Aerospace personnel operation quality evaluation method based on multi-modal data fusion
Through multimodal data fusion technology, the EEG signals, operation records and error feedback data of astronauts are obtained, and feature extraction and fusion are performed using a neural network model. This solves the problem of insufficient accuracy in the assessment of astronaut operation quality in traditional evaluation methods and realizes scientific and real-time evaluation of space operations.
Patent Information
- Application Number
- CN202510917266.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional methods for evaluating the operational quality of aerospace personnel rely on single-dimensional data, ignoring the fatigue and psychological stress factors of aerospace personnel under long-term and complex tasks, and lack the ability to identify dynamic behaviors, resulting in low evaluation accuracy.
A multimodal data fusion method is used to obtain EEG signal data, operation record data and error feedback data. Features are extracted through long short-term memory networks, bidirectional gated recurrent unit networks and multi-head self-attention mechanisms, and fused to generate shared features. Dynamic adjustments are made based on the forward neural network to construct a target evaluation model.
It has achieved scientific and accurate evaluation of space operations, improved the ability to identify and evaluate the operational quality of space personnel, and provided a scientific basis for evaluation.
Smart Images

Figure CN120706993A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of aerospace operation technology, and in particular to a method for evaluating the operation quality of aerospace personnel based on multimodal data fusion. Background Art
[0002] Space operations refer to the operations performed by astronauts on spacecraft and related equipment. Human reliability is a comprehensive measure of the operational accuracy, stability, responsiveness, and adaptability of astronauts when performing various space missions. Human reliability is usually assessed through the operational quality of astronauts. The operational quality of astronauts refers to the quality of operators in space systems. When astronauts perform complex tasks, their operational quality is affected by many factors, including the operator's degree of distraction, fatigue, psychological state, and ability to respond to emergencies. Because space missions typically require high precision, complex operational processes, and irreversible operational consequences, the operational quality of astronauts is directly related to the safety and success rate of space missions.
[0003] Traditional methods for evaluating aerospace crew performance rely primarily on single-dimensional data from the spacecraft, such as feedback on spacecraft operational errors or mission logs. These methods ignore the impact of factors such as fatigue and psychological stress on crew performance during long, complex missions. Furthermore, most evaluation models employ static analysis methods, lacking the ability to process time series data and identify dynamic behavioral changes. Consequently, most evaluation techniques focus on post-mission summaries and assessments, lacking feedback on the dynamic behavior of aerospace crew members.
[0004] In summary, there is currently a lack of evaluation methods for aerospace personnel's operational quality, and the accuracy and scientificity of the evaluation need to be improved. Summary of the Invention
[0005] The purpose of this application is to provide a method for evaluating the operational quality of aerospace personnel based on multimodal data fusion, in order to solve the problem of low accuracy in the quality evaluation of aerospace operators.
[0006] To achieve the above objectives, the present application provides, in a first aspect, a method for evaluating the operational quality of aerospace personnel based on multimodal data fusion, comprising: Acquiring multiple modal training data relevant to spaceflight operations, the training data including EEG signal data, operation record data, and error feedback data; Performing a data normalization operation on the training data of multiple modalities to obtain a plurality of time-aligned single-modal data; Based on a long short-term memory network, a bidirectional gated recurrent unit network, and a multi-head self-attention mechanism, feature extraction is performed on the plurality of unimodal data to obtain a plurality of unimodal features; Fusing a plurality of the unimodal features to obtain a cross-modal fusion feature, and generating a shared feature based on the unimodal features and the cross-modal fusion feature; Inputting the shared features into an initial forward neural network, learning the attention weights of the shared features according to the task objectives, calculating the loss values of the shared features under different attention weights, and dynamically adjusting the weight coefficients of the training data of multiple modalities based on the loss values to obtain a target forward neural network; The online data of the space operation is input into the target forward neural network to obtain an evaluation index of the space operation.
[0007] A second aspect of the present application provides a computer-readable storage medium, in which a program is stored. The program can be loaded by a processor and execute the above-mentioned method for evaluating the operational quality of aerospace personnel based on multimodal data fusion.
[0008] The beneficial effects of this application are: This application first acquires training data from multiple modalities relevant to space operations, encompassing the multi-dimensional information of astronauts performing space operations. Data normalization aligns the data from different modalities, facilitating subsequent analysis and processing, and improving data availability and accuracy. Then, a long-short-term memory network is used to capture the long-term dependencies of the training data. A bidirectional gated recurrent unit network extracts features from both forward and reverse directions to obtain more comprehensive temporal information. A multi-head self-attention mechanism is used to mine the interactions between the training data, capturing richer features from the training data. This allows for feature extraction from multiple unimodal data, resulting in multiple unimodal features. Next, the weighted results of bidirectional attention are fused to generate shared features. An initial forward neural network is then trained based on the shared features. Attention weights for the shared features are learned according to the mission objectives. The loss values for the shared features at different attention weights are calculated. Based on these loss values, the weight coefficients of the training data from multiple modalities are dynamically adjusted to construct a target forward neural network. Finally, the online space operation data is input into the trained target forward neural network to obtain real-time evaluation metrics for space operations, enabling a more scientific and accurate assessment of the operational quality of astronauts.
[0009] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A flowchart of a method for evaluating aerospace personnel operation quality based on multimodal data fusion provided in an embodiment of the present application is provided; Figure 2A flow chart of a feature extraction method provided in an embodiment of the present application; Figure 3 This is a flowchart of a method for evaluating the operational quality of aerospace personnel based on multimodal data fusion provided in a specific embodiment of the present application. DETAILED DESCRIPTION
[0011] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0012] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and are not to be construed as indicating or implying relative importance or implicitly specifying the number of the technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of this application, "plurality" means two or more, unless otherwise specifically qualified. In this application, the word "exemplary" is used to mean "serving as an example, illustration, or illustration." Any embodiment described in this application as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is provided to enable anyone skilled in the art to implement and use the present application. In the following description, details are listed for illustrative purposes. It should be understood that one of ordinary skill in the art will recognize that the present application can be implemented without these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0013] Figure 1 This is a flow chart of a method for evaluating the operational quality of aerospace personnel based on multimodal data fusion provided in an embodiment of the present application. Figure 1 As shown, the method may include steps 101-106, which are described in detail below.
[0014] Step 101: Acquire multiple modal training data related to space operations, where the training data includes EEG signal data, operation record data, and error feedback data.
[0015] In related technologies, the evaluation of aerospace personnel's operational quality is usually based on a single dimension, which easily overlooks the impact that factors such as fatigue and psychological stress caused by long-term and complex missions may have on aerospace operations. Based on this, the embodiments of the present application integrate and analyze training data from multiple modalities through multimodal fusion technology, and perform inference based on training data from multiple modalities, thereby compensating for the shortcomings of single-modal data and improving the ability to recognize complex behaviors and the accuracy of decision-making.
[0016] In the embodiment of the present application, the training data is the data required for training the model for evaluating the operational quality of astronauts. As an example, the training data can be selected from historical data of astronauts performing space operations, for example, historical data within one month.
[0017] Multimodal training data refers to training data of different types, each of which can reflect a factor that influences a spacecraft's operational behavior. For example, EEG data can reflect a spacecraft's mental state and cognitive load. Operation log data can include detailed operational behavior and process information. Error feedback data can reveal errors made by spacecraft during operations.
[0018] Multimodal training data can cover multiple aspects of information such as astronauts' physiology, behavior, and errors. Compared with single-modal data, it can more comprehensively analyze the comprehensive status of astronauts, explore the potential relationship between the astronauts' own status and operational quality, and lay a data foundation for evaluating the operational quality of astronauts.
[0019] Step 102: Perform data standardization on the training data of multiple modalities to obtain multiple single-modal data that are time-aligned.
[0020] To eliminate differences in training data between different modalities and the format of the same training data, and to ensure comparability and fusion of training data from different modalities, multiple training data can be standardized. For example, EEG signal data may record real-time neural activity using high-frequency sampling, with dense and continuous time points. Operational recording data, on the other hand, may be event-triggered, with sparse and discontinuous time points. Error feedback data may be manually annotated through retrospective analysis, and time tags may have lags or errors. Therefore, it is necessary to time-align the training data of multiple modalities to obtain standardized single-modal data.
[0021] By standardizing multiple training data sets, we can eliminate differences in dimensions and scales, facilitating unified processing. Time alignment ensures that training data across modalities correspond in time, improving data quality and standardizing and unifying data from diverse sources and characteristics. This enhances the data's usability and accuracy in subsequent feature extraction and model training, reducing errors and mistakes caused by data discrepancies.
[0022] Step 103: Based on a long short-term memory (LSTM) network, a bidirectional gated recurrent unit (BiGRU) network, and a multi-head self-attention (MHSA) mechanism, feature extraction is performed on multiple unimodal data to obtain multiple unimodal features.
[0023] LSTM networks can process long sequences of data through a gating mechanism, resolving gradient issues and effectively capturing long-term dependencies in training data. BiGRU networks can process sequences in both forward and reverse directions, acquiring more comprehensive temporal information about training data. Multi-head self-attention mechanisms can calculate correlations across multiple subspaces, capturing multi-level relationships in training data.
[0024] Therefore, the embodiment of the present application uses an LSTM network to extract local information of unimodal data and maps the input dimension to the same dimension to obtain the unimodal time series features of the unimodal data. Then, the unimodal time series features obtained after processing by the LSTM network are input into the BiGRU network to extract the bidirectional hidden state time series features of the unimodal data. Then, a multi-head self-attention network mechanism is used to extract the information connection between the unimodal time series features, and finally the unimodal features of each unimodal data are obtained.
[0025] The comprehensive use of LSTM networks, BiGRU networks and multi-head self-attention mechanisms can give full play to the advantages of each model in processing time series data and capturing sequence relationships. It can deeply explore the characteristics of single-modal data from different angles and levels, obtain richer and more effective feature representations, and provide high-quality features for subsequent feature fusion and model training.
[0026] Step 104: fuse multiple unimodal features to obtain cross-modal fusion features, and generate shared features based on the unimodal features and the cross-modal fusion features.
[0027] EEG data can reflect a space crew member's cognitive state, but it cannot directly link them to specific operational behaviors. Operation log data can record the precise timing of operational actions, but it struggles to explain the neural substrates underlying these behaviors. Error feedback data can indicate the correctness of operational results, but lacks procedural data support. Therefore, it is necessary to fuse unimodal features extracted from multiple unimodal data sets, link the causal chains between training data from multiple modalities, and construct a complete analytical closed loop to reduce misjudgments caused by single modalities.
[0028] In an embodiment of the present application, the feature obtained by fusing multiple single-modal features is a cross-modal fusion feature, and the cross-modal fusion feature is not limited to the fusion between two modalities and multiple modalities. The feature composed of single-modal features and cross-modal features is a shared feature. In this way, the features extracted from the training data include both single-modal features reflecting the features of each modality and cross-modal features containing the correlation relationship between at least two modalities, providing richer data support for model training. Through fusion, the feature advantages of different modalities can be integrated to achieve information complementarity, so that shared features can more completely and accurately describe the complex situation of aerospace operations, which helps to more accurately analyze the complex relationships and potential problems in the aerospace operation process.
[0029] Step 105: Input the shared features into the initial forward neural network, learn the attention weights of the shared features according to the task objectives, calculate the loss values of the shared features under different attention weights, and dynamically adjust the weight coefficients of the training data of multiple modalities based on the loss values to obtain the target forward neural network.
[0030] Since the evaluation of the operational quality of aerospace personnel involves a complex mapping of multimodal features to reliability indicators, the mapping is usually nonlinear and highly complex. Based on this, the embodiment of the present application selects a feedforward neural network, which realizes automatic learning of high-order interactions between multimodal features through multi-layer nonlinear transformations, and can fit arbitrarily complex functions. Compared with models such as recurrent neural networks, feedforward neural networks have a non-circular structure, fast forward propagation speed, and higher computational efficiency, making them suitable for implementing the evaluation of the operational quality of aerospace personnel. At the same time, feedforward neural networks reduce the common gradient vanishing and explosion problems in recurrent neural networks, and the training process is more controllable. In addition, multi-task learning and dynamic weight adjustment can be simultaneously met.
[0031] In the embodiments of this application, the initial feedforward neural network is an untrained feedforward neural network. The task objective is the specific evaluation task of the feedforward neural network and the optimization direction set to achieve this task. For example, the skill mastery of astronauts, the error correction efficiency of astronauts, etc. This is essentially a goal-oriented feature selection. Through the task-driven optimization process, it is possible to determine which modal features are more critical to the specific evaluation task, thereby achieving accurate modeling of the evaluation of astronaut operational quality.
[0032] By inputting the shared features into the initial forward neural network, learning attention weights based on the mission objectives, calculating the loss values under different attention weights, and then dynamically adjusting the weight coefficients of the training data for each modality based on the loss values, a trained target forward neural network can be obtained. By training the forward neural network, the target forward neural network can focus on important features and adaptively adjust its emphasis on different modal data, improving the relevance and accuracy of model training, optimizing the performance of the target neural network, and thus improving its adaptability and generalization capabilities to different space operation scenarios.
[0033] Step 106: Input the online data of the space operation into the target forward neural network to obtain the evaluation index of the space operation.
[0034] In the embodiments of the present application, online data refers to operational data related to aerospace operations collected in real time. By inputting the online data of aerospace operations into the trained target forward neural network in real time, an evaluation index for the current aerospace operation can be obtained. This evaluation index is based on the results output after processing and training in the above steps, and can be used to evaluate the quality and reliability of aerospace operations, and provide a scientific and objective basis for operation evaluation, personnel training, process optimization, etc. in the aerospace field. For example, the weak links of aerospace personnel can be discovered through evaluation indicators, and targeted training can be carried out. For another example, the rationality of the operating process can be evaluated to provide a reference for process improvement, thereby improving the safety and efficiency of aerospace operations, which has significant application value.
[0035] In summary, the present embodiment first acquires training data from multiple modalities relevant to spaceflight operations, encompassing the multi-dimensional information of spaceflight personnel involved in spaceflight operations. Data standardization aligns the data from different modalities, facilitating subsequent analysis and processing, and improving data availability and accuracy. Then, a long-short-term memory network is used to capture the long-term dependencies of the training data. A bidirectional gated recurrent unit network extracts features from both positive and negative directions to obtain more comprehensive temporal information. Furthermore, a multi-head self-attention mechanism is used to mine the interactions between the training data, capturing richer features from the training data. This allows for feature extraction from multiple unimodal data, resulting in multiple unimodal features. Next, the weighted results of the bidirectional attention are fused to generate shared features. An initial forward neural network is then trained based on the shared features. Attention weights for the shared features are learned according to the mission objectives. The loss values for the shared features at different attention weights are calculated. Based on the loss values, the weight coefficients of the training data from multiple modalities are dynamically adjusted to construct a target forward neural network. Finally, the online spaceflight data is input into the trained target forward neural network to obtain real-time evaluation metrics for spaceflight operations, enabling a more scientific and accurate assessment of spaceflight operations.
[0036] In step 101, the multimodal training data may include EEG signal data , operation record data , and error feedback data .
[0037] For EEG signal data Initial EEG signal data reflecting attention levels during spaceflight operations can be collected using an EEG sensor, such as a ThinkGear Access Module (TGAM). Because the initial EEG signal data has a low amplitude, it can be amplified using an amplifier and then converted to a digital signal using an analog-to-digital converter. Principal Component Analysis (PCA) can then be performed on the collected EEG. PCA is a dimensionality reduction technique that reconstructs the data by identifying the principal components (i.e., directions) in the data. This determines the contribution weights of the digital signal and then reduces the dimensionality of the initial EEG signal data based on the contribution weights to obtain the EEG data. This transforms a set of potentially correlated variables into a set of linearly uncorrelated variables, making the data more concise in the new coordinate system and making the principal components independent of each other.
[0038] Specifically defined as: set up for -dimensional random variable, mean vector , the covariance matrix . No. principal components satisfy: (1) ; (2) When hour, ; (3) .
[0039] Will After the linear space is rotated, a new space is obtained, and the subspace in the new space has the largest sample variance. Specifically, is through the rotation matrix The calculated principal components, is with The corresponding eigenvector. The variance of the principal components By solving the covariance matrix Therefore, we can use the The proportion of the principal component variance in all variances is used to define its contribution weight in the total variance. : ; in, For the The eigenvalues corresponding to the principal components are From 1 to The eigenvalues corresponding to the principal components are summed.
[0040] By processing EEG signal data through PCA, the original high-dimensional data can be approximated by fewer principal components while minimizing information loss, thereby achieving data dimensionality reduction.
[0041] Operation record data in the embodiment of the present application It can include three types of indicators, namely test plan requirement coverage , Test plan completion on time and test requirement execution completeness ,Right now The test plan requirement coverage is used to evaluate the degree to which the operation training test activities cover the actual requirements. The test plan completion on time is used to evaluate the time management ability of the operation training personnel. The test requirement execution completeness is used to evaluate the quality and effect of the operator's execution of the training test activities. , Test plan completion on time and test requirement execution completeness , you can get the operation record data.
[0042] In an example, the test plan requirement coverage can be calculated based on the number of requirements covered by the actual test of aerospace operations and the total number of requirements, that is, = Number of requirements actually covered by testing / Total number of requirements. This metric covers two types of data required for spaceflight operations: all operational requirements required during spaceflight operations (including but not limited to mission launch, equipment commissioning, system checks, etc.) and the number of operational requirements actually completed and successfully executed by spaceflight crews. Specifically, the following steps are performed: List all required operations by spaceflight crews based on the spacecraft operating manual, mission planning documents, or system operating specifications; query mission logs to record the operations performed by spaceflight crews and mark which operations have been completed; and calculate the ratio of the number of operational requirements actually completed to the total number of operational requirements.
[0043] In one example, the on-time completion rate of the test plan can be calculated based on the number of tests completed on time and the total number of tests for space operations, that is, = Number of tests completed on time / Total number of tests. Specifically, you can first list all tasks, clearly defining the steps and scheduled time for each task. Then, by querying the task log to record the actual completion time of each task, you can check whether it was completed on schedule. Finally, calculate the ratio of the number of tasks completed on schedule to the total number of tasks.
[0044] In one example, the test requirement execution completeness can be calculated based on the number of operations that meet the operation standards and the total number of operations of the aerospace operation, that is, = Number of operations that meet the operational standards / Total number of operations. Specifically, all operations that require crew members to perform can be listed based on spacecraft operating specifications and mission requirements. Each operation can be checked for compliance with the standard requirements using the crew member's status monitoring system and mission feedback. The ratio of successfully completed and standard-compliant operations to the total number of operations can be calculated.
[0045] In the embodiment of the present application, the error feedback data can be calculated based on the error operation, limit-exceeding operation and total number of operations of the aerospace operation, that is, = (Error Operations + Out-of-Limit Operations) / Total Number of Operations. Specifically, determine the total number of operations that astronauts must perform. Use the space operation monitoring system to record the astronauts' operation logs to determine whether errors (such as incorrect operations or incorrect input) or out-of-limit operations (such as timeouts or overloads) occur. Record the number of error operations and out-of-limit operations. Add the number of error operations and out-of-limit operations together and compare it with the total number of operations to determine the error feedback rate.
[0046] Because EEG signal data may contain various types of noise during acquisition, such as blinking, the amplitude of EEG signal data is relatively small, while the signal amplitude caused by blinking varies greatly. The larger the blinking movement, the larger the amplitude, which appears as a large peak. Therefore, the collected data packets need to be processed for artifact removal and verification. If the verification number does not match, the packet is directly ignored and the data verification failure is considered packet loss.
[0047] Since the target detection signal is two-dimensional, the EEG signal decomposition method of multivariate empirical mode decomposition (MEMD) is introduced to simultaneously decompose multi-channel data to ensure that the number of intrinsic mode function (IMF) components obtained from the decomposition of each channel is the same and the corresponding order frequency components are consistent. This overcomes the mode calibration problem of traditional methods such as empirical mode decomposition (EMD) when processing multivariate data, and realizes the synchronous analysis of multi-channel signals.
[0048] The MEMD method was first used in dimensional space to establish a uniformly distributed projection vector set, then calculate the projection envelope of the signal in each direction, and finally define the local mean function of the multidimensional signal by calculating the envelope mean. Dimensional input signal , the specific process of the MEMD algorithm is: 1. Set substitution variables , assuming that the number of MIMF layers of the current multidimensional pattern model is ; 2. Search on the sphere Evenly distributed points, find the projection ,in ; 3. Calculate the projection ,in ; 4. For all , by taking the maximum value to get the local maximum , that is, calculate all the moments of projection ; 5. Using cubic spline interpolation, Interpolate each dimension to get a multidimensional envelope , and calculate The mean of the dimension envelope ; 6. Subtract the mean from the original signal to get the current model signal ; 7. If If the end condition is met, the current MIMF transmission is stopped and the current MIMF amount is obtained. , go to the next step. Otherwise, set , and continue to the next step; 8. From the current signal Get the amount of the previous MIMF component If the current signal If the decomposition condition (such as monotonic function or other user-defined conditions) is met, the remaining components are transmitted; otherwise, set , and start calculating the next layer of MIMF; 9. Final output model results, including MIMF components and the remaining trend items ,Right now .
[0049] The correctly parsed data contains the EEG signal features obtained by the TGAM module using its own embedded algorithm , including Attention Meditation , ranging from 1 to 100, with upset, agitated, or abnormal states indicated by lower values.
[0050] After preprocessing the EEG signal, data standardization of each modality is an important step after data preprocessing. Standardization can eliminate the dimensionality effect, accelerate the convergence speed, improve the model performance and avoid numerical instability, etc., to prepare for subsequent model training and testing. For example, after standardizing the EEG signal data, we can get . EEG signal after preprocessing and normalization , spacecraft operation records , Error Feedback Is the length of Each modality corresponds to a real number evaluation representing the emotional state.
[0051] For multimodal data, various modalities have different time point lengths, sampling rates, or time delays. In order to effectively fuse multimodal data, in one example, dynamic time warping (DTW) and interpolation methods can be introduced.
[0052] Specifically, in step 102, for each single modal data ,in , we can first extract the original timestamp sequence of the training data of each modality separately , and align multiple original timestamp sequences to the same starting time , and obtain multiple normalized time series ,The original timestamp sequence includes training data corresponding to multiple time points.
[0053] Then, the normalized time series with the widest time coverage among multiple normalized time series (assuming it is the normalized time series of modality A) is selected as the alignment reference sequence, and the time axis of the alignment reference sequence is used as the reference time axis. The training data included in the normalized time series to be aligned is mapped to the reference time axis, where the reference time axis includes multiple reference time points and can be defined as: Then, the training data corresponding to other single modalities B and C are mapped to the reference time axis superior.
[0054] Next, for any mode to be aligned The normalized time series is constructed by constructing a distance metric function: , calculate the Euclidean distance of the normalized time series to be aligned at each time point, and construct the cost matrix based on the Euclidean distance , and calculate the minimum cumulative distance of the normalized time series to be aligned at each time point. The cumulative distance is recursively defined as follows: .
[0055] By backtracking, starting from the last element of the cost matrix, tracing back in the direction corresponding to the minimum cumulative distance, we can obtain the target alignment path, i.e. the shortest alignment path. , to align the training data of the normalized time series to be aligned with the training data of the reference time axis, that is, each pair Indicates modality The The time point and the reference time axis The time points are relatively aligned.
[0056] If there is a reference time point in the target alignment path If there is no corresponding normalized time series training data to be aligned, the linear interpolation method is used to reconstruct the training data corresponding to the reference time point, that is: .
[0057] The above alignment method is performed on the data of each modality. Finally, the training data corresponding to each time point of the normalized time series of multiple modalities are spliced to obtain multiple aligned single-modal data. The reconstruction on the unified reference timeline is represented as: .
[0058] All modal data at each time point Splicing to form a joint feature representation after multimodal alignment : .
[0059] Output uniformly aligned multimodal time series data , as the input for subsequent single-modal modeling and multi-modal fusion.
[0060] EEG signals after preprocessing and normalization , operation record data , error feedback data Is the length of The time series of , each modal data corresponds to a real number evaluation representing an emotional state. In order to mine the unique features within a single modality, in step 103, multiple single modal data can be first input into the LSTM network to obtain multiple single modal time series features. Then, the multiple single modal time series features are input into the BiGRU network to obtain the bidirectional hidden state time series features corresponding to multiple time points of each single modal time series feature. Next, each bidirectional hidden state time series feature is input into the multi-head self-attention network to obtain the single modal features of each single modal data.
[0061] In the embodiments of the present application, the LSTM network may include a forget gate, an input gate, and an output gate. An LSTM network is a special type of recurrent neural network that uses three gating mechanisms to control the flow of information while maintaining a long-term memory unit (state). This solves the vanishing and exploding gradient problems of ordinary recurrent neural networks when processing long sequences.
[0062] Specifically, each unimodal data is input into the forget gate respectively to determine the forgetting ratio of the unimodal data from the previous time point to the current time point.
[0063] The decision gate determines how much of the cell state from the previous time point's unimodal data is forgotten .
[0064] ; in, Controls the proportion of forgetting (close to 1 means retention, close to 0 means forgetting).
[0065] The input gate determines how much new information is written to the state at the current time point. By feeding unimodal data with a value greater than the set forgetting ratio into the input gate, we can determine the write ratio of new unimodal data at the current time point and generate candidate states for new unimodal data based on the write ratio.
[0066] For example, first calculate the activation value of the input gate Determine the writing ratio of new unimodal data and then generate new unimodal data candidate states.
[0067] .
[0068] Next, the state of the unimodal data at the current time point is updated based on the forgetting ratio and the candidate state, namely: .
[0069] The output gate determines which parts of the state are used to output the hidden state at the current time point. By inputting the state of the unimodal data at the current time point into the output gate, the hidden state of the unimodal data at the current time point can be generated, that is: .
[0070] Finally, based on the hidden state of the unimodal data, the unimodal time series features reflecting the long-term dependency between the unimodal data are generated, namely: .
[0071] In this embodiment, the BiGRU network includes a forward gated recurrent unit network and a backward gated recurrent unit network. The Gated Recurrent Unit (GRU) neural network is an improvement on the recurrent neural network. It uses two gating mechanisms (a reset gate and an update gate) to address the vanishing gradient problem and accelerate training. The features processed by the LSTM network are input into the BiGRU network. By continuously updating the hidden state, the corresponding bidirectional hidden state at each moment is extracted as a new feature.
[0072] The calculation process of GRU is as follows: 1. Update Gate : Control the current hidden state From the previous point in time How much information to inherit: .
[0073] 2. Reset the door : Control the previous time point The information in the current time point Impact: .
[0074] 3. Candidate hidden states :Based on the current input and the previous hidden state after reset Compute candidate states: .
[0075] 4. Update hidden state :According to the update gate Balance the contribution of the previous hidden state and the candidate hidden state: .
[0076] In the BiGRU network, the hidden states of the forward GRU and the reverse GRU are calculated separately. The calculation process is as follows: 1. Set the hidden state dimension: Set the hidden state dimension to , extract the bidirectional hidden state corresponding to each moment as a new feature.
[0077] 2. Input multiple unimodal time series features into the forward GRU neural network and calculate the forward hidden state of the unimodal time series features in order from front to back: ;in, is the current input, is the forward hidden state.
[0078] 3. Input multiple unimodal time series features into the reverse GRU neural network and calculate the reverse hidden state of the unimodal time series features in reverse order from back to front: ;in, is the inverted hidden state.
[0079] 4. Concatenate the forward hidden state and the reverse hidden state to obtain a bidirectional hidden state temporal feature including multiple time points. The output is the concatenation of the forward hidden state and the reverse hidden state: .
[0080] The final output for the entire sequence is the set of hidden states at all time steps: ; The shape of the data after BiGRU processing is .
[0081] In the embodiments of this application, the multi-head self-attention mechanism is an important component of the Transformer model. Compared to the single-head self-attention mechanism, it uses multiple attention heads working in parallel to more comprehensively capture the multi-level relationships between different positions in the input sequence. The multi-head self-attention mechanism is introduced at the output of the BiGRU to extract richer temporal connections by calculating the similarity between the query vector and the index vector in different subspaces.
[0082] Specifically, the input sequence is represented as: the bidirectional hidden state temporal characteristics of the input sequence are ,in , is the first elements, is the sequence length, Is the dimension of the input vector. For each bidirectional hidden state time series feature, three sets of weight matrices are used to transform the bidirectional hidden state time series feature Mapped to query vector Query , key vector Key Sum value vector Value .in, , , , It is the vector dimension of Query and Key in each header.
[0083] For each attention head , the correlation between the query vector Query and the key vector Key is calculated by dot product, and the attention scores of multiple attention heads are calculated in parallel and scaled, which is expressed as: .
[0084] in, is the number of attention heads, , , The corresponding Query 、Key 、 Value Mapping matrix, mapping the original data to different low-dimensional spaces; The calculation is the similarity between Query and Key. is a scaling factor to avoid gradient instability caused by large inner product values.
[0085] Next, the attention scores of multiple attention heads are concatenated to obtain the concatenated output: .
[0086] Finally, layer normalization is used to avoid gradient explosion caused by excessive values. , the concatenated output is mapped to the original dimension through the fully connected layer to obtain the unimodal features: .
[0087] In step 104, the scores of the attention heads between any two unimodal features can be calculated separately, and the scores of the attention heads between any two unimodal features can be spliced together to obtain a unidirectional fusion feature and a reverse unidirectional fusion feature between any two unimodal features. Then, the unidirectional fusion feature and the reverse unidirectional fusion feature are spliced together to obtain a cross-modal fusion feature between any two unimodal features. Next, average pooling is used to integrate the unimodal features and cross-modal fusion features at multiple time points, and the unimodal features are projected into a space of the same dimension as the cross-modal fusion features through linear mapping. Finally, multiple unimodal features and multiple cross-modal fusion features are spliced together to obtain a shared feature.
[0088] The multimodal Transformer model can achieve efficient fusion of temporal multimodal data through the attention mechanism. The model uses the data of each time point of modality A as a reference, calculates its similarity with the data of all time points of modality B, and embeds the relevant information in modality B into modality A, thereby completing the one-way feature integration from modality A to modality B (A→B). This method is highly flexible in processing sequences of different lengths and can maintain good fusion effects even when the sequences are not aligned. , spacecraft operation records , Error Feedback Perform pairwise fusion to use EEG signals Spacecraft Operation Records Taking fusion as an example, the calculation process is as follows.
[0089] 1. Calculation from EEG signal data To operation record data Each attention head during fusion is: .
[0090] 2. Concatenate all attention heads to obtain unidirectional fusion features: .
[0091] 3. Calculate and record data To EEG signal data Fusion features: .
[0092] 4. Concatenate the bidirectional fusion results to obtain a complete cross-modal fusion result: .
[0093] 5. Use multi-head self-attention to further extract features and obtain EEG signal data and operation record data The final result of fusion.
[0094] .
[0095] Finally, three unimodal feature representations are obtained 、 、 And three cross-modal fusion features 、 、 They are all two-dimensional matrices, using average pooling to integrate data at all moments, and using linear mapping to project unimodal features into a space of the same dimension as the cross-modal features, and splice them into complete shared features.
[0096] Figure 2 Schematic diagram of a feature extraction method provided in an embodiment of the present application. Figure 2 As shown in the figure, the preprocessed data is sequentially fed into an LSTM network, a BiGRU network, and a self-attention mechanism to generate unimodal features. The unimodal features are then fused to generate cross-modal fusion features. Finally, shared features are generated based on the unimodal and cross-modal fusion features.
[0097] In step 105, the shared features can be input into the initial forward neural network, and the weight coefficients of the training data of multiple modalities are set as initial weight values. At the beginning of training, the weight coefficients of multiple modalities are initialized to 1, that is, the average weight. The initial forward neural network is then iteratively trained based on the initial weight values, and the loss value corresponding to the training data of each modality is calculated separately. Then, based on the loss value, convergence status, and evaluation index corresponding to the training data of each modality, the initial weight value of the training data of each modality is adjusted according to the preset adjustment rules, and the total loss value of each round of training is calculated. When the total loss value is less than the set loss value or the number of iterations reaches the set number, the training is stopped, and the trained initial forward neural network is used as the target forward neural network.
[0098] Specifically, EEG signal data (i.e., the EEG signal in the figure) , Operation record data (i.e., operation record in the figure) and error feedback data (i.e., error feedback during the process) They can all deliver a certain level of operation, but their contribution varies when performing different tasks, and the importance of each modality or feature for different task goals is also different. Therefore, it is necessary to measure the importance of each modality information according to the task goal. Based on this, the embodiment of the application proposes a dynamic weight allocation method that combines task goals and modality features. , operation record data and error feedback data Together they work to comprehensively evaluate the quality of the operation. These tasks can be viewed as subtasks in multi-task learning, and the corresponding task loss function is expressed as ,in Indicates the loss of EEG-based operational quality assessment; Indicates the loss of operational quality assessment based on operational record data (such as task completion, time management, etc.); Represents the evaluation loss based on error feedback data (such as wrong operation, out-of-limit operation, etc.). The overall operation quality evaluation loss function It consists of the weighted loss of each task: ; in, is the dynamic weight coefficient of each task, which indicates the importance of each task in the evaluation process.
[0099] In order to adapt to the dynamic requirements of tasks and the training progress of each task, the following dynamic adjustment strategies are used to update the weight coefficients in real time: (1) Dynamic weight adjustment based on task loss: During the training process, the loss values of different modalities will change as the learning progress of the task. For tasks with large losses (such as when the evaluation error of EEG signal data is large), the weight of the task can be increased so that the model can pay more attention to the learning of the task. Conversely, for tasks with small losses, their weight can be reduced. The adjustment method is as follows: ; in, It's a task The corresponding loss, It is a hyperparameter that controls the degree of influence of loss on weight adjustment.
[0100] (2) Weight adjustment based on task convergence: For tasks that have converged relatively quickly (such as tasks where the loss of operation record data is gradually decreasing), their weights can be reduced to avoid over-optimizing the task and affecting the learning of other tasks. For tasks that converge more slowly (such as tasks that evaluate error feedback data), their weights can be increased to promote faster convergence. This adjustment can be achieved by calculating the rate of change of task loss: .
[0101] (3) Adjustment based on task evaluation indicators: In some cases, the modality of EEG signal data may provide higher accuracy, while other modalities (such as error feedback data) may provide higher robustness. By monitoring the evaluation indicators of each task and dynamically adjusting the weight coefficient according to the performance of the task in each iteration. For example, if the contribution of EEG signal data to the operation quality is significantly improved, its weight can be increased accordingly. Specifically: ; in, is a hyperparameter that controls the sensitivity of the evaluation score to weight adjustments. It's a task evaluation score.
[0102] In one example, the steps for adaptively assigning weight coefficients are: 1. At the beginning of training, the initial weight coefficients of all tasks are set to 1, that is, ; 2. Calculate the loss value for each task (such as EEG signals, spacecraft operation records, error feedback) ; 3. Using the above dynamic adjustment rules, adjust the corresponding weight coefficients based on the loss, convergence and evaluation indicators of each task ; 4. In each round of training, the total loss is calculated based on the dynamically adjusted weights ; 5. Stop when the task loss converges or reaches the preset training rounds.
[0103] The evaluation of aerospace personnel's operational quality is a regression task. The number of neurons in the output layer is 1. The Mean Absolute Error (MAE) is used as the loss function according to the specific training project. The training of the entire model is based on the sum of the losses of each task: ; in, To adjust the hyperparameters of different task training levels, the larger the parameter value, the faster the corresponding task converges. Finally, the network outputs the operation quality score of the astronauts. .
[0104] The following is an example of a method for evaluating the operational quality of aerospace personnel based on multimodal data fusion according to an embodiment of the present application, using a specific embodiment. Figure 3 This is a flow chart of a method for evaluating the operational quality of aerospace personnel based on multimodal data fusion provided in a specific embodiment of this application. Figure 3 As shown, in a specific embodiment, the method includes steps 301-307.
[0105] Step 301: Data collection. Obtain EEG signal data , operation record data , error feedback data As input data, the operation record data contains three types of indicators .
[0106] Step 302: Data preprocessing. Preprocess the input data. The correctly parsed data contains EEG signal data obtained by the TGAM module using its own embedded algorithm. .
[0107] Step 303: Data standardization. After standardization, we get , the three types of unimodal data after processing are .
[0108] Step 304: LSTM+BiGRU network captures the internal temporal information of the unimodal data. The three unimodal data are input into the LSTM network, and the flow of information is controlled by three gating mechanisms. At the same time, a long-term memory unit is maintained to obtain the processed input features. , is the sequence length, Represents the feature dimension. The features processed by LSTM are input into BiGRU, and the hidden state is continuously updated by calculating the forward GRU and reverse GRU using two gating mechanisms. The bidirectional hidden state corresponding to each moment is extracted as a new feature. The output of the entire sequence is the set of hidden states of all time steps. , the shape of the data after BiGRU processing is A multi-head self-attention mechanism is introduced on the output of BiGRU. By calculating the similarity between the query vector and the index vector in different subspaces, richer temporal connections can be extracted. The correlation between the query and the key of each attention head is calculated separately, and the attention scores of multiple attention heads are calculated in parallel. The results are scaled to obtain the attention scores: .
[0109] Then concatenate the outputs of all attention heads to get the complete output: .
[0110] Finally, a fully connected layer is used to map it back to the original dimension to obtain the unimodal features of the final sequence: .
[0111] Step 305: The multimodal Transformer network fuses multiple single-modal features. , operation record data , error feedback data Perform pairwise fusion to use EEG signal data and operation record data Take fusion as an example: Brain Telecommunications Data To spacecraft operation record data Each attention head during fusion is: .
[0112] Concatenate all attention heads to get unidirectional fusion features: .
[0113] Similarly, get the operation record data To EEG signal data Fusion features: , stitching the bidirectional fusion results: , use multi-head self-attention to further extract features .
[0114] Finally, three unimodal feature representations are obtained 、 、 And three cross-modal fusion features 、 、 They are all two-dimensional matrices. Average pooling is used to integrate data at all moments, and linear mapping is used to project unimodal features into a space of the same dimension as the cross-modal features, and spliced into complete shared features: .
[0115] Step 306: The feedforward neural network adaptively assigns weight coefficients. The feedforward neural network is used to learn the attention weights of each feature representation according to the task objectives: .in, and are the weight parameters of the two fully connected layers, and is the corresponding bias.
[0116] The attention weights corresponding to each feature representation constitute a task-specific attention vector: .
[0117] The evaluation operation is a regression task. The number of neurons in the output layer is 1. The mean absolute error (MAE) is used as the loss function according to the specific training project. The training of the entire model is based on the sum of the losses of each task. ; in, To adjust the hyperparameters of different task training levels, the larger the parameter value, the faster the corresponding task converges.
[0118] Step 307: Real-time scoring. First, train the model. During the training process, use the Adam optimizer, set the learning rate to 1E-4, the number of batch training samples to 256, the number of BiGRU hidden neurons to 100, and set the hyperparameters on the loss function. All are 1. After the model is trained, online data is obtained to calculate the operation quality score of the aerospace personnel. .
[0119] In the embodiment of the present application, EEG signal data, operation record data, and error feedback are used as input data, which is more scientific and accurate than the traditional use of single-modal data for evaluation. The single-modal extraction network based on long short-term memory network and bidirectional gated recurrent unit network can automatically extract features from the time series data of EEG signal data, operation record data, and error feedback data, which is more effective than the static analysis method used in traditional evaluation models. Traditional evaluation techniques are mostly summarized and evaluated after the mission is completed, while the evaluation method of the embodiment of the present application can provide real-time feedback during the training of aerospace personnel, improving the efficiency of evaluating the operational quality of aerospace personnel.
[0120] An embodiment of the present application also provides a computer-readable storage medium, which stores a program that can be loaded and executed by a processor, such as a method for evaluating the operational quality of aerospace personnel based on multimodal data fusion, as described in any one of the embodiments of the present application.
[0121] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer program. When all or part of the functions in the above embodiments are implemented by computer program, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above functions can be implemented. In addition, when all or part of the functions in the above embodiments are implemented by computer program, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and saved in the memory of the local device by downloading or copying, or the system of the local device is updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be implemented.
[0122] The above specific examples are used to illustrate the present application, which is only used to help understand the present application and is not intended to limit the present application. For those skilled in the art of the present application, based on the concept of the present application, they can also make some simple deductions, modifications or substitutions.
Claims
1. A method for evaluating the operational quality of aerospace personnel based on multimodal data fusion, characterized in that: include: Acquiring multiple modal training data relevant to spaceflight operations, the training data including EEG signal data, operation record data, and error feedback data; Performing a data normalization operation on the training data of multiple modalities to obtain a plurality of time-aligned single-modal data; Based on a long short-term memory network, a bidirectional gated recurrent unit network, and a multi-head self-attention mechanism, feature extraction is performed on the plurality of unimodal data to obtain a plurality of unimodal features; fusing a plurality of the unimodal features to obtain a cross-modal fusion feature, and generating a shared feature based on the unimodal features and the cross-modal fusion feature; Inputting the shared features into an initial forward neural network, learning the attention weights of the shared features according to the task objectives, calculating the loss values of the shared features under different attention weights, and dynamically adjusting the weight coefficients of the training data of multiple modalities based on the loss values to obtain a target forward neural network; The online data of the space operation is input into the target forward neural network to obtain an evaluation index of the space operation.
2. The method according to claim 1, characterized in that The acquiring of multiple modal training data relevant to aerospace operations includes: collecting initial EEG signal data reflecting the attention level during the spaceflight operation through an EEG sensor, amplifying the initial EEG signal data, and converting the data into a digital signal through an analog-to-digital converter; performing principal component analysis on the converted digital signal, reconstructing data based on the principal components in the digital signal to determine contribution weights in the digital signal, and reducing the dimension of the initial EEG signal data based on the contribution weights to obtain the EEG signal data; Calculate the test plan requirement coverage based on the number of requirements actually covered by the test of the space operation and the total number of requirements; Calculate the on-time completion rate of the test plan based on the number of tests completed on time and the total number of tests for the space operation; Calculate the test requirement execution completeness based on the number of operations that meet the operation standard and the total number of operations of the aerospace operation; Obtaining the operation record data based on the test plan requirement coverage, the test plan completion on time, and the test requirement execution completeness; Error feedback data is calculated based on the erroneous operations, out-of-limit operations and total number of operations of the aerospace operation.
3. The method according to claim 1, characterized in that The operation of performing data normalization on the training data of multiple modalities to obtain a plurality of time-aligned single-modal data includes: Extracting the original timestamp sequence of the training data of each modality respectively, and aligning multiple original timestamp sequences to the same starting time to obtain multiple normalized time series, wherein the original timestamp sequence includes training data corresponding to multiple time points; Selecting the normalized time series with the widest time coverage among the multiple normalized time series as an alignment reference sequence, and using the time axis of the alignment reference sequence as a reference time axis, mapping the training data included in the normalized time series to be aligned to the reference time axis, wherein the reference time axis includes multiple reference time points; For any of the normalized time series to be aligned, a distance metric function is constructed to calculate the Euclidean distance of the normalized time series to be aligned at each time point, a cost matrix is constructed based on the Euclidean distance, and the minimum cumulative distance of the normalized time series to be aligned at each time point is calculated; By backtracking, starting from the last element of the cost matrix, tracing back in the direction corresponding to the minimum cumulative distance, a target alignment path is obtained to align the training data of the normalized time series to be aligned with the training data of the reference time axis; If there is a reference time point in the target alignment path that has no corresponding training data of the normalized time series to be aligned, the training data corresponding to the reference time point is reconstructed using a linear interpolation method; The training data corresponding to each time point of the normalized time series of multiple modalities are spliced to obtain a plurality of aligned single-modal data.
4. The method according to claim 1, wherein The method is based on the long short-term memory network, the bidirectional gated recurrent unit network and the multi-head self-attention mechanism to extract features from the plurality of unimodal data to obtain a plurality of unimodal features, including: Inputting the plurality of unimodal data into the long short-term memory network to obtain a plurality of unimodal time series features; Inputting the plurality of unimodal time series features into the bidirectional gated recurrent unit network to obtain bidirectional hidden state temporal features corresponding to a plurality of time points of each of the unimodal time series features; Each of the bidirectional hidden state temporal features is input into a multi-head self-attention network to obtain a unimodal feature of each of the unimodal data.
5. The method according to claim 4, characterized in that The unimodal data includes data at multiple time points, the long short-term memory network includes a forget gate, an input gate, and an output gate, and the multiple unimodal data are input into the long short-term memory network to obtain multiple unimodal time series features, including: Input each of the unimodal data into the forget gate respectively to determine the forgetting ratio of the unimodal data from the previous time point to the current time point; Inputting the unimodal data greater than a set forgetting ratio into the input gate, obtaining a write ratio of new unimodal data at a current time point, and generating a candidate state of the new unimodal data based on the write ratio; updating the state of the unimodal data at the current time point based on the forgetting ratio and the candidate state; Inputting the state of the unimodal data at the current time point into the output gate to generate the hidden state of the unimodal data at the current time point; Based on the hidden state of the unimodal data, a unimodal time series feature reflecting the long-term dependency relationship between the unimodal data is generated.
6. The method according to claim 4, characterized in that The bidirectional gated recurrent unit network includes a forward gated recurrent unit network and a backward gated recurrent unit network. The inputting of the plurality of unimodal time series features into the bidirectional gated recurrent unit network to obtain bidirectional hidden state temporal features corresponding to a plurality of time points of each unimodal time series feature includes: Set the hidden state dimension; Inputting a plurality of the unimodal time series features into the forward gated recurrent unit network, and calculating the forward hidden states of the unimodal time series features in a front-to-back order; Inputting a plurality of the unimodal time series features into the reverse gated recurrent unit network, and calculating the reverse hidden states of the unimodal time series features in reverse order from back to front; The forward hidden state and the reverse hidden state are concatenated to obtain a bidirectional hidden state temporal feature including a plurality of the time points.
7. The method according to claim 4, characterized in that The step of inputting each of the bidirectional hidden state temporal features into the multi-head self-attention network to obtain a unimodal feature of each of the unimodal data comprises: For each of the bidirectional hidden state temporal features, mapping the bidirectional hidden state temporal features into a query vector, a key vector, and a value vector using three sets of weight matrices; For each attention head, the correlation between the query vector and the key vector is calculated by dot product, and the attention scores of multiple attention heads are calculated in parallel and scaled; Concatenating the attention scores of the multiple attention heads to obtain a concatenated output; The concatenated output is mapped to the original dimension through a fully connected layer using layer normalization to obtain the unimodal features.
8. The method according to claim 1, characterized in that The fusing of the plurality of unimodal features to obtain a cross-modal fusion feature, and generating a shared feature based on the unimodal features and the cross-modal fusion feature, includes: Calculate the score of the attention head between any two of the unimodal features respectively, Splicing the scores of the attention heads between any two of the unimodal features to obtain a unidirectional fusion feature and a reverse unidirectional fusion feature between any two of the unimodal features; Splicing the unidirectional fusion feature and the reverse unidirectional fusion feature to obtain a cross-modal fusion feature between any two unimodal features; Using average pooling to integrate the unimodal features and the cross-modal fusion features at multiple time points, and projecting the unimodal features to a space of the same dimension as the cross-modal fusion features through linear mapping; The multiple single-modal features and the multiple cross-modal fusion features are spliced together to obtain the shared features.
9. The method according to claim 1, characterized in that The method includes inputting the shared features into an initial forward neural network, learning the attention weights of the shared features according to the task objectives, calculating the loss values of the shared features under different attention weights, and dynamically adjusting the weight coefficients of the training data of multiple modalities based on the loss values to obtain a target forward neural network, including: Inputting the shared features into the initial forward neural network, and setting the weight coefficients of the training data of multiple modalities as initial weight values; Iteratively training the initial feedforward neural network based on the initial weight values, and respectively calculating the loss value corresponding to the training data of each modality; Adjust the initial weights of the training data for each modality based on the loss value, convergence status, and evaluation indicators corresponding to the training data for each modality using preset adjustment rules, and calculate the total loss value for each round of training; When the total loss value is less than the set loss value or the number of iterations reaches the set number, the training is stopped, and the trained initial forward neural network is used as the target forward neural network.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which can be loaded by a processor and executed by a method for evaluating the operational quality of aerospace personnel based on multimodal data fusion according to any one of claims 1 to 9.
Citation Information
Cited By
Disaster emergency rescue nursing skill training system based on virtual reality
CN121148210A