Multi-mode sailor emotion recognition method and device and electronic equipment
By employing a multimodal crew emotion recognition method, which utilizes latent space representation and causal graph model to decouple interference factors, high-precision crew emotion recognition is achieved, solving the problem of emotion recognition in complex maritime environments and improving navigation safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle to achieve high-precision, interpretable crew emotion recognition in complex maritime environments, increasing the risk of operational errors and accidents.
By constructing a multimodal crew emotion recognition method, physiological, behavioral and environmental data streams are acquired. Latent space representation, counterfactual generator and latent space causal graph model are used to decouple interference factors and causal alignment, generate physiological emotion representation and behavioral representation with high signal-to-noise ratio, and finally predict emotion state.
Achieving high-precision and interpretable crew emotion recognition in complex maritime environments reduces the probability of mission errors and improves navigation safety.
Smart Images

Figure CN121901708A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of ship control technology, and in particular relates to a multimodal crew emotion recognition method, device and electronic equipment. Background Technology
[0002] With the deepening development of intelligent ships and human factors engineering in the maritime field, the real-time and accurate identification of crew members' emotional state has become a key technical requirement for improving navigation safety, preventing human error, optimizing human-machine collaboration, and ensuring the mental health of crew members.
[0003] In high-risk, high-load scenarios such as ocean voyages, maneuvering in narrow waterways, crossing fishing areas, or night watch, crew members are exposed to complex environmental disturbances such as ship rolling, main engine noise, and visual obstruction, as well as high-intensity cognitive tasks. This can easily lead to negative emotions such as fatigue, anxiety, and inattention, increasing the risk of operational errors and accidents.
[0004] Therefore, there is an urgent need to develop a method for high-precision and interpretable crew emotion recognition in complex maritime environments, in order to provide early warning of high-risk psychological states, reduce the probability of mission errors, and prevent major accidents such as collisions and groundings. Summary of the Invention
[0005] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a multimodal crew emotion recognition method, device, and electronic device, which can achieve high-precision and interpretable crew emotion recognition in complex maritime environments.
[0006] Firstly, this application provides a multimodal crew emotion recognition method, which includes: Acquire physiological and behavioral data streams from crew members, as well as environmental data streams from the ship; The physiological data stream, the behavioral data stream, and the environmental data stream are mapped to the latent space to obtain the initial joint representation of the crew member. The initial joint representation includes physiological emotional representation, behavioral emotional representation, and environmental representation. The initial joint representation is input into the counterfact generator to decouple the initial joint representation from interference factors, and the counterfact behavior representation output by the counterfact generator is obtained. The physiological emotional representation and the counterfactual behavioral representation are input into the latent space causal graph model to perform modal intervention and causal alignment on the physiological emotional representation and the counterfactual behavioral representation, so as to obtain the emotional latent variables output by the latent space causal graph model. The behavioral emotion representation and the emotion latent variables are input into the emotion classifier to obtain the emotional state of the crew member; The counterfact generator is obtained by training a condition generator, the latent space causal graph model is constructed based on a causal structure algorithm, and the emotion classifier is used to predict the emotions of the crew members.
[0007] According to one embodiment of this application, the physiological emotion representation is obtained based on the following steps: Based on the anchor points of maritime events, the physiological data stream is segmented by boundaries to obtain physiological data segments; HRV time-domain and HRV frequency-domain indices are extracted from the physiological data segments to calculate long-range autonomic nervous system dynamic characteristics. The long-range autonomic neural dynamic features are input into a physiological coding network, which maps them to the latent space. The vestibular-related frequency band weights of the physiological coding network are adjusted by yaw cycle attention gating to obtain the physiological emotion representation.
[0008] According to one embodiment of this application, the behavioral emotion representation is obtained based on the following steps: Based on the anchor points of maritime events, the behavioral data stream is segmented by boundaries to obtain behavioral data segments; The facial action unit intensity sequence and discrete operation behavior sequence are extracted from the behavioral data segment, and prosodic features and eye movement entropy are calculated. The facial motion unit intensity sequence, the discrete operational behavior sequence, the prosodic features, and the eye movement entropy are input into the behavior representation encoder, which includes a visual branch, a speech branch, and an operational branch. The visual branch is used to concatenate and encode the facial action unit intensity sequence and the eye movement entropy to obtain a visual behavior sub-representation. Through the speech branch, the prosodic feature input is encoded with acoustic and semantic features to obtain a speech behavior sub-representation; Through the operation branch, the discrete operation behavior sequence and the watch type prior embedding vector corresponding to the navigation event anchor point are encoded to obtain the operation behavior sub-representation; The visual behavior sub-representation, the speech behavior sub-representation, and the operational behavior sub-representation are aligned across modalities in the latent space to obtain the behavioral emotion representation.
[0009] According to one embodiment of this application, the environmental characterization is obtained based on the following steps: Based on the maritime event anchor points, the environmental data stream is segmented by boundaries to obtain environmental data segments; The noise spectrum of the cabin audio signal in the environmental data segment is extracted, and the peak frequency and Mel spectrum entropy are calculated. Based on the messages from other vessels in the environmental data, the AIS density is obtained; The noise spectrum, the ship's roll parameters, the AIS density, and the watchkeeping type label corresponding to the navigation event anchor point are input into the environmental coding network, encoded by the environmental coding network, and mapped to the latent space to obtain the environmental characterization.
[0010] According to one embodiment of this application, the counterfactual generator is obtained based on the following steps: Acquire paired training samples and sampled reconstruction noise. The paired training samples include multimodal observation data of crew members in multiple navigation environments. The paired training samples include positive samples and negative samples. The multimodal observation data includes sample physiological data, sample behavioral data, and sample environmental data. Generate sample behavior and emotion representations for the positive and negative samples respectively, and obtain corresponding clean and perturbation behavior representations; A conditional variational autoencoder is constructed as a generator, and an adversarial discriminator is also constructed. The disturbance behavior characterization is input into the encoder of the conditional variational autoencoder to obtain latent variables; The sampled reconstruction noise and the latent variables are input into the decoder of the conditional variational autoencoder to output a counterfactual behavior representation; The reconstruction loss is calculated by combining the counterfactual behavior representation with the clean behavior representation, and the adversarial loss is calculated by inputting the counterfactual behavior representation into the adversarial discriminator to obtain the total training loss. By minimizing the total training loss, the parameters of the conditional variational autoencoder and the adversarial discriminator are optimized to obtain the trained counterfactual generator.
[0011] According to one embodiment of this application, inputting the physiological emotional representation and the counterfactual behavioral representation into a latent space causal graph model includes: Obtain the task semantic tags corresponding to the maritime event anchor points, wherein the task semantic tags include task type, risk level and required focus; Based on the task semantic labels, determine the average physiological response pattern and emotional expression behavior pattern of the group that match the current task; Based on the average physiological response pattern of the group and the emotional expression behavior pattern, the standardized deviation of the physiological emotional representation relative to the average physiological response pattern of the group, and the degree of matching between the counterfactual behavioral representation and the emotional expression behavior pattern are calculated to determine the credibility weights of the physiological emotional representation and the counterfactual behavioral representation. The physiological emotion representation and its corresponding credibility weight, as well as the counterfactual behavior representation and its corresponding credibility weight, are input into the latent space causal graph model.
[0012] According to one embodiment of this application, the latent space causal graph model includes emotional latent variable nodes, physiological emotional representation nodes, and counterfactual behavior representation nodes. The emotional latent variable nodes point to the physiological emotional representation nodes to form a direct causal path, and the emotional latent variable nodes point to the counterfactual behavior representation nodes to form an emotional expression path. The latent space causal graph model also includes a physiological decoder, a behavioral decoder, and a variational inference network. The physiological decoder is used to parameterize the direct causal path; The behavior decoder is used to parameterize the emotion expression path; The variational inference network is used to estimate the posterior distribution of the emotion latent variable from the observable input; The latent space causal graph model is trained through differentiable directed acyclic graph constraint optimization. It takes physiological emotion representation and counterfactual behavior representation as observation inputs and learns the adjacency matrix between the latent emotion variable, the physiological emotion representation and the counterfactual behavior representation. The adjacency matrix is used to represent the causal connection strength.
[0013] According to one embodiment of this application, the modal intervention and causal alignment of the physiological emotional representation and the counterfactual behavioral representation to obtain the emotional latent variables output by the latent space causal graph model includes: The physiological emotion representation and the counterfactual behavior representation are concatenated along the feature dimension to obtain a fusion vector; The fusion vector is input into the variational inference network, which includes a shared encoder and a linear projection head, and the linear projection head includes a mean head and a log-variance head. The shared encoder encodes the common information of physiology and behavior in the fusion vector under the current causal hypothesis to obtain a joint context vector. The joint context vector is input into the mean header and the log-variance header respectively to obtain the posterior distribution parameters of the emotion latent variable; The posterior distribution parameters are sampled to obtain the emotional latent variables, which include the corresponding valence dimension and arousal dimension.
[0014] Secondly, this application provides a multimodal crew member emotion recognition device, the device comprising: The acquisition module is used to acquire physiological data streams, behavioral data streams, and environmental data streams of the crew. The first processing module is used to map the physiological data stream, the behavioral data stream and the environmental data stream to the latent space to obtain the initial joint representation of the crew member. The initial joint representation includes physiological emotional representation, behavioral emotional representation and environmental representation. The second processing module is used to input the initial joint representation into the counterfact generator to decouple the initial joint representation from interfering factors and obtain the counterfact behavior representation output by the counterfact generator. The third processing module is used to input the physiological emotional representation and the counterfactual behavioral representation into the latent space causal graph model to perform modal intervention and causal alignment on the physiological emotional representation and the counterfactual behavioral representation, so as to obtain the emotional latent variables output by the latent space causal graph model. The fourth processing module is used to input the behavioral emotion representation and the emotion latent variables into the emotion classifier to obtain the emotional state of the crew member; The counterfact generator is obtained by training a condition generator, the latent space causal graph model is constructed based on a causal structure algorithm, and the emotion classifier is used to predict the emotions of the crew members.
[0015] Thirdly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal crew emotion recognition method as described in the first aspect above.
[0016] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal crew emotion recognition method as described in the first aspect above.
[0017] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the multimodal crew emotion recognition method as described in the first aspect.
[0018] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the multimodal crew emotion recognition method as described in the first aspect above.
[0019] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application.
[0020] The multimodal crew emotion recognition method, device, and electronic device provided in this application have the following advantages over the prior art: (1) By constructing a latent space representation of physiology, behavior and environment, the internal state, external behavior and navigation context information of crew members are effectively integrated. Through the counterfactual generator, non-emotional behavioral components such as cognitive overload are actively stripped away when the task load is high or the environmental interference is strong, so as to retain the true emotional expression. Through the latent space causal graph model, the physiological response and the behavior after interference removal are forced to share the same emotional latent variable as a common cause, so as to achieve cross-modal causal alignment and robust reasoning. The latent emotional variables that are consistent with the behavioral context and causality are classified, so as to achieve high-precision and interpretable crew emotion recognition in complex navigation environment.
[0021] (2) By extracting physiological data segments with task semantics from the anchor point of the nautical event, and generating HRV features that reflect the long-term dynamic changes of the autonomic nervous system based on multi-window sliding statistics, the physiological trends related to emotions are effectively captured rather than instantaneous noise. At the same time, combined with the real-time acquired roll cycle information, a frequency band adaptive attention gating mechanism is constructed to actively suppress the vestibular interference frequency band caused by ship swaying during the physiological encoding process, thereby reducing the pollution of heart rate variability analysis by environmental motion artifacts. The resulting physiological emotion representation retains the essence of the emotion-driven autonomic nervous system response and has robustness to nautical-specific interference, providing high signal-to-noise ratio and causally interpretable physiological input for subsequent multimodal emotion recognition. Attached Figure Description
[0022] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart illustrating the multimodal crew emotion recognition method provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of the multimodal crew member emotion recognition device provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0024] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0025] The multimodal crew emotion recognition method, multimodal crew emotion recognition device, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0026] Among them, the multimodal crew emotion recognition method can be applied to the terminal, specifically by the hardware or software in the terminal.
[0027] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).
[0028] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.
[0029] The multimodal crew emotion recognition method provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the multimodal crew emotion recognition method. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras and wearable devices. The multimodal crew emotion recognition method provided in this application embodiment will be described below using an electronic device as the execution subject.
[0030] like Figure 1 As shown, this multimodal crew emotion recognition method includes: Step 110: Obtain the physiological data stream, behavioral data stream, and environmental data stream of the crew.
[0031] Understandably, physiological data streams are time series of physiological signals of crew members continuously collected by wearable sensors such as ECG chest straps and photoplethysmography (PPG) wristbands. These signals include ECG signals, PPG pulse waves, skin conductance, and body temperature, and are used to reflect the activity state of the autonomic nervous system.
[0032] Behavioral data streams are signals of crew behavior collected by sensing devices deployed on the ship's bridge or living quarters. These include video streams, audio streams, and operation logs recorded by the ship's automation systems. They are used to analyze facial expressions, eye movements, speech rhythms, and operational behaviors.
[0033] The environmental data stream is environmental and operational condition sensor data collected in real time from the ship's integrated platform. It includes noise audio from the engine room or bridge, roll angle sequences output by the ship's attitude instrument, messages from other ships parsed by the Automatic Identification System (AIS) receiver, and watch type labels provided by the electronic scheduling system. It is used to characterize external disturbances and task load in the current navigation situation.
[0034] In step 110, the physiological data stream, behavioral data stream, and environmental data stream of the crew are simultaneously acquired through wearable devices, cameras, microphones, and the ship information system. All data are time-aligned based on navigational event anchor points such as entering and leaving ports and crossing fishing areas.
[0035] Step 120: Map the physiological data stream, the behavioral data stream, and the environmental data stream to the latent space to obtain the initial joint representation of the crew member. The initial joint representation includes physiological emotional representation, behavioral emotional representation, and environmental representation.
[0036] It is understandable that the latent space is a low-dimensional continuous vector space obtained by mapping the original high-dimensional observation data from the deep neural network. In the latent space, semantically similar emotional states are closer in geometric distance, which is convenient for subsequent modeling of causal relationships and decoupling interference.
[0037] Physiological emotion representation is a latent space vector obtained by extracting and encoding physiological data streams. It is 128-dimensional and reflects the autonomic neural response driven by emotions. Furthermore, the vestibular system artifact interference caused by ship roll has been suppressed through the roll cycle attention gating mechanism.
[0038] Behavioral emotion representation is a 128-dimensional latent space vector obtained by passing behavioral data through multi-branch encoding and cross-modal alignment. It integrates information from visual, speech, and operational modalities and embeds prior knowledge of the current duty type. The visual modality includes facial action unit intensity and eye movement entropy, the speech modality includes fundamental frequency, speech rate, and pause duration, and the operational modality includes steering command response delay and alarm confirmation frequency.
[0039] Environmental characterization is a latent space vector obtained by encoding the noise spectrum, roll parameter, AIS density and duty type label in the environmental data stream. It is used to quantify the potential interference intensity of the current environment on emotion recognition.
[0040] In step 120, heart rate variability (HRV) time-frequency analysis and roll attention gating encoding are performed on the physiological data stream to obtain a physiological emotion representation; visual, speech, and operational three-branch encoding and cross-modal alignment are performed on the behavioral data stream to obtain a behavioral emotion representation; the environmental data stream is processed by spectrum, extracted by motion parameters, and embedded in context to obtain an environmental representation; the physiological emotion representation, behavioral emotion representation, and environmental representation are combined into an initial joint representation in the form of structured tuples.
[0041] Step 130: Input the initial joint representation into the counterfact generator to decouple the interference factors from the initial joint representation and obtain the counterfact behavior representation output by the counterfact generator; It is understandable that a counterfactual generator is a conditional generative model trained on paired data. It is built on a conditional variational autoencoder and supplemented by an adversarial discriminator, and is used to generate behavioral representations without interference from behavioral representations that are affected by environmental interference.
[0042] Counterfactual behavioral representations are the output of the counterfactual generator. They have been stripped of external interference factors such as swaying, host noise, and shift rhythm, and only retain behavioral components driven by real emotions, such as downturned corners of the mouth and low tone of voice. They suppress non-emotional features such as hand tremors or fragmented sentences caused by high cognitive load.
[0043] In step 130, the behavioral and emotional representations and environmental representations in the initial joint representations are input into the counterfactual generator. The generator uses a task load gating mechanism to automatically suppress cognitive overload-related behavioral features when environmental interference is strong, and outputs counterfactual behavioral representations that retain only emotion-specific components.
[0044] Step 140: Input the physiological emotional representation and the counterfactual behavioral representation into the latent space causal graph model to perform modal intervention and causal alignment on the physiological emotional representation and the counterfactual behavioral representation, and obtain the latent emotional variables output by the latent space causal graph model; Understandably, the latent space causal graph model is a neural network implementation based on the structural causal model, including emotional latent variable nodes, physiological emotional representation nodes, and counterfactual behavioral representation nodes. Among them, the emotional latent variable serves as a common cause, pointing to the physiological and behavioral nodes respectively, forming a direct causal path and an emotional expression path. The model is trained using a differentiable directed acyclic graph constrained optimization algorithm to ensure that the causal structure conforms to prior knowledge.
[0045] Emotional latent variables are two-dimensional latent variables output by the latent space causal graph model. The first dimension is the emotional valence with positive or negative tendencies, and the second dimension is arousal, which is used to characterize the level of activation and represents the crew's true emotional state, which is the common cause of physiological and behavioral changes.
[0046] In step 140, the physiological emotional representation and the counterfactual behavioral representation are input into the latent space causal graph model. The latent space causal graph model estimates the posterior distribution of the emotional latent variable through a variational inference network. Under the reconstruction constraints of the physiological decoder and the behavioral decoder, the emotional latent variable is forced to become the common causal source of both, thereby achieving modal intervention and causal alignment, and outputting a two-dimensional emotional latent variable.
[0047] Step 150: Input the behavioral emotion representation and the emotion latent variables into the emotion classifier to obtain the emotional state of the crew member; The counterfact generator is obtained by training a condition generator, the latent space causal graph model is constructed based on a causal structure algorithm, and the emotion classifier is used to predict the emotions of the crew members.
[0048] Understandably, emotion classifiers are built on lightweight classification networks to jointly map behavioral emotion representations and latent emotion variables to discrete emotion categories, such as calm, fatigue, anxiety, alertness, or anger, thereby completing the final emotion state determination.
[0049] The condition generator is built on the conditional variational autoencoder (VAE) or conditional generative adversarial network (GAN). The loss function of the counterfactual generator is built on the reconstruction loss, decoupling loss and sentiment consistency loss. The decoupling loss is used to maximize the mutual information lower bound to ensure that the counterfactual representation is independent of the environment.
[0050] In step 150, the behavioral emotion representation with complete behavioral context is concatenated with the emotion latent variables and input into the emotion classifier. The supervised training classification network outputs the crew member's current emotion state category and its confidence level.
[0051] The multimodal crew emotion recognition method provided in this application effectively integrates crew internal state, overt behavior, and navigation context information by constructing a latent space representation of physiology, behavior, and environment. Through a counterfactual generator, it actively removes non-emotional behavioral components such as cognitive overload when the task load is high or environmental interference is strong, preserving genuine emotional expression. By using a latent space causal graph model, it forces physiological responses and post-interference behaviors to share the same latent emotional variable as a common cause, achieving cross-modal causal alignment and robust inference. Finally, it classifies latent emotional variables that are consistent with behavioral context and causality, thereby achieving high-precision and interpretable crew emotion recognition in complex maritime environments.
[0052] In some embodiments, the physiological emotional representation is obtained based on the following steps: Based on the anchor points of maritime events, the physiological data stream is segmented by boundaries to obtain physiological data segments; HRV time-domain and HRV frequency-domain indices are extracted from the physiological data segments to calculate long-range autonomic nervous system dynamic characteristics. The long-range autonomic neural dynamic features are input into a physiological coding network, which maps them to the latent space. The vestibular-related frequency band weights of the physiological coding network are adjusted by yaw cycle attention gating to obtain the physiological emotion representation.
[0053] Understandably, navigational event anchor points are key operational nodes with clear semantics and time markers during a ship's navigation process, such as the start of port entry, the end of port departure, entry into fishing areas, and night shift handover. Navigational event anchor points are automatically recorded by the ship's integrated bridge system or electronic log and serve as a time reference for aligning physiological data streams, behavioral data streams, and environmental data streams.
[0054] Physiological data segments are segments of physiological data of a fixed duration that are taken forward or backward from the anchor point of a navigation event, ensuring that each segment of data corresponds to a specific navigation mission scenario.
[0055] HRV time-domain indices are statistical features calculated based on the interval between adjacent heartbeats, including but not limited to the mean RR interval, standard deviation, and root mean square of the difference between adjacent RR intervals, used to reflect the short-term and long-term regulatory capacity of parasympathetic nerve activity.
[0056] HRV frequency domain index refers to the energy distribution characteristics obtained after converting the RR interval sequence to the frequency domain through fast Fourier transform or autoregressive model. It mainly includes low-frequency component LF, high-frequency component HF, and the ratio of low-frequency component to high-frequency component LF / HF. Among them, LF reflects the mixed regulation of sympathetic and parasympathetic, while HF mainly reflects the activity of parasympathetic (vagus nerve).
[0057] Long-range autonomic nervous system dynamic features are high-dimensional feature vectors formed by sliding calculation and aggregation of HRV time-domain indicators and HRV frequency-domain indicators within a 15-minute window. These features are used to characterize the overall state of the crew's autonomic nervous system evolution over time during this voyage phase.
[0058] The physiological coding network is built on a lightweight deep neural network, including one-dimensional convolutional layers and fully connected layers. It is used to map high-dimensional long-range autonomic neural dynamic features to a low-dimensional latent space and can run efficiently on shipborne edge computing devices.
[0059] Roll cycle attention gating is a learnable attention mechanism. It takes the roll angle time series output by the ship's attitude instrument as input, extracts the dominant roll cycle through spectral analysis, and constructs a gating weight vector aligned with the physiological signal frequency band. This weight vector dynamically adjusts the importance of each frequency band channel in the frequency domain processing layer of the physiological coding network. For vestibular related frequency bands that overlap with the harmonics of the roll cycle, their weights are reduced to suppress non-emotional physiological artifacts caused by ship rolling, while for emotionally sensitive frequency bands such as the HF band, their weights are maintained or increased.
[0060] Physiological emotional representation is a 128-dimensional latent space vector that retains the characteristics of emotion-driven autonomic neural responses while effectively reducing the contamination of physiological signals by environmental disturbances such as ship rolling, thus more accurately reflecting the crew's true emotional state.
[0061] In actual operation, the wearable device receives raw PPG or ECG signals in real time and simultaneously acquires the timestamps of the navigation event anchor points from the ship information system. When a new navigation event anchor point is detected, the system enters the narrow channel, traces back and intercepts the physiological signals of the 15 minutes prior to the anchor point, forming a segment of physiological data.
[0062] The physiological data segment was preprocessed by denoising, R-wave detection, and RR interval extraction. Then, the HRV time-domain index and HRV frequency-domain index were calculated in each window with a sliding window of 5 minutes and a step size of 1 minute. The coefficient of variation of all window results was statistically aggregated to generate a long-range autonomic neural dynamic feature vector.
[0063] Long-range autonomic neural dynamic feature vectors are input into the physiological coding network. Simultaneously, a roll angle sequence is obtained from the ship's inertial measurement unit (IMU), and the main roll frequency is identified through Fast Fourier Transform (FFT), generating roll cycle attention gating weights accordingly. These weights are applied to the intermediate frequency domain representation layer of the physiological coding network to attenuate vestibular interference frequencies. Finally, the physiological coding network outputs a 128-dimensional vector, which is the physiological emotion representation. This representation is then fed into a subsequent joint representation construction module for multimodal emotion recognition.
[0064] In this embodiment, physiological data segments with task semantics are extracted by using nautical event anchor points as trigger points, and HRV features reflecting long-term dynamic changes of the autonomic nervous system are generated based on multi-window sliding statistics. This effectively captures emotion-related physiological trends rather than instantaneous noise. At the same time, combined with real-time acquired roll cycle information, a frequency band adaptive attention gating mechanism is constructed to actively suppress vestibular interference frequency bands caused by ship swaying during the physiological encoding process, reducing the contamination of heart rate variability analysis by environmental motion artifacts. The resulting physiological emotion representation retains the essence of emotion-driven autonomic nervous system response and is robust to nautical-specific interference, providing high signal-to-noise ratio and causally interpretable physiological input for subsequent multimodal emotion recognition.
[0065] In some embodiments, the behavioral emotion representation is obtained based on the following steps: Based on the anchor points of maritime events, the behavioral data stream is segmented by boundaries to obtain behavioral data segments; The facial action unit intensity sequence and discrete operation behavior sequence are extracted from the behavioral data segment, and prosodic features and eye movement entropy are calculated. The facial motion unit intensity sequence, the discrete operational behavior sequence, the prosodic features, and the eye movement entropy are input into the behavior representation encoder, which includes a visual branch, a speech branch, and an operational branch. The visual branch is used to concatenate and encode the facial action unit intensity sequence and the eye movement entropy to obtain a visual behavior sub-representation. Through the speech branch, the prosodic feature input is encoded with acoustic and semantic features to obtain a speech behavior sub-representation; Through the operation branch, the discrete operation behavior sequence and the watch type prior embedding vector corresponding to the navigation event anchor point are encoded to obtain the operation behavior sub-representation; The visual behavior sub-representation, the speech behavior sub-representation, and the operational behavior sub-representation are aligned across modalities in the latent space to obtain the behavioral emotion representation.
[0066] Understandably, the behavioral data stream consists of crew overt behavioral signals continuously collected by multimodal sensing devices deployed on the ship's bridge or living quarters, including video, audio, and ship operation logs. Navigational event anchors are time markers with clear semantics during ship navigation, used to align behavioral data streams with task context. A behavioral data segment is a fixed-duration subsequence of behavioral data extracted from the anchor point.
[0067] The Facial Action Unit Intensity Sequence is a time series of intensity values of standard action units of the Facial Action Coding System (FACS) extracted frame by frame from video frames by the lightweight Facial Action Unit Recognition Model AU-Net. For example, the standard action units of FACS include AU4 frowning and AU12 raising the corners of the mouth.
[0068] Discrete operational behavior sequences are structured operational event flows parsed from the logs of ship automation systems, such as time-stamped discrete actions like confirming alarms, adjusting rudder angles, and marking radar targets.
[0069] Prosodic features include acoustic parameters that reflect emotional states, such as speech fundamental frequency, speech rate, pause duration, and energy changes; eye movement entropy is an information uncertainty index calculated based on pupil trajectory, used to quantify the degree of gaze distraction or cognitive load.
[0070] In actual execution, the original behavioral data stream is segmented based on the anchor points of the navigation events to obtain behavioral data segments corresponding to specific navigation tasks.
[0071] Facial motion unit intensity sequences are extracted from the video stream of the behavior data segment and eye-tracking entropy is calculated by combining them with eye-tracking data. Prosodic features are extracted from the audio stream of the behavior data segment, and discrete operation behavior sequences are parsed from the operation log of the behavior data segment.
[0072] The facial motion unit intensity sequence, the discrete operational behavior sequence, the prosodic features, and the eye movement entropy are respectively input into three dedicated branches of the behavior representation encoder.
[0073] Among them, the visual branch uses lightweight ViT-Tiny to concatenate the action unit sequence with eye movement entropy and then encodes it into a visual behavior sub-representation through a one-dimensional convolutional network.
[0074] The speech branch uses a fine-tuned WavLM acoustic model to jointly model the acoustic and high-level semantic features of prosody, and outputs a speech behavior sub-representation.
[0075] The operation branch embeds the discrete operation behavior sequence into a vector through a multilayer perceptron (MLP), and concatenates it with the learnable prior embedding vector corresponding to the current duty type label obtained from the electronic scheduling system. The result is then encoded into an operation behavior sub-representation by the multilayer perceptron.
[0076] The three sub-representations are mapped to a unified latent space, and cross-modal contrastive learning is used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs to align different modalities at the emotional semantic level, outputting a fused 128-dimensional behavioral emotion representation.
[0077] In this embodiment, task-anchor-driven behavior segmentation ensures data semantic consistency. A three-branch dedicated encoder is used to fully mine emotional cues from visual, speech, and operational modalities, and a duty type prior is introduced to enhance the contextual interpretation capability of operational behaviors. Through a cross-modal comparison alignment mechanism, different modalities are forced to cluster around common emotional semantics in the latent space, effectively mitigating the performance degradation caused by single modal failures such as inaccurate facial expression recognition due to insufficient light and noise interference with speech. This generates robust, comprehensive, and context-aware behavioral and emotional representations, laying a solid foundation for subsequent high-precision emotion recognition.
[0078] In some embodiments, the environmental characterization is obtained based on the following steps: Based on the maritime event anchor points, the environmental data stream is segmented by boundaries to obtain environmental data segments; The noise spectrum of the cabin audio signal in the environmental data segment is extracted, and the peak frequency and Mel spectrum entropy are calculated. Based on the messages from other vessels in the environmental data, the AIS density is obtained; The noise spectrum, the ship's roll parameters, the AIS density, and the watchkeeping type label corresponding to the navigation event anchor point are input into the environmental coding network, encoded by the environmental coding network, and mapped to the latent space to obtain the environmental characterization.
[0079] Understandably, maritime event anchors are used to contextualize environmental data. An environmental data segment is a fixed-duration subsequence of environmental data extracted based on a maritime event anchor point.
[0080] Noise spectrum refers to the frequency-energy distribution obtained by performing a short-time Fourier transform on the audio signal collected by the microphone in the cockpit or cockpit. It is used to characterize the acoustic interference generated by the operation of main engine, auxiliary engine and other equipment.
[0081] The peak frequency is the strongest frequency component of energy identified from the noise spectrum, corresponding to the host speed harmonics.
[0082] Mel spectral entropy is the information entropy calculated by mapping the spectrum to the Mel scale. It reflects the complexity and non-stationarity of noise. High entropy values often correspond to sudden mechanical noise or mixed speech from multiple people.
[0083] AIS density refers to the number of other vessels received via AIS within a 10-nautical-mile radius of the current maritime event, used to quantify the traffic density of the navigation area.
[0084] Roll parameters are key statistics calculated from the roll angle time series obtained from the ship's IMU or attitude instrument, such as roll period and root mean square value of angular acceleration, which reflect the intensity of the ship's sway caused by the sea waves.
[0085] The duty type label is the current shift attribute read from the electronic scheduling system, such as lookout, steering, or rest period, serving as a priori indication of the workload.
[0086] In actual implementation, the original environmental data stream is time-aligned and boundary-segmented based on the anchor points of the navigation events to form environmental data segments corresponding to specific navigation scenarios.
[0087] The audio signal in the environmental data segment is pre-emphasized, framed, windowed, and subjected to Fast Fourier Transform to extract the noise spectrum in the range of 0–500Hz, and the main frequency peak frequency and Mel spectrum entropy are further calculated.
[0088] The AIS density is obtained by counting the number of other ships per unit time from AIS messages.
[0089] The roll period and root mean square value of roll angular acceleration are calculated from the roll angle time series output by the ship's attitude instrument as roll parameters. The current duty type label is obtained from the scheduling system. The noise spectrum compressed by the Mel filter bank, the roll parameters, and the AIS density are concatenated into a continuous feature vector. The duty type label is converted into a learnable vector representation through an embedding layer. The continuous feature vector and the learnable vector representation are input into an environment coding network. The environment coding network includes a spectrum processing subnetwork based on one-dimensional convolution, a motion parameter subnetwork based on fully connected layers, and a context embedding subnetwork. Through the spectrum processing subnetwork, the noise spectrum... After compression by the Mel filter bank, the input is given to a one-dimensional convolutional layer, which outputs a noise feature vector. The roll parameter and AIS density are concatenated through the motion parameter sub-network and then input to a fully connected layer, outputting a navigation interference feature vector. The watch type label is mapped to a learnable watch type embedding vector through the context embedding sub-network. The noise feature vector, navigation interference feature vector, and watch type embedding vector are concatenated and input to a shared multilayer perceptron to obtain a 64-dimensional intermediate environmental representation. The outputs of each word network are concatenated, fused through the shared multilayer perceptron, and linearly projected onto a 128-dimensional latent space to output the environmental representation.
[0090] For example, let the timestamp of the kth maritime event anchor point be... The environmental data stream is a continuous time series. The corresponding environmental data segment Defined as:
[0091] in, The preset window length represents the observation duration before the anchor point.
[0092] Assume the cabin audio signal is The complex spectrum is obtained after short-time Fourier transform (STFT). Its amplitude spectrum is The noise spectrum is then defined as the time-averaged power spectrum. :
[0093] in, For frequency, This represents the number of time frames.
[0094] Peak frequency The frequency at which the energy is maximum in the noise spectrum:
[0095] power spectrum The Mel spectral coefficients are obtained by mapping the Mel filter bank to the Mel frequency domain. After normalization, the probability distribution is obtained. Then the Mel spectrum entropy is:
[0096] Set in time window Within this vessel, the number of AIS messages received from other vessels within a radius centered on this vessel is: Then the AIS density is:
[0097] Let the roll angle sequence be... Calculate the root mean square value of its angular acceleration as an indicator of roll intensity:
[0098] Simultaneously, the dominant roll cycle is obtained through spectrum analysis. And convert it into a roll frequency. The roll parameter vector is denoted as:
[0099] Let the watchkeeping type label corresponding to the current maritime event anchor point be a discrete variable. Through learnable embedding matrices Mapped to an embedding vector:
[0100] The above features are concatenated into an input vector:
[0101] This can be further compressed into Mel spectrum vectors or key frequency band energy.
[0102] Environment coding network based on multilayer perceptron Will Mapped to D - Potential Space:
[0103] Output This is the environmental characterization.
[0104] In this embodiment, environmental data is contextualized by using maritime event anchors to accurately capture multidimensional environmental factors such as mechanical noise characteristics, ship motion intensity, traffic density in the navigation area, and task load type, which are strongly correlated with emotional interference. Heterogeneous environmental signals are uniformly mapped to the latent space through a dedicated coding structure to generate highly expressive environmental representations. This not only comprehensively characterizes the physical and task dimensions of external interference but also provides key conditional inputs for the subsequent counterfactual generator, enabling it to effectively decouple the mixed effects of the environment on behavior and improve its anti-interference capability and situational adaptability in complex ocean scenarios.
[0105] In some embodiments, the counterfactual generator is obtained based on the following steps: Acquire paired training samples and sampled reconstruction noise. The paired training samples include multimodal observation data of crew members in multiple navigation environments. The paired training samples include positive samples and negative samples. The multimodal observation data includes sample physiological data, sample behavioral data, and sample environmental data. Generate sample behavior and emotion representations for the positive and negative samples respectively, and obtain corresponding clean and perturbation behavior representations; A conditional variational autoencoder is constructed as a generator, and an adversarial discriminator is also constructed. The disturbance behavior characterization is input into the encoder of the conditional variational autoencoder to obtain latent variables; The sampled reconstruction noise and the latent variables are input into the decoder of the conditional variational autoencoder to output a counterfactual behavior representation; The reconstruction loss is calculated by combining the counterfactual behavior representation with the clean behavior representation, and the adversarial loss is calculated by inputting the counterfactual behavior representation into the adversarial discriminator to obtain the total training loss. By minimizing the total training loss, the parameters of the conditional variational autoencoder and the adversarial discriminator are optimized to obtain the trained counterfactual generator.
[0106] Understandably, paired training samples are time-aligned multimodal observation data pairs collected from the same crew member under different navigation environmental conditions, used to learn the impact of environmental disturbances on behavior.
[0107] Positive examples are data collected from crew members under low-interference, low-task-load conditions (such as the quiet period at port). The crew members' behavior is mainly driven by real emotions and is a clean reference.
[0108] Negative examples are data collected from the same crew member under high-interference, high-cognitive-load conditions (such as during fishing seasons on the same route). Their behavior is a mixture of environmental stress and mission overload, and thus constitutes interference.
[0109] Multimodal observation data includes physiological signals, behavioral data such as video / audio / operation logs, and environmental data such as noise, roll, and AIS for the corresponding time period.
[0110] Clean behavioral representations are latent space vectors obtained by performing the aforementioned behavioral emotion representation generation process on the behavioral data of positive examples. They represent emotional behavioral patterns without significant environmental interference.
[0111] The representation of interfering behavior is a representation obtained after the same processing of negative samples, which includes a mixture of emotional and interfering components.
[0112] Conditional Variational Autoencoders (CVAEs) are generative models in which the encoder maps perturbation behavior representations to latent variables, and the decoder reconstructs the behavior representations by introducing environmental conditions and random noise.
[0113] Adversarial discriminators are binary classification neural networks used to distinguish between generated counterfactual behavioral representations and real clean behavioral representations, thereby improving the quality of generation.
[0114] Sampling reconstruction noise is a vector randomly sampled from a standard normal distribution, used to enhance generative diversity and support reparameterized training.
[0115] The reconstruction loss uses mean squared error to measure the distance between the generated representation and the clean representation; the adversarial loss is calculated based on the discriminator output to make the generated distribution approximate the true clean distribution. The total training loss is a weighted sum of the reconstruction loss and the adversarial loss, used to optimize the entire generative adversarial framework end-to-end.
[0116] In actual implementation, paired samples that meet the criteria are selected from the historical navigation database, including synchronous multimodal data of positive examples of the same crew member during the quiet period at port and negative examples of the same fishing area during the high-load period.
[0117] The behavioral sentiment representation generation process is executed on the behavioral data streams of positive and negative examples respectively to obtain the corresponding clean behavioral representations and interfering behavioral representations.
[0118] A conditional variational autoencoder is constructed as a generator, taking the representation of the disturbance behavior as input and the representation of the environment as conditions, and a discriminator is designed to accompany it.
[0119] During training, the representation of interfering behaviors is input into the encoder to obtain latent variables. The latent variables are then concatenated with reconstructed noise sampled from a standard normal distribution and fed into the decoder. Under the guidance of environmental conditions, the decoder outputs counterfactual behavior representations. This output is used to calculate the L2 reconstruction loss with the clean behavior representation and is also fed into the discriminator to calculate the adversarial loss. The two losses are weighted and summed to form the total loss, which is used to jointly update the parameters of the generator and the discriminator through backpropagation. After training, the generator becomes a counterfactual generator, which can receive any interfering behavior representations and environmental representations during the inference phase and output counterfactual behavior representations that have removed environmental interferences and retain only the emotion-specific components.
[0120] In this embodiment, by constructing positive and negative paired samples, the mapping relationship between interfering and non-interfering behaviors is explicitly modeled. A collaborative training mechanism between a conditional variational autoencoder and an adversarial discriminator is used to enable the generator to learn to remove non-emotional factors such as swaying, noise, and high cognitive load from contaminated behaviors. An environmental condition and task load perception mechanism is introduced in the decoding stage to specifically suppress cognitive overload features such as hand tremors and fragmented sentences, while retaining core emotional expressions such as downturned lips and low tone of voice. The resulting counterfactual behavioral representation effectively eliminates environmental confounding effects, providing clean behavioral input for subsequent causal modeling and improving the accuracy, robustness, and causal interpretability of emotion recognition in complex maritime scenarios.
[0121] In some embodiments, inputting the physiological emotional representation and the counterfactual behavioral representation into a latent space causal graph model includes: Obtain the task semantic tags corresponding to the maritime event anchor points, wherein the task semantic tags include task type, risk level and required focus; Based on the task semantic labels, determine the average physiological response pattern and emotional expression behavior pattern of the group that match the current task; Based on the average physiological response pattern of the group and the emotional expression behavior pattern, the standardized deviation of the physiological emotional representation relative to the average physiological response pattern of the group, and the degree of matching between the counterfactual behavioral representation and the emotional expression behavior pattern are calculated to determine the credibility weights of the physiological emotional representation and the counterfactual behavioral representation. The physiological emotion representation and its corresponding credibility weight, as well as the counterfactual behavior representation and its corresponding credibility weight, are input into the latent space causal graph model.
[0122] Understandably, task semantic tags are structured task descriptions automatically extracted from the ship's electronic task management system or navigation log and associated with the current maritime event anchor point. They include task types such as lookout, steering, and emergency response, risk levels, and required focus, and are used to characterize the expected impact of the current situation on the crew's cognition and emotions.
[0123] The group average physiological response pattern refers to the statistical mean vector of physiological and emotional representations calculated from all crew samples with the same task semantic label in a historical database. It reflects the typical autonomic nervous response baseline under the task and is obtained by querying a pre-built task-physiological baseline library.
[0124] Emotional expression behavior patterns are cluster centers or mean vectors of counterfactual behavior representations of all samples under the corresponding task. They represent typical explicit behavioral paradigms driven by emotions in the task and are obtained by querying the task-behavior expression norm library.
[0125] Standardized deviation is calculated by dividing the Euclidean distance between the current physiological emotional representation and the average physiological response pattern of the group by the standard deviation of the physiological representation under the task. It is used to quantify whether an individual's physiological response deviates significantly from the reasonable range of the task. If the deviation is too large, it may indicate that the signal is affected by artifacts or dominated by non-emotional factors.
[0126] The degree of matching is measured by cosine similarity or Mahalanobis distance to determine how close the current counterfactual behavior representation is to the emotional expression behavior pattern. A high degree of matching represents behavior that conforms to the typical emotional expression logic under the task.
[0127] The credibility weight is a value between 0 and 1 obtained by mapping the deviation and matching degree mentioned above. It is used to characterize the reliability of the modality input in the current task context. For example, a low physiological deviation results in a high physiological credibility weight, and a high behavioral matching degree results in a high behavioral credibility weight.
[0128] In actual execution, the task semantic tag corresponding to the current maritime event anchor point is obtained from the ship mission management system.
[0129] Query the group average physiological response pattern and its covariance matrix under the same label in the pre-built task-physiological baseline library, and query the corresponding emotion expression behavior pattern in the pre-built task-behavioral norm library.
[0130] The Mahalanobis distance of the current physiological emotional representation relative to the group mean is calculated and converted into a standardized deviation. Then, it is mapped to physiological credibility weights using a sigmoid function. The smaller the deviation, the higher the weight. At the same time, the cosine similarity between the counterfactual behavioral representation and the behavioral pattern is calculated and also mapped to behavioral credibility weights. The higher the similarity, the higher the weight.
[0131] Physiological emotional representations and their credibility weights, as well as counterfactual behavioral representations and their credibility weights, are input into the latent space causal graph model as the basis for regulating modal confidence in subsequent causal inference.
[0132] Furthermore, a cross-module causal consistency constraint term is introduced into the loss function of the latent space causal graph model. The cross-module causal consistency constraint term dynamically adjusts the gradient contribution according to the credibility weights of the physiological emotion representation and the counterfactual behavior representation. This ensures that when both representations have high credibility, they jointly participate in the reconstruction process of the latent emotion variable. However, when the credibility of one representation is low, it mainly relies on another high-credibility representation to reconstruct the latent emotion variable, thereby reducing the impact of potential noise or inconsistent information.
[0133] In this embodiment, by introducing a task semantic-driven context-aware mechanism, external navigation mission knowledge is integrated into the multimodal fusion process, enabling the system to dynamically judge the rationality and credibility of physiological and behavioral signals in the current context. When a certain modality is distorted due to individual differences, equipment noise, or mission specificity, its credibility weight is automatically reduced to avoid erroneous information dominating emotion inference. Under the dominance of high-quality modalities, the causal graph model can more accurately reconstruct shared latent emotion variables, enhancing the model's adaptability to complex ocean conditions and improving the contextual consistency and decision credibility of emotion recognition.
[0134] In some embodiments, the latent space causal graph model includes latent emotion variable nodes, physiological emotion representation nodes, and counterfactual behavior representation nodes. The latent emotion variable nodes point to the physiological emotion representation nodes to form a direct causal path, and the latent emotion variable nodes point to the counterfactual behavior representation nodes to form an emotion expression path. The latent space causal graph model also includes a physiological decoder, a behavioral decoder, and a variational inference network. The physiological decoder is used to parameterize the direct causal path; The behavior decoder is used to parameterize the emotion expression path; The variational inference network is used to estimate the posterior distribution of the emotion latent variable from the observable input; The latent space causal graph model is trained through differentiable directed acyclic graph constraint optimization. It takes physiological emotion representation and counterfactual behavior representation as observation inputs and learns the adjacency matrix between the latent emotion variable, the physiological emotion representation and the counterfactual behavior representation. The adjacency matrix is used to represent the causal connection strength.
[0135] Understandably, the latent emotion node is used to represent the unobservable real emotional state. It is a two-dimensional latent variable root node in the latent space causal graph. The dimensions correspond to the valence that represents the positive or negative tendency of emotions in psychology, and the arousal that represents the activation level.
[0136] The physiological emotion representation node and the counterfactual behavior representation node are two observable leaf nodes that represent the physiological response and behavioral expression after interference decoupling, respectively, and serve as input evidence for the model.
[0137] Direct causal paths refer to directed edges from latent emotional variables to physiological emotional representations, describing the driving effect of emotions on the activity of the autonomic nervous system.
[0138] The emotion expression path is a directed edge from latent emotional variables to counterfactual behavioral representations, reflecting the process of emotion expression through overt behaviors such as facial expressions and tone of voice.
[0139] The physiological decoder is built on a multilayer perceptron and is used for the parameterized implementation of structural causal equations, that is, to reconstruct physiological emotional representations from emotional latent variables.
[0140] The behavioral decoder has a similar structure to the physiological decoder, and is used to reconstruct counterfactual behavioral representations starting from the same emotional latent variables; The variational inference network is the encoder part. It takes physiological and counterfactual behavioral representations as input and outputs the mean and variance of the posterior distribution of the emotion latent variable. It can support reparameterized sampling to achieve end-to-end training.
[0141] Differentiable directed acyclic graph constrained optimization is a continuous optimization structure learning algorithm based on the NOTEARS algorithm. By applying a smooth acyclicity penalty term to the adjacency matrix, the model automatically learns a connection structure that conforms to causal priors during training. The adjacency matrix is a 3×3 real-valued matrix whose elements represent the causal strength from one node to another. It is constrained to only allow the latent emotional variables to point to the other two nodes, and the connection weight of the environment to physiology approaches zero.
[0142] In actual implementation, the structure of the latent space causal graph model assumes that physiology is not directly affected by the environment, and that environmental factors do not establish direct causal edges with the physiological emotional representations, so as to ensure that physiological signals are not affected by false causal influences from environmental interference. During the training phase, paired physiological emotion representations and counterfactual behavior representations are used as observational inputs. The posterior distribution of the emotion latent variables is estimated and sampled through a variational inference network.
[0143] The latent emotion variables are input into the physiological decoder and the behavioral decoder respectively to generate reconstructions of the two inputs. At the same time, the model maintains a learnable adjacency matrix, whose off-diagonal elements are updated through optimization of the differentiable directed acyclic graph (DAG) constraints to ensure that the causal graph is acyclic and conforms to the prior structure that emotion is a common cause. The loss function includes reconstruction error, KL divergence regularization term and DAG penalty term, which jointly optimize the parameters of the variational inference network, the two decoders and the adjacency matrix.
[0144] After training, the adjacency matrix converges to a sparse structure, clearly demonstrating two effective causal paths, with the emotional latent variable becoming the only shared cause for physiology and behavior.
[0145] During the inference phase, only variational inference networks are used to output two-dimensional emotional latent variables by inputting new physiological and counterfactual behavioral representations.
[0146] In addition, variables in the latent space can be set to include physiological emotional representations, counterfactual behavioral representations, and implicit emotional latent variables, and a causal structure prior can be presupposed. Among them, the emotional latent variables have a direct causal influence on the physiological emotional representations, the emotional latent variables drive the counterfactual behavioral representations to form an emotional expression path, and environmental factors do not directly affect the physiological emotional representations. Using physiological emotional representations and counterfactual behavioral representations as observational inputs, a differentiable directed acyclic graph constrained optimization algorithm is employed to learn the adjacency matrix between the above variables. This adjacency matrix is used to represent the causal strength between variables in the latent space. During the learning process of the adjacency matrix, a sparse regularization term is introduced to encourage the model to retain key causal paths and suppress redundant connections, generating a concise causal graph and forcing the weight of the causal influence of the environment on physiology to approach zero. In the final output causal structure, the causal weight of the environment on behavior is significantly greater than that of the environment on physiology, which is consistent with the prior knowledge in the nautical scenario.
[0147] A structural causal model is constructed based on the adjacency matrix after training convergence, in which physiological emotion representation is jointly determined by latent emotion variables and independent noise, and counterfactual behavior representation is also generated by latent emotion variables and independent noise. A variational inference network is constructed, which takes physiological emotional representation and counterfactual behavioral representation as inputs, estimates the posterior distribution of the emotional latent variable, and outputs a two-dimensional emotional latent variable including valence and arousal dimensions.
[0148] During the backpropagation process of model training, the parameter gradient of the physiological decoder is proportional to the physiological confidence weight, and the parameter gradient of the behavior decoder is proportional to the behavior confidence weight. When the confidence weight of a certain physiological or behavioral representation is low, the gradient magnitude of its corresponding decoder is linearly compressed, thereby reducing the contribution of that modality in parameter updates; when both confidence weights are high, the two gradients jointly dominate the optimization direction of the emotional latent variable, achieving strong causal alignment under high confidence.
[0149] In this embodiment, by deeply fusing a structural causal model with a deep generative network, not only are the causal mechanisms of emotion → physiology and emotion → behavior explicitly modeled, but differentiable DAG learning is also used to ensure the rationality and sparsity of the causal structure. The introduction of physiological and behavioral decoders enables the causal relationship to have a computable functional form, while the variational inference network realizes inverse reasoning from observed data to latent causal variables. During training, the entire model forces physiology and behavior to be explained by the same latent emotional variable, naturally achieving cross-modal causal alignment and effectively suppressing spurious associations caused by environmental confounding or modal noise. It can output causally interpretable and physiological-behavioral consistent latent emotional variables in complex maritime scenarios.
[0150] In some embodiments, the modal intervention and causal alignment of the physiological emotional representation and the counterfactual behavioral representation to obtain the latent emotional variables output by the latent space causal graph model includes: The physiological emotion representation and the counterfactual behavior representation are concatenated along the feature dimension to obtain a fusion vector; The fusion vector is input into the variational inference network, which includes a shared encoder and a linear projection head, and the linear projection head includes a mean head and a log-variance head. The shared encoder encodes the common information of physiology and behavior in the fusion vector under the current causal hypothesis to obtain a joint context vector. The joint context vector is input into the mean header and the log-variance header respectively to obtain the posterior distribution parameters of the emotion latent variable; The posterior distribution parameters are sampled to obtain the emotional latent variables, which include the corresponding valence dimension and arousal dimension.
[0151] It is understandable that physiological emotional representations are latent space vectors that reflect the autonomic nervous activity driven by emotions after being suppressed by sway interference.
[0152] Counterfactual behavioral representations are latent space vectors that have been stripped of environmental interferences such as noise, sway, and high cognitive load, retaining only the emotionally relevant overt behavioral components.
[0153] Modal intervention involves dynamically adjusting the influence of each modality on sentiment estimation based on the credibility or contextual plausibility of each modality during causal reasoning.
[0154] Causal alignment forces both physiological and behavioral modalities to share the same latent emotional variable as a common cause, ensuring that they are consistent in emotional semantics.
[0155] The fusion vector is a joint input vector formed by concatenating physiological emotional representations and counterfactual behavioral representations along the feature dimension.
[0156] Variational inference network is a parametric encoder used to approximate the posterior distribution of the real emotional latent variable. The shared encoder consists of two fully connected layers to extract common semantic information between physiology and behavior. The linear projection head includes two independent linear layers: the mean head outputs the mean vector of the posterior distribution of the emotional latent variable, and the log-variance head outputs its log-variance vector, which together define a two-dimensional Gaussian distribution.
[0157] The joint context vector is the output of the shared encoder, representing the emotional context that physiology and behavior jointly point to under the causal structure assumption that the current emotion is the common cause.
[0158] The emotional latent variable is a two-dimensional latent variable in the final output. The first dimension corresponds to valence in psychology, representing the positive or negative tendency of emotion, and the second dimension corresponds to arousal, representing the level of activation, used to characterize the crew's true emotional state.
[0159] In actual implementation, the 128-dimensional physiological emotion representation and the 128-dimensional counterfactual behavior representation are spliced along the feature dimensions to form a 256-dimensional fusion vector.
[0160] The fused vector is input into the shared encoder of the variational inference network, which passes through fully connected layers of 256→128→64 dimensions and is activated by ReLU, outputting a 64-dimensional joint context vector.
[0161] The joint context vector is fed into the mean header and the log-variance header respectively to obtain the mean of the posterior distribution. Sum of logarithmic variance .
[0162] During the training phase, noise is sampled from a standard normal distribution. Calculate latent emotion variables using reparameterization techniques ,in, In the reasoning stage, we directly take... As a deterministic output, this two-dimensional vector represents the latent emotion variable, with its two dimensions guided by the monitoring signal as valence and arousal, respectively.
[0163] In this embodiment, variational inference networks jointly encode physiological and behavioral information under the constraints of a structural causal model. This enables the model to not only integrate multimodal information but also extract shared emotional semantics under the prior guidance of "emotion as a common cause," thereby achieving true causal alignment. Simultaneously, since the entire process is embedded in a learnable generative framework, when a certain modality is disturbed or inconsistent with the task context, its contribution to the joint context vector is naturally weakened, implicitly achieving modal intervention. The final output two-dimensional emotional latent variables possess both psychological interpretability and cross-modal consistency, providing a reliable and causally reasonable intermediate representation for subsequent high-precision and robust emotion classification.
[0164] The multimodal crew emotion recognition method provided in this application can be executed by a multimodal crew emotion recognition device. This application uses the execution of the multimodal crew emotion recognition method by a multimodal crew emotion recognition device as an example to illustrate the multimodal crew emotion recognition device provided in this application.
[0165] This application also provides a multimodal crew member emotion recognition device.
[0166] like Figure 2 As shown, the multimodal crew emotion recognition device includes: The acquisition module 210 is used to acquire the physiological data stream, behavioral data stream, and environmental data stream of the crew members; The first processing module 220 is used to map the physiological data stream, the behavioral data stream and the environmental data stream to the latent space to obtain the initial joint representation of the crew member. The initial joint representation includes physiological emotional representation, behavioral emotional representation and environmental representation. The second processing module 230 is used to input the initial joint representation to the counterfact generator to decouple the interference factors of the initial joint representation and obtain the counterfact behavior representation output by the counterfact generator. The third processing module 240 is used to input the physiological emotional representation and the counterfactual behavioral representation into the latent space causal graph model, so as to perform modal intervention and causal alignment on the physiological emotional representation and the counterfactual behavioral representation, and obtain the emotional latent variables output by the latent space causal graph model. The fourth processing module 250 is used to input the behavioral emotion representation and the emotion latent variables into the emotion classifier to obtain the emotional state of the crew member; The counterfact generator is obtained by training a condition generator, the latent space causal graph model is constructed based on a causal structure algorithm, and the emotion classifier is used to predict the emotions of the crew members.
[0167] The multimodal crew emotion recognition device provided in this application effectively integrates crew internal state, overt behavior, and navigation context information by constructing a latent space representation of physiology, behavior, and environment. Through a counterfactual generator, it actively removes non-emotional behavioral components such as cognitive overload when the task load is high or the environmental interference is strong, thus preserving the true emotional expression. By using a latent space causal graph model, it forces physiological responses and behaviors after interference removal to share the same latent emotional variable as a common cause, achieving cross-modal causal alignment and robust inference. It classifies latent emotional variables that are consistent with behavioral context and causality, thereby achieving high-precision and interpretable crew emotion recognition in complex maritime environments.
[0168] The multimodal crew emotion recognition device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a Mobile Internet Device (MID), an Ultra-Mobile Personal Computer (UMPC), a server, or a Personal Computer (PC), etc., and this application embodiment does not impose specific limitations.
[0169] The multimodal crew emotion recognition device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0170] The multimodal crew emotion recognition device provided in this application embodiment can realize the various processes implemented in the multimodal crew emotion recognition method embodiment as described above. To avoid repetition, it will not be described again here.
[0171] In some embodiments, such as Figure 3 As shown, this application embodiment also provides an electronic device 300, including a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the program is executed by the processor 301, it implements the various processes of the above-described multimodal crew emotion recognition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0172] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0173] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described multimodal crew emotion recognition method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0174] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0175] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described multimodal crew emotion recognition method.
[0176] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0177] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described multimodal crew emotion recognition method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0178] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0179] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0180] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the multimodal crew emotion recognition method of the various embodiments of this application.
[0181] In the description of this application, "first feature" and "second feature" may include one or more of the features.
[0182] In the description of this application, "multiple" means two or more.
[0183] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0184] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0185] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
Claims
1. A multimodal crew emotion recognition method, characterized in that, include: Acquire physiological and behavioral data streams from crew members, as well as environmental data streams from the ship; The physiological data stream, the behavioral data stream, and the environmental data stream are mapped to the latent space to obtain the initial joint representation of the crew member. The initial joint representation includes physiological emotional representation, behavioral emotional representation, and environmental representation. The initial joint representation is input into the counterfact generator to decouple the initial joint representation from interference factors, and the counterfact behavior representation output by the counterfact generator is obtained. The physiological emotional representation and the counterfactual behavioral representation are input into the latent space causal graph model to perform modal intervention and causal alignment on the physiological emotional representation and the counterfactual behavioral representation, so as to obtain the emotional latent variables output by the latent space causal graph model. The behavioral emotion representation and the emotion latent variables are input into the emotion classifier to obtain the emotional state of the crew member; The counterfact generator is obtained by training a condition generator, the latent space causal graph model is constructed based on a causal structure algorithm, and the emotion classifier is used to predict the emotions of the crew members.
2. The multimodal crew emotion recognition method according to claim 1, characterized in that, The physiological emotional representation is obtained based on the following steps: Based on the anchor points of maritime events, the physiological data stream is segmented by boundaries to obtain physiological data segments; HRV time-domain and HRV frequency-domain indices are extracted from the physiological data segments to calculate long-range autonomic nervous system dynamic characteristics. The long-range autonomic neural dynamic features are input into a physiological coding network, which maps them to the latent space. The vestibular-related frequency band weights of the physiological coding network are adjusted by yaw cycle attention gating to obtain the physiological emotion representation.
3. The multimodal crew emotion recognition method according to claim 1, characterized in that, The behavioral emotion representation is obtained based on the following steps: Based on the anchor points of maritime events, the behavioral data stream is segmented by boundaries to obtain behavioral data segments; The facial action unit intensity sequence and discrete operation behavior sequence are extracted from the behavioral data segment, and prosodic features and eye movement entropy are calculated. The facial motion unit intensity sequence, the discrete operation behavior sequence, the prosodic features, and the eye movement entropy are input into the behavior representation encoder, which includes a visual branch, a speech branch, and an operation branch. The visual branch is used to concatenate and encode the facial action unit intensity sequence and the eye movement entropy to obtain a visual behavior sub-representation. Through the speech branch, the prosodic feature input is encoded with acoustic and semantic features to obtain a speech behavior sub-representation; Through the operation branch, the discrete operation behavior sequence and the watch type prior embedding vector corresponding to the navigation event anchor point are encoded to obtain the operation behavior sub-representation; The visual behavior sub-representation, the speech behavior sub-representation, and the operational behavior sub-representation are aligned across modalities in the latent space to obtain the behavioral emotion representation.
4. The multimodal crew emotion recognition method according to claim 1, characterized in that, The environmental characterization is obtained based on the following steps: Based on the maritime event anchor points, the environmental data stream is segmented by boundaries to obtain environmental data segments; The noise spectrum of the cabin audio signal in the environmental data segment is extracted, and the peak frequency and Mel spectrum entropy are calculated. Based on the messages from other vessels in the environmental data, the AIS density is obtained; The noise spectrum, the ship's roll parameters, the AIS density, and the watchkeeping type label corresponding to the navigation event anchor point are input into the environmental coding network, encoded by the environmental coding network, and mapped to the latent space to obtain the environmental characterization.
5. The multimodal crew emotion recognition method according to claim 1, characterized in that, The counterfact generator is obtained based on the following steps: Acquire paired training samples and sampled reconstruction noise. The paired training samples include multimodal observation data of crew members in multiple navigation environments. The paired training samples include positive samples and negative samples. The multimodal observation data includes sample physiological data, sample behavioral data, and sample environmental data. Generate sample behavior and emotion representations for the positive and negative samples respectively, and obtain corresponding clean and perturbation behavior representations; A conditional variational autoencoder is constructed as a generator, and an adversarial discriminator is also constructed. The disturbance behavior characterization is input into the encoder of the conditional variational autoencoder to obtain latent variables; The sampled reconstruction noise and the latent variables are input into the decoder of the conditional variational autoencoder to output a counterfactual behavior representation; The reconstruction loss is calculated by combining the counterfactual behavior representation with the clean behavior representation, and the adversarial loss is calculated by inputting the counterfactual behavior representation into the adversarial discriminator to obtain the total training loss. By minimizing the total training loss, the parameters of the conditional variational autoencoder and the adversarial discriminator are optimized to obtain the trained counterfactual generator.
6. The multimodal crew emotion recognition method according to claim 1, characterized in that, The step of inputting the physiological emotional representation and the counterfactual behavioral representation into the latent space causal graph model includes: Obtain the task semantic tags corresponding to the maritime event anchor points, wherein the task semantic tags include task type, risk level and required focus; Based on the task semantic labels, determine the average physiological response pattern and emotional expression behavior pattern of the group that match the current task; Based on the average physiological response pattern of the group and the emotional expression behavior pattern, the standardized deviation of the physiological emotional representation relative to the average physiological response pattern of the group, and the degree of matching between the counterfactual behavioral representation and the emotional expression behavior pattern are calculated to determine the credibility weights of the physiological emotional representation and the counterfactual behavioral representation. The physiological emotion representation and its corresponding credibility weight, as well as the counterfactual behavior representation and its corresponding credibility weight, are input into the latent space causal graph model.
7. The multimodal crew emotion recognition method according to claim 1, characterized in that, The latent space causal graph model includes latent emotion variable nodes, physiological emotion representation nodes, and counterfactual behavior representation nodes. The latent emotion variable nodes point to the physiological emotion representation nodes, forming a direct causal path, and the latent emotion variable nodes point to the counterfactual behavior representation nodes, forming an emotion expression path. The latent space causal graph model also includes a physiological decoder, a behavioral decoder, and a variational inference network. The physiological decoder is used to parameterize the direct causal path; The behavior decoder is used to parameterize the emotion expression path; The variational inference network is used to estimate the posterior distribution of the emotion latent variable from the observable input; The latent space causal graph model is trained through differentiable directed acyclic graph constraint optimization. It takes physiological emotion representation and counterfactual behavior representation as observation inputs and learns the adjacency matrix between the latent emotion variable, the physiological emotion representation and the counterfactual behavior representation. The adjacency matrix is used to represent the causal connection strength.
8. The multimodal crew emotion recognition method according to claim 7, characterized in that, The modal intervention and causal alignment of the physiological emotional representation and the counterfactual behavioral representation to obtain the emotional latent variables output by the latent space causal graph model include: The physiological emotion representation and the counterfactual behavior representation are concatenated along the feature dimension to obtain a fusion vector; The fusion vector is input into the variational inference network, which includes a shared encoder and a linear projection head, and the linear projection head includes a mean head and a log-variance head. The shared encoder encodes the common information of physiology and behavior in the fusion vector under the current causal hypothesis to obtain a joint context vector. The joint context vector is input into the mean header and the log-variance header respectively to obtain the posterior distribution parameters of the emotion latent variable; The posterior distribution parameters are sampled to obtain the emotional latent variables, which include the corresponding valence dimension and arousal dimension.
9. A multimodal crew member emotion recognition device, characterized in that, include: The acquisition module is used to acquire physiological data streams, behavioral data streams, and environmental data streams of the crew. The first processing module is used to map the physiological data stream, the behavioral data stream, and the environmental data stream to the latent space to obtain the initial joint representation of the crew member. The initial joint representation includes physiological emotional representation, behavioral emotional representation, and environmental representation. The second processing module is used to input the initial joint representation into the counterfact generator to decouple the initial joint representation from interference factors and obtain the counterfact behavior representation output by the counterfact generator. The third processing module is used to input the physiological emotional representation and the counterfactual behavioral representation into the latent space causal graph model to perform modal intervention and causal alignment on the physiological emotional representation and the counterfactual behavioral representation, and obtain the emotional latent variables output by the latent space causal graph model. The fourth processing module is used to input the behavioral emotion representation and the emotion latent variables into the emotion classifier to obtain the emotional state of the crew member; The counterfact generator is obtained by training a condition generator, the latent space causal graph model is constructed based on a causal structure algorithm, and the emotion classifier is used to predict the emotions of the crew members.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal crew emotion recognition method as described in any one of claims 1-8.