Motion sickness recognition system based on multi-modal contrast learning

A motion sickness recognition system based on multimodal contrastive learning constructs a connectivity graph representation of EEG, video, and motion data, solving the problem of insufficient cross-modal structural alignment in existing technologies and improving the security of virtual devices.

CN120899179APending Publication Date: 2025-11-07BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511103486.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing multimodal fusion recognition methods lack cross-modal structure alignment mechanisms, resulting in real-time response delays and inaccurate intervention measures. This fails to effectively improve users' motion sickness and affects the security of virtual devices.

Method used

A motion sickness recognition system based on multimodal contrastive learning is adopted. The system collects EEG data, video data and motion data from virtual reality devices through a data processing chip to construct an initial brain connectivity representation field. Through connection constraint contrast fusion, a multimodal brain connectivity representation is generated. The system uses a preset difference metric function to identify motion sickness state and adjusts the screen of the virtual reality device to improve motion sickness.

Benefits of technology

It improves the safety of virtual reality device use by reducing latency and blur, accurately adjusting virtual reality device parameters, and alleviating motion sickness in users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120899179A_ABST
    Figure CN120899179A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a motion sickness recognition system based on multi-mode contrast learning. According to one specific embodiment, the system comprises a virtual reality device, a data processing chip, a camera, an electroencephalogram signal receiver and a motion signal receiver, and the data processing chip is used for collecting multi-modal data from the virtual reality device based on a virtual reality scene; the data processing chip is used for constructing an initial brain connection graph representation field; the data processing chip is used for executing a training step on the initial brain connection graph representation field; the data processing chip is used for inputting to-be-recognized multi-modal input data into the multi-modal contrast learning brain connection diagram representation field to obtain a motion sickness recognition result; and the data processing chip is used for adjusting the picture of the virtual reality equipment according to the motion sickness recognition result. Therefore, by reducing delay and motion blur of the virtual reality equipment, the motion sickness state of the user is improved, and therefore the safety of the user in the using process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of virtual reality and brain-computer interface, and particularly relate to a motion sickness recognition system based on multi-modal contrast learning. BACKGROUND

[0002] With the development of the technical field of virtual reality and brain-computer interface, the problem of motion sickness of users in an immersive environment is increasingly prominent, which seriously affects the quality of experience of users. In the existing technology, common motion sickness recognition methods include motion analysis based on video content, feature extraction based on physiological signals, and multi-modal data fusion processing methods that fuse multiple modal data. Among them, the multi-modal fusion method is widely concerned by introducing more dimensional data.

[0003] However, it is found in practice that when the above multi-modal fusion recognition method is used, the following technical problems often exist:

[0004] Using a simple feature splicing or stacking method for multi-modal fusion, the real-time response delay and inaccurate intervention caused by the lack of cross-modal structure alignment mechanism, resulting in the virtual device being unable to improve the user's motion sickness state by adjusting the virtual reality device, thereby making the safety of the user when using the virtual device is low.

[0005] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore, it can contain information that does not form the prior art known to those of ordinary skill in the art in the country. SUMMARY

[0006] The summary section of the present disclosure is used to introduce the concepts in a brief form, which will be described in detail in the specific embodiments section. The summary section of the present disclosure is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0007] Some embodiments of the present disclosure propose a motion sickness recognition system based on multi-modal contrast learning to solve one or more of the technical problems mentioned in the background section.

[0008] Some embodiments of the present disclosure provide a motion sickness recognition system based on multi-modal contrast learning, comprising: a virtual reality device, a data processing chip, a camera, an electroencephalogram signal receiver, and a motion signal receiver, wherein: the data processing chip is configured to collect multi-modal data based on a virtual reality scene from the virtual reality device, wherein the multi-modal data comprises electroencephalogram data, video data, and motion data; the data processing chip is configured to construct an initial brain connection graph representation field according to the electroencephalogram data, the video data, and the motion data, wherein the initial brain connection graph representation field comprises: an electroencephalogram connection graph representation, a video and motion connection graph representation, and a standardized brain connection graph representation; the data processing chip is configured to perform the following training steps on the initial brain connection graph representation field: the data processing chip performs connection constraint contrast fusion on the video and motion connection graph and the electroencephalogram connection graph representation based on the standardized brain connection graph representation to obtain a fused multi-modal brain connection graph representation; the data processing chip determines a target difference value between the multi-modal brain connection graph representation and labeled motion sickness state data based on a preset difference measure function; the data processing chip determines the initial brain connection graph representation field as a multi-modal contrast learning brain connection graph representation field in response to determining that the target difference value is less than a preset difference threshold; the data processing chip is configured to input multi-modal input data to be recognized into the multi-modal contrast learning brain connection graph representation field to obtain a motion sickness recognition result; and the data processing chip is configured to adjust a picture of the virtual reality device according to the motion sickness recognition result.

[0009] The above various embodiments of the present disclosure have the following beneficial effects: the motion sickness recognition system based on multi-modal contrast learning of some embodiments of the present disclosure improves the safety of users in using virtual devices. The reason for the low safety of users in using virtual devices is that the multi-modal fusion is performed in a simple feature splicing or stacking manner, and the real-time response delay and inaccurate intervention caused by the lack of cross-modal structure alignment mechanism, which leads to the virtual device being unable to improve the motion sickness state of the user by adjusting the virtual reality device, thereby reducing the safety of the user in using the virtual device. Based on this, some embodiments of the present disclosure provide a motion sickness recognition system based on multi-modal contrast learning, which comprises a virtual reality device, a data processing chip, a camera, an electroencephalogram receiver, and a motion signal receiver, wherein: the data processing chip is configured to collect multi-modal data based on a virtual reality scene from the virtual reality device, wherein the multi-modal data comprises electroencephalogram data, video data, and motion data; the data processing chip is configured to construct an initial brain connection graph representation field according to the electroencephalogram data, the video data, and the motion data, wherein the initial brain connection graph representation field comprises an electroencephalogram connection graph representation, a video and motion connection graph representation, and a standardized brain connection graph representation; the data processing chip is configured to perform the following training steps on the initial brain connection graph representation field: the data processing chip performs connection constraint contrast fusion on the video and motion connection graph and the electroencephalogram connection graph representation based on the standardized brain connection graph representation to obtain a fused multi-modal brain connection graph representation; the data processing chip determines a target difference value between the multi-modal brain connection graph representation and labeled motion sickness state data based on a preset difference measure function; the data processing chip determines the initial brain connection graph representation field as a multi-modal contrast learning brain connection graph representation field in response to determining that the target difference value is less than a preset difference threshold; the data processing chip is configured to input multi-modal input data to be recognized into the multi-modal contrast learning brain connection graph representation field to obtain a motion sickness recognition result; and the data processing chip is configured to adjust the picture of the virtual reality device according to the motion sickness recognition result. Thus, by reducing the virtual reality device delay and motion blur, the motion sickness state of the user is improved, thereby improving the safety of the user during use. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:

[0011] Figure 1is a structural schematic diagram of some embodiments of a motion sickness recognition system based on multi-modal contrast learning according to the present disclosure;

[0012] Figure 2 is a feature space distribution diagram generated by a connection constraint contrast fusion method according to the present disclosure. DETAILED DESCRIPTION

[0013] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0014] It should also be noted that, for ease of description, only parts related to the present application are shown in the drawings. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0015] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0016] It should be noted that the adjectives "one", "multiple" mentioned in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0017] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of these messages or information.

[0018] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0019] First, please refer to Figure 1 , Figure 1 The structural schematic diagram 100 of some embodiments of a motion sickness recognition system based on multi-modal contrast learning according to the present disclosure is shown. The motion sickness recognition system based on multi-modal contrast learning includes a virtual reality device 101, a data processing chip 102, a camera 103, an electroencephalogram signal receiver 104, and a motion signal receiver 105. The virtual reality device 101 is configured to obtain a video signal from the camera 103, collect an electroencephalogram signal from the electroencephalogram signal receiver 104, and obtain a motion signal from the motion signal receiver 105.

[0020] The data processing chip 102 is configured to collect multi-modal data from the virtual reality device 101 based on a virtual reality scene.

[0021] In some embodiments, the data processing chip 101 described above can be used to collect multi-modal data from the virtual reality device 101 described above based on a virtual reality scene. The multi-modal data can include electroencephalogram data, video data, and motion data. The virtual reality scene is constructed by a 3D engine, supports dynamic scenes where users can adjust the moving path, motion speed, acceleration, and turning rate, and supports triggering motion sickness in the exploration process. The virtual reality device 101 can include, but is not limited to, a head-mounted electroencephalogram acquisition device (such as Brain Products ActiCHamp) worn on the user's head and a virtual reality head-mounted device (such as Pico 4 Ultra).

[0022] In practice, the data processing chip 102 described above is configured to collect multi-modal data by the following steps:

[0023] First, the electroencephalogram data is collected and preprocessed. Here, the virtual reality device 101 can collect high-density electroencephalogram signals as electroencephalogram data through the electroencephalogram signal receiver 104 in the head-mounted electroencephalogram acquisition device (such as Brain Products ActiCHamp) worn on the user's head. The electroencephalogram acquisition device can include, but is not limited to, 32 electrode channels, an electrode layout conforming to the internationally accepted 10-20 system topology specification, covering key brain area locations such as frontal lobe, parietal lobe, central lobe, and occipital lobe. In practice, the electroencephalogram data sampling rate can be no less than 500 Hz, and the electrode impedance can be controlled to be below 5kΩ during the experiment. During the collection of electroencephalogram data, the data processing chip 102 can establish a communication connection with the recording terminal through a low-latency transmission protocol to monitor and record real-time data. At the same time, a high-precision time trigger module can be configured in the electroencephalogram acquisition device to synchronize the video and motion modalities at the millisecond level to ensure the temporal consistency of multi-modal data across devices.

[0024] In practice, the data processing chip 102 described above is configured to preprocess the electroencephalogram data using an open-source neural signal processing tool (such as EEGLAB) by the following steps:

[0025] Step one, resample the electroencephalogram data to obtain the resampled electroencephalogram data. Here, the data processing chip 102 can uniformly resample the original electroencephalogram data to 128 Hz to ensure that the key data features are preserved while reducing the complexity of the calculation.

[0026] Step two, re-reference the above sampled electroencephalogram data as electroencephalogram data. In practice, the above data processing chip 102 can take the double-mastoid electrode as the reference channel, determine the average potential difference between all active electrodes and the two reference electrodes, subtract the reference potential from the original signal to obtain the reference electroencephalogram data, so as to improve the consistency and spatial resolution between data channels.

[0027] Step three, remove the interference segments in the above electroencephalogram data to obtain electroencephalogram data without signal interference. In practice, the above data processing chip 102 can remove the power frequency interference and high frequency artifacts in the above electroencephalogram data through a 0.1-30 Hz band-pass filter. Here, non-physiological activity signals with an amplitude greater than or equal to 100 μV can be marked and removed. As an example, the above non-physiological activity signals can include but are not limited to electrode pop-out signals, cable movement signals, and electromagnetic interference signals.

[0028] Step four, perform independent component analysis (ICA) on the above signal interference-free electroencephalogram data to obtain electroencephalogram data for constructing multi-modal data. In practice, first, the above data processing chip 102 can perform ICA decomposition on the above electroencephalogram data to obtain a plurality of independent components and their corresponding spatial distribution maps and time series. Second, the spatial distribution characteristics of each component (such as the blink artifact reflected by the forehead region), the spectral characteristics (such as the high-frequency energy highlighted indicating electromyographic interference), and the time waveform characteristics (such as sudden high-amplitude fluctuations) can be comprehensively judged. Third, components identified as non-neurogenic signals (such as blink, electromyographic artifact components) can be removed, and the remaining components can be reconstructed into electroencephalogram data for constructing multi-modal data.

[0029] Second step, collect and preprocess the video data. In practice, the above virtual reality device 101 can collect the first perspective video data of the user in the use process through the camera 103 in the image acquisition module of the virtual reality head-mounted device (such as Pico 4 Ultra), to ensure complete capture of the speed, depth change and color difference details in the visual stimulation scene. The above video data frame rate can include but is not limited to 60 frames per second. The above process of collecting video data can rely on the VR SDK interface to realize seamless integration through the scene rendering module to record the real visible area of the user at each moment. At the same time, the scene rendering module can ensure that the video frames collected are consistent with the time stamps of the electroencephalogram and motion data by accessing an external synchronization controller, to realize accurate alignment between modalities.

[0030] In practice, the above data processing chip 102 is configured to perform the following preprocessing steps on the above video data:

[0031] Step 1, slice the video data by seconds. In practice, each video data can be divided into 1-second segments to ensure consistency with the EEG data and motion data.

[0032] Step 2, standardize the frame rate and resolution of the video data. In practice, the frame rate of the video data can be standardized to 30 frames per second, and the resolution can be standardized to 256x256 pixels.

[0033] Step 3, remove invalid pictures in the video data. The invalid pictures can include but are not limited to black screen, system setting interface and loading buffer frame.

[0034] Step 4, modal alignment of the video data. In practice, the data processing chip 102 can be configured to perform frame-level modal alignment on the time-offset segments of the video data to ensure the timing consistency of the multi-modal data.

[0035] Step 3, collect and preprocess the motion data. In practice, the data processing chip 102 can record the user's motion parameters in real time through the motion signal receiver 105 in the motion capture module built-in the virtual reality device 101 as motion data. The motion data can include but are not limited to: position data (Position, Vector3), linear velocity (Velocity, Vector3), acceleration (Acceleration, Vector3), angular velocity (AngularVelocity, Vector4) and immersion time (Immersion Time, Double). The motion data sampling frequency can include but are not limited to 30Hz.

[0036] In practice, the data processing chip 102 can be configured to perform the following preprocessing steps on the motion data:

[0037] Step 1, align the video data timestamp and downsample. In practice, the data processing chip 102 can divide the original 30Hz motion data by seconds to align with the EEG data and video data.

[0038] Step 2, remove abnormal data in the motion data. The abnormal data can include but are not limited to: missing data, garbled characters and illegal floating point value motion data.

[0039] Step three, re-encoding the motion data. Here, the data processing chip 102 can re-encode the motion data into UTF-8 encoding format while removing special characters and tag contents. The special characters can be characters in the string that do not belong to the regular character set. The tag contents can be ANSI control sequence characters used to mark the CLI color. As an example, the special characters can include but are not limited to control characters (\x00), escape characters, and non-printable characters.

[0040] Step four, identifying and removing extreme values in the motion data. In practice, the data processing chip 102 can use the interquartile range (IQR) method to identify extreme values in the motion data. First, calculate the first quartile (Q1) and the third quartile (Q3) of the motion data. Then, set the normal value range as [Q1-1.5xIQR, Q3+1.5xIQR]. Wherein, the IQR is the difference between the first quartile (Q1) and the third quartile (Q3). Finally, the motion data outside the range is determined as extreme value and removed.

[0041] Step four, time aligning the EEG data, the video data and the motion data to obtain multi-modal data. In practice, the data processing chip 102 can synchronize the three modal data according to the global timestamp carried by the collected video data, video data and motion data, as a standardized sample unit. The standardized sample unit can include but is not limited to: a 1-second EEG signal segment, a video frame sequence in the same time window, a synchronized motion behavior parameter and a corresponding subjective motion sickness label.

[0042] In practice, the data processing chip 102 can record the time when the user perceives the onset of motion sickness symptoms by triggering the button pressing operation through the virtual reality device 101, as the subjective feedback information. Here, the data processing chip 102 can generate a pair of 8-bit markers according to the subjective feedback information to record the start time and end time of the motion sickness symptoms. According to the start time and end time, the EEG data, video data and motion data can be synchronously marked. At the same time, the data processing chip 102 can millisecond-level calibrate the video data, video data and motion data in the time dimension through the synchronous module of software and hardware cooperation, to ensure the strict alignment relationship in the subsequent steps. In order to ensure the sample quality, the segment with missing or noise pollution in any modality will be discarded as a whole. In practice, the data processing chip 102 can divide the obtained training samples into fixed time windows according to the hierarchical acquisition strategy, and accurately calibrate the multi-modal data according to the start time of motion sickness, and then adjust the inter-class balance of the calibrated multi-modal data, and count the class ratio in the annotation result. When the non-motion sickness sample data (0 class) is more than 3 times of the motion sickness sample data (1 class), the inter-class balance adjustment mechanism is triggered, and the non-motion sickness sample is down-sampled and the motion sickness sample data is data-augmented, to ensure that the positive and negative sample ratio is maintained between 1:1 and 1:2 during training, so as to balance the label distribution of motion sickness events and non-motion sickness samples.

[0043] The data processing chip 102 is configured to construct an initial brain connection graph representation field according to the EEG data, video data and motion data.

[0044] In some embodiments, the data processing chip 102 can construct an initial brain connection graph representation field according to the EEG data, video data and motion data. The initial brain connection graph representation field can include an EEG connection graph representation, a video and motion connection graph representation, and a standardized brain connection graph representation. The data processing chip 102 can be an intelligent processing chip for converting EEG data, video data and motion data into motion sickness recognition results.

[0045] In some optional implementations of some embodiments, the data processing chip 102 can construct an initial brain connection graph representation field by the following steps:

[0046] Optionally, the data processing chip 102 is configured to construct an EEG connection graph representation and a video and motion connection graph representation by the following steps:

[0047] First, an EEG connection graph representation is constructed based on the EEG data.

[0048] In practice, the data processing chip 102 can determine the connection strength between brain regions and group the electroencephalogram data according to gender. The electroencephalogram data is taken as input and the motion sickness label is taken as output. The graph convolutional neural network corresponding to different genders is trained. The graph weight parameters in the trained and converged graph convolutional neural network are extracted as the connection strength of the electroencephalogram connection graph representation. Here, the formula for constructing the electroencephalogram connection graph representation of the data processing chip 102 can be as follows:

[0049] G E =(V, E, W E ).

[0050] Wherein, V represents the corresponding brain region node set of different functional regions of the brain. E represents the set of edges constructed according to the nodes. W E represents the connection strength between brain regions related to motion sickness. G E represents the constructed electroencephalogram connection graph representation.

[0051] Secondly, the visual features of the video data are extracted. The visual features include the encoding of static frames and optical flow images. In practice, the data processing chip 102 can use a convolutional neural network and a variational autoencoder to encode the original video frames and the optical flow images corresponding thereto, and extract static visual features and dynamic visual features.

[0052] Thirdly, the motion data is modeled by motion parameters to obtain a structured representation of the motion data. In practice, the data processing chip 102 can encode the motion data by a feature extractor to output a structured representation D m of the motion data with a dimension matching the electroencephalogram connection graph representation. These features can be fused by a neural prior attention encoder to output a weighted adjacency matrix of the connection strength between the video and the motion connection graph representation. The extraction process can be represented by the following mapping function:

[0053] f NAPE :(D v , D o , D m )→W MV .

[0054] Wherein, D v represents the original video frame. D o represents the optical flow image corresponding to the original video frame. f NAPE represents the neural prior attention encoder. D m represents the structured representation of the motion data. W MV represents the connection strength matrix between the video data and the motion data related to motion sickness.

[0055] Fourthly, the visual features and the structured representation of the motion data are mapped into a graph space corresponding to the brain region topology to construct a video-motion connectivity graph representation associated with brain function. In practice, the data processing chip 102 can map the visual features and the structured representation of the motion data into the brain region topology space through graph position coding, align the node positions with the brain function, and adjust the connection strength between the nodes in the video-motion connectivity graph representation guided by the edge weight of the electroencephalogram connectivity graph representation. Then, the brain function grouping can be integrated through the topology calibration constraint and the sparse regularization term, and the video-motion connectivity graph representation associated with the brain function can be output.

[0056] Fifthly, the key regions in the video-motion connectivity graph representation are enhanced through a neural prior attention module to form an enhanced video-motion connectivity graph representation as the video-motion connectivity graph representation. In practice, the data processing chip 102 can enhance the video and motion data by taking the brain function distribution in neuroscience as the prior feature, and map the feature to the video-motion connectivity graph representation conforming to the 10-20 electroencephalogram system. This can be done by the following formula:

[0057] G MV =(V, E, W NV ).

[0058] Wherein, V represents a set of corresponding brain region nodes of different functional regions of the brain. E represents a set of edges constructed according to the nodes. W MV represents the connection strength matrix of the video data and the motion data related to motion sickness. G MV represents the constructed video-motion connectivity graph representation.

[0059] In some optional implementations of the embodiments, the data processing chip 102 is configured to further perform the following steps on the video-motion connectivity graph:

[0060] Step one, the video-motion connectivity graph representation is optimized and a loss function is determined. In practice, the data processing chip 102 can constrain the constructed video-motion connectivity graph structure and the electroencephalogram connectivity graph representation by introducing a joint loss function of mean square error and sparse regularization term. This can be done by the following formula:

[0061] L NPAE =L MSE +λL1.

[0062]

[0063] Wherein, i represents the first to the Mth sample. represents the electroencephalogram connectivity graph representation weight corresponding to the i-th sample. represents the video-motion connectivity graph representation weight corresponding to the i-th sample. λ represents the sparse regularization factor. N represents the number of samples. M represents the total number of samples. L NPAE represents the joint loss function term. L MSE represents the mean square error term. W MV represents the connection strength matrix between the video data and the motion data related to motion sickness. W E represents the connection strength between brain regions related to motion sickness. L1 represents the sparse regularization term of the norm. F represents the Frobenius norm. ||1 represents the L1 norm. represents the square of the Frobenius norm.

[0064] Step two, standardize the video-motion connectivity graph representation and output to obtain the standardized video-motion connectivity graph representation in the unified modality as the video-motion connectivity graph representation. In practice, the data processing chip 102 can perform topology calibration and node embedding dimension alignment operations on the video-motion connectivity graph representation. The topology calibration can be achieved by adjusting the node connection mode of the video-motion connectivity graph to be consistent with the topology of the electroencephalogram connectivity graph representation (such as the distribution of brain regions in the 10-20 system). The node embedding dimension alignment can be achieved by linear transformation to map the features of all modalities to the same vector space, ensuring the comparability of data from different sources. The final output video-motion connectivity graph representation can be determined as the standardized video-motion connectivity graph representation in the unified modality.

[0065] The first step-step two and its related content as one of the invention points of the embodiments of the present disclosure, combined with the following step "step 105", solves the technical problem "by independently processing video content analysis (such as frame rate / optical flow) and motion sensor data, lack of association with neurophysiological mechanisms, so that the virtual device cannot coordinate the conflicting signals of the visual and vestibular neural pathways, resulting in the user still having motion sickness in the picture adjusted by the virtual device, causing the user's safety during the experience process to be low", which leads to the factors that the virtual device cannot adjust the device parameters by coordinating the conflicting signals of the visual and vestibular neural pathways. The factors are often as follows: by independently processing video content analysis (such as frame rate / optical flow) and motion sensor data, lack of association with neurophysiological mechanisms, so that the virtual device cannot coordinate the conflicting signals of the visual and vestibular neural pathways, resulting in the user still having motion sickness in the picture adjusted by the virtual device, causing the user's safety during the experience process to be low. If the above factors are solved, the virtual device can highly match the actual discomfort state of the user, and then adjust the picture parameters, thereby improving the safety of the user when using the virtual device. In order to achieve this effect, first, based on the above EEG data, an EEG connection graph representation is constructed. Second, the visual features of the above video data are extracted. Third, the motion data is modeled for motion parameters to obtain a structured representation of the motion data. Fourth, the visual features and the structured representation of the motion data are mapped into a graph space corresponding to the brain region topology to construct a video and motion connection graph representation associated with the function of the brain region. Fifth, through a neural prior attention module, the key areas in the video and motion connection graph representation are enhanced in features to form an enhanced video and motion connection graph representation as the video and motion connection graph representation. Sixth, the video and motion connection graph representation is optimized and a loss function is determined. Seventh, the video and motion connection graph representation is standardized and output to obtain a standardized video and motion connection graph representation in a unified modality as the video and motion connection graph representation. By constructing the EEG connection graph representation and the video and motion connection graph representation, combined with feature extraction, the motion sickness state recognition is further completed. Combined with step "step 105", according to the motion sickness recognition result, the virtual reality device is adjusted. Thus, by adjusting the virtual reality device parameters, the user can avoid the situation of motion sickness, thereby improving the safety of the user when using the virtual device.

[0066] Third, based on the individual's EEG connection graph representation, a standardized brain connection graph representation is generated.

[0067] In practice, the above data processing chip 102 is configured to generate a standardized brain connection graph representation based on the individual's EEG connection graph representation by the following steps:

[0068] Step one, constructing population characteristic oriented electroencephalogram connectivity graph representation. In practice, the above data processing chip 102 can process the electroencephalogram data by individual difference factors affecting motion sickness reaction, and obtain the data of each group. The data of each group can be learned by dynamic graph convolutional neural network to generate gender-specific electroencephalogram connectivity graph representation. Wherein, the individual difference factors affecting the motion sickness reaction can include but not limited to user gender and age. Here, the following formula can be realized:

[0069]

[0070] Wherein, m represents the gender grouping identifier. G E represents the constructed electroencephalogram connectivity graph representation. represents the constructed electroencephalogram connectivity graph representation of the corresponding gender group. W E represents the connection strength between brain regions related to motion sickness. represents the adjacency matrix of the corresponding gender group. V represents the set of corresponding brain region nodes of different functional regions of the brain. E represents the set of edges constructed according to the nodes. represents the adjacency matrix of the corresponding gender group of the ith sample. E i represents the adjacency matrix of the ith sample. i represents the traversal of the first to the N m sample. N represents the number of samples. N m represents the number of samples in this gender group. T represents the transpose operation of the matrix. i represents the traversal of the first to the N m sample.

[0071] Step two, generating shared connectivity graph structure by introducing standardization decomposition algorithm. In practice, the above data processing chip 102 can extract electroencephalogram data with consistent characteristics in the electroencephalogram connectivity graph representation between different individuals, and can separate the shared component W from by standardization decomposition algorithm, while eliminating individual-specific noise. Here, the above data processing chip 102 can realize by solving the optimization problem:

[0072]

[0073] Wherein, K = Φ(W E ) represents the kernel matrix after mapping the original adjacency matrix to the reproducing kernel Hilbert space. Φ() represents the Gaussian kernel function. W represents the shared connectivity graph structure. H represents the gender-specific component. γ represents the regularization factor, which is used to suppress the feature coupling interference between different individuals. tr represents the trace operation, which is used to measure the mutual orthogonality of the feature space. W E represents the connection strength between brain regions related to motion sickness. T represents the transpose operation of the matrix. denotes the square of the Frobenius norm. F denotes the Frobenius norm. H1 denotes the gender-specific connectivity pattern matrix of the first gender group (e.g. male group). H2 denotes the gender-specific connectivity pattern matrix of the second gender group (e.g. female group). K1 denotes the kernel matrix of the first gender group (e.g. male group) mapped by the Gaussian kernel function Φ(). K2 denotes the kernel matrix of the second gender group (e.g. female group) mapped by the Gaussian kernel function Φ().

[0074] Step three, optimizing the standardized brain connectivity graph representation based on the multiplication iterative update rule. Here, the data processing chip 102 can be implemented by the following formula:

[0075] updating the shared standardized brain connectivity graph representation matrix:

[0076]

[0077] updating the individual-specific feature component:

[0078]

[0079] wherein K = Φ(W E ) denotes the kernel matrix of mapping the original adjacency matrix to the reproducing kernel Hilbert space. Φ() denotes the Gaussian kernel function. W denotes the shared connectivity graph structure. H denotes the gender-specific component. γ denotes the regularization factor. ∈ denotes the numerical stability parameter to prevent division by zero. T denotes the transpose operation of the matrix. H1 denotes the gender-specific connectivity pattern matrix of the first gender group (e.g. male group). H2 denotes the gender-specific connectivity pattern matrix of the second gender group (e.g. female group). K1 denotes the kernel matrix of the first gender group (e.g. male group) mapped by the Gaussian kernel function Φ(). K2 denotes the kernel matrix of the second gender group (e.g. female group) mapped by the Gaussian kernel function Φ().

[0080] Step four, outputting the standardized brain connectivity graph representation. In practice, the data processing chip 102 can output the standardized brain connectivity graph representation by the following formula:

[0081]

[0082] G S = (V, E, W S ).

[0083] wherein W denotes the shared connectivity graph structure. T denotes the transpose operation of the matrix. V denotes the corresponding set of brain region nodes of different functional regions of the brain. E denotes the set of edges constructed according to the nodes. W S denotes the connection strength related to motion sickness of the standardized brain connectivity graph representation. G S denotes the constructed standardized brain connectivity graph representation.

[0084] The data processing chip 102 is configured to perform the following training steps on the initial brain connectivity graph representation field:

[0085] The data processing chip 102 performs connection constraint comparison fusion on the video and motion connectivity graph and the electroencephalogram connectivity graph based on the standardized brain connectivity graph representation to obtain a fused multi-modal brain connectivity graph representation.

[0086] In some embodiments, the data processing chip 102 described above is configured to perform connection constraint comparison fusion on the video and motion connectivity graph and the electroencephalogram connectivity graph based on the standardized brain connectivity graph representation to obtain a fused multi-modal brain connectivity graph representation. In practice, this can be achieved through the following sub-steps:

[0087] Sub-step one, based on the standardized brain connectivity graph representation, data augmentation is performed on the electroencephalogram connectivity graph and the video and motion connectivity graph to form a plurality of graph structure views. In practice, first, the data processing chip 102 can randomly discard part of the connection edges of the adjacency matrix of the electroencephalogram connectivity graph and the video and motion connectivity graph through a mask mechanism with a probability of 10%-20%, while adding random noise disturbance conforming to a Gaussian distribution (mean value of 0, standard deviation of 0.1) to the edge weights of the retained edges. Second, at the feature space level, the data processing chip 102 can use a graph convolution network to extract node-level feature embedding and use a Mixup strategy for feature mixing, and perform linear interpolation on the graph representations of the two modalities through the mixing coefficients sampled from the Beta distribution. Third, the data processing chip 102 can implement oversampling on key connection regions (such as the prefrontal-parietal connection region) with high centrality in the graph based on the topological prior knowledge provided by the standardized brain connectivity graph representation to strengthen important features, and can perform random subgraph sampling on non-key connection regions (such as the ipsilateral temporal lobe low-frequency connection region) to increase data diversity and generate enhanced views with continuous transition characteristics.

[0088] Sub-step two, classification processing is performed on the plurality of graph structure views to generate a positive sample set and a negative sample set. Each positive sample in the positive sample set is the electroencephalogram connectivity graph representation and the video and motion connectivity graph representation corresponding to the same sample. Each negative sample in the negative sample set is the electroencephalogram connectivity graph representation and the video and motion connectivity graph representation corresponding to different samples.

[0089] Sub-step three, based on the positive sample set and the negative sample set, a similarity matrix is constructed to obtain a fused multi-modal brain connectivity graph representation.

[0090] In practice, the data processing chip 102 can map the electroencephalogram connectivity graph representation and the video and motion connectivity graph representation to a shared latent space based on the InfoNCE objective function to obtain a fused connectivity graph representation, and construct a similarity matrix for the positive and negative sample pairs in the classified graph structure view. Here, the data processing chip 102 can be implemented by the following formula:

[0091]

[0092] G * = (V, E, W * ).

[0093] Where S represents the similarity matrix. τ represents the temperature coefficient for adjusting the sensitivity of the difference between samples. V represents the corresponding brain region node set of different functional regions of the brain. E represents the set of edges constructed according to the nodes. G * represents the fused normalized brain connectivity graph representation as the final feature basis for motion sickness discrimination. W * represents the generated fused connectivity graph structure.

[0094] In some optional implementations of the embodiments, the data processing chip 102 is configured to visualize the effect before and after fusion by a dimensionality reduction algorithm. This can be achieved by the following steps:

[0095] First, analyze the distribution of the sample set before and after fusion without connection constraint, and obtain the sample set distribution before fusion. The sample set includes positive sample set and negative sample set. In practice, the data processing chip 102 can extract the sample set before and after fusion without connection constraint, map it to a two-dimensional coordinate system, and obtain the distribution of the positive sample set (the same sample corresponds to the electroencephalogram connectivity graph representation and the video and motion connectivity graph representation, i.e. motion sickness sample) and the negative sample set (different samples correspond to the electroencephalogram connectivity graph representation and the video and motion connectivity graph representation, i.e. non-motion sickness sample) before fusion.

[0096] Second, analyze the distribution of the sample set before and after fusion with connection constraint, and obtain the sample set distribution after fusion. Here, the data processing chip 102 can extract the sample set before and after fusion with connection constraint, map it to a two-dimensional coordinate system, and obtain the distribution of the positive sample set and the negative sample set after fusion. At the same time, the data processing chip 102 can mark the optimization effect of the sample set after fusion by marking false positive samples (incorrectly identifying normal state as motion sickness) and false negative samples (incorrectly identifying motion sickness state as normal).

[0097] Third, the distribution of the sample set before fusion and the distribution of the sample set after fusion are visualized to obtain a two-dimensional feature distribution map. As an example, the data processing chip 102 can visualize the obtained distribution to obtain a two-dimensional feature distribution map to represent the distribution of the positive sample set and the negative sample set. Among them, the left side of the two-dimensional feature distribution map is the sample distribution before fusion, and the right side is the sample distribution after fusion. The distribution shows that the boundary of the feature distribution of the positive sample set and the negative sample set before fusion is blurred and overlaps a lot. The boundary of the feature distribution of the positive sample set and the negative sample set after fusion is clear, the distance between classes is significantly increased, and the false positive and false negative sample points are significantly reduced. The feature space distribution map generated by the connection constraint comparison fusion method is shown in FIG. 8. Figure 2

[0098] The data processing chip 102 determines the target difference value between the multi-modal brain connection map representation and the labeled motion sickness state data based on a preset difference measure function.

[0099] In some embodiments, the data processing chip 102 can determine the target difference value between the multi-modal brain connection map representation and the labeled motion sickness state data based on a preset difference measure function. In practice, the data processing chip 102 can calculate the sum of the neural prior attention encoding loss, the contrastive learning loss, the classification loss, the connection centrality loss, and the binary structure loss as the total loss function value. Among them, the neural prior attention encoding loss can calculate the difference between the video and motion connection map representation and the electroencephalogram connection map representation through the weighted sum of the mean square error and the sparse regularization term. The contrastive learning loss can be obtained by calculating the similarity difference of positive and negative samples in the shared latent space. The classification loss can measure the difference between the prediction result and the true label through the cross-entropy function. The connection centrality loss can calculate the difference between the fusion graph and the standardized graph in the node centrality index. The binary structure loss can measure the similarity of the two graph structures in the existence of edges through the binary cross-entropy.

[0100] The data processing chip 102 determines the target difference value between the multi-modal brain connection map representation and the labeled motion sickness state data based on a preset difference measure function.

[0101] ​In some embodiments, the data processing chip 102 can determine the initial brain connectivity graph representation field as the multi-modal contrastive learning brain connectivity graph representation field in response to determining that the target difference value is less than the preset difference threshold. In practice, the data processing chip 102 can determine whether the training process is terminated by judging the size relationship between the target difference value and the preset difference threshold. When the target difference value is less than the preset difference threshold, the training process is terminated. The initial brain connectivity graph representation field at this time can be determined as the final multi-modal contrastive learning brain connectivity graph representation field. The multi-modal contrastive learning brain connectivity graph representation field can be used to determine the motion sickness state, and the fusion connectivity graph structure generated thereby retains the uniqueness of each modality while establishing a cross-modal correlation pattern.

[0102] Optionally, the data processing chip 102 can be further configured to:

[0103] In the first step, in each training iteration, the connection centrality loss and the binary structure loss between the fused multi-modal brain connectivity graph representation and the standardized brain connectivity graph representation are determined, and the connection centrality loss and the binary structure loss are taken as the structure constraint term. In practice, the data processing chip 102 can be implemented by the following formula:

[0104]

[0105] L bin = BCE(I(G * > 0), I(G S > 0)).

[0106] wherein, represents the weighted degree centrality of node i in graph G. L deg represents the connection centrality loss. L bin represents the binary structure loss. i represents the target node for which the centrality is calculated. i represents the traversal of the first to the Nth sample. j represents the traversal of all nodes connected to i. A ij represents the connection strength between node i and node j in the adjacency matrix. N represents the number of samples. C i is an indicator function. represents the node centrality of the real graph. represents the node centrality of the standardized brain connectivity graph. BCE() represents the binary cross-entropy loss, which is used to measure the binary similarity of two graph structures in terms of the existence of edges. G S represents the constructed standardized brain connectivity graph representation. G * represents the fused standardized brain connectivity graph representation, which is the final feature basis for motion sickness determination. I() represents a binary function that outputs 1 when the condition is met, and otherwise outputs 0.

[0107] Secondly, the structural constraint term is combined with the contrastive learning loss, the classification loss and the neural prior attention encoding loss to form a joint loss function. In practice, the joint loss function can be realized by the following formula:

[0108] The total loss function (L total ) can be realized by the following formula:

[0109] L total = L NPAE + L NCE + L CE + L deg + L bin .

[0110] Wherein, L total represents the total loss. L NPAE represents the neural prior attention encoding loss. L NCE represents the contrastive learning loss. L CE represents the classification loss. L deg represents the connection centrality loss. L bin represents the binary structure loss.

[0111] The neural prior attention encoding loss (L NPAE ) can be realized by the following formula:

[0112] L NPAE = L MSE + λL1.

[0113]

[0114] Wherein, L NPAE represents the neural prior attention encoding loss. L1 represents the sparse regularization term of norm. L MSE represents the mean square error term. λ represents the sparse regularization factor. N represents the sample quantity. M represents the total sample quantity. i represents the first to the Mth sample. represents the video and motion connection weight matrix of the ith sample. represents the electroencephalogram connection weight matrix of the ith sample. F represents the Frobenius norm. ||1 represents the L1 norm. represents the square of the Frobenius norm.

[0115] The contrastive learning loss (L NCE ) can be realized by the following formula:

[0116]

[0117] Wherein, L NCE represents the contrastive learning loss. L CErepresents a classification loss. S represents a similarity matrix. S T represents a transpose of S. y represents a label vector.

[0118] The classification loss (L CE ) can be implemented by the following formula:

[0119]

[0120] wherein, L CE represents a classification loss. i represents traversing the first to the Nth sample. N represents the number of samples. y represents a label vector. y i represents a label vector of the ith sample. S represents a similarity matrix. S i represents a similarity matrix of the ith sample. softmax() represents a normalized exponential function.

[0121] The third step, based on an end-to-end manner, updates the parameters of the entire network according to the above joint loss function. In practice, the above data processing chip 102 can calculate the gradient and synchronously update all parameters from the data encoding layer, the graph comparison fusion layer to the classification output layer by using a unified optimizer (such as Adam) to calculate the gradient of the above constituted joint loss function through a back propagation algorithm, and realize multi-module collaborative optimization.

[0122] The fourth step, through a global maximum pooling operation, extracts the node features of the above multi-modal brain connection graph representation. In practice, the above data processing chip 102 can perform a global maximum pooling operation on the feature vector of each node in the above multi-modal brain connection graph representation, take the maximum value of all channels along the node dimension, and compress the variable size multi-modal brain connection graph representation into a fixed length node feature.

[0123] The fifth step, the above node features are input into a multi-layer perceptron to complete the motion sickness state recognition. In practice, the above data processing chip 102 can input the above node features into a multi-layer perceptron to complete the motion sickness state recognition. A binary cross-entropy loss can be used to calculate the difference between the predicted value and the true label y, and the weights of the multi-layer perceptron are adjusted through gradient descent. It can be implemented by the following formula:

[0124]

[0125] wherein, represents the final predicted label of whether motion sickness occurs. Pooling() represents a pooling operation on the input matrix. MLP() represents a classification operation on the pooled features. W * represents a generated fusion connection graph structure.

[0126] The first step to the fifth step and its related content as one of the invention points of the embodiment of the present disclosure, combined with the following step "step 105", solves the technical problem "the physiological signal and behavior data are misaligned when multi-modal feature fusion occurs, which makes the adaptive adjustment ability of the virtual reality device insufficient due to individual differences, so that the picture adjusted by the virtual device still causes the user to have motion sickness, resulting in low safety of the user during the experience process", which leads to the factors that the virtual device cannot adjust the device parameters according to the individual differences of the user, which are often as follows: the physiological signal and behavior data are misaligned when multi-modal feature fusion occurs, which makes the adaptive adjustment ability of the virtual device insufficient due to individual differences, so that the picture adjusted by the virtual device still causes the user to have motion sickness, resulting in low safety of the user during the experience process. If the above factors are solved, the virtual device can be highly matched to the actual discomfort state of the user, and then the picture parameters are adjusted to improve the safety of the user when using the virtual device. In order to achieve this effect, first, in each training iteration, the connection centrality loss and the binary structure loss between the fused multi-modal brain connection graph representation and the standardized brain connection graph representation are calculated, and the connection centrality loss and the binary structure loss are used as structure constraint terms. Second, the structure constraint term, the contrastive learning loss, the classification loss and the neural prior attention encoding loss constitute a joint loss function. Third, the parameters of the entire network can be updated based on an end-to-end manner according to the joint loss function. Fourth, the node features of the multi-modal brain connection graph representation can be extracted through a global maximum pooling operation. Fifth, the node features can be input into a multi-layer perceptron to complete motion sickness state recognition. Through the joint loss function and the end-to-end updating method, combined with feature extraction, the recognition of motion sickness state is further completed. Combined with step "step 105", the virtual reality device is adjusted according to the motion sickness recognition result. Thus, by adjusting the parameters of the virtual reality device, the user can be prevented from having motion sickness, thereby improving the safety of the user when using the virtual device.

[0127] The data processing chip 102 is configured to input the multi-modal input data to be recognized into the multi-modal contrastive learning brain connection graph representation field to obtain a motion sickness recognition result.

[0128] In some embodiments, the data processing chip 102 can input the multi-modal input data to be identified into a multi-modal contrastive learning brain connectivity graph representation field to obtain a motion sickness identification result. In practice, the data processing chip 102 can collect the brain electrical data, video data and motion data (such as acceleration and angular velocity) of the user in real time through a virtual reality device (such as a head-mounted display). After preprocessing the collected brain electrical data, video data and motion data, the multi-modal input data is obtained. The multi-modal input data is input into the multi-modal contrastive learning brain connectivity graph representation field, and a motion sickness probability score between 0 and 1 is output. If the score exceeds a preset threshold (such as 0.5), it can be determined that the user is currently in a motion sickness state.

[0129] In some optional implementations of some embodiments, the data processing chip 102 can analyze and process the motion sickness identification result to generate an analysis result. The analysis result includes identification accuracy, identification precision, identification sensitivity, identification specificity, and area under the curve. The identification sensitivity is used to measure the identification ability of the positive sample set, and the identification specificity is used to measure the identification ability of the negative sample set. In practice, the data processing chip 102 can determine the convergence performance of the multi-modal contrastive learning brain connectivity graph representation field based on an early stopping strategy. The early stopping strategy includes triggering a condition that the validation set loss does not improve within a preset number of rounds. Here, the data processing chip 102 can continuously monitor the validation set loss based on the convergence performance evaluation of the early stopping strategy. The patience threshold (such as 10 consecutive rounds) can be set at initialization and the counter can be reset. If the current validation loss is lower than the historical best value during training, the best loss is updated and the counter is reset. Otherwise, the counter is incremented by 1. When the counter reaches the preset patience value, early stopping is triggered, the parameters of this round are saved as the final model, and the performance indicators at convergence are output.

[0130] The data processing chip 102 is configured to adjust the picture of the virtual reality device 101 according to the motion sickness identification result.

[0131] In some embodiments, the data processing chip 102 can adjust the picture of the virtual reality device 101 according to the motion sickness identification result obtained by the data processing chip.

[0132] In some optional implementations of some embodiments, the data processing chip 102 can adjust the parameters of the virtual reality device based on the identified degree of motion sickness (e.g., mild, severe) to optimize the user’s experience after obtaining the motion sickness identification result. In practice, the data processing chip 102 can employ a hierarchical response mechanism by strategy to take different intervention measures for different degrees of motion sickness. As an example, if the identification result is mild motion sickness (threshold between 0.5 and 0.7), the data processing chip 102 can reduce the dynamic complexity of the virtual scene in the virtual reality device 101, such as reducing the number of moving objects or reducing the speed of movement, while appropriately reducing the user’s field of view to reduce the interference of visual flow to the vestibular system. If the identification result is severe motion sickness (threshold > 0.7), the data processing chip 102 can take intervention measures such as automatically switching the virtual scene in the virtual reality device 101 to a static scene (e.g., a virtual rest room) and reducing the screen refresh rate of the virtual reality device 101 (e.g., from 90 Hz to 72 Hz) to relieve visual fatigue. At the same time, the data processing chip 102 can optimize in real time based on the identification result. For example, when the user continues to have motion sickness symptoms after adjustment, the displacement of the virtual camera can be adjusted in the opposite direction to offset the user’s head movement, or dynamic blur frames can be inserted between frames to smooth the visual transition. At the same time, the virtual reality device 101 will remind the user to take a break through voice prompts or tactile feedback (e.g., handle vibration), and record the effects of different adjustment strategies for subsequent optimization of the intervention scheme.

[0133] The above various embodiments of the present disclosure have the following beneficial effects: the motion sickness recognition system based on multi-modal contrast learning of some embodiments of the present disclosure improves the safety of users using virtual devices. The reason for the low safety of users using virtual devices is that the multi-modal fusion is performed in a simple feature splicing or stacking manner, and the real-time response delay and inaccurate intervention caused by the lack of cross-modal structure alignment mechanism, which leads to the virtual device being unable to improve the motion sickness state of the user by adjusting the virtual reality device, thereby reducing the safety of the user using the virtual device. Based on this, some embodiments of the present disclosure provide a motion sickness recognition system based on multi-modal contrast learning, which comprises a virtual reality device, a data processing chip, a camera, an electroencephalogram receiver, and a motion signal receiver, wherein: the data processing chip is configured to collect multi-modal data based on a virtual reality scene from the virtual reality device, wherein the multi-modal data comprises electroencephalogram data, video data, and motion data; the data processing chip is configured to construct an initial brain connection graph representation field according to the electroencephalogram data, the video data, and the motion data, wherein the initial brain connection graph representation field comprises an electroencephalogram connection graph representation, a video and motion connection graph representation, and a standardized brain connection graph representation; the data processing chip is configured to perform the following training steps on the initial brain connection graph representation field: the data processing chip performs connection constraint contrast fusion on the video and motion connection graph and the electroencephalogram connection graph representation based on the standardized brain connection graph representation to obtain a fused multi-modal brain connection graph representation; the data processing chip determines a target difference value between the multi-modal brain connection graph representation and labeled motion sickness state data based on a preset difference measure function; the data processing chip determines the initial brain connection graph representation field as a multi-modal contrast learning brain connection graph representation field in response to determining that the target difference value is less than a preset difference threshold; the data processing chip is configured to input multi-modal input data to be recognized into the multi-modal contrast learning brain connection graph representation field to obtain a motion sickness recognition result; and the data processing chip is configured to adjust the picture of the virtual reality device according to the motion sickness recognition result. Thus, by reducing the virtual reality device delay and motion blur, the motion sickness state of the user is improved, thereby improving the safety of the user during use.

Claims

1. A motion sickness recognition system based on multi-modal contrastive learning, wherein, The motion sickness recognition system comprises a virtual reality device, a data processing chip, a camera, an electroencephalogram signal receiver and a motion signal receiver, wherein: The data processing chip is configured to collect multi-modal data based on a virtual reality scene from the virtual reality device, wherein the multi-modal data comprises electroencephalogram data, video data and motion data; The data processing chip is configured to construct an initial brain connection graph representation field according to the electroencephalogram data, the video data and the motion data, wherein the initial brain connection graph representation field comprises an electroencephalogram connection graph representation, a video and motion connection graph representation and a standardized brain connection graph representation; The data processing chip is configured to perform the following training steps on the initial brain connection graph representation field: The data processing chip is configured to perform connection constraint comparative fusion on the video and motion connection graph and the electroencephalogram connection graph based on the standardized brain connection graph representation, to obtain a fused multi-modal brain connection graph representation; The data processing chip is configured to determine a target difference value between the multi-modal brain connection graph representation and labeled motion sickness state data based on a preset difference measurement function; The data processing chip is configured to determine the initial brain connection graph representation field as a multi-modal comparative learning brain connection graph representation field in response to determining that the target difference value is less than a preset difference threshold; The data processing chip is configured to input multi-modal input data to be recognized into the multi-modal comparative learning brain connection graph representation field, to obtain a motion sickness recognition result; The data processing chip is configured to adjust a picture of the virtual reality device according to the motion sickness recognition result.

2. Motion sickness recognition system according to claim 1, wherein The data processing chip is configured to: perform data enhancement on the electroencephalogram connection graph representation and the video and motion connection graph representation based on the standardized brain connection graph representation, to form a plurality of graph structure views; perform classification processing on the plurality of graph structure views, to generate a positive sample set and a negative sample set, wherein each positive sample in the positive sample set is a same sample corresponding electroencephalogram connection graph representation and video and motion connection graph representation, and each negative sample in the negative sample set is different sample corresponding electroencephalogram connection graph representation and video and motion connection graph representation; construct a similarity matrix based on the positive sample set and the negative sample set, to obtain a fused multi-modal brain connection graph representation.

3. Motion sickness recognition system according to claim 2, wherein The data processing chip is further configured to: analyze a distribution of a sample set that has not been subjected to connection constraint comparative fusion, to obtain a distribution of a sample set before fusion, wherein the sample set comprises the positive sample set and the negative sample set; analyze a distribution of a sample set that has been subjected to connection constraint comparative fusion, to obtain a distribution of a sample set after fusion; perform visual processing on the distribution of the sample set before fusion and the distribution of the sample set after fusion, to obtain a two-dimensional feature distribution graph.

4. The motion sickness recognition system of claim 1, wherein, The data processing chip is configured to: The motion sickness recognition result is analyzed and processed to generate an analysis result, wherein the analysis result includes recognition accuracy, recognition precision, recognition sensitivity, recognition specificity, and area under the curve, the recognition sensitivity is used to measure the recognition ability of the positive sample set, and the recognition specificity is used to measure the recognition ability of the negative sample set; Based on the early stopping strategy, the convergence performance of the multi-modal contrast learning brain connection graph representation field is determined, wherein the early stopping strategy includes a condition that the validation set loss does not improve within a preset number of rounds.

5. The motion sickness recognition system of claim 1, wherein, The data processing chip is configured to: Collect high-density electroencephalogram signals as electroencephalogram data based on a head-mounted electroencephalogram acquisition device from the virtual reality device; Collect video data of the first perspective of the user through the virtual reality head-mounted device; Real-time record motion parameters of the user as motion data through the virtual reality device, wherein the motion data includes position data, linear velocity, acceleration, angular velocity, and immersion time; Time-align the electroencephalogram data, the video data, and the motion data to obtain multi-modal data.

6. Motion sickness recognition system according to claim 5, wherein The data processing chip is further configured to: Resample the electroencephalogram data to obtain sampled electroencephalogram data; Re-reference the sampled electroencephalogram data as electroencephalogram data; Remove interference segments in the electroencephalogram data to obtain electroencephalogram data without signal interference; Perform independent component analysis on the electroencephalogram data without signal interference to obtain electroencephalogram data used to construct the multi-modal data.