A vehicle control method, system, device and medium based on electroencephalogram signals and a multi-modal large model

By establishing a thought-action mapping dictionary in both stationary and autonomous driving states and decoding driving intention features using a multimodal large model, combined with visual evoked potential signals for fusion and conflict arbitration, the low accuracy of intention recognition and the time lag of physical takeover methods in brain-computer interface-assisted driving are solved, achieving high-precision vehicle control and smooth human-machine collaboration.

CN122443483APending Publication Date: 2026-07-24BEIJING ELECTRIC VEHICLE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ELECTRIC VEHICLE
Filing Date
2026-06-05
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, brain-computer interface-assisted driving suffers from low intent recognition accuracy, leading to driver cognitive fatigue, insufficient control resolution, inconsistencies between autonomous driving decisions and human intentions, time lag in physical takeover methods, weak proactive correction capabilities, and poor smoothness of human-machine collaboration.

Method used

By establishing a thought-action mapping dictionary when the vehicle is stationary, the driving intention features are decoded in autonomous driving mode using EEG signals and a multimodal large model. Combined with visual evoked potential signals, multiple alternative adjustment operations are generated, and the final operation is determined through conflict arbitration. Error signals are analyzed in real time for correction and adjustment.

Benefits of technology

It improves the accuracy of driver intent recognition and spatial trajectory control, realizes dynamic collaboration between human and machine decision-making, enhances the vehicle's active adaptive correction capability under abnormal operating conditions, and improves driving safety and continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122443483A_ABST
    Figure CN122443483A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of vehicle control method, system, equipment and medium based on electroencephalogram and multimodal big model, wherein, method includes: in stationary state, establish thought action mapping dictionary;In the automatic driving state, the second state electroencephalogram is decoded into driving intention feature in conjunction with the dictionary;Through vehicle-mounted visual guidance device, virtual thought center of gravity is projected, and the comprehensive driving intention is obtained by fusing the evoked visual evoked potential signal and driving intention feature;Real-time vehicle environment perception data and comprehensive driving intention are used to multimodal big model Cross-modal collaborative reasoning and conflict collaborative arbitration to determine final adjustment operation;After execution, third state electroencephalogram is analyzed, if it is judged that there is error related negative potential signal exceeding threshold, then adaptive generation and execution of rectification adjustment operation are carried out. Therefore, the present application reduces the cognitive load of driver, improves trajectory control precision, improves man-machine cooperation smoothness and driving safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent vehicle control technology, and in particular to a vehicle control method, system, device, and medium based on electroencephalogram (EEG) signals and a multimodal large model. Background Technology

[0002] Human-machine collaborative control technology for automobiles is currently a research hotspot in the field of intelligent transportation. However, traditional brain-computer interface-assisted driving technology usually requires drivers to engage in high-intensity, action-level mental imagery. Due to the significant dynamic differences in EEG waveforms among individuals and the non-stationarity of the environment, existing technologies have low accuracy and poor fault tolerance in recognizing driver control intentions during actual operation, which can easily lead to severe cognitive fatigue in drivers during long-term driving.

[0003] Meanwhile, in the complex electromagnetic and vibration environment of vehicles, EEG signals are susceptible to noise interference and have limited spatial and temporal resolution. Relying solely on macroscopic thinking and imagination is insufficient to guarantee stable control of the vehicle on a fine spatial trajectory. Therefore, the system cannot meet the high-precision driving control requirements of vehicles under complex road conditions.

[0004] At the decision-making level, conventional autonomous driving systems typically plan paths based on fixed mechanical rules or algorithms. Because their decision-making logic is severely disconnected from the real cognition and driving style of human drivers, when the autonomous driving decision does not align with the driver's subjective intentions, the vehicle control output often appears too abrupt, resulting in a low overall comfort level of human-machine collaboration.

[0005] Finally, existing vehicle control methods primarily rely on traditional physical intervention methods, such as the driver forcefully applying the brakes or turning the steering wheel, to intervene in abnormal or erroneous driving situations. Because physical intervention involves significant neural transmission delays and mechanical action lags, it cannot achieve millisecond-level instantaneous corrections. Furthermore, forced physical intervention easily disrupts the continuity of vehicle movement, ultimately resulting in a significantly insufficient active correction capability of the system, severely impacting vehicle safety and ride comfort. Summary of the Invention

[0006] (a) Technical problems to be solved

[0007] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a vehicle control method, system, device and medium based on EEG signals and multimodal large model. It solves the technical problems of low intention recognition accuracy leading to driver fatigue, limited control resolution leading to insufficient accuracy, disconnect between autonomous driving decision and human will leading to stiff control, and reliance on physical takeover methods with time delay leading to weak active correction ability and poor smoothness of human-machine collaboration.

[0008] (II) Technical Solution

[0009] To achieve the above objectives, the main technical solutions adopted by the present invention include:

[0010] In a first aspect, embodiments of the present invention provide a vehicle control method based on electroencephalogram (EEG) signals and a multimodal large model, comprising:

[0011] When the vehicle is stationary, multiple sets of vehicle driving scene images are played to the driver to obtain the driver's first state EEG signal in order to establish a thought-motor mapping dictionary.

[0012] When the vehicle is in autonomous driving mode, the driver's second-state EEG signal is acquired and decoded into driving intention features with probability distribution by combining the thought-action mapping dictionary;

[0013] The vehicle-mounted visual guidance device projects a virtual intention crosshair onto the driver and acquires the visual evoked potential signal triggered by the virtual intention crosshair. The visual evoked potential signal is then fused with the driving intention characteristics to obtain a comprehensive driving intention.

[0014] Using a pre-defined multimodal large model, the acquired real-time vehicle environment perception data and comprehensive driving intention are used to perform cross-modal collaborative reasoning to generate multiple alternative adjustment operations. Conflict arbitration is then performed between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation.

[0015] After the vehicle performs the final adjustment operation, the acquired third-state EEG signal is analyzed in the time domain. If it is determined that there is an error-related negative potential signal that represents the driver's expected deviation and the amplitude exceeds the preset threshold, the error-related negative potential signal and the real-time vehicle environment perception data are fed into the multimodal large model to adaptively generate and execute the correction adjustment operation.

[0016] Optionally, while the vehicle is stationary, multiple sets of vehicle driving scene images are played to the driver to obtain the driver's first-state EEG signals, in order to establish a thought-motor mapping dictionary, including:

[0017] Multiple sets of vehicle driving scene images were played to the driver, and the on-board EEG acquisition device was used simultaneously to obtain the driver's first-state EEG signal under the stimulation of multiple sets of vehicle driving scene images.

[0018] The first-state EEG signal is input into a preset variational autoencoder to be projected onto a continuous latent space. The mean and variance parameters corresponding to the continuous latent space are calculated. Based on the mean and variance parameters, the probability density distribution of the latent variables of the first-state EEG signal in the continuous latent space is constructed.

[0019] Feature decoupling is performed based on the probability density distribution of latent variables to separate physiological noise components, and decoupled feature vectors are generated through reparameterized sampling.

[0020] The decoupled feature vectors are reassembled into a feature sequence according to time sequence and input into a preset Transformer network. The multi-head self-attention mechanism is used to calculate the dot product similarity between different time steps within the feature sequence to obtain attention weights. The attention weights are then used to perform a weighted summation of the feature sequence to obtain context features.

[0021] Global temporal average pooling is performed on the context features to transform them into context vectors, and then linear projection is used to map the context vectors to a semantic space to obtain EEG semantic vectors.

[0022] Based on the EEG semantic vector and the driving condition action labels corresponding to multiple sets of vehicle driving scene images, the conditional probability distribution of each driving condition action label under a given EEG semantic vector is calculated, and a thought action mapping dictionary is generated based on the conditional probability distribution.

[0023] Optionally, when the vehicle is in autonomous driving mode, the driver's second-state EEG signal is acquired and decoded into driving intention features with probability distribution using a thought-motion mapping dictionary, including:

[0024] Using an in-vehicle EEG acquisition device, the second-state EEG signals of the driver in autonomous driving mode are collected in real time under the current driving conditions.

[0025] The second-state EEG signal is sequentially processed by a variational autoencoder and a Transformer network to output an EEG semantic vector corresponding to the second-state EEG signal.

[0026] The EEG semantic vector corresponding to the second state EEG signal is input into the thought-action mapping dictionary for conditional probability retrieval. The target conditional probability distribution of each driving condition action label in the thought-action mapping dictionary under the given EEG semantic vector is extracted, and the target conditional probability distribution is used as the driving intention feature with probability distribution.

[0027] Optionally, a virtual intention crosshair is projected onto the driver via an in-vehicle visual guidance device, and visual evoked potential signals excited by the virtual intention crosshair are acquired. These visual evoked potential signals are then fused with driving intention characteristics to obtain a comprehensive driving intention, including:

[0028] The vehicle-mounted visual guidance device projects a virtual mind-guided crosshair interface containing multiple flashing coded areas onto the driver's field of vision, which is dynamically adjusted according to the vehicle's driving status. Each flashing coded area corresponds to a different driving condition action label and flashes visually at a different preset characteristic flashing frequency.

[0029] The vehicle-mounted EEG acquisition device was used to acquire the visual evoked potential signal of the driver caused by the virtual intention crosshair interface. The visual evoked potential signal was subjected to multi-band feature mapping, and the feature response components corresponding to each preset feature flashing frequency were extracted.

[0030] Calculate the waveform correlation strength between each characteristic response component and the preset sine and cosine standard reference signals corresponding to each preset characteristic flashing frequency, and perform normalization processing on the waveform correlation strength to construct a focusing probability distribution feature that corresponds one-to-one with each driving condition action label.

[0031] The probability distribution features of the focus are multiplied with the driving intention features by corresponding bits, and the result of the multiplication is normalized to calculate the joint probability distribution of the action labels of each driving condition.

[0032] The driving condition action label corresponding to the highest probability value is selected from the joint probability distribution to serve as the comprehensive driving intention.

[0033] Optionally, a pre-defined multimodal large model is used to perform cross-modal collaborative reasoning on the acquired real-time vehicle environment perception data and the comprehensive driving intention, generating multiple alternative adjustment operations. Conflict arbitration is then performed between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation, including:

[0034] Real-time vehicle environment perception data is acquired through on-board environmental sensors. The real-time vehicle environment perception data is then spatially aligned with the multimodal coding layer of the comprehensive driving intention input multimodal big model to generate a cross-modal collaborative feature sequence containing environmental semantics and intention semantics.

[0035] The cross-modal collaborative feature sequence is passed to the spatiotemporal autoregressive decoding layer of the multimodal large model for spatiotemporal autoregressive inference, generating multiple sets of alternative adjustment operations representing different control trajectories, and generating the model prediction probability distribution corresponding to each alternative adjustment operation.

[0036] By comparing the overall driving intention with the action direction of each alternative adjustment operation, if the overall driving intention conflicts with the action direction of the alternative adjustment operation corresponding to the highest probability value in the model's predicted probability distribution, then a human-machine conflict is determined to exist.

[0037] Based on real-time vehicle environment perception data, identify the safety boundaries around the vehicle body to determine the safety boundary weights corresponding to each alternative adjustment operation.

[0038] Using the joint probability distribution as the prior probability distribution and the model prediction probability distribution as the observation likelihood distribution, a pre-defined intention arbitrator based on the Bayesian probability algorithm is used, combined with the safety boundary weight, to perform dynamic weighted balancing calculations on the prior probability distribution and the observation likelihood distribution for intention coordination and conflict resolution, and to construct the posterior probability space of each alternative adjustment operation.

[0039] The candidate adjustment operation corresponding to the maximum posterior probability value is selected from the posterior probability space and determined as the final adjustment operation.

[0040] Optionally, the multimodal large model includes:

[0041] The multimodal coding layer includes an environmental perception feature sub-network for extracting spatiotemporal feature tensors of the vehicle's surroundings from real-time vehicle environmental perception data, a semantic embedding layer for projecting feature vectors onto comprehensive driving intention or error-related negative potential signals, and a cross-modal fusion module for aligning spatiotemporal feature tensors and feature vectors with spatial features and outputting cross-modal collaborative feature sequences.

[0042] The spatiotemporal autoregressive decoding layer includes a cross-modal cross-attention module for fusing cross-modal collaborative feature sequences in a heterogeneous feature space, and a Transformer decoder for generating multiple sets of alternative adjustment operations representing different control trajectories and the model prediction probability distributions corresponding to each alternative adjustment operation according to a preset time step autoregression based on the heterogeneous feature space fusion results.

[0043] Optionally, after the vehicle performs the final adjustment operation, the acquired third-state EEG signal is subjected to time-domain feature analysis. If it is determined that there is an error-related negative potential signal representing a violation of the driver's expected behavior and the amplitude exceeds a preset threshold, the error-related negative potential signal and real-time vehicle environmental perception data are fed into a multimodal large model to adaptively generate and execute a correction adjustment operation, including:

[0044] After the vehicle performs the final adjustment operation, the third-state EEG signal is acquired using the on-board EEG acquisition device, and the time-domain characteristic waveform of the third-state EEG signal is extracted.

[0045] The system detects whether there is an error-related negative potential signal that meets the preset waveform shape in the time-domain characteristic waveform, and determines that the driver has violated the expected behavior when it is confirmed that there is an error-related negative potential signal and the amplitude of the error-related negative potential signal exceeds the preset threshold.

[0046] In response to the expected violation, the error-related negative potential signal and the real-time vehicle environment perception data obtained by the on-board environment sensor are input into the multimodal coding layer of the multimodal large model for spatial feature alignment, generating a cross-modal correction feature sequence containing correction semantics and environmental semantics.

[0047] The cross-modal correction feature sequence is handed over to the spatiotemporal autoregressive decoding layer of the multimodal large model for spatiotemporal autoregressive inference to adaptively generate correction adjustment operations;

[0048] Control the vehicle to perform correction and adjustment operations, so as to dynamically close the loop to correct the vehicle's driving status after the final adjustment operation.

[0049] Secondly, embodiments of the present invention provide a vehicle control system based on electroencephalogram (EEG) signals and a multimodal large model, comprising:

[0050] The dictionary building module is used to play multiple sets of vehicle driving scene images to the driver when the vehicle is stationary, and to obtain the driver's first state EEG signal in order to build a thought-action mapping dictionary.

[0051] The intention feature decoding module is used to acquire the driver's second-state EEG signal when the vehicle is in autonomous driving mode, and decode it into driving intention features with probability distribution by combining the intention action mapping dictionary.

[0052] The integrated intent fusion module is used to project a virtual intention crosshair to the driver through the in-vehicle visual guidance device, and to obtain the visual evoked potential signal excited by the virtual intention crosshair. The integrated driving intent is obtained by fusing the visual evoked potential signal with the driving intention characteristics.

[0053] The reasoning and arbitration module is used to perform cross-modal collaborative reasoning on the acquired real-time vehicle environment perception data and comprehensive driving intention using a preset multimodal large model, generate multiple alternative adjustment operations, and perform conflict collaborative arbitration between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation.

[0054] The adaptive correction control module is used to perform time-domain feature analysis on the acquired third-state EEG signal after the vehicle performs the final adjustment operation. If it is determined that there is an error-related negative potential signal that represents the driver's expected deviation and the amplitude exceeds the preset threshold, the error-related negative potential signal and real-time vehicle environment perception data are fed into the multimodal large model to adaptively generate and execute the correction adjustment operation.

[0055] Thirdly, embodiments of the present invention provide a vehicle control device based on electroencephalogram (EEG) signals and a multimodal large model, comprising: at least one controller; and a memory communicatively connected to the at least one controller; wherein the memory stores instructions executable by the at least one controller, the instructions being executed by the at least one controller to enable the at least one controller to execute the vehicle control method based on EEG signals and a multimodal large model as described above.

[0056] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a controller, implement the vehicle control method based on EEG signals and a multimodal large model as described above.

[0057] (III) Beneficial Effects

[0058] The beneficial effects of this invention are:

[0059] First, a thought-action mapping dictionary is established by playing driving scene images when the vehicle is stationary. Then, in autonomous driving mode, the second-state EEG signal is decoded into driving intention features with probability distribution. This reduces the impact of dynamic differences in EEG waveforms among different individuals and the lack of rapid calibration methods on the accuracy of intention conversion. It realizes the mapping of non-stationary EEG signals to features with probability distribution, reducing the cognitive load on the driver caused by action-level thinking and imagination.

[0060] Based on this, a virtual intention crosshair is projected through an in-vehicle visual guidance device, and visual evoked potential signals are obtained and fused with driving intention characteristics to obtain a comprehensive driving intention. This reduces the impact of limited spatial resolution and susceptibility to noise interference of single-modal intention EEG signals in the complex in-vehicle environment. By deeply integrating the driver's microscopic real-time visual focusing dynamics with driving thinking, the precision of capturing control intentions and the accuracy of spatial trajectory control are improved.

[0061] Subsequently, a multimodal large model is used to perform cross-modal collaborative reasoning between real-time vehicle environment perception data and comprehensive driving intentions, generating multiple alternative adjustment operations. Conflict arbitration is then performed between the two to determine the final adjustment operation. This alleviates the problem of inconsistency between the decision logic of conventional path planning algorithms and driver cognition, achieving collaborative reasoning between external environment semantics and driver driving intentions. Furthermore, through a dynamic conflict arbitration mechanism, a dynamic weighted balance of human-machine decision-making is achieved while satisfying vehicle safety boundaries, improving the smoothness of vehicle control output and human-machine collaboration.

[0062] Finally, after the vehicle performs the final adjustment operation, the EEG signal of the third state is analyzed in real time in the time domain. If it is determined that there is an erroneous negative potential signal that represents a violation of the driver's expected behavior and whose amplitude exceeds the preset threshold, it is fed back into the multimodal large model along with the environmental perception data to perform adaptive correction. This reduces the impact of neural conduction delay and mechanical action lag on driving continuity caused by traditional physical takeover methods, realizes closed-loop trajectory correction, enhances the vehicle's active adaptive correction capability under abnormal conditions, and improves the safety of vehicle driving and the continuity of human-vehicle cooperation. Attached Figure Description

[0063] Figure 1 This is a schematic diagram of the overall process of the method provided in the embodiments of the present invention;

[0064] Figure 2 This is a schematic diagram illustrating the specific process of step S1 of the method provided in this embodiment of the invention;

[0065] Figure 3 This is a detailed flowchart illustrating step S2 of the method provided in this embodiment of the invention;

[0066] Figure 4 This is a detailed flowchart illustrating step S3 of the method provided in this embodiment of the invention;

[0067] Figure 5 This is a detailed flowchart illustrating step S4 of the method provided in this embodiment of the invention;

[0068] Figure 6 A schematic diagram of the specific process of step S5 of the method provided in the embodiment of the present invention. Detailed Implementation

[0069] To better explain and facilitate understanding of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0070] like Figure 1 As shown in the embodiment of the present invention, a vehicle control method based on electroencephalogram (EEG) signals and a multimodal large model is proposed, which includes: when the vehicle is stationary, playing multiple sets of vehicle driving scene images to the driver and acquiring the driver's first-state EEG signals to establish a thought-action mapping dictionary; when the vehicle is in autonomous driving mode, acquiring the driver's second-state EEG signals and decoding them into driving intention features with probability distribution by combining them with the thought-action mapping dictionary; projecting a virtual thought crosshair to the driver through an onboard visual guidance device and acquiring the visual evoked potential signals excited by the virtual thought crosshair; and fusing the visual evoked potential signals with the driving intention features. The system integrates real-time vehicle environment perception data with the comprehensive driving intention using a pre-defined multimodal model. This results in cross-modal collaborative reasoning to generate multiple alternative adjustment operations. Conflict arbitration is then performed between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation. After the vehicle executes the final adjustment operation, the acquired third-state EEG signal undergoes temporal feature analysis. If an error-related negative potential signal indicating a violation of the driver's expected behavior is detected and its amplitude exceeds a pre-defined threshold, the error-related negative potential signal and the real-time vehicle environment perception data are fed into the multimodal model to adaptively generate and execute a correction adjustment operation. This invention achieves the following technical efficiencies:

[0071] First, a thought-action mapping dictionary is established by playing driving scene images when the vehicle is stationary. Then, in autonomous driving mode, the second-state EEG signal is decoded into driving intention features with probability distribution. This reduces the impact of dynamic differences in EEG waveforms among different individuals and the lack of rapid calibration methods on the accuracy of intention conversion. It realizes the mapping of non-stationary EEG signals to features with probability distribution, reducing the cognitive load on the driver caused by action-level thinking and imagination.

[0072] Based on this, a virtual intention crosshair is projected through an in-vehicle visual guidance device, and visual evoked potential signals are obtained and fused with driving intention characteristics to obtain a comprehensive driving intention. This reduces the impact of limited spatial resolution and susceptibility to noise interference of single-modal intention EEG signals in the complex in-vehicle environment. By deeply integrating the driver's microscopic real-time visual focusing dynamics with driving thinking, the precision of capturing control intentions and the accuracy of spatial trajectory control are improved.

[0073] Subsequently, a multimodal large model is used to perform cross-modal collaborative reasoning between real-time vehicle environment perception data and comprehensive driving intentions, generating multiple alternative adjustment operations. Conflict arbitration is then performed between the two to determine the final adjustment operation. This alleviates the problem of inconsistency between the decision logic of conventional path planning algorithms and driver cognition, achieving collaborative reasoning between external environment semantics and driver driving intentions. Furthermore, through a dynamic conflict arbitration mechanism, a dynamic weighted balance of human-machine decision-making is achieved while satisfying vehicle safety boundaries, improving the smoothness of vehicle control output and human-machine collaboration.

[0074] Finally, after the vehicle performs its final adjustment, the EEG signal in the third state is analyzed in real time using temporal characteristics. If an erroneous negative potential signal is detected, indicating a deviation from the driver's expected behavior and exceeding a preset threshold in amplitude, it is fed back into the multimodal large model along with environmental perception data to perform adaptive correction. This reduces the impact of neural conduction delays and mechanical action lags on driving continuity caused by traditional physical takeover methods, achieves closed-loop trajectory correction, enhances the vehicle's active adaptive correction capability under abnormal conditions, and improves vehicle driving safety and the continuity of human-vehicle collaboration.

[0075] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and that the scope of the present invention can be fully conveyed to those skilled in the art.

[0076] Specifically, embodiments of the present invention provide a vehicle control method based on electroencephalogram (EEG) signals and a multimodal large model, comprising:

[0077] S1. When the vehicle is stationary, play multiple sets of vehicle driving scene images to the driver and obtain the driver's first state EEG signal to establish a thought-action mapping dictionary.

[0078] Furthermore, such as Figure 2 As shown, step S1 includes:

[0079] S11. Play multiple sets of vehicle driving scene images to the driver, and simultaneously use the on-board EEG acquisition device to obtain the driver's first-state EEG signal under the stimulation of multiple sets of vehicle driving scene images.

[0080] Specifically, with the vehicle stationary, a calibration process lasting 5-10 minutes is initiated: multiple sets of vehicle driving scene video clips covering different driving conditions are played to the driver through the in-vehicle display device. The standard video clips include acceleration, deceleration, steering, lane changing, overtaking, conventional braking, and active obstacle avoidance. While playing the vehicle driving scene video clips, the in-vehicle EEG acquisition device is used to collect and acquire the driver's first-state EEG signals under the stimulation of multiple sets of vehicle driving scene video clips in real time.

[0081] Specifically, the vehicle-mounted EEG acquisition device employs a non-invasive high-density EEG array or a fNIRS (functional near-infrared spectroscopy) array, integrated into an intelligent autonomous driving helmet that meets collision safety standards. During signal acquisition, the non-invasive high-density EEG array is used to acquire microvolt-level postsynaptic potential electrical activity signals of neuronal groups in the cerebral cortex in real time, while the fNIRS array is used to simultaneously acquire local blood oxygen dynamics response signals in the cerebral cortex (including changes in the concentrations of oxyhemoglobin and deoxyhemoglobin). Accordingly, the first-state EEG signal obtained thereby specifically includes at least one of the following: spontaneous EEG rhythm fluctuation signals, sensorimotor rhythm signals, and motor imagery (MI) characteristic signals spontaneously generated by the driver's cerebral cortex when facing the aforementioned image stimuli of different driving conditions.

[0082] S12. Input the first-state EEG signal into a preset variational autoencoder to project it onto a continuous latent space, calculate the mean parameter and variance parameter corresponding to the continuous latent space, and construct the latent variable probability density distribution of the first-state EEG signal in the continuous latent space based on the mean parameter and variance parameter.

[0083] Specifically, the first-state EEG signal, after digital preprocessing, is input into a preset variational autoencoder (VAE). The encoding network of this VAE uses its nonlinear mapping structure to project the first-state EEG signal onto a preset low-dimensional (preset to be 64-dimensional, 128-dimensional, or 256-dimensional) continuous latent space. In the continuous latent space, the output of the encoding network calculates two sets of low-dimensional parameter vectors that are completely consistent with the dimensions of the preset latent space, namely a mean vector (representing the central location of the distribution) and a variance vector (representing the degree of dispersion of the distribution).

[0084] Next, the mean vector is used as the center position parameter, and the variance vector is reduced to the variance parameters of each feature dimension as the scale parameter. These are then substituted into the Gaussian distribution formula to explicitly establish a multidimensional independent Gaussian latent variable probability density distribution in the latent space. This successfully transforms discrete and unstable EEG samples into a continuous and regular probability space representation. The Gaussian distribution formula is:

[0085] ;

[0086] In the formula, p(z) is the probability density distribution of multidimensional independent Gaussian latent variables established in the continuous latent space. The symbol represents the normal distribution, μ is the mean vector composed of the mean parameters of each feature dimension, and σ is the mean vector. 2 It is a variance vector composed of the variance parameters of each feature dimension.

[0087] S13. Feature decoupling is performed based on the latent variable probability density distribution to separate physiological noise components, and decoupled feature vectors are generated through reparameterized sampling. Considering that the original brain signal is mixed with a large amount of irrelevant interference, the system performs feature decoupling based on the constructed latent variable probability density distribution, effectively separating physiological noise components, including eye movement flicker, facial electromyography, and baseline drift, from the core driving intention components, and removing the physiological noise components, thereby extracting the latent variable probability density distribution representing pure driving intention.

[0088] After feature decoupling is completed, lossless sampling is performed from the probability density distribution of latent variables representing pure driving intentions through reparameterization sampling technology. Specifically, a random noise sample is first extracted from the preset standard normal distribution. Then, the random noise sample is multiplied by the standard deviation parameter calculated based on the variance parameter, and then summed with the corresponding mean parameter. Through this linear combination sampling method, a decoupled feature vector with both high noise resistance and robustness is generated.

[0089] S14. The decoupled feature vectors are recombined into a feature sequence according to time sequence and input into a preset Transformer network. A multi-head self-attention mechanism is used to calculate the dot product similarity between different time steps within the feature sequence to obtain attention weights. These attention weights are then used to perform a weighted summation of the feature sequence to obtain contextual features. To deeply explore the temporal evolution of driver thinking, the obtained decoupled feature vectors are recombined into a continuous feature sequence according to time sequence and input into a preset Transformer network. Utilizing the network's built-in multi-head self-attention mechanism, the dot product similarity between different time steps within the feature sequence is calculated in parallel. The dot product similarity is then normalized using the Softmax function to obtain attention weights on the time axis. This allows us to obtain the dynamic correlation and causal correspondence of the driver's thinking at different time points when the brain is stimulated by scene images. Then, the attention weights are used to perform a weighted summation of the feature sequence over the entire time period to obtain contextual features containing long-term temporal dependencies in thinking.

[0090] S15. Global temporal average pooling is performed on the context features to transform them into a context vector. This context vector is then mapped to a semantic space via linear projection to obtain the EEG semantic vector. To eliminate length differences in the temporal dimension and extract the core global information, global temporal average pooling is applied to the obtained context features to reduce their dimensionality and transform them into a fixed-dimensional context vector. Subsequently, a linear projection layer maps this context vector to a preset high-dimensional semantic space (preset to 512 or 1024 dimensions), ultimately outputting a standardized EEG semantic vector. In this process, the aforementioned variational autoencoder and Transformer network jointly construct a neural decoding engine. Its core value lies in parsing the originally disordered and chaotic raw EEG signals into EEG semantic vectors containing vehicle control features such as deceleration, overtaking, confirmation, and negation.

[0091] S16. Based on the EEG semantic vector and the driving condition action labels corresponding to multiple sets of vehicle driving scene images, calculate the conditional probability distribution of each driving condition action label under a given EEG semantic vector, and generate a thought-action mapping dictionary based on the conditional probability distribution. In this step, the mapped EEG semantic vector is associated with the driving condition action labels (such as acceleration, deceleration, steering, lane changing, overtaking, conventional braking, or active obstacle avoidance labels) pre-bound when playing multiple sets of vehicle driving scene images. By calculating the conditional probability distribution corresponding to each driving condition action label under a given EEG semantic vector, the matching degree between the current EEG features and each driving action is evaluated, following the formula: P(a|v), where v is the input EEG semantic vector and a is the candidate driving condition action label; finally, based on this conditional probability distribution, the matched EEG semantic vector and the corresponding driving condition action label are mapped and bound in the form of key-value pairs, thereby establishing and generating a personalized thought-action mapping dictionary exclusive to the current driver.

[0092] S2. When the vehicle is in autonomous driving mode, acquire the driver's second-state EEG signal and decode it into driving intention features with probability distribution by combining the thought-action mapping dictionary.

[0093] Furthermore, such as Figure 3 As shown, step S2 includes:

[0094] S21. Using an onboard EEG acquisition device, the driver's second-state EEG signals under the current driving conditions are acquired in real time while in autonomous driving mode. Specifically, after the vehicle enters the actual autonomous driving state, the non-invasive high-density EEG array or fNIRS array in the intelligent autonomous driving helmet remains continuously operational, acquiring the second-state EEG signal stream of the driver under the current real driving conditions online in real time. This signal stream is continuously input on the time axis. Accordingly, the acquired second-state EEG signals specifically include at least one of the spontaneous EEG rhythm fluctuation signals, sensorimotor rhythm signals, and motor imagery (MI) characteristic signals spontaneously generated by the driver's cerebral cortex when facing the current real driving conditions during dynamic driving.

[0095] S22. The second-state EEG signal is sequentially processed by a variational autoencoder and a Transformer network to output an EEG semantic vector corresponding to the second-state EEG signal. Specifically, the real-time acquired second-state EEG signal stream is continuously input into the neural decoding engine, and following the same steps as in steps S12 to S15, forward inference processing is performed through the previously trained and constructed variational autoencoder and Transformer network to output an EEG semantic vector corresponding to the second-state EEG signal.

[0096] S23. Input the EEG semantic vector corresponding to the second-state EEG signal into the thought-action mapping dictionary for conditional probability retrieval. Extract the target conditional probability distribution of each driving condition action label in the thought-action mapping dictionary under the given EEG semantic vector, and use the target conditional probability distribution as the driving intention feature with probability distribution. In this step, the current EEG semantic vector output in real time is used as the retrieval input and is passed in real time to the personalized thought-action mapping dictionary established and stored in the aforementioned calibration stage for online conditional probability retrieval. By quickly comparing the matching degree between the current EEG semantic vector and the preset features in the dictionary, the target conditional probability distribution of each driving condition action label (such as acceleration, deceleration, steering, lane changing, overtaking, conventional braking, or active obstacle avoidance labels) in the dictionary under the currently given EEG semantic vector is extracted and output as the driving intention feature with probability distribution. Through this step, the driver's implicit thought fluctuations are successfully quantified and decoded in real time into an explicit multi-dimensional action trigger probability distribution.

[0097] S3. A virtual intention crosshair is projected to the driver through the vehicle-mounted visual guidance device, and the visual evoked potential signal excited by the virtual intention crosshair is obtained. The visual evoked potential signal is fused with the driving intention characteristics to obtain a comprehensive driving intention.

[0098] Furthermore, such as Figure 4 As shown, step S3 includes:

[0099] S31. A virtual thought-guided crosshair interface, containing multiple flashing coded areas and dynamically adjusted according to the vehicle's driving status, is projected onto the driver's field of vision via an in-vehicle visual guidance device. Each flashing coded area corresponds to a different driving condition action label and flashes visually at a different preset characteristic flashing frequency. Specifically, during vehicle operation, a lightweight augmented reality interactive interface, namely the virtual thought-guided crosshair interface, is projected onto the driver's dynamic forward field of vision using an in-vehicle visual guidance device (which may be an augmented reality (AR) mask integrated into a smart autonomous driving helmet, an in-vehicle augmented reality head-up display (AR-HUD), or digital augmented reality smart glasses). To avoid interfering with the driver's line of sight and improve comfort, this virtual thought-guided crosshair interface can perform real-time position tracking and adaptive scaling adjustments based on the vehicle's current speed, steering angle, and vehicle posture, ensuring it remains within the driver's natural gaze area.

[0100] In this virtual mind-guided crosshair interface, multiple independent flashing coded regions are pre-defined and arranged. Each flashing coded region is bound in the system backend to a specific driving condition action label (such as acceleration, deceleration, steering, lane changing, overtaking, conventional braking, or active obstacle avoidance labels). At the same time, each flashing coded region flashes continuously with its own preset characteristic flashing frequency (for example, set to 11Hz, 12Hz, 13Hz, etc., which are highly recognizable steady-state visual evoked frequencies) to specifically modulate the driver's visual pathway.

[0101] S32. Use an on-board EEG acquisition device to acquire the visual evoked potential signal triggered by the driver's virtual intention crosshair interface, perform multi-band feature mapping on the visual evoked potential signal, and extract the feature response components corresponding to each preset feature flashing frequency.

[0102] Specifically, when the driver's gaze is fixed on a flashing coded area of ​​the virtual mind-guided crosshair interface, the occipital visual cortex of the brain will generate an electrophysiological synchronous resonance response at the preset characteristic flashing frequency and the corresponding harmonic frequency of that area. Using the onboard EEG acquisition device in the intelligent autonomous driving helmet, the steady-state visual evoked potential (SSVEP) signal stream induced by the visual stimulation of this interface can be acquired in real time.

[0103] Subsequently, the acquired visual evoked potential signal stream is input into a preset multi-band filter group for multi-band feature mapping (specifically, multiple overlapping bandpass filters are used to decompose the signal into their respective frequency bands) to filter out baseline drift and electromyographic noise caused by vehicle bumps, thereby extracting the feature response components corresponding to each preset feature flicker frequency in each corresponding frequency band.

[0104] S33. Calculate the waveform correlation strength between each characteristic response component and the sine and cosine standard reference signals corresponding to each preset characteristic flashing frequency, and perform normalization processing on the waveform correlation strength to construct a focusing probability distribution feature that corresponds one-to-one with each driving condition action label.

[0105] Specifically, for each preset characteristic flicker frequency, a set of sine and cosine standard reference signals composed of that frequency and its corresponding harmonic components are constructed. Then, the waveform correlation strength between each characteristic response component and each set of sine and cosine standard reference signals is calculated (specifically by calculating the correlation coefficient or similarity score between time series waveforms to measure the degree of matching between the two in waveform rhythm).

[0106] After obtaining the waveform correlation intensity corresponding to each preset feature flashing frequency, the Softmax normalization algorithm is used to normalize each waveform correlation intensity to eliminate the influence of individual skin impedance differences and absolute signal amplitude. Finally, a set of focusing probability distribution features corresponding one-to-one with each driving condition action label is constructed. The sum of each component in the focusing probability distribution feature is 1, which is used to characterize the driver's current visual gaze focus distribution.

[0107] S34. Multiply the focus probability distribution features with the driving intention features by corresponding bits, and normalize the result of the multiplication to calculate the joint probability distribution of the action labels for each driving condition.

[0108] Specifically, in order to integrate driving intention features with focus probability distribution features, the probability score of the action label belonging to a certain driving condition in the focus probability distribution features is multiplied bitwise with the probability score of the action label of the same driving condition in the driving intention features (i.e., element-wise dot product calculation), thereby synchronizing the probabilities of the two modalities.

[0109] Subsequently, the product results obtained after multiplication are normalized as a whole to restore the total probability sum after fusion to 1, thereby calculating the joint probability distribution of the action labels for each driving condition; this reduces the erroneous operation caused by single-modal false triggering and improves the robustness of control commands.

[0110] S35. Select the driving condition action label corresponding to the highest probability value from the joint probability distribution to serve as the comprehensive driving intention.

[0111] Furthermore, a maximum value search operation is performed on the calculated joint probability distribution to select the target probability value with the largest value, and the driving condition action label associated with this target probability value is extracted as the comprehensive driving intention output. This comprehensive driving intention integrates the highest-level decision-making information from "eye-brain coordination" and is subsequently used to trigger corresponding vehicle control actions.

[0112] S4. Using a pre-set multimodal large model, the acquired real-time vehicle environment perception data and comprehensive driving intention are used to perform cross-modal collaborative reasoning to generate multiple alternative adjustment operations. Conflict arbitration is then performed between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation.

[0113] Furthermore, such as Figure 5 As shown, step S4 includes:

[0114] S41. Acquire real-time vehicle environment perception data through on-board environmental sensors, align the real-time vehicle environment perception data with the multimodal coding layer of the comprehensive driving intention input multimodal big model for spatial feature alignment, and generate a cross-modal collaborative feature sequence containing environmental semantics and intention semantics.

[0115] Specifically, during vehicle operation, the vehicle acquires real-time vehicle environment perception data (including spatial topology information such as surrounding vehicle trajectories, pedestrian positions, lane line boundaries, obstacle distances, and traffic signal status) through on-board environmental sensors (specifically including at least one of on-board LiDAR, surrounding high-definition cameras, millimeter-wave radar, and ultrasonic sensors).

[0116] Subsequently, real-time vehicle environmental perception data will be combined with comprehensive driving intention data. Figure 1 In the multimodal coding layer of the same input multimodal large model: inside the multimodal coding layer, real-time vehicle environment perception data is input into the environment perception feature sub-network, and processed using a three-dimensional convolutional structure or a long short-term memory network to filter out environmental noise and extract spatiotemporal feature tensors that represent the physical space state around the vehicle; at the same time, the comprehensive driving intention is input into the semantic embedding layer, and transformed into a feature vector with uniform dimensional specifications through word vector lookup table projection.

[0117] Finally, the cross-modal fusion module receives the aforementioned spatiotemporal feature tensors and feature vectors, and uses spatial coordinate system transformation and cross-projection techniques to align these two heterogeneous features in a unified spatial geometric coordinate system, thereby outputting a cross-modal collaborative feature sequence that includes external environment semantics and internal driving intention semantics.

[0118] S42. The cross-modal collaborative feature sequence is passed to the spatiotemporal autoregressive decoding layer of the multimodal large model for spatiotemporal autoregressive inference, generating multiple sets of alternative adjustment operations representing different control trajectories, and generating the model prediction probability distribution corresponding to each alternative adjustment operation.

[0119] Specifically, the generated cross-modal collaborative feature sequence is fed into the spatiotemporal autoregressive decoding layer of the multimodal large model for trajectory prediction inference. Inside the spatiotemporal autoregressive decoding layer, the cross-modal cross-attention module is used to map the cross-modal collaborative feature sequence into query vectors, key vectors, and value vectors, and the cross-modal self-attention association scores in the heterogeneous feature space are calculated in parallel to achieve the fusion of environmental features and intent features. Then, the Transformer decoder is used to perform autoregressive forward inference (i.e., predict the control trajectory point of the next moment based on the current moment) according to the preset time steps based on the fusion result.

[0120] Through this autoregressive inference process, the Transformer decoder generates multiple sets of alternative adjustment operations representing different control trajectories (e.g., alternative adjustment operation 1 is accelerating to the left to change lanes, alternative adjustment operation 2 is decelerating and maintaining the original lane, and alternative adjustment operation 3 is emergency braking to avoid obstacles). Simultaneously, the Transformer decoder is used to calculate the model prediction probability distribution corresponding to each alternative adjustment operation. This model prediction probability distribution is used to characterize the predicted probability of each alternative adjustment operation.

[0121] S43. Compare the overall driving intention with the action direction of each alternative adjustment operation. If the overall driving intention and the alternative adjustment operation corresponding to the maximum probability value in the model prediction probability distribution conflict in action direction, then it is determined that there is a human-machine conflict.

[0122] In this step, the action type indicated by the overall driving intention is compared with the action direction of the alternative adjustment operation corresponding to the highest probability value in the model's predicted probability distribution. For example, if the action direction of the overall driving intention is "accelerate to the left to change lanes", while the alternative adjustment operation corresponding to the highest probability value in the model's predicted probability distribution, calculated using a multimodal large model based on real-time vehicle environment perception data, is "maintain the original lane, decelerate and follow", then a conflict exists between the two in terms of action direction or spatial trajectory guidance, and a trigger signal is output to determine that a human-machine conflict exists.

[0123] S44. Identify the safety boundaries around the vehicle body based on real-time vehicle environment perception data to determine the safety boundary weights corresponding to each alternative adjustment operation.

[0124] Specifically, the safety boundary around the vehicle is constructed based on real-time vehicle environmental perception data, which includes the following joint construction steps for longitudinal and lateral safety boundaries:

[0125] (1) Establish a two-dimensional vehicle body coordinate system with the current vehicle body center as the origin, where the forward direction of the vehicle head is the longitudinal positive axis and the horizontal direction of the left side of the vehicle body is the horizontal positive axis.

[0126] (2) Extract the current vehicle's longitudinal speed, as well as the relative distance and speed of each surrounding obstacle relative to the current vehicle body, from the real-time vehicle environment perception data. For each obstacle, calculate the corresponding dynamic collision time using the following formula:

[0127] TTC i =(ΔX i ) / (ΔV i );

[0128] In the formula, TTC i ΔX represents the dynamic collision time between the current vehicle and the i-th obstacle; iΔV represents the current relative distance component between the vehicle and the i-th obstacle; i This represents the current relative velocity component between the vehicle and the i-th obstacle.

[0129] (3) If the calculated TTC i If the distance is less than or equal to the preset conflict time threshold, then the dynamic longitudinal safety distance threshold for the i-th obstacle is calculated using the following formula based on the current longitudinal speed of the vehicle and the preset braking deceleration:

[0130] L safe,i =V ego *t delay +(V ego 2 -V i 2 ) / (2*a max )+d buffer ;

[0131] In the formula, L safe,i V represents the dynamic longitudinal safety distance threshold for the i-th obstacle; ego V represents the current longitudinal speed of the vehicle. i Let t be the absolute velocity of the i-th obstacle; delay The preset vehicle braking system response delay time; a max d is the preset absolute value of the maximum braking deceleration; buffer This is the preset minimum safety redundancy distance; if the calculated TTC i If the time exceeds the preset conflict time threshold, the dynamic longitudinal safety distance threshold for the i-th obstacle will be set to the minimum safety redundancy distance d. buffer The maximum value among the dynamic longitudinal safety distance thresholds corresponding to all obstacles is determined as the longitudinal safety boundary along the positive longitudinal axis at the current moment.

[0132] (4) Extract the topology data of the current lane line boundary where the current vehicle is located from the real-time vehicle environment perception data. Based on the known vehicle width, the first absolute distance between the vehicle centerline and the left lane line boundary, and the second absolute distance between the vehicle centerline and the right lane line boundary, subtract the preset lateral safety boundary blank distance along the corresponding left lateral positive axis direction and the corresponding right lateral negative axis direction, respectively, to construct a lateral safety boundary composed of the left lane constraint boundary and the right lane constraint boundary.

[0133] Next, each alternative adjustment operation is projected onto the safety boundary, and the degree of overlap between the control trajectory corresponding to each alternative adjustment operation and the safety boundary is evaluated: if the corresponding control trajectory is within the safe area of ​​the safety boundary, its corresponding safety boundary weight is assigned a first preset weight value (e.g., set to 1.0); if the corresponding control trajectory approaches or intrudes into the safety boundary, a safety penalty is imposed, and its corresponding safety boundary weight is assigned a second preset weight value (e.g., set to 0.1 or 0.01), wherein the second preset weight value is less than the first preset weight value. Through this step, the external physical constraints are converted into safety boundary weights corresponding to each alternative adjustment operation.

[0134] S45. Using the joint probability distribution as the prior probability distribution and the model prediction probability distribution as the observation likelihood distribution, a pre-set intention arbitrator based on the Bayesian probability algorithm is used to perform dynamic weighted balancing calculations on the prior probability distribution and the observation likelihood distribution in combination with the safety boundary weights to carry out intention coordination and conflict resolution, and to construct the posterior probability space of each alternative adjustment operation.

[0135] The joint probability distribution output in step S34 is used as the prior probability distribution, and the model prediction probability distribution output in step S42 is used as the observation likelihood distribution. Using a preset intention arbitrator based on the Bayesian probability algorithm, and in conjunction with the safety boundary weights determined in step S44, a weighted balance calculation is performed on the prior probability distribution and the observation likelihood distribution.

[0136] In this process, the multimodal large model receives the comprehensive driving intent as explicit semantic features at the encoding layer, while the intent arbitrator retains the full-space uncertainty of the joint probability distribution as the prior probability distribution, thereby realizing multi-level human-machine collaborative control. Regardless of whether a human-machine conflict is determined to exist, the intent arbitrator routinely performs dynamic weighted balancing calculations; when a human-machine conflict is determined to exist, the intent arbitrator adjusts the posterior probability space by introducing safety boundary weights to resolve the human-machine conflict and control the vehicle's driving state.

[0137] By using an intention arbitrator, the posterior probability space of each alternative adjustment operation is constructed by calculating the numerical product and normalized sum of each independent action component. The posterior probability calculation formula is shown below:

[0138]

[0139] In the formula, A k Let M be the k-th candidate adjustment operation; M be the total number of candidate adjustment operations generated; m be the loop variable for summing all candidate adjustment operations; P post (A k P represents the posterior probability value finally calculated for the k-th alternative adjustment operation. prior (A kP represents the prior probability value corresponding to the k-th alternative adjustment operation; like (A k ) represents the observed likelihood probability value corresponding to the k-th alternative adjustment operation; w k P represents the safety boundary weights for the k-th alternative adjustment operation. prior (A m Let P be the prior probability value corresponding to the m-th alternative adjustment operation. like (A m Let w be the observed likelihood probability value corresponding to the m-th alternative adjustment operation. m The safety boundary weights are for the m-th alternative adjustment operation.

[0140] This formula is used to construct the posterior probability space for each alternative adjustment operation by multiplying the probabilities and weights of each feature dimension and component in numerical form and summing the results with the denominator normalized.

[0141] S46. Select the candidate adjustment operation corresponding to the maximum posterior probability value from the posterior probability space, and determine it as the final adjustment operation.

[0142] In this step, the aforementioned intent arbitrator is used to perform a numerical search on the constructed posterior probability space, from which the maximum posterior probability value is selected, and the candidate adjustment operation corresponding to the maximum posterior probability value is determined as the final adjustment operation.

[0143] Before sending the final adjustment operation to the vehicle's drive-by-wire chassis actuator, the driver's mental state is verified using a preset drive safety monitoring loop: by calculating and analyzing the real-time frequency ratio of Alpha waves to Theta waves in the brainwaves acquired by the aforementioned non-invasive high-density array, it is determined whether the real-time frequency ratio is within the preset safety range.

[0144] If the real-time wavy ratio is within the preset safety range, the final adjustment operation is output, and the driver receives audio feedback via the bone conduction headphones built into the helmet. This final adjustment operation includes not only mechanical motion control commands sent to the vehicle's drive-by-wire chassis actuator, but also eco-linkage control commands sent to the vehicle body control domain and the cabin comfort control domain. The mechanical motion control commands in the final adjustment operation are then sent to the vehicle's drive-by-wire chassis actuator to dynamically adjust the vehicle's driving state. When the final adjustment operation indicates that the driver has a tendency to relax and enjoy the scenery, the drive-by-wire chassis actuator performs deceleration and fine-tunes the vehicle's front end. At the same time, the cabin comfort control domain controls the windows to lower and triggers the vehicle's odor modulation system to release a fragrant gas that matches the current environmental semantics.

[0145] If the real-time tidal ratio is not within the preset safety range, the chassis output path of the final adjustment operation is cut off, and the vehicle's drive-by-wire chassis actuator is controlled to forcibly execute the corresponding emergency braking obstacle avoidance operation in the aforementioned alternative adjustment operation.

[0146] S5. After the vehicle performs the final adjustment operation, perform time-domain feature analysis on the acquired third-state EEG signal. If it is determined that there is an error-related negative potential signal that represents the driver's expected deviation and the amplitude exceeds the preset threshold, then feed the error-related negative potential signal and the real-time vehicle environment perception data into the multimodal large model to adaptively generate and execute the correction adjustment operation.

[0147] Furthermore, such as Figure 6 As shown, step S5 includes:

[0148] S51. After the vehicle performs the final adjustment operation, the vehicle-mounted EEG acquisition device is used to acquire the third-state EEG signal and extract the time-domain characteristic waveform of the third-state EEG signal.

[0149] Specifically, as the chassis actuators respond and begin the final adjustment operation, the EEG input is switched to dynamic feedback monitoring mode. Using an onboard EEG acquisition device, the third-state EEG signal generated by the driver during the execution of this final adjustment operation is simultaneously acquired. Accordingly, the acquired third-state EEG signal specifically includes at least one of the following: error-related negative potential (ErrP) signal spontaneously generated by the driver's cerebral cortex after the vehicle performs the final adjustment operation and faces the current actual driving state of the vehicle; spontaneous EEG rhythmic fluctuation signal; sensorimotor rhythm signal; and motor imagery (MI) characteristic signal.

[0150] When performing temporal feature extraction, the starting time of the final adjustment operation is first used as the time zero point to divide the temporal region into slices, and a segment of data within a preset time window (e.g., 0 milliseconds to 1000 milliseconds after the action occurs) is extracted. Then, the random fluctuations of spontaneous background EEG are eliminated by using a baseline correction algorithm and temporal moving average filtering, thereby extracting the temporal feature waveform of the third-state EEG signal on the time axis.

[0151] S52. Detect whether there is an error-related negative potential signal in the time-domain characteristic waveform that meets the preset waveform shape, and determine that the driver has committed an expected violation when it is confirmed that there is an error-related negative potential signal and the amplitude of the error-related negative potential signal exceeds the preset threshold.

[0152] In this step, it is detected whether the time-domain characteristic waveform exhibits a characteristic negative peak that conforms to the preset intrinsic error-related negative potential characteristics within a window period of 200 to 400 milliseconds after the vehicle action occurs.

[0153] To eliminate random amplitude noise caused by normal blinking or gaze shifting by the driver, a double check is performed: First, the latency and slope of the negative peak are checked to see if they are within the preset standard morphological range; second, the amplitude component of the negative peak is extracted and compared with a preset amplitude threshold (e.g., a voltage threshold between -5 microvolts and -10 microvolts).

[0154] Only when it is confirmed that there is an error-related negative potential signal that meets the preset waveform shape in the time-domain characteristic waveform, and the absolute amplitude of the error-related negative potential signal exceeds the preset threshold, a conflict alarm is generated, and it is determined that the driver has violated the expected driving state of the vehicle.

[0155] If there is no negative potential signal related to the error, or if the absolute amplitude does not exceed the preset threshold, it is determined that no expected violation has occurred, and the process returns to step S51 for continuous monitoring.

[0156] S53. In response to the expected violation, the error-related negative potential signal and the real-time vehicle environment perception data obtained by the on-board environment sensor are input into the multimodal coding layer of the multimodal large model for spatial feature alignment, generating a cross-modal correction feature sequence containing correction semantics and environmental semantics.

[0157] Once a driver is determined to have violated the expected behavior, in response to the expected behavior trigger signal, the vehicle's environmental sensors acquire the latest real-time vehicle environmental perception data to obtain the latest spatiotemporal dynamic changes that have occurred around the vehicle after it has performed the final adjustment operation.

[0158] The latest real-time vehicle environment perception data and the extracted temporal feature waveform are input together into the multimodal coding layer of the multimodal large model: the spatiotemporal feature tensor of the current environment is extracted using the environment perception feature sub-network, and the temporal feature waveform is transformed into a correction feature vector representing the error correction intention using the semantic embedding layer. Subsequently, the spatiotemporal feature tensor and the correction feature vector are spatially aligned in a unified spatial geometric coordinate system through the cross-modal fusion module, thereby adaptively generating a cross-modal error correction feature sequence containing error correction semantics and the latest environment semantics.

[0159] S54. The cross-modal correction feature sequence is handed over to the spatiotemporal autoregressive decoding layer of the multimodal large model for spatiotemporal autoregressive inference to adaptively generate correction adjustment operations.

[0160] In this step, the generated cross-modal correction feature sequence is fed into the spatiotemporal autoregressive decoding layer of the multimodal large model. Inside this decoding layer, the cross-modal cross-attention module is used as a weight constraint to focus on the spatiotemporal features of the environment in order to identify the danger sources that cause the expected violation.

[0161] Next, the Transformer decoder is used to perform spatiotemporal autoregressive inference based on this feature fusion result. The decoder is guided by multi-objective optimization, which aims to reduce the amplitude of the error-related negative potential signal and keep the vehicle trajectory within the safety boundary. At the output, it adaptively generates correction adjustment operations for adaptive correction. These correction adjustment operations include emergency lane change abort and return to the original lane, or adding pressure compensation trajectory for emergency braking on the basis of the original deceleration.

[0162] S55. Control the vehicle to perform a correction adjustment operation to dynamically close the loop and correct the vehicle's driving status after the final adjustment operation.

[0163] Finally, after the multimodal large model outputs the correction adjustment operation, the correction control parameters are converted into electronic control low-level drive commands. These correction control parameters include emergency avoidance angle and braking deceleration value. The electronic control low-level drive commands are sent to the vehicle's drive-by-wire chassis actuator, which interrupts or corrects the final adjustment operation being executed, so that the vehicle can adaptively execute the correction adjustment operation in physical space.

[0164] The online reinforcement learning mechanism of the multimodal large model is triggered synchronously. While adaptively generating the correction adjustment operation, the chassis's driving trajectory control strategy network is adaptively iteratively optimized. In this way, the vehicle's driving state after the final adjustment operation is dynamically corrected in a closed loop. This step utilizes error-related negative potential signals to establish an adaptive error correction closed-loop control loop, reducing the driving risks caused by trajectory control deviations.

[0165] Furthermore, embodiments of the present invention provide a vehicle control system based on electroencephalogram (EEG) signals and a multimodal large model, comprising: a dictionary construction module, used to play multiple sets of vehicle driving scene images to the driver when the vehicle is stationary, and acquire the driver's first-state EEG signals to establish a thought-action mapping dictionary; an intention feature decoding module, used to acquire the driver's second-state EEG signals when the vehicle is in autonomous driving mode, and decode them into driving intention features with probability distribution by combining them with the thought-action mapping dictionary; and a comprehensive intention fusion module, used to project a virtual intention crosshair to the driver through an onboard visual guidance device, and acquire the visual evoked potential signals excited by the virtual intention crosshair, and based on the visual evoked potential signals and driving intention features... The system integrates various data points to obtain a comprehensive driving intention. A reasoning and arbitration module uses a pre-defined multimodal model to perform cross-modal collaborative reasoning with the acquired real-time vehicle environment perception data and the comprehensive driving intention, generating multiple alternative adjustment operations. It then performs conflict arbitration between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation. An adaptive correction control module analyzes the temporal features of the acquired third-state EEG signal after the vehicle executes the final adjustment operation. If it determines that there is an error-related negative potential signal representing a violation of the driver's expected behavior and that the amplitude exceeds a preset threshold, it feeds the error-related negative potential signal and the real-time vehicle environment perception data into the multimodal model to adaptively generate and execute the correction adjustment operation.

[0166] Furthermore, embodiments of the present invention provide a vehicle control device based on electroencephalogram (EEG) signals and a multimodal large model, comprising: at least one controller; and a memory communicatively connected to the at least one controller; wherein the memory stores instructions executable by the at least one controller, the instructions being executed by the at least one controller to enable the at least one controller to perform the vehicle control method based on EEG signals and a multimodal large model as described above.

[0167] Then, this embodiment of the invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a controller, implement the vehicle control method based on EEG signals and a multimodal large model as described above.

[0168] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0169] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.

[0170] It should be noted that any reference numerals placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims that enumerate several means, several of these means may be embodied by the same hardware. The use of the terms first, second, third, etc., is merely for convenience of expression and does not indicate any order. These terms can be understood as part of the component names.

[0171] Furthermore, it should be noted that in the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0172] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the claims should be interpreted to include both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0173] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention should also include these modifications and variations.

Claims

1. A vehicle control method based on electroencephalogram (EEG) signals and a multimodal large model, characterized in that, include: When the vehicle is stationary, multiple sets of vehicle driving scene images are played to the driver to obtain the driver's first state EEG signal in order to establish a thought-motor mapping dictionary. When the vehicle is in autonomous driving mode, the driver's second-state EEG signal is acquired and decoded into driving intention features with probability distribution by combining the thought-action mapping dictionary; The vehicle-mounted visual guidance device projects a virtual intention crosshair onto the driver and acquires the visual evoked potential signal triggered by the virtual intention crosshair. The visual evoked potential signal is then fused with the driving intention characteristics to obtain a comprehensive driving intention. Using a pre-defined multimodal large model, the acquired real-time vehicle environment perception data and comprehensive driving intention are used to perform cross-modal collaborative reasoning to generate multiple alternative adjustment operations. Conflict arbitration is then performed between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation. After the vehicle performs the final adjustment operation, the acquired third-state EEG signal is analyzed in the time domain. If it is determined that there is an error-related negative potential signal that represents the driver's expected deviation and the amplitude exceeds the preset threshold, the error-related negative potential signal and the real-time vehicle environment perception data are fed into the multimodal large model to adaptively generate and execute the correction adjustment operation.

2. The vehicle control method based on EEG signals and a multimodal large model as described in claim 1, characterized in that, With the vehicle stationary, multiple sets of video footage of the vehicle in motion are played to the driver to obtain the driver's initial EEG signals, in order to establish a thought-motor mapping dictionary, including: Multiple sets of vehicle driving scene images were played to the driver, and the on-board EEG acquisition device was used simultaneously to obtain the driver's first-state EEG signal under the stimulation of multiple sets of vehicle driving scene images. The first-state EEG signal is input into a preset variational autoencoder to be projected onto a continuous latent space. The mean and variance parameters corresponding to the continuous latent space are calculated. Based on the mean and variance parameters, the probability density distribution of the latent variables of the first-state EEG signal in the continuous latent space is constructed. Feature decoupling is performed based on the probability density distribution of latent variables to separate physiological noise components, and decoupled feature vectors are generated through reparameterized sampling. The decoupled feature vectors are reassembled into a feature sequence according to time sequence and input into a preset Transformer network. The multi-head self-attention mechanism is used to calculate the dot product similarity between different time steps within the feature sequence to obtain attention weights. The attention weights are then used to perform a weighted summation of the feature sequence to obtain context features. Global temporal average pooling is performed on the context features to transform them into context vectors, and then linear projection is used to map the context vectors to a semantic space to obtain EEG semantic vectors. Based on the EEG semantic vector and the driving condition action labels corresponding to multiple sets of vehicle driving scene images, the conditional probability distribution of each driving condition action label under a given EEG semantic vector is calculated, and a thought action mapping dictionary is generated based on the conditional probability distribution.

3. The vehicle control method based on EEG signals and a multimodal large model as described in claim 2, characterized in that, When the vehicle is in autonomous driving mode, the driver's second-state EEG signal is acquired and decoded into driving intention features with probability distribution using a thought-motion mapping dictionary, including: Using an in-vehicle EEG acquisition device, the second-state EEG signals of the driver in autonomous driving mode are collected in real time under the current driving conditions. The second-state EEG signal is sequentially processed by a variational autoencoder and a Transformer network to output an EEG semantic vector corresponding to the second-state EEG signal. The EEG semantic vector corresponding to the second state EEG signal is input into the thought-action mapping dictionary for conditional probability retrieval. The target conditional probability distribution of each driving condition action label in the thought-action mapping dictionary under the given EEG semantic vector is extracted, and the target conditional probability distribution is used as the driving intention feature with probability distribution.

4. The vehicle control method based on EEG signals and a multimodal large model as described in claim 3, characterized in that, A virtual intention crosshair is projected onto the driver through an in-vehicle visual guidance device, and visual evoked potential signals excited by the virtual intention crosshair are acquired. These visual evoked potential signals are then fused with driving intention characteristics to obtain a comprehensive driving intention, including: The vehicle-mounted visual guidance device projects a virtual mind-guided crosshair interface containing multiple flashing coded areas onto the driver's field of vision, which is dynamically adjusted according to the vehicle's driving status. Each flashing coded area corresponds to a different driving condition action label and flashes visually at a different preset characteristic flashing frequency. The vehicle-mounted EEG acquisition device was used to acquire the visual evoked potential signal of the driver caused by the virtual intention crosshair interface. The visual evoked potential signal was subjected to multi-band feature mapping, and the feature response components corresponding to each preset feature flashing frequency were extracted. Calculate the waveform correlation strength between each characteristic response component and the preset sine and cosine standard reference signals corresponding to each preset characteristic flashing frequency, and perform normalization processing on the waveform correlation strength to construct a focusing probability distribution feature that corresponds one-to-one with each driving condition action label. The probability distribution features of the focus are multiplied with the driving intention features by corresponding bits, and the result of the multiplication is normalized to calculate the joint probability distribution of the action labels of each driving condition. The driving condition action label corresponding to the highest probability value is selected from the joint probability distribution to serve as the comprehensive driving intention.

5. The vehicle control method based on EEG signals and a multimodal large model as described in claim 4, characterized in that, Using a pre-defined multimodal large model, the acquired real-time vehicle environment perception data and comprehensive driving intention are used for cross-modal collaborative reasoning to generate multiple alternative adjustment operations. Conflict arbitration is then performed between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation, including: Real-time vehicle environment perception data is acquired through on-board environmental sensors. The real-time vehicle environment perception data is then spatially aligned with the multimodal coding layer of the comprehensive driving intention input multimodal big model to generate a cross-modal collaborative feature sequence containing environmental semantics and intention semantics. The cross-modal collaborative feature sequence is passed to the spatiotemporal autoregressive decoding layer of the multimodal large model for spatiotemporal autoregressive inference, generating multiple sets of alternative adjustment operations representing different control trajectories, and generating the model prediction probability distribution corresponding to each alternative adjustment operation. By comparing the overall driving intention with the action direction of each alternative adjustment operation, if the overall driving intention conflicts with the action direction of the alternative adjustment operation corresponding to the highest probability value in the model's predicted probability distribution, then a human-machine conflict is determined to exist. Based on real-time vehicle environment perception data, identify the safety boundaries around the vehicle body to determine the safety boundary weights corresponding to each alternative adjustment operation. Using the joint probability distribution as the prior probability distribution and the model prediction probability distribution as the observation likelihood distribution, a pre-defined intention arbitrator based on the Bayesian probability algorithm is used, combined with the safety boundary weight, to perform dynamic weighted balancing calculations on the prior probability distribution and the observation likelihood distribution for intention coordination and conflict resolution, and to construct the posterior probability space of each alternative adjustment operation. The candidate adjustment operation corresponding to the maximum posterior probability value is selected from the posterior probability space and determined as the final adjustment operation.

6. The vehicle control method based on EEG signals and a multimodal large model as described in claim 5, characterized in that, Multimodal large models include: The multimodal coding layer includes an environmental perception feature sub-network for extracting spatiotemporal feature tensors of the vehicle's surroundings from real-time vehicle environmental perception data, a semantic embedding layer for projecting feature vectors onto comprehensive driving intention or error-related negative potential signals, and a cross-modal fusion module for aligning spatiotemporal feature tensors and feature vectors with spatial features and outputting cross-modal collaborative feature sequences. The spatiotemporal autoregressive decoding layer includes a cross-modal cross-attention module for fusing cross-modal collaborative feature sequences in a heterogeneous feature space, and a Transformer decoder for generating multiple sets of alternative adjustment operations representing different control trajectories and the model prediction probability distributions corresponding to each alternative adjustment operation according to a preset time step autoregression based on the heterogeneous feature space fusion results.

7. The vehicle control method based on EEG signals and a multimodal large model as described in claim 5 or 6, characterized in that, After the vehicle performs its final adjustment, the acquired third-state EEG signal is analyzed for temporal features. If an error-related negative potential signal indicating a violation of the driver's expected behavior is detected and its amplitude exceeds a preset threshold, the error-related negative potential signal and real-time vehicle environmental perception data are fed into a multimodal large model to adaptively generate and execute corrective adjustment operations, including: After the vehicle performs the final adjustment operation, the third-state EEG signal is acquired using the on-board EEG acquisition device, and the time-domain characteristic waveform of the third-state EEG signal is extracted. The system detects whether there is an error-related negative potential signal that meets the preset waveform shape in the time-domain characteristic waveform, and determines that the driver has violated the expected behavior when it is confirmed that there is an error-related negative potential signal and the amplitude of the error-related negative potential signal exceeds the preset threshold. In response to the expected violation, the error-related negative potential signal and the real-time vehicle environment perception data obtained by the on-board environment sensor are input into the multimodal coding layer of the multimodal large model for spatial feature alignment, generating a cross-modal correction feature sequence containing correction semantics and environmental semantics. The cross-modal correction feature sequence is handed over to the spatiotemporal autoregressive decoding layer of the multimodal large model for spatiotemporal autoregressive inference to adaptively generate correction adjustment operations; Control the vehicle to perform correction and adjustment operations, so as to dynamically close the loop to correct the vehicle's driving status after the final adjustment operation.

8. A vehicle control system based on electroencephalogram (EEG) signals and a multimodal large model, characterized in that, include: The dictionary building module is used to play multiple sets of vehicle driving scene images to the driver when the vehicle is stationary, and to obtain the driver's first state EEG signal in order to build a thought-action mapping dictionary. The intention feature decoding module is used to acquire the driver's second-state EEG signal when the vehicle is in autonomous driving mode, and decode it into driving intention features with probability distribution by combining the intention action mapping dictionary. The integrated intent fusion module is used to project a virtual intention crosshair to the driver through the in-vehicle visual guidance device, and to obtain the visual evoked potential signal excited by the virtual intention crosshair. The integrated driving intent is obtained by fusing the visual evoked potential signal with the driving intention characteristics. The reasoning and arbitration module is used to perform cross-modal collaborative reasoning on the acquired real-time vehicle environment perception data and comprehensive driving intention using a preset multimodal large model, generate multiple alternative adjustment operations, and perform conflict collaborative arbitration between the comprehensive driving intention and each alternative adjustment operation to determine the final adjustment operation. The adaptive correction control module is used to perform time-domain feature analysis on the acquired third-state EEG signal after the vehicle performs the final adjustment operation. If it is determined that there is an error-related negative potential signal that represents the driver's expected deviation and the amplitude exceeds the preset threshold, the error-related negative potential signal and real-time vehicle environment perception data are fed into the multimodal large model to adaptively generate and execute the correction adjustment operation.

9. A vehicle control device based on electroencephalogram (EEG) signals and a multimodal large model, characterized in that, include: At least one controller; and a memory that is communicatively connected to at least one controller; The memory stores instructions that can be executed by at least one controller, which enables the at least one controller to perform the vehicle control method based on EEG signals and a multimodal large model as described in any one of claims 1-7.

10. A computer-readable storage medium storing computer-executable instructions thereon, characterized in that, When the executable instructions are executed by the controller, the vehicle control method based on EEG signals and a multimodal large model as described in any one of claims 1-7 is implemented.