Virtual human and virtual space computing brain-computer interface system
Through the electroencephalogram signal processing of deep convolutional neural networks and long-term memory networks, combined with wavelet transformation and independent component analysis, virtual scene optimization of voxel lighting and instantiated rendering, and virtual human behavior generation of reinforcement learning and emotional computing, the problem of low signal noise and rendering frame rate in the virtual human interaction system and achieves an efficient and accurate interactive experience.
Patent Information
- Application Number
- CN202510265053.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing virtual human-virtual space interaction system, EEG signal acquisition is susceptible to noise interference, poor signal quality, low accuracy of traditional algorithms, and low frame rate of virtual scene rendering, which cannot meet the needs of immersive experience.
The EEG signal intention recognition algorithm is adopted that fuses deep convolutional neural networks and long-term memory networks, combines wavelet transformation and denoising technology of independent component analysis, combines virtual scene optimization algorithms with voxel lighting and instantiated rendering, and integrates virtual human behavior generation algorithms with reinforcement learning and emotional computing, and uses a personalized EEG template construction method with deep autoencoder and cluster analysis.
It significantly improves the signal-to-noise ratio and intention recognition accuracy of EEG signals, improves the rendering frame rate and immersion of virtual scenes, enhances the naturalness and considerate interaction between virtual people and users, and improves the interactive quality.
Smart Images

Figure CN120295458A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and in particular to a virtual human and virtual space computing brain-computer interface system. Background Art
[0002] In today's digital age, the integration of virtual reality (VR) and brain-computer interface (BCI) technologies has become a frontier exploration field, giving rise to a highly potential research direction of virtual human and virtual space interaction. Traditional human-computer interaction methods, such as keyboards, mice, and even touch screen operations, appear clumsy and restricted in virtual scenarios and cannot meet the deep interaction needs of immersive experiences. Although existing brain-computer interface technologies have made progress, they still face numerous difficulties. When collecting electroencephalogram (EEG) signals, it is extremely easy to mix in electromyogram (EMG), electrooculogram (EOG), and environmental electromagnetic noise, resulting in a significant reduction in signal quality and interfering with subsequent accurate interpretation. In terms of EEG signal interpretation, traditional algorithms have low accuracy and slow speed, making it difficult to timely and accurately understand the user's intentions, greatly restricting the timeliness and rationality of virtual human interaction responses. For virtual scene rendering, conventional rendering technologies either sacrifice frame rate in the pursuit of image quality, causing frame drops, or produce rough results to ensure frame rate, losing the true sense of light and shadow and unable to give users a sense of immersion. Generally speaking, the current virtual human and virtual space interaction system urgently needs a comprehensive innovation to overcome these thorny problems and improve the interaction experience. Summary of the Invention
[0003] The present invention provides a virtual human and virtual space computing brain-computer interface system, including an EEG signal acquisition module, a signal preprocessing module, a feature extraction and recognition module, a virtual scene construction and rendering module, and an interaction control module. Using at least one EEG signal denoising algorithm, an EEG signal intention recognition algorithm, a virtual scene real-time rendering optimization algorithm, a virtual human behavior generation algorithm, and a personalized EEG template construction algorithm, it realizes the efficient interaction between virtual humans and virtual spaces; the EEG signal intention recognition algorithm has a fusion architecture of a deep convolutional neural network and a long short-term memory network, relying on the feature formula output by the convolutional layer of the deep convolutional neural network: And the long short-term memory network unit formula: i t =σ(W ii x t +b ii +W hi h t-1 +b hi ), to extract and learn the features of EEG signals changing with time and space; at the same time, first use transfer learning to complete pre-training in a publicly available large-scale EEG database to obtain a general EEG feature representation, and then use an extreme learning machine to quickly determine the calculation formula of the output weight: To adapt to the intention recognition needs of specific users or tasks.
[0004] Furthermore, the EEG signal denoising algorithm covers a scheme combining wavelet transform and independent component analysis. First, the EEG signal is decomposed into different frequency bands by wavelet transform to highlight the noise frequency band, and then independent component analysis is used to separate independent signal components, and the noise is removed according to a preset EEG signal source template. It also includes the combination of empirical mode decomposition and adaptive filtering. First, the collected original EEG signal is subjected to empirical mode decomposition to become multiple intrinsic mode functions and a residual term. Then, for the high-frequency intrinsic mode functions, an adaptive filtering algorithm is adopted to dynamically adjust the filter weights to remove noise.
[0005] Furthermore, the virtual scene real-time rendering optimization algorithm includes a hybrid rendering method based on level of detail and ray tracing. The level of detail of scene objects is determined by distance calculation, and ray tracing is carried out according to the ray propagation equation to optimize the rendering. It also simulates the global illumination effect through voxel lighting calculation, and combines instance rendering to reduce the redundant data transmission and calculation of repeated objects in the scene, improving the rendering efficiency. Furthermore, the virtual human behavior generation algorithm includes the integration of reinforcement learning and emotion computing. An action set A, a state space S, and a reward function R(s, a) of the virtual human are set up to build a reinforcement learning environment, and emotion computing is integrated at the same time. According to the emotion characteristics analyzed from the user's EEG signal, the behavior performance of the virtual human is adjusted in real time.
[0006] Furthermore, the personalized EEG template construction algorithm combines a deep autoencoder and clustering analysis. The deep autoencoder is used to extract the unique features of the user's EEG signal, and then clustering analysis is used to generate a personalized EEG template.
[0007] Furthermore, in the joint denoising of empirical mode decomposition and adaptive filtering, the denoising effect is optimized by adjusting the step size parameter of adaptive filtering and the key parameter of the termination condition of empirical mode decomposition.
[0008] Furthermore, in the intention recognition combining transfer learning and extreme learning machine, by reasonably selecting the source data of transfer learning and accurately setting the number of hidden layer nodes parameter of the extreme learning machine, the intention recognition accuracy in small and complex task scenarios is improved, and the time consumed for model training is shortened.
[0009] Furthermore, in the optimization of voxel-based global illumination and instance rendering, by adjusting the relevant parameters of the spherical harmonic function in voxel lighting calculation and optimizing the geometric data reuse strategy of instance rendering, the rendering frame rate of the virtual scene is enhanced, making the scene more realistic.
[0010] Furthermore, in the behavior of virtual humans based on the integration of reinforcement learning and emotion computing, by finely designing the weights of each item in the reward function and optimizing the parameters of the emotion feature quantization method, the satisfaction of users during the interaction between virtual humans and users is improved.
[0011] Beneficial effects:
[0012] In the processing of electroencephalogram (EEG) signals, denoising algorithms such as the combination of empirical mode decomposition and adaptive filtering, and the combination of wavelet transform and independent component analysis, compared with traditional filtering, the signal-to-noise ratio is greatly improved, laying a solid foundation for subsequent accurate recognition. The intention recognition algorithms that combine transfer learning and extreme learning machine, and the integration of deep convolutional neural network and long short-term memory network, significantly increase the accuracy rate and greatly shorten the response time, and can accurately understand the user's thoughts even in complex tasks. In terms of virtual scene rendering, voxel-based global illumination and instanced rendering, and hybrid rendering based on level of detail and ray tracing, significantly improve the frame rate, with a realistic lighting effect and a sudden increase in immersion. The generation of virtual human behavior incorporates reinforcement learning and emotion computing, and the integration of generative adversarial network and emotion-based reinforcement learning, strengthening the emotional resonance between virtual humans and users and making the interaction more natural and considerate. The construction of personalized EEG templates uses the combination of deep belief network and hierarchical clustering, which is more accurate and stable than traditional methods, can accurately capture the user's intention in the long term, and comprehensively improve the quality of interaction in the virtual space. Brief description of the drawings
[0013] Figure 1 Schematic diagram of the algorithm flow. Detailed implementation manners
[0014] Example 1:
[0015] Intention recognition algorithm for EEG signals - Modeling and solving process of the integration of deep convolutional neural network and long short-term memory network. Data preprocessing: Collect a large amount of EEG signal data from different individuals performing various tasks (such as motor imagery, visual perception tasks, etc.). Standardize the original EEG signals so that their mean is 0 and variance is 1, reducing the differences brought by different individuals and acquisition devices. Divide the data set into a training set, a validation set, and a test set according to a specific ratio, such as 70%, 20%, 10%.
[0016] Network architecture construction: The front end uses a deep convolutional neural network (CNN), with multiple convolutional layers set, such as 3 - 5 layers. The convolutional kernel size is selected as 3x3 or 5x5. Through convolutional operations, the spatial correlation features between different electrode positions of the EEG signals are automatically extracted. For example, the first convolutional layer can capture the local features of adjacent electrode signals. As the number of layers deepens, the receptive field expands, and more macroscopic spatial features are extracted. The pooling layer follows the convolutional layer, adopting the maximum pooling or average pooling strategy to reduce the data dimension and computational amount. The back end is connected to a long short - term memory network (LSTM). The input gate, forget gate, and output gate inside the LSTM unit work together. Based on the feature sequence output by the CNN, it learns the dynamic changes of the EEG signals over time.
[0017] Training process: The training set data is input into the network in sequence. The feature map output by the CNN is fed into the LSTM. The final hidden state output by the LSTM is connected to the fully - connected layer and the Softmax classifier. The cross - entropy loss function is used to measure the difference between the predicted intention category and the true label. Stochastic gradient descent (SGD) and its variants, such as Adagrad, Adam, etc., are used as optimization algorithms to update the network weights through backpropagation. During the training process, according to the loss value and accuracy on the validation set, hyperparameters such as the learning rate and batch size are adjusted to prevent overfitting.
[0018] Model evaluation: After training is completed, the test set is used to evaluate the model performance. Metrics such as accuracy, recall rate, and F1 - value are calculated to judge the accuracy of the model in identifying different intentions.
[0019] The traditional support vector machine (SVM) method is selected for comparison. On a test data set containing 10 common EEG intentions, the accuracy of the SVM method is about 60%, the recall rate is around 55%, and the F1 - value is 0.57. Under this fusion architecture, the accuracy is increased to 80%, the recall rate reaches 75%, and the F1 - value increases to 0.77. In terms of response time, the SVM takes an average of 400 milliseconds to process each sample, while this architecture only requires 200 milliseconds, significantly improving the accuracy and real - time performance of intention recognition.
[0020] Example 2:
[0021] EEG signal denoising algorithm - the solution process of the combined modeling of empirical mode decomposition and adaptive filtering. Empirical mode decomposition (EMD): After obtaining the original EEG signal, the EMD algorithm gradually decomposes it into multiple intrinsic mode functions (IMFs) and a residue term. Specifically, by finding the local extreme points of the signal, using cubic spline interpolation to fit the upper and lower envelopes, calculating the mean of the envelopes, and iterating continuously until the stopping criterion is met to obtain the IMFs. For example, for an EEG signal contaminated with electromyogram interference, the high - frequency IMF components usually contain more noise, while the low - frequency IMFs and the residue term are more inclined to retain the main components of the EEG signal.
[0022] Adaptive filtering: An adaptive filter is constructed for high-frequency IMFs. Taking the least mean square (LMS) algorithm as an example, the initial weight vector of the filter is set. The input noise reference signal (which can be obtained additionally from the acquisition environment or estimated based on signal characteristics) is input, and the error between the filter output and the desired output (i.e., the pure noise estimate) is calculated. The weights are dynamically adjusted according to the update formula. For example, the update formula is w(n + 1) = w(n) + 2μe(n)x(n), where μ is the step size parameter that controls the convergence speed, e(n) is the error signal, and x(n) is the input signal. Iterations are continued until the error converges to a certain threshold. The filtered high-frequency IMFs are recombined with other low-frequency IMFs and the residual term to obtain the denoised EEG signal.
[0023] Parameter tuning: The influence of different step size parameters μ on the denoising effect is tested through experiments. At the same time, the stopping criterion of EMD is determined according to factors such as signal length and noise intensity to optimize the overall denoising performance.
[0024] Compared with the traditional Butterworth low-pass filtering, in the tests of 20 groups of EEG signals with different noise intensities, the signal-to-noise ratio of the signals after Butterworth filtering is increased by an average of 2 dB, while the proposed joint denoising algorithm can increase the signal-to-noise ratio by an average of 5 dB. In terms of preserving signal details, Butterworth filtering will blur some high-frequency EEG rhythm features, and this algorithm can more accurately restore the subtle fluctuations of the original signal, providing a better signal basis for subsequent accurate intention recognition.
[0025] Example 3:
[0026] Optimization algorithm for real-time rendering of virtual scenes - Solving process of voxel-based global illumination and instanced rendering modeling. Voxelization and global illumination calculation: The geometric model of the virtual scene is converted into a three-dimensional voxel grid, and each voxel serves as the basic unit for illumination calculation. According to the radiosity transfer equation, the spherical harmonic function is used for low-frequency approximation of illumination. First, the initial illumination energy emitted by the scene light source is calculated, and then the illumination intensity of each voxel is updated by tracing the scattering and reflection paths of light between voxels. For example, in an indoor scene, light starts from a light bulb and is reflected and scattered by the surfaces of objects such as walls and furniture, causing each voxel to accumulate corresponding illumination energy. The calculation formula In this formula, the illumination calculation accuracy is controlled by adjusting the order N of the spherical harmonic function.
[0027] Instanced rendering: For repeated objects in the scene, such as trees in forest scenes and street lights in urban landscapes, only one copy of geometric data is generated, including vertex, patch and other information. Then, the transformation matrix is used to give each instance different position, rotation and scaling attributes. When rendering, the GPU quickly draws multiple instances based on these attributes, reducing data transmission bandwidth and calculation. At the same time, the geometric data reuse strategy of instanced rendering is optimized, such as grouping instances by distance from the viewpoint, giving priority to high-precision instances near and low-precision instances far away.
[0028] Integration and optimization: Combine voxel lighting calculation results with scene objects after instanced rendering, adjust rendering pipeline parameters such as depth buffer and stencil buffer settings, optimize rendering order, reduce rendering redundancy, and improve overall rendering efficiency.
[0029] A large virtual city landscape scene test was built with a large number of repeated buildings and vegetation. The traditional direct lighting rendering method has a frame rate of only 12fps, and the scene lighting effects are stiff and lack of realism. After adopting this optimization algorithm, the frame rate is increased to 30fps, the lighting transition is natural, and the scene realism is greatly enhanced, which can bring users a more immersive virtual space experience.
[0030] Embodiment 4:
[0031] Virtual human behavior generation algorithm - based on the fusion modeling and solution process of reinforcement learning and emotional computing, reinforcement learning environment construction: define the action set A of the virtual human, covering behaviors such as movement, dialogue, and expression changes; the state space S integrates the characteristics of the user's EEG signal, the current state of the virtual scene, and the virtual human's own state (such as fatigue and emotional value). Design a reward function R(s′a), and give positive rewards when the virtual human's behavior is in line with the user's potential intention, completes the task efficiently, or triggers the user's positive emotions, and negative rewards otherwise. For example, if the user's EEG shows that he is highly focused on a virtual object, and the virtual human actively approaches and introduces relevant information, a positive reward will be given. Using reinforcement learning algorithms such as deep Q network (DQN) or proximal policy optimization (PPO), the intelligent agent (virtual human) explores and learns in the environment based on the strategy π(a|s), and regularly updates the policy network parameters.
[0032] Emotional computing integration: Real-time analysis of emotion-related rhythms in the user's EEG signals, such as changes in alpha, beta, and theta waves, and use machine learning models, such as support vector regression or neural networks, to map EEG features to emotional states (happy, sad, anxious, etc.). Integrate emotional states into the state space of reinforcement learning and adjust the weight of the reward function so that the behavior of the virtual human not only considers task completion, but also takes into account the user's emotional experience. For example, if the user's anxiety is detected, the virtual human slows down the speech speed and switches to soothing background music.
[0033] Model Training and Optimization: In a simulated interaction environment, let the virtual human interact with the virtual user multiple times, collect interaction data, continuously optimize the policy network based on the reward feedback, adjust the parameters of the sentiment analysis model, and improve the rationality and emotional adaptability of the virtual human's behavior generation.
[0034] In the test of the simulated customer service consultation scenario, compared with the virtual human that generates behavior based on fixed rules, the user satisfaction of the virtual human with fixed rules is 60%, and the interaction process is rigid and mechanical. After adopting this fusion algorithm, the user satisfaction is increased to 85%, the interaction between the virtual human and the user is natural and smooth, and it can flexibly adjust the response method according to the user's emotions, significantly improving the interaction experience.
[0035] Example 5:
[0036] Personalized EEG Template Construction Algorithm - The Solving Process of Joint Modeling of Deep Autoencoder and Cluster Analysis
[0037] Feature Extraction of Deep Autoencoder: Collect the EEG signal data of a certain user for a long time, construct a deep autoencoder with multiple hidden layers, such as 3 - 4 layers. The encoder gradually compresses the high - dimensional EEG data into a low - dimensional feature representation, and the decoder tries to reconstruct the original data from the low - dimensional features by minimizing the reconstruction error Train the network. During the training process, adopt a strategy of combining layer - by - layer pre - training and overall fine - tuning. First, train each layer separately, and then jointly fine - tune to ensure that the user's personalized and stable EEG features are extracted.
[0038] Template Generation by Cluster Analysis: Input the low - dimensional feature vectors output by the autoencoder into the clustering algorithm. Taking K - means clustering as an example, preset the number of clusters, and allocate the data points to different clusters according to the Euclidean distance between the feature vectors. Continuously update the cluster centers Until the clusters are stable. Each cluster center represents a typical EEG pattern, and combining these patterns forms a personalized EEG template. When identifying the user's intention later, it is compared and matched with the features of the real - time EEG signal.
[0039] Template Update and Optimization: As the user continues to use the system, new EEG data accumulates continuously. Regularly input the new data into the autoencoder and the clustering model to update the personalized template and adapt to the long - term changes of the EEG signal.
[0040] For 15 users who have long participated in the test, compared with the method of constructing a template simply relying on mean statistics, in the intention recognition task, the average accuracy of the mean - statistics template method is 65%, and the misjudgment rate is about 20%. After adopting this joint solution, the accuracy is increased to 80%, and the misjudgment rate is reduced to 10%, which can more accurately capture the individual EEG signal features and improve the accuracy of the interaction system's understanding of the user's intention.
[0041] Example 6:
[0042] EEG Signal Intention Recognition Algorithm - Modeling and Solving Process Combining Transfer Learning and Extreme Learning Machine, Data Collection and Preprocessing:
[0043] EEG signal data of different people and under different experimental conditions are widely collected from multiple public data sources and the in - house - accumulated EEG database. These data cover basic movement instructions, simple perception tasks, etc. For the collected raw EEG signals, obvious outliers are first removed. These outliers may stem from momentary failures of the acquisition device or sudden interference actions of the subjects. Then, the signals are standardized so that the mean of all samples becomes 0 and the variance becomes 1, thereby eliminating the influence brought by different acquisition devices and individual physiological differences. The dataset is randomly divided into a training set, a validation set, and a test set according to the ratio of 70%, 20%, and 10%.
[0044] Transfer Learning Pretraining:
[0045] A fully - connected neural network architecture with 3 - 4 layers is selected. The number of nodes in the input layer is determined according to the number of electrodes and the feature dimensions of the EEG signals. In the pre - training stage, large - scale general EEG data are input into the network. During forward propagation, linear weighted summation is performed between neurons according to the weight matrix, and then a non - linear transformation is introduced through an activation function (such as the ReLU function). The mean square error loss between the output result and the true label is calculated, and the weights are updated by backpropagation using the stochastic gradient descent algorithm. As the number of training rounds increases, the network gradually learns the common spatio - temporal feature patterns of different EEG tasks. After the pre - training is completed, the weight parameters of each layer of the network are saved.
[0046] Extreme Learning Machine Adaptation:
[0047] When facing a specific user or a complex new task, the output layer of the pre - trained model is removed and an extreme learning machine is connected. The extreme learning machine randomly initializes the connection weight matrix between the input layer and the hidden layer, and sets the number of neurons in the hidden layer according to the complexity of the new task and the number of expected output categories. The training set data are input into this new architecture. The extreme learning machine quickly calculates the output of the hidden layer and directly determines the weights between the hidden layer and the output layer using the method of the generalized inverse matrix, so that the model output matches the true label as much as possible. In this process, with only a small number of labeled new - task samples, the model can be fine - tuned to adapt to the new scenario.
[0048] Model Fusion and Optimization:
[0049] Organically combine the feature extraction ability obtained by transfer learning with the fast learning ability of the extreme learning machine, and repeatedly test different combinations of hyperparameters on the validation set. For example, adjust parameters such as the sampling ratio and learning rate of the pre-training data in the transfer learning stage, as well as the number of hidden layer nodes and regularization coefficient of the extreme learning machine. Observe the changes in indicators such as accuracy and recall on the validation set, select the optimal parameter combination, and enable the model to achieve the best performance on the test set of the new task.
[0050] For 10 extremely niche and complex EEG intention tasks such as identifying inspiration flashes and complex artistic conceptions in the creative writing process, traditional neural networks start training from scratch. Due to scarce data, the model is difficult to converge, with an accuracy of only about 35% and a training duration of up to 8 hours. By adopting the combination scheme of transfer learning and extreme learning machine, and using the knowledge transfer of pre-training with general data, the model can quickly adapt to the new task, with the accuracy improved to 60% and the training time shortened to 1.5 hours, greatly enhancing the practicality and efficiency of the model.
[0051] Example 7:
[0052] EEG signal denoising algorithm - the modeling and solution process of the combination of wavelet transform and independent component analysis, data preparation and wavelet decomposition:
[0053] After collecting the EEG signal, first perform low-pass filtering on the signal to initially remove extreme noises such as high-frequency electromagnetic interference and ensure the basic smoothness of the signal. Select a suitable wavelet basis function, such as the db4 wavelet of the Daubechies series, which is more suitable for the EEG signal in terms of time-frequency localization characteristics. Use wavelet transform to convert the EEG signal from the time domain to the wavelet domain of different scales and frequencies. In this process, by continuously changing the scale factor and translation factor, the original signal is decomposed into wavelet coefficients of multiple different frequency bands. The wavelet coefficients in the high-frequency band correspond to the rapidly changing part of the signal, that is, the area where noise is likely to concentrate, while the low-frequency band retains more of the main rhythm of the EEG signal.
[0054] Independent component analysis separation:
[0055] Regard the wavelet coefficients of each frequency band obtained by wavelet decomposition as a mixed signal and input it into the independent component analysis model. ICA assumes that the mixed signal is linearly mixed by multiple independent source signals. Through iterative algorithms such as FastICA, continuously adjust the demixing matrix and try to separate the independent components. In this process, use the prior statistical characteristics of the EEG signal in the frequency domain and time domain, such as the energy distribution of different frequency bands and the rhythm periodicity, to construct a reference template, and use it to identify and eliminate those independent components that conform to the noise characteristics, leaving the components that are closer to the pure EEG signal.
[0056] Signal reconstruction and optimization:
[0057] The components of each frequency band retained after ICA processing are recombined back into the time-domain signal through inverse wavelet transform to obtain the pre-denoised EEG signal. To further optimize the denoising effect, a cross-validation method is adopted to fine-tune parameters such as the decomposition level of wavelet transform and the convergence threshold of ICA algorithm on a part of the samples. By comparing indicators such as the signal-to-noise ratio of the reconstructed signal and the correlation with the original clean signal under different parameter settings, the optimal parameter combination is determined, and a high-quality denoised EEG signal is output.
[0058] Thirty groups of EEG signal samples containing high-intensity environmental noise and human physiological noise (such as electromyogram, electrooculogram) interference are selected. After denoising by the traditional median filtering method, the signal-to-noise ratio of the signal is increased by an average of 1.8 dB, and the tiny EEG rhythms in the signal are still blurred. While using the method combining wavelet transform and independent component analysis, the signal-to-noise ratio can be increased by an average of up to 4.5 dB, and the weak EEG rhythm changes can be clearly restored, providing a more reliable input for subsequent accurate intention recognition.
[0059] Example 8:
[0060] Virtual scene real-time rendering optimization algorithm - The solution process of hybrid rendering modeling based on level of detail (LOD) technology and ray tracing. Scene analysis and LOD division: Load the 3D model data of the virtual scene, analyze the type, importance of each object in the scene, and the potential interaction frequency with the viewpoint. Calculate the distance from the object to the virtual viewpoint according to the distance formula, and set multiple distance thresholds according to the scene complexity. For example, the objects in the near view (distance from the viewpoint less than 5 meters) are set to high detail level, the middle view (5 - 20 meters) is set to medium detail, and the far view (more than 20 meters) is set to low detail. For different levels, simplify the geometric structure of the model, such as reducing the number of polygons and lowering the texture resolution, to generate the corresponding LOD models.
[0061] Ray tracing lighting calculation:
[0062] In the key visual area, start the ray tracing process according to the principle of light propagation. Emit a large number of rays from the viewpoint. When the rays intersect with the surface of the objects in the scene, calculate the reflection, refraction, scattering and other behaviors of the rays according to the material properties of the objects (such as reflectivity, refractive index, diffuse reflection color, etc.), trace the multiple propagation paths of the rays in the scene, accumulate the lighting energy, and accurately simulate the lighting effects in the real world, including shadows, indirect lighting, etc. Use spatial acceleration structures, such as octree, BVH tree, to accelerate the intersection calculation of rays and objects, and reduce the calculation amount.
[0063] Rendering integration and performance optimization:
[0064] In the rendering pipeline, first determine which LOD models need to be rendered based on the viewpoint position and viewing direction, and send low, medium, and high detail level objects to the rendering process in order from far to near. Superimpose the lighting effects calculated by ray tracing on the corresponding objects, use the depth buffer to ensure the correct occlusion relationship, and the template buffer to handle special effects (such as transparent object layering). Throughout the process, continuously monitor performance indicators such as frame rate and memory usage, dynamically adjust parameters such as the number of ray tracing samples and LOD switching thresholds, and balance rendering quality and real-time performance.
[0065] Constructing a large virtual ancient town scene with massive vegetation and complex buildings, using traditional raster rendering, the frame rate is only 8fps, the lighting effect is stiff, the shadows are unnatural, and the distant scenes are blurred. After using hybrid rendering based on LOD and ray tracing, the frame rate is increased to 28fps, the lighting transition is natural, the shadows are delicate and realistic, and the distant and near scenes are presented with clear and rich details, which significantly enhances the user's immersion in the virtual scene.
[0066] Embodiment 9:
[0067] Virtual human behavior generation algorithm - based on the fusion modeling and solution process of generative adversarial network (GAN) and emotional reinforcement learning, GAN construction and pre-training:
[0068] Build a generative adversarial network. The generator consists of multiple layers of fully connected layers and transposed convolution layers. The input is a fusion vector of random noise and user EEG features. It is gradually upsampled through transposed convolution to generate the initial value of the virtual human's behavior sequence, such as the coordinates of key points of posture and the word vector of the dialogue text. The discriminator is also a multi-layer convolutional network structure. It receives real behavior samples and generator outputs, extracts features through the convolution layer, and uses the fully connected layer to output the judgment result (true or false). First, use a large amount of public virtual human behavior data to pre-train GAN, so that the generator can learn the basic behavior generation mode, the discriminator can master the ability to distinguish between true and false, and use the adversarial loss function to adjust the weights of both parties.
[0069] Emotional reinforcement learning environment construction:
[0070] Define the state space of reinforcement learning, including the user's current EEG emotional characteristics (such as the quantitative values of emotional states such as excitement, calmness, anxiety, etc.), the current atmosphere of the virtual scene (lively, quiet, etc.), and the virtual person's own state (energy value, fatigue). The action space contains all possible actions of the virtual person, such as movement, expression changes, conversation initiation, etc. Design a reward function. If the virtual person's behavior triggers positive emotions in the user, fits the scene atmosphere, and conforms to its own state logic, give positive rewards; otherwise, give negative rewards. Based on the policy gradient algorithm, the intelligent agent continuously tries different actions in the simulation environment and collects reward feedback.
[0071] Fusion and collaborative training:
[0072] Integrate the optimization objective of emotional reinforcement learning into GAN training. Use the high-reward behavior samples generated during the reinforcement learning process to train the GAN discriminator, enabling the discriminator to more accurately distinguish high-quality behaviors. At the same time, utilize the new behaviors output by the GAN generator to enrich the exploration space of the reinforcement learning agent, prompting it to discover more potential high-reward behavior paths. During multiple rounds of alternating training, continuously adjust the hyperparameters of the GAN and the reinforcement learning module, such as the learning rates of the GAN generator and discriminator, the discount factor of the reinforcement learning, etc., to optimize the virtual human behavior generation effect. In the simulation test of a multi-person social gathering scenario, for a virtual human that generates behaviors solely based on preset rules, the user's emotional resonance is only 25%. The interaction process is mechanical and rigid. After adopting the fusion scheme of GAN and emotional reinforcement learning, the emotional resonance is increased to 65%. The virtual human can adjust the interaction style in real time according to the user's emotions, creating lively, interesting, and atmosphere-fitting interaction content, greatly enhancing the interaction experience.
[0073] Example 10:
[0074] Personalized EEG template construction algorithm - the joint modeling and solution process of deep belief network (DBN) and hierarchical clustering. DBN feature extraction: Input the massive EEG signal data collected from a user over a long period into the deep belief network. The DBN is composed of multiple restricted Boltzmann machines (RBMs) stacked together. The first RBM takes the EEG signal vector as the input of the visible layer. The number of neurons in the hidden layer is set according to an empirical formula and is trained through the contrastive divergence algorithm to adjust the connection weights between the visible layer and the hidden layer, enabling the RBM to learn the first-order feature representation of the input data. The subsequent RBMs take the output of the previous hidden layer as the input and repeat the training process to gradually extract higher-level and more abstract EEG features, mining deep features closely related to individual habits and thinking patterns from the original signals.
[0075] Hierarchical clustering template generation:
[0076] Input the feature vectors output by the last hidden layer of the DBN into the hierarchical clustering algorithm, such as agglomerative hierarchical clustering. At the beginning, each feature vector forms a separate class. Calculate the distance between classes, and metrics such as Euclidean distance and cosine similarity can be selected. Gradually merge the two closest classes, update the distance between the new class and other classes, and continuously iterate this process to form a clustering tree. There is no need to preset the number of clusters. Let the data's own structure determine the final clustering result. The feature vectors corresponding to each clustering center are combined to construct a personalized EEG template.
[0077] Dynamic update and adaptation:
[0078] As new EEG data are regularly collected, the new data are merged with historical data, and the DBN is retrained. Using incremental learning techniques, only the network parts affected by the new data need to be fine-tuned to quickly update the feature extraction results. Then, the new feature vectors are input into the existing hierarchical clustering model to update the cluster centers and class structures, enabling the personalized EEG template to adapt dynamically to the long-term changes in the user's EEG signals and maintaining high accuracy in user intention recognition.
[0079] An experiment was conducted on 25 long-term tracked users. Compared with the method of constructing templates using simple mean statistics, the accuracy of the mean statistic templates fluctuates greatly in the long-term intention recognition task, averaging about 55%, and is easily affected by short-term emotional fluctuations and changes in physical states of the users. After adopting the combined DBN and hierarchical clustering scheme, the long-term accuracy is stabilized above 78%, which can accurately capture the subtle long-term evolution trends of the users' EEG signals and maintain a stable high recognition accuracy.
Claims
1. A virtual human and virtual space computing brain-computer interface system, characterized in that, It includes an electroencephalogram (EEG) signal acquisition module, a signal preprocessing module, a feature extraction and recognition module, a virtual scene construction and rendering module, and an interaction control module. By using at least one EEG signal denoising algorithm, an EEG signal intention recognition algorithm, a virtual scene real-time rendering optimization algorithm, a virtual human behavior generation algorithm, and a personalized EEG template construction algorithm, it realizes the efficient interaction between the virtual human and the virtual space. The EEG signal intention recognition algorithm has a fusion architecture of a deep convolutional neural network and a long short-term memory network, relying on the feature formula output by the convolutional layer of the deep convolutional neural network: And the long short-term memory network unit formula: i t =σ(W ii x t +b ii +W hi h t-1 +b hi ), to extract the features of the learning EEG signal changing with time and space. At the same time, first use transfer learning to complete pre-training in a publicly available large-scale EEG database to obtain a general EEG feature representation, and then use an extreme learning machine to quickly determine the calculation formula of the output weight: To adapt to the intention recognition requirements of specific users or tasks.
2. The system according to claim 1, characterized in that The EEG signal denoising algorithm first decomposes the EEG signal into different frequency bands by means of wavelet transform to highlight the noise frequency band, then uses independent component analysis to separate independent signal components, and eliminates noise according to a preset EEG signal source template; it also includes the combination of empirical mode decomposition and adaptive filtering. First, the collected original EEG signal is subjected to empirical mode decomposition to become multiple intrinsic mode functions and a residual term. Then, for the high-frequency intrinsic mode functions, an adaptive filtering algorithm is used to dynamically adjust the filter weights to remove noise.
3. The system according to claim 1, wherein The virtual scene real-time rendering optimization algorithm includes a hybrid rendering method based on level of detail and ray tracing. The level of detail of scene objects is determined by distance calculation, and ray tracing optimization rendering is carried out according to the ray propagation equation; the global illumination effect is simulated by voxel lighting calculation, and instanced rendering is combined to reduce redundant data transmission and calculation of repeated objects in the scene.
4. The system according to claim 1, characterized in that, The virtual human behavior generation algorithm includes the integration of reinforcement learning and emotion computing. The action set A, state space S, and reward function R(s′a) of the virtual human are set to build a reinforcement learning environment, and emotion computing is integrated at the same time. According to the emotion characteristics analyzed from the user's EEG signal, the behavior performance of the virtual human is adjusted in real time.
5. The system according to claim 1, wherein The personalized EEG template construction algorithm combines a deep autoencoder and cluster analysis. The deep autoencoder is used to minimize the reconstruction error to extract the unique features of the user's EEG signal, and then cluster analysis is used to generate a personalized EEG template.
6. The system according to claim 2, wherein In the joint denoising of empirical mode decomposition and adaptive filtering, the denoising effect is optimized by adjusting the step size parameter of adaptive filtering and the key parameter of the termination condition of empirical mode decomposition.
7. The system according to claim 1, characterized in that In the intention recognition combining transfer learning and extreme learning machine, by reasonably selecting the source data of transfer learning and accurately setting the number of hidden layer nodes parameter of the extreme learning machine, the intention recognition accuracy in a niche and complex task scenario is improved, and the time consumed for model training is shortened.
8. The system according to claim 3, characterized in that, In the optimization of voxel-based global illumination and instanced rendering, by adjusting the relevant parameters of the spherical harmonic function in voxel lighting calculation and optimizing the geometric data reuse strategy of instanced rendering, the rendering frame rate of the virtual scene is enhanced, making the scene more realistic.
9. The system according to claim 4, characterized in that, In the virtual human behavior based on the integration of reinforcement learning and emotion computing, by carefully designing the weights of each item in the reward function and optimizing the parameters of the emotion feature quantization method, the user satisfaction when the virtual human interacts with the user is improved.
Citation Information
Patent Citations
Original electroencephalogram deep learning classification method and application
CN112790774A
Multi-agent training method and system based on reinforcement learning
CN117009811A
Rendering optimization method, electronic equipment and computer readable storage medium
CN117557703A
Multi-user real-time interaction method based on electroencephalogram signals and computer equipment
CN118131917A
Immersive space virtual-real interaction method and system based on meta universe
CN119311124A