VR scene automatic generation method based on deep learning

Through the automatic generation method of VR scenes based on deep learning, the existing VR scene generation methods are solved, and an efficient, real and changeable virtual environment is realized, and the user's immersion and interactivity are enhanced.

CN120107520AInactive Publication Date: 2025-06-06THE FIRST AFFILIATED HOSPITAL OF XINXIANG MEDICAL UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510055709.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing VR scene generation methods rely on a large number of manual modeling and texture maps, which are time-consuming and labor-intensive and difficult to achieve highly realistic effects. They ignore the experience of the olfactory sense, resulting in insufficient immersion, inflexible viewing angle switching, users cannot freely observe the virtual environment, and the scene lacks the ability to change in real time.

Method used

The automatic generation method of VR scenes based on deep learning is adopted, and multimodal feature data is automatically generated through data collection and preprocessing, scene perspective capture, real-time weather collection, physiological information collection and analysis, and generation of adversarial network (GAN) generation models, so as to realize multi-view switching, odor linkage, real-time weather system and emotional state adjustment.

Benefits of technology

It reduces the demand for manual modeling and texture maps, improves productivity, enhances user immersion and interactivity, and provides a more realistic and changeable virtual environment to help users adjust their emotions and improve overall satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107520A_ABST
    Figure CN120107520A_ABST
Patent Text Reader

Abstract

The invention discloses a VR scene automatic generation method based on deep learning, and relates to the related technical field of computer vision, and the method comprises the following steps: collecting image, sound and smell multi-modal feature data in a scene, and carrying out the preprocessing and feature extraction; different visual angles in the scene are collected, and a user freely switches the visual angles for observation; collecting real-time weather data, and generating a VR scene to reflect the current weather condition; evaluating the emotional state of the user; and constructing a scene generation model of the generative adversarial network, and automatically generating a VR scene according to the scene data, the visual angle, the weather and the emotional state. According to the invention, the deep learning can automatically learn and generate a VR scene, the requirements for artificial modeling and texture mapping are reduced, the production efficiency is improved, the GANs can continuously optimize the generated scene, the quality and diversity of the scene are improved, various leading-edge technologies such as smell linkage, multi-view switching, emotional state and real-time weather system are fused together, and the method is suitable for popularization and application. And a unique VR scene generation system is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to computer vision, and in particular to a method for automatically generating VR scenes based on deep learning. Background Art

[0002] With the continuous advancement of virtual reality (VR) technology, its applications in games, entertainment, education, medical care and other fields are becoming increasingly widespread. However, the existing VR scene generation methods still have many shortcomings in terms of realism and user experience.

[0003] Traditional VR scene generation methods mainly rely on a large amount of manual modeling and texture mapping, which is not only time-consuming and labor-intensive, but also difficult to achieve highly realistic effects. In addition, the manual modeling method limits the diversity and complexity of the scene, making the generated scene often lack the richness and details of the real world; existing VR technology mainly focuses on the simulation of vision and hearing, while ignoring the experience of other senses such as smell, resulting in insufficient immersion; the perspective switching of most VR systems is not flexible enough, and users cannot freely observe the virtual environment from different angles, which limits interactivity and user experience. Existing VR scenes often lack the ability to change in real time, such as the dynamic adjustment of the weather system, making the scene appear static and unreal.

[0004] Therefore, it is necessary to propose a deep learning-based VR scene automatic generation method to solve the above problems. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a method for automatically generating VR scenes based on deep learning, which solves the problem that the existing VR scene generation method mainly relies on a large amount of manual modeling and texture mapping, which is not only time-consuming and labor-intensive, but also difficult to achieve highly realistic effects, ignores the experience of the olfactory sense, and leads to insufficient immersion; the perspective switching is not flexible enough, and users cannot freely observe the virtual environment from different angles.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0007] A method for automatically generating a VR scene based on deep learning, comprising the following steps:

[0008] Step 1: Data collection and preprocessing: Collect multimodal feature data for VR scenes, including images, sounds, and smells, and perform preprocessing and feature extraction;

[0009] Step 2: Scene perspective capture: multiple cameras are set up in the scene to collect different perspectives of the scene, and users can freely switch perspectives for observation;

[0010] Step 3: Real-time weather collection: collect real-time weather data as one of the input parameters and generate a VR scene to reflect the current weather conditions;

[0011] Step 4: Physiological information collection and analysis: Evaluate the user's emotional state by analyzing the user's physiological signals (such as heart rate, skin conductivity) and using natural language processing technology to analyze the user's voice or text input;

[0012] Step 5: Scene generation: Build a scene generation model based on generative adversarial network (GAN) to automatically generate VR scenes based on scene data, perspective, weather, and emotional state.

[0013] Optionally, the feature extraction described in step 1 includes the following steps:

[0014] S1: Use a convolutional neural network to extract features from image data. The convolutional neural network includes a convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer extracts local features through a convolution operation.

[0015] The pooling layer reduces the feature map size by maximum pooling or average pooling;

[0016] The fully connected layer flattens the feature map into a vector for classification or regression;

[0017] S2: extracting features from the sound data using a recurrent neural network, wherein the recurrent neural network (RNN) is used to extract features from the sound data and includes a recurrent layer and an output layer;

[0018] The recurrent layer processes sequence data and captures temporal dependencies;

[0019] The output layer designs output according to task requirements;

[0020] S3: Feature extraction of odor data using chemical sensors;

[0021] The chemical sensor response model is as follows: the concentration of the chemical substance is converted into a numerical representation, and feature extraction is performed to extract features from the sensor response.

[0022] Optionally, the scene perspective capturing described in step 2 includes the following steps:

[0023] S1: Data collection: In the environment, multiple cameras are deployed to capture image data from different perspectives. The cameras are installed on mobile devices such as drones or robots.

[0024] S2: Synchronization and calibration: synchronize and calibrate multiple cameras to ensure that the image data captured by multiple cameras can be accurately integrated. Synchronization means ensuring that all cameras start and stop capturing images at the same time, while calibration means adjusting the parameters of the cameras so that the images they capture are geometrically consistent.

[0025] S3: Data processing: Extract features from image data from multiple perspectives, including edges and corners. The features are used to establish correspondence between different images for subsequent 3D modeling. Through feature matching algorithms, the correspondence between the same features in different images is found. This is one of the key steps in multi-perspective image reconstruction.

[0026] Optionally, the real-time weather collection described in step 3 includes the following steps:

[0027] S1: API real-time weather collection: Use the real-time weather API to obtain current weather conditions, including temperature, humidity, wind speed, and precipitation probability. The API provides data in JSON or XML format for easy program parsing and processing;

[0028] S2: Data preprocessing: Clean, format, and convert weather data obtained from the API for use in VR scenarios, including removing unnecessary information, unifying data formats, and converting data into a coordinate system suitable for VR rendering;

[0029] S3: Real-time update: Use WebSocket real-time communication technology to push processed weather data to the VR client, and the client dynamically updates the weather elements in the VR scene, such as clouds, raindrops, and wind direction, based on the received data;

[0030] S4: Data visualization: Present abstract weather data to users in an intuitive form by creating 3D models, textures, and animations, so that users can intuitively experience scene changes under different weather conditions in a VR environment.

[0031] Optionally, the A-FRAME framework is used to build a VR environment in the real-time weather collection step. The framework is based on HTML syntax and is used to simplify the creation process of 3D and VR content. Through custom tags and components of A-FRAME, scenes, entities, lighting, camera VR elements are defined, and real-time weather data is integrated.

[0032] Optionally, the physiological information collection and analysis described in step 4 includes the following steps:

[0033] S1: Data collection: Use wearable devices (such as smart watches and heart rate monitors) to collect users’ physiological data in real time, including heart rate, skin conductivity, and breathing rate; record users’ voice input through microphones; and collect text data from users’ interactions with the system through text input (such as chat boxes);

[0034] S2: Data preprocessing: Physiological signals are filtered and normalized; speech signals are denoised, features (such as MFCC) are extracted, and they are segmented into frames; text data is segmented, stop words are removed, and stems are extracted;

[0035] S3: Feature extraction: Use time series analysis methods (such as Fourier transform) to extract features of physiological signals; use Mel-frequency cepstral coefficients (MFCC) to extract fundamental frequency, pitch, and formant features of speech signals, and use word embedding models (such as Word2Vec, BERT) to convert text into vector representation;

[0036] S4: Physiological signal emotion classification: Support vector machine (SVM) is used to classify physiological signal features, recurrent neural network (RNN) is used to classify speech features, and deep learning models (such as CNN, RNN, Transformer) are used to classify text features. Combined with information from multiple modalities of physiological signals, speech and text, multimodal fusion algorithms (such as decision tree integration) are used to improve the accuracy of emotion recognition. The emotion classification results of the above parts are comprehensively analyzed using the weighted average method to obtain the final user emotional state (such as happiness, sadness, anger).

[0037] Optionally, the generative adversarial network (GAN) described in step 5 mainly comprises a generator and a discriminator. The goal of the generator is to generate realistic data, and the goal of the discriminator is to distinguish the generated data from the real data. The two parts compete with each other through adversarial training to improve the quality of the generated data.

[0038] The generator includes an input layer, a hidden layer, a view embedding layer, a weather embedding layer, a feature fusion layer, a generator network and an output layer;

[0039] Input layer: input image, sound, smell multimodal features F, view parameters V, real-time weather data W, and sentiment analysis results E; hidden layer: including multiple fully connected layers or convolutional layers for learning data distribution; view embedding layer: embed the view parameters V into a high-dimensional space to combine with multimodal features, real-time weather data and sentiment analysis results; weather embedding layer: embed the real-time weather data W into a high-dimensional space to combine with multimodal features, view parameters and sentiment analysis results; feature fusion layer: embed the multimodal features, sentiment analysis results, view parameters and real-time weather data for fusion; generator network: including multiple fully connected layers or convolutional layers for generating VR scene data from the fused features; output layer: generated VR scene data G (F, E, V, W);

[0040] The discriminator includes an input layer, a hidden layer, an output layer and a discriminator network; the input layer: inputs real VR scene data Xreal and generated VR scene data G(F, E, V, W); the hidden layer: includes multiple fully connected layers or convolutional layers, used to extract features and determine whether the input is real data; the output layer: probability values ​​D(Xreal) and D(G(F, E, V, W)), indicating the probability that the input is real data; the discriminator network: includes multiple fully connected layers or convolutional layers, used to determine whether the input VR scene data is real data.

[0041] Optionally, the adversarial training process of the generative adversarial network (GAN) is:

[0042] S1: Initialization: Randomly initialize the parameters of the generator and discriminator;

[0043] S2: Discriminator update: Sample a batch of real data from the real data distribution and generate a batch of fake data from the generator, and use these data to update the parameters of the discriminator to maximize its loss function;

[0044] S3: Generator update: Update the parameters of the generator using only the fake data generated by the generator to minimize its loss function;

[0045] S4: View parameter learning: During the training process, the view parameter V is optimized by gradient descent to learn the best view representation;

[0046] S5: Weather parameter learning: The real-time weather data W is optimized via gradient descent to learn how to adjust the generated scenes according to different weather conditions;

[0047] S6: Emotional parameter learning: The sentiment analysis result E is optimized via gradient descent to learn how to adjust the generated scenes according to different emotional states:

[0048] S7: Alternating iteration: repeat steps S2, S3, S4, S5 and S6 until a predetermined number of training rounds is reached or the loss function converges;

[0049] S8: Application scenario generation: After training is completed, the trained generator is used to generate VR scenarios based on new multimodal features, sentiment analysis results, viewing angle parameters, and real-time weather data; mathematically expressed as:

[0050] S=G(F new , E new , V new , W new )

[0051] Among them, F new is a new multimodal feature, E new is the new sentiment analysis result, V new is the new viewing angle parameter, W new is the new real-time weather data.

[0052] Optionally, a data management system is provided in the scene generation step, and the data management system is used to store and manage scene data. The data management system is a relational database (such as MySQL, PostgreSQL), and the database is used to store scene models, materials, textures, audio, video elements, as well as interactive elements, physical properties, and animation effect information in the scene.

[0053] The present invention provides a method for automatically generating VR scenes based on deep learning, which has the following beneficial effects:

[0054] 1. The deep learning of the present invention can automatically learn and generate VR scenes, reducing the need for manual modeling and texture mapping and improving production efficiency. Traditional VR scene generation requires a lot of manual work, while GANs can automatically complete these tasks, saving a lot of time and cost. Through continuous training and iteration, GANs can continuously optimize the generated scenes and improve their quality and diversity.

[0055] 2. Through the odor linkage system, the present invention adds olfactory stimulation to the user on the basis of vision and hearing. Multi-sensory stimulation can significantly enhance the user's sense of immersion. Therefore, integrating the odor linkage system into the VR scene and releasing corresponding odors according to changes in the virtual environment can further enhance the user's immersive experience and make the user feel that he is really in the environment.

[0056] 3. The multi-perspective switching technology of the present invention allows users to freely move their perspectives in VR scenes, enhancing interactivity and realism. Users can freely observe the surrounding environment and discover details from different angles just like in the real world. Multi-perspective technology can reduce motion sickness caused by rapid movement or rotation, and improve user participation and satisfaction.

[0057] 4. The present invention introduces a real-time weather system to enable VR scenes to change dynamically, thereby enhancing the vividness and realism of the scenes. Real-time weather scene generation allows users to interact with the virtual environment more naturally. Users can observe the movement of clouds, feel the flow of wind, and even influence weather changes through specific operations, allowing users to feel the variability and richness of the environment.

[0058] 5. The present invention can help users regulate their emotions by analyzing their emotional state and adjusting the VR scene accordingly. When the user feels angry or anxious, a calm and relaxing scene is generated, which helps to alleviate negative emotions. This technology can also be applied to the field of psychotherapy. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a flowchart of the method for automatically generating a VR scene according to the present invention;

[0060] Figure 2 A schematic diagram of the adversarial training process of the generative adversarial network of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] Example 1

[0063] See also Figure 1 , a method for automatically generating VR scenes based on deep learning, comprising the following steps:

[0064] Step 1: Data collection and preprocessing: Collect multimodal feature data for VR scenes, including images, sounds, and smells, and perform preprocessing and feature extraction;

[0065] Feature extraction includes the following steps:

[0066] S1: Use convolutional neural network to extract features from image data. The convolutional neural network includes convolution layer, pooling layer and fully connected layer. The convolution layer extracts local features through convolution operation, which can be expressed mathematically as:

[0067]

[0068] Among them, I is the input image, K is the convolution kernel, and C(i,j) is the position of the feature map;

[0069] The pooling layer reduces the feature map size by maximum pooling or average pooling, which can be expressed mathematically as:

[0070] P(i,j)=max(C(i,j)) (Maximum pooling)

[0071] (average pooling);

[0072] The fully connected layer flattens the feature map into a vector for classification or regression, which is mathematically expressed as:

[0073] y=W·x+b,

[0074] Among them, W is the weight matrix, x is the input vector, and b is the bias term;

[0075] S2: Using a recurrent neural network to extract features from sound data. The recurrent neural network (RNN) is used to extract features from sound data, including a recurrent layer and an output layer.

[0076] The recurrent layer processes sequential data and captures temporal dependencies, which can be expressed mathematically as:

[0077] h t =σ(W·[h t-1 , x t ]+b),

[0078] Among them, h t is the hidden state, x t is the input, W and b are parameters, and σ is the activation function;

[0079] The output layer designs the output according to the task requirements, which can be expressed mathematically as:

[0080] y t =g(V·h t +c),

[0081] Among them, V and c are parameters, and g is the activation function;

[0082] S3: Feature extraction of odor data using chemical sensors;

[0083] The sensor response model of the chemical sensor is: convert the chemical concentration into a numerical representation, mathematically expressed as: R = f(C), where R is the sensor response, C is the chemical concentration vector, and f is the conversion function; and perform feature extraction to extract features from the sensor response, mathematically expressed as: F = extract_features(R), where F is the extracted feature vector

[0084] Step 2: Scene perspective capture: multiple cameras are set up in the scene to collect different perspectives of the scene, and users can freely switch perspectives for observation;

[0085] Step 3: Real-time weather collection: collect real-time weather data as one of the input parameters and generate a VR scene to reflect the current weather conditions;

[0086] Step 4: Physiological information collection and analysis: Evaluate the user's emotional state by analyzing the user's physiological signals (such as heart rate, skin conductivity) and using natural language processing technology to analyze the user's voice or text input;

[0087] Step 5: Scene generation: Build a scene generation model based on generative adversarial network (GAN) to automatically generate VR scenes based on scene data, perspective, weather, and emotional state.

[0088] In this embodiment, through the odor linkage system, the user adds olfactory stimulation on the basis of vision and hearing, which greatly enhances the sense of immersion. For example, smelling the fragrance of grass and trees in a virtual forest, or smelling the smell of car exhaust in a virtual city, makes the user feel that he is really in that environment. The multi-view switching technology allows the user to freely move the perspective in the VR scene, which enhances the interactivity and realism. The user can freely observe the surrounding environment as in the real world and discover details from different angles, which improves the user's sense of participation and satisfaction. The real-time weather system is introduced to enable the VR scene to change dynamically, such as rain, snow, wind and other natural phenomena, which enhances the vividness and realism of the scene. For example, experiencing a sudden rainstorm in a virtual city, or seeing snowflakes falling in a virtual mountain area, makes the user feel the variability and richness of the environment.

[0089] Using deep learning models to automatically learn and generate high-quality 3D models and textures significantly improves the efficiency and quality of VR scene generation. For example, by training deep neural networks, complex 3D scenes and high-resolution textures can be quickly generated, reducing the time and cost of manual modeling.

[0090] Example 2

[0091] This embodiment is a further optimization based on the embodiment 1. Specifically, the scene perspective capture in step 2 includes the following steps:

[0092] S1: Data collection: In the environment, multiple cameras are deployed to capture image data from different perspectives. The cameras are installed on mobile devices such as drones or robots.

[0093] S2: Synchronization and calibration: Synchronize and calibrate multiple cameras to ensure that the image data captured by multiple cameras can be accurately integrated. Synchronization means ensuring that all cameras start and stop capturing images at the same time, while calibration means adjusting the parameters of the cameras so that the images they capture are geometrically consistent. In addition to image data, depth information can also be recorded using structured light scanners or depth camera devices. Depth information is crucial for subsequent 3D reconstruction and model optimization.

[0094] S3: Data processing: Extract features from image data from multiple perspectives, including edges and corners. The features are used to establish correspondence between different images for subsequent 3D modeling. Through feature matching algorithms, the correspondence between the same features in different images is found. This is one of the key steps in multi-perspective image reconstruction.

[0095] In this embodiment, providing multiple perspectives through scene perspective capture can enhance immersion and improve user experience; through multi-perspective scene generation, users can observe the virtual environment from different perspectives. This all-round visual experience can greatly enhance the user's immersion. Users can explore different parts of the scene by changing the observation angle. This dynamic interaction method makes the virtual environment more vivid and realistic.

[0096] Users can choose the viewing angle according to their preferences. This personalized experience can meet the needs of different users and improve overall satisfaction. Multi-perspective technology can reduce motion sickness caused by rapid movement or rotation because it allows users to explore virtual space in a more natural way.

[0097] Example 3

[0098] This embodiment is a further optimization based on the embodiment 1. Specifically, the real-time weather collection in step 3 includes the following steps:

[0099] S1: API real-time weather collection: Use the real-time weather API to obtain current weather conditions, including temperature, humidity, wind speed, and precipitation probability. The API provides data in JSON or XML format for easy program parsing and processing;

[0100] S2: Data preprocessing: Clean, format, and convert weather data obtained from the API for use in VR scenarios, including removing unnecessary information, unifying data formats, and converting data into a coordinate system suitable for VR rendering;

[0101] S3: Real-time update: Use WebSocket real-time communication technology to push processed weather data to the VR client, and the client dynamically updates the weather elements in the VR scene, such as clouds, raindrops, and wind direction, based on the received data;

[0102] S4: Data visualization: Present abstract weather data to users in an intuitive form by creating 3D models, textures, and animations, so that users can intuitively experience scene changes under different weather conditions in a VR environment;

[0103] In the real-time weather collection step, the A-FRAME framework is used to build a VR environment. The framework is based on HTML syntax and is used to simplify the creation process of 3D and VR content. Through the custom tags and components of A-FRAME, scenes, entities, lighting, camera VR elements are defined, and real-time weather data is integrated.

[0104] In this embodiment, by integrating real-time weather information, the VR scene can be dynamically adjusted according to changes in the external environment, allowing users to experience a more realistic and immersive experience. For example, when it rains outside, the VR scene will also simulate the falling of raindrops and a wet environment. Real-time weather scene generation allows users to interact with the virtual environment more naturally. Users can observe the movement of clouds, feel the flow of wind, and even affect weather changes through specific operations, thereby obtaining a richer interactive experience.

[0105] For applications in certain specific areas, such as agriculture and travel planning, the integration of real-time weather information can help users better assess the impact of weather changes on crop growth or travel arrangements, thereby improving the efficiency and accuracy of decision-making.

[0106] Example 4

[0107] This embodiment is a further optimization based on the embodiment 1. Specifically, the physiological information collection and analysis in step 4 includes the following steps:

[0108] S1: Data collection: Use wearable devices (such as smart watches and heart rate monitors) to collect users’ physiological data in real time, including heart rate, skin conductivity, and breathing rate; record users’ voice input through microphones; and collect text data from users’ interactions with the system through text input (such as chat boxes);

[0109] S2: Data preprocessing: Physiological signals are filtered and normalized; speech signals are denoised, features (such as MFCC) are extracted, and they are segmented into frames; text data is segmented, stop words are removed, and stems are extracted;

[0110] S3: Feature extraction: Use time series analysis methods (such as Fourier transform) to extract features of physiological signals; use Mel-frequency cepstral coefficients (MFCC) to extract fundamental frequency, pitch, and formant features of speech signals, and use word embedding models (such as Word2Vec, BERT) to convert text into vector representation;

[0111] The mathematical formula for physiological signal feature extraction is: Assuming the heart rate signal is H(t), its frequency domain features can be extracted through Fourier transform;

[0112]

[0113] S4: Physiological signal emotion classification: Support vector machine (SVM) is used to classify physiological signal features, recurrent neural network (RNN) is used to classify speech features, and deep learning models (such as CNN, RNN, Transformer) are used to classify text features. Combined with information from multiple modalities of physiological signals, speech and text, multimodal fusion algorithms (such as decision tree integration) are used to improve the accuracy of emotion recognition. The emotion classification results of the above parts are comprehensively analyzed using the weighted average method to obtain the final user emotional state (such as happiness, sadness, anger);

[0114] The mathematical formula for sentiment classification using support vector machine is: train an SVM classifier to classify the physiological signal features into sentiments: y = SVM (X), where X is the physiological signal feature vector and y is the predicted sentiment label;

[0115] The mathematical formula for emotion recognition results that use weighted average method to fuse multiple modalities is:

[0116]

[0117] Among them, E final is the final emotional state, E i is the emotion recognition result of the i-th modality, w i is the weight of the i-th mode.

[0118] In this embodiment, by analyzing the user's emotional state, the scene can be automatically adjusted to match the user's emotional state, which can enhance the user's emotional resonance. By analyzing the user's emotional state, the system can provide a more personalized VR experience. For example, when the user feels happy, a vibrant and colorful scene can be generated; when the user feels sad, a warm and comforting environment can be provided. Automatically adjusting the scene to match the user's emotional state can enhance the user's emotional resonance and make the user feel understood and cared for, thereby improving overall satisfaction and immersion.

[0119] By analyzing the user's emotional state, the VR scene can be adjusted accordingly to help the user regulate their emotions. For example, when the user feels angry or anxious, a calm and relaxing scene can be generated to help alleviate negative emotions. This technology can also be applied to the field of psychotherapy to help patients deal with emotional problems such as phobias and depression by creating specific virtual environments.

[0120] Example 5

[0121] See also Figure 2 This embodiment is a further optimization based on the embodiment 1. Specifically, the main components of the generative adversarial network (GAN) in step 5 include a generator and a discriminator. The goal of the generator is to generate realistic data, and the goal of the discriminator is to distinguish the generated data from the real data. The two parts compete with each other through adversarial training to improve the quality of the generated data.

[0122] The generator includes input layer, hidden layer, view embedding layer, weather embedding layer, feature fusion layer, generator network and output layer;

[0123] Input layer: input image, sound, smell multimodal features F, view parameters V (viewing angle and position), real-time weather data W, and sentiment analysis results E;

[0124] Hidden layer: includes multiple fully connected layers or convolutional layers, which are used to learn the distribution of data;

[0125] View embedding layer: embeds the view parameter V into a high-dimensional space to combine with multimodal features, real-time weather data, and sentiment analysis results; mathematical expression is:

[0126] V embed = embedding(V);

[0127] Weather embedding layer: embeds real-time weather data W into a high-dimensional space to combine with multimodal features, view parameters, and sentiment analysis results; mathematical expression is:

[0128] W embed = embedding(W);

[0129] Feature fusion layer: multimodal features, sentiment analysis results, view parameters and real-time weather data are embedded and fused; mathematical expression is:

[0130] F combined =concatenate(F,E,V embed ,V embed );

[0131] Generator network: includes multiple fully connected layers or convolutional layers, which are used to generate VR scene data from the fused features; mathematically expressed as:

[0132] G(F combined )=generator(F combined );

[0133] Output layer: generated VR scene data G (F, E, V, W);

[0134] The discriminator includes an input layer, a hidden layer, an output layer, and a discriminator network; the input layer: inputs the real VR scene data Xreal and the generated VR scene data G(F, E, V, W); the hidden layer: includes multiple fully connected layers or convolutional layers, which are used to extract features and determine whether the input is real data; the output layer: probability values ​​D(Xreal) and D(G(F, E, V, W)), which represent the probability that the input is real data; the discriminator network: includes multiple fully connected layers or convolutional layers, which are used to determine whether the input VR scene data is real data; the mathematical expression is:

[0135] D(X) = discriminator(X);

[0136] Loss function of the generator: The goal of the generator is to maximize the probability that the discriminator mistakenly identifies the generated data as real data, that is, to minimize the following loss function; mathematically expressed as:

[0137]

[0138] Where z is a random noise vector, usually from a standard normal distribution or other distributions;

[0139] Loss function of the discriminator: The goal of the discriminator is to correctly distinguish between real data and generated data, that is, to maximize the following loss function; mathematically expressed as:

[0140]

[0141] Among them, x is the real data sample;

[0142] Overall loss function: The overall loss function of GAN is the sum of the generator loss and the discriminator loss; mathematically expressed as:

[0143] L GAN =L G +L D .

[0144] The adversarial training process of the Generative Adversarial Network (GAN) is:

[0145] S1: Initialization: Randomly initialize the parameters of the generator and discriminator;

[0146] S2: Discriminator update: Sample a batch of real data from the real data distribution and generate a batch of fake data from the generator, and use these data to update the parameters of the discriminator to maximize its loss function;

[0147] S3: Generator update: Update the parameters of the generator using only the fake data generated by the generator to minimize its loss function;

[0148] S4: View parameter learning: During the training process, the view parameter V is optimized by gradient descent to learn the best view representation;

[0149] S5: Weather parameter learning: The real-time weather data W is optimized via gradient descent to learn how to adjust the generated scenes according to different weather conditions;

[0150] S6: Emotional parameter learning: The sentiment analysis result E is optimized via gradient descent to learn how to adjust the generated scenes according to different emotional states:

[0151] S7: Alternating iteration: repeat steps S2, S3, S4, S5 and S6 until a predetermined number of training rounds is reached or the loss function converges;

[0152] S8: Application scenario generation: After training is completed, the trained generator is used to generate VR scenarios based on new multimodal features, sentiment analysis results, viewing angle parameters, and real-time weather data; mathematically expressed as:

[0153] S=G(F new , E new , V new , W new )

[0154] Among them, F new is a new multimodal feature, E new is the new sentiment analysis result, V new is the new viewing angle parameter, W new is the new real-time weather data;

[0155] Suppose the user is currently in an angry state. Once the system detects this emotion, it will automatically adjust the light in the VR scene to dark red, add gloomy background music, and simulate a thunderstorm effect to match the user's current emotional state. In addition, the system will adjust the scene's perspective and details in real time according to the user's position changes, providing a more personalized immersive experience.

[0156] The scene generation step also includes the odor linkage mechanism and real-time weather system scene generation;

[0157] Smell linkage mechanism: according to the specific elements in the scene (such as flowers, grass), the corresponding smell is released through the smell generator to enhance the sense of immersion; Smell linkage scene generation: such as instantaneous point source release, using Gaussian distribution function as analysis, in the two-dimensional case, the smell concentration C(x,y,t) is expressed as:

[0158]

[0159] , where: M is the total amount of odor released, (x 0 ,y 0 ) is the location of the odor source, D is the diffusion coefficient, and t is the time; assuming that in a closed room of 10×10 meters, the odor is released from the center point (5,5) meters away, the diffusion coefficient D = 0.01 square meters per second, and the total amount released M = 1 gram, then at t = 10 seconds, the odor concentration distribution can be calculated by the above formula;

[0160] Real-time weather system scene generation mechanism: Use fluid dynamics equations to simulate atmospheric movement and generate real-time weather effects. The Navier-Stokes equations are used to describe the movement of incompressible fluids:

[0161]

[0162] Among them, u is the fluid velocity field, p is the pressure field, μ is the dynamic viscosity coefficient, and f is the external force field.

[0163] A data management system is set up in the scene generation step. The data management system is used to store and manage scene data. The data management system is a relational database (such as MySQL, PostgreSQL). The database is used to store scene models, materials, textures, audio, video elements, as well as interactive elements, physical properties, and animation effect information in the scene.

[0164] In this embodiment, the generative adversarial network (GANs) can generate high-quality three-dimensional models and textures through the adversarial process of training the generator and the discriminator, making the generated VR scene more realistic. For example, during the training process, the generator continuously tries to generate more realistic images, while the discriminator continuously learns how to distinguish between real images and generated images, thereby improving the quality of the generated images. GANs can capture subtle details in the scene, such as light and shadow effects and material texture, making the generated scene more vivid and realistic.

[0165] GANs can automatically learn and generate VR scenes, reducing the need for manual modeling and texture mapping, and improving production efficiency. Traditional VR scene generation requires a lot of manual work, while GANs can automatically complete these tasks, saving a lot of time and cost. Through continuous training and iteration, GANs can continuously optimize the generated scenes and improve their quality and diversity. For example, with the continuous increase of training data and the continuous adjustment of model parameters, the generated scenes will become closer and closer to the complexity and diversity of the real world.

[0166] By simulating and generating smells, users can not only see and hear in the virtual world, but also smell various smells, such as the fragrance of flowers, thus providing a more realistic sensory experience. Multi-perspective scene generation technology allows users to observe and experience scenes from different angles and positions, increasing flexibility and selectivity. The real-time weather system can simulate different climatic conditions, such as sunny days, rainy days, and snowy days, making the virtual environment more varied and realistic. By analyzing the user's emotional state, the system can automatically adjust the scene, such as generating a sunny scene when the user feels happy, and a warm environment when the user is sad.

[0167] It integrates multiple cutting-edge technologies such as odor linkage, multi-perspective switching, emotional state and real-time weather system to form a unique VR scene generation system, demonstrating the innovation of the technology; it uses deep learning models to learn and simulate complex phenomena, which improves the intelligence level of scene generation. In the field of education, this technology can be used to create a more vivid and realistic learning environment and improve learning effects; by providing a more realistic and immersive VR experience, it helps to improve people's quality of life, especially in entertainment and relaxation. In the medical field, by simulating different environments and situations, it can help patients with rehabilitation training or psychotherapy.

[0168] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A method for automatically generating VR scenes based on deep learning, characterized in that: The following steps are involved: Step 1: Data collection and preprocessing: Collect multimodal feature data for VR scenes, including images, sounds, and smells, and perform preprocessing and feature extraction; Step 2: Scene perspective capture: multiple cameras are set up in the scene to collect different perspectives of the scene, and users can freely switch perspectives for observation; Step 3: Real-time weather collection: collect real-time weather data as one of the input parameters and generate a VR scene to reflect the current weather conditions; Step 4: Physiological information collection and analysis: Evaluate the user's emotional state by analyzing the user's physiological signals and using natural language processing technology to analyze the user's voice or text input; Step 5: Scene generation: Build a scene generation model based on a generative adversarial network to automatically generate VR scenes based on scene data, perspective, weather, and emotional state.

2. The method for automatically generating VR scenes based on deep learning according to claim 1, characterized in that: Feature extraction as described in step 1 The following steps are involved: S1: Use a convolutional neural network to extract features from image data. The convolutional neural network includes a convolution layer, a pooling layer, and a fully connected layer. The convolution layer extracts local features through convolution operations; the pooling layer reduces the size of the feature map through maximum pooling or average pooling; the fully connected layer flattens the feature map into a vector for classification or regression; S2: Extract features from the sound data using a recurrent neural network, wherein the recurrent neural network is used to extract features from the sound data and includes a recurrent layer and an output layer; the recurrent layer processes sequence data and captures time dependencies; and the output layer designs outputs according to task requirements; S3: Using a chemical sensor to extract features from the odor data; the sensor response model of the chemical sensor is: converting the concentration of the chemical substance into a numerical representation, and performing feature extraction to extract features from the sensor response.

3. The method for automatically generating VR scenes based on deep learning according to claim 1, characterized in that: The scene perspective capture described in step 2 includes the following steps: S1: Data collection: In the environment, multiple cameras are deployed to capture image data from different perspectives. The cameras are installed on mobile devices. S2: Synchronization and calibration: synchronize and calibrate multiple cameras to ensure that the image data captured by multiple cameras can be accurately integrated. Synchronization means ensuring that all cameras start and stop capturing images at the same time. Calibration means adjusting the parameters of the camera so that the images it captures are geometrically consistent. S3: Data processing: Extract features from image data from multiple perspectives, including edges and corners. The features are used to establish correspondence between different images. Through feature matching algorithms, the correspondence between the same features in different images is found.

4. The method for automatically generating VR scenes based on deep learning according to claim 1, characterized in that: The real-time weather collection described in step 3 includes the following steps: S1: API real-time weather collection: Use the real-time weather API to obtain current weather conditions, including temperature, humidity, wind speed, and precipitation probability. The API provides data in JSON or XML format; S2: Data preprocessing: Clean, format and convert weather data obtained from the API, including removing unnecessary information, unifying data formats, and converting data into a coordinate system suitable for VR rendering; S3: Real-time update: Use WebSocket real-time communication technology to push processed weather data to the VR client, and the client dynamically updates the weather elements in the VR scene based on the received data; S4: Data visualization: Present abstract weather data to users in an intuitive form by creating 3D models, textures, and animations, so that users can intuitively experience scene changes under different weather conditions in a VR environment.

5. The method for automatically generating VR scenes based on deep learning according to claim 4, characterized in that: In the real-time weather collection step, the A-FRAME framework is used to build a VR environment. The framework is based on HTML syntax and is used to simplify the creation process of 3D and VR content. Through the custom tags and components of A-FRAME, scenes, entities, lighting, camera VR elements are defined, and real-time weather data is integrated.

6. The method for automatically generating VR scenes based on deep learning according to claim 1, characterized in that: The physiological information collection and analysis described in step 4 includes the following steps: S1: Data collection: Use wearable devices to collect users' physiological data in real time, including heart rate, skin conductivity, and breathing rate, record users' voice input through microphones, and collect text data from users' interactions with the system through text input; S2: Data preprocessing: Physiological signals are filtered and normalized; speech signals are denoised, features are extracted, and they are segmented into frames; text data are segmented, stop words are removed, and stems are extracted; S3: Feature extraction: Use time series analysis methods to extract features of physiological signals; use Mel frequency cepstral coefficients to extract fundamental frequency, pitch, and formant features of speech signals, and use word embedding models to convert text into vector representations; S4: Physiological signal emotion classification: Support vector machine is used to classify physiological signal features, recurrent neural network is used to classify speech features, and deep learning model is used to classify text features. By combining information from multiple modalities such as physiological signals, speech and text, a multimodal fusion algorithm is used to improve the accuracy of emotion recognition. The emotion classification results of the above parts are comprehensively analyzed using the weighted average method to obtain the final user emotional state.

7. The method for automatically generating VR scenes based on deep learning according to claim 1, characterized in that: The main components of the generative adversarial network described in step 5 include a generator and a discriminator. The goal of the generator is to generate realistic data, and the goal of the discriminator is to distinguish between generated data and real data. The two parts compete with each other through adversarial training; The generator includes an input layer, a hidden layer, a view embedding layer, a weather embedding layer, a feature fusion layer, a generator network and an output layer; Input layer: input image, sound, smell multimodal features F, view parameters V, real-time weather data W, and sentiment analysis results E; Hidden layer: includes multiple fully connected layers or convolutional layers for learning data distribution; View embedding layer: embeds the view parameter V into a high-dimensional space to combine with multimodal features, real-time weather data and sentiment analysis results; Weather embedding layer: embeds the real-time weather data W into a high-dimensional space to combine with multimodal features, view parameters and sentiment analysis results; Feature fusion layer: fuses multimodal features, sentiment analysis results, view parameters and real-time weather data embedding; Generator network: includes multiple fully connected layers or convolutional layers, which are used to generate VR scene data from the fused features; Output layer: generated VR scene data G (F, E, V, W); The discriminator includes an input layer, a hidden layer, an output layer and a discriminator network; Input layer: inputs real VR scene data Xreal and generated VR scene data G (F, E, V, W); Hidden layer: includes multiple fully connected layers or convolutional layers, which are used to extract features and determine whether the input is real data; Output layer: probability values ​​D(Xreal) and D(G(F, E, V, W)), indicating the probability that the input is real data; Discriminator network: including multiple fully connected layers or convolutional layers, used to determine whether the input VR scene data is real data.

8. The method for automatically generating VR scenes based on deep learning according to claim 7, characterized in that: The adversarial training process of the generative adversarial network is: S1: Initialization: Randomly initialize the parameters of the generator and discriminator; S2: Discriminator update: Sample a batch of real data from the real data distribution and generate a batch of fake data from the generator, and use these data to update the parameters of the discriminator to maximize its loss function; S3: Generator update: Update the parameters of the generator using only the fake data generated by the generator to minimize its loss function; S4: View parameter learning: During the training process, the view parameter V is optimized by gradient descent to learn the best view representation; S5: Weather parameter learning: The real-time weather data W is optimized via gradient descent to learn how to adjust the generated scenes according to different weather conditions; S6: Emotional parameter learning: The sentiment analysis result E is optimized via gradient descent to learn how to adjust the generated scenes according to different emotional states: S7: Alternating iteration: repeat steps S2, S3, S4, S5 and S6 until a predetermined number of training rounds is reached or the loss function converges; S8: Application scenario generation: After training is completed, the trained generator is used to generate VR scenarios based on new multimodal features, sentiment analysis results, viewing angle parameters, and real-time weather data; mathematically expressed as: S=G(F new ,E new ,V new W new ), Among them, F new is a new multimodal feature, E new is the new sentiment analysis result, V new is the new viewing angle parameter, W new is the new real-time weather data.

9. The method for automatically generating VR scenes based on deep learning according to claim 7, characterized in that: A data management system is provided in the scene generation step, and the data management system is used to store and manage scene data. The data management system is a relational database, and the database is used to store scene models, materials, textures, audio, video elements, as well as interactive elements, physical properties, and animation effect information in the scene.

Citation Information

Cited By

  • Aviation special situation simulation training scene automatic generation method based on generative adversarial network

    CN122368222A

  • Air force emergency situation training scene automatic generation method based on generative adversarial network

    CN122368222B