A method and implementation system for voice-driven virtual scene generation and switching

Through voice recognition and edge computing technology, the delay and operation inconvenience of VR devices in scene generation and switching are solved, user privacy is guaranteed, and the dizziness during exit is reduced through frame rate adjustment and scene fading technology, and the overall VR interactive experience is improved.

CN119512372BActive Publication Date: 2025-07-01NANJING INST OF MECHATRONIC TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411569111.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-07-01
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

In the scene generation and switching of existing VR devices, there are problems such as delayed response, inconvenient operation, insufficient privacy protection, and dizziness when exiting.

Method used

Quickly parse user instructions through voice recognition, use edge computing to reduce latency and protect user privacy, and optimize exit experience through frame rate smooth adjustment and scene fade out.

Benefits of technology

It realizes efficient and convenient interaction between users in metacosmic VR devices, improves operational convenience and experience immersion, ensures user privacy and security, and reduces dizziness when exiting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119512372B_ABST
    Figure CN119512372B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice-driven virtual scene generation and switching method and implementation system, which is applied to a metaverse VR virtual reality glasses device and includes: identifying a user's instruction through speech recognition and natural language processing technologies, generating or switching a virtual scene in real time, and realizing personalized scene creation and object generation. An edge computing node is used to process high-computation tasks, significantly reducing the latency of scene generation and switching, and improving the data transmission efficiency by dynamically selecting the optimal node, ensuring user privacy and security. The present invention optimizes the user's experience of exiting the virtual scene through frame rate smoothing adjustment and scene fade-out mechanism, reduces dizziness, and improves the fluency and comfort of interaction. The present invention has the characteristics of rapid response, convenient operation and strong privacy protection, and is widely applicable to virtual reality application fields such as games, education and medical treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of virtual reality (VR), and particularly to a method and implementation system for generating and switching virtual scenes driven by voice, which are applicable to VR devices and related metaverse environment interaction systems. Background Art

[0002] With the rapid development of virtual reality (VR) technology and speech recognition technology, VR devices have gradually been applied in fields such as gaming, education, and healthcare. However, existing VR applications usually rely on manual operations or traditional remote control devices to achieve interaction, making it difficult for users to operate conveniently while wearing VR devices. The limitation of this operation mode is that when the wearer performs operations such as scene switching and object generation, the overall experience may be affected due to inconvenient operation or response delay. In addition, the traditional manual operation method is difficult to achieve efficient interaction in the scenario of using VR head-mounted devices, especially in metaverse applications, where quickly generating and switching virtual scenes becomes particularly important.

[0003] In current technologies, some voice-controlled VR devices have emerged, but there are still problems such as high response delay and insufficient operation stability, mainly manifested as unsatisfactory speeds of scene generation and switching. In complex scenarios, the accuracy and parsing speed of speech recognition affect the coherence of the user experience to a certain extent. Especially when users need to generate specific virtual items or adjust their appearance parameters, the actual experience is often insufficient due to the insufficient computing power and response speed of existing technologies. At the same time, there are also obvious deficiencies in current technologies in terms of user privacy protection. For example, users' voice data and interaction information often need to be transmitted to remote servers for processing, which not only may pose a risk of privacy leakage, but also increases the response time of scene generation and switching due to network latency. In addition, users often experience a slight sense of dizziness when exiting virtual scenes, and existing VR devices have not effectively addressed this problem, resulting in a decline in the user experience.

[0004] Therefore, there is an urgent need in the market for a voice-driven VR system that can overcome the above deficiencies to achieve the efficiency and stability of scene generation and switching, and further ensure the privacy security of users and improve the smoothness of the exit experience. Summary of the Invention

[0005] In view of the problems of delay, inconvenient operation, insufficient privacy protection, and dizziness when exiting in the existing VR devices during scene generation and switching, the present invention provides a voice-driven virtual scene generation and switching method and an implementation system. The present invention quickly parses user instructions through voice recognition to achieve the generation, switching of virtual scenes, and creation of objects; uses edge computing to reduce latency and protect user privacy; optimizes the exit experience through frame rate smoothing adjustment and scene fade-out, thereby improving the fluency and comfort of VR interaction.

[0006] To achieve the above object, the present invention is implemented through the following technical solutions:

[0007] A voice-driven virtual scene generation and switching method, applied to a metaverse VR virtual reality glasses device, includes the following steps:

[0008] Step S1: The wearer wears the metaverse VR virtual reality glasses on the head, aligns the eyes directly in front of the glasses, and turns on the voice recognizer;

[0009] Step S2: Obtain the voice instructions of the wearer, parse the instructions through voice recognition technology, and identify the instructions for scene generation, item creation, or scene switching;

[0010] Step S3: According to the recognized voice instructions, use virtual reality technology to generate corresponding virtual scenes or virtual items, allowing users to build a personalized metaverse in the virtual world;

[0011] Step S4: Respond to the voice instructions of the user in real time, switch or update the virtual scene, or create new virtual items according to the user's needs;

[0012] Step S5: When the user inputs an exit instruction, the system performs an exit operation, or the user manually closes the virtual reality environment by pressing the button in the middle of the glasses;

[0013] Step S6: Use optimized computing processing technology to reduce the latency in scene generation and switching, ensure the privacy protection of user data, and at the same time reduce the slight dizziness feeling accompanied by the user during the exit process.

[0014] Preferably, in step S1, when the wearer wears the metaverse VR virtual reality glasses on the head, the device automatically detects the wearing position and strength through the built-in inertial measurement unit IMU and pressure sensor, and the sensor judges whether the wearing is correct according to the set threshold, and prompts the wearer to adjust the wearing state through voice or vibration feedback.

[0015] Preferably, in step S2, the voice recognition technology uses a voice recognition model based on a deep neural network DNN, and specifically includes the following steps:

[0016] Perform multi-layer neural network inference and calculation on the user's voice input through the voice recognition model to generate an output text; train the voice recognition model using the following formula:

[0017]

[0018] where P(y|x) represents the probability of outputting the corresponding text y given the input voice signal x; y t represents the output text or text segment corresponding to time step t; x t represents the input voice signal at time step t; is the product symbol, indicating the cumulative product of the probabilities P(y t |x t ,h t-1 ) for each time step t; h t-1 is the hidden state at the previous moment; T represents the number of time steps of the voice signal sequence; the model generates the output text based on the hidden state at the previous moment in the feedforward neural network for subsequent scene generation or switching operations.

[0019] Preferably, the parsing of the voice command includes the following steps:

[0020] Convert the user's voice input into text through an end-to-end voice-to-text ASR model;

[0021] Use a natural language processing NLP model to perform intent classification on the generated text, and determine the user's intent through the following formula:

[0022] C(y|x) = softmax(W * h + b)

[0023] where W is the classification weight matrix, h is the feature vector of the text input, b is the bias term, and the output generates the classification result through the softmax function to determine the scene generation, item creation, or scene switching operation corresponding to the text.

[0024] Preferably, in step S3, the virtual reality technology adopts a generative adversarial network GAN and a neural radiance field NeRF technology to generate a virtual scene, and the specific steps include:

[0025] Use a GAN composed of a generator G and a discriminator D to perform adversarial generation and optimization on the image elements in the virtual scene, where the generator receives a random noise input z and generates a scene, and the calculation formula is:

[0026] G(z) = f(W g * z + b g )

[0027] Among them, G(z) is the output image of the generator; f is the activation function, which is used to introduce non-linear factors to make the generated images more diverse and complex; z is the random noise input, which is used to generate new images; W g is the weight matrix in the generator, which controls the transformation of the input noise by the generator; b g is the bias vector of the generator, which is used to adjust the output result;

[0028] The discriminator judges the authenticity of the generated image through adversarial training. The expression is:

[0029] D(x) = σ(W d *x + b d )

[0030] Among them, D(x) is the output of the discriminator, which represents the authenticity score of the generated image; σ is the activation function; x is the generated image or real image, which is input to the discriminator for judging true or false; W d is the weight matrix of the discriminator, which determines the feature extraction method of the discriminator for the input image; b d is the bias vector of the discriminator, which is used to adjust the discriminant output;

[0031] Based on the NeRF technology, three-dimensional scene reconstruction and ray tracing are used to achieve high-precision three-dimensional model rendering. The calculation formula is:

[0032]

[0033] Among them, C(r) represents the color value along the ray r, which is the color information of the finally generated three-dimensional image; tn and tf respectively represent the starting and ending positions of the light ray, which determine the range of ray tracing; T(t) is the light ray transmission function; σ(t) is the volume density function, which is used to describe the density of the object at position t and determines the absorption degree of the light ray; c(t) is the color vector, which represents the color information of the object at position t; dt represents the tiny distance between each point and the next point;

[0034] The generation process is used to achieve realistic and adjustable scenes and items in a virtual reality environment.

[0035] Preferably, the size, shape, color, and position attributes of the generated corresponding virtual scene or virtual item are adjusted in real time through the specific parameters input by the user's voice. The adjustment process includes:

[0036] Using the generator G to dynamically generate items and modify the three-dimensional coordinates (x, y, z) of the items in real time. The expression for controlling the size adjustment is:

[0037] S = α * S0

[0038] Among them, S is the adjusted size of the item, S0 is the initial size, and α is the adjustment ratio parameter input by the user.

[0039] Preferably, in step S4, the switching or updating of the virtual scene is optimized by reinforcement learning technology. The system learns the user's scene switching history and preferences through the deep Q-network DQN algorithm. The specific steps include:

[0040] The state s during the scene switching process is the user's historical operation record, and the action a is the user's scene switching instruction;

[0041] Use the following Q function to calculate the recommended scene:

[0042] Q(s,a) = r + γm a a ′ xQ(s′,a′)

[0043] Among them, Q(s,a) is the Q value of taking action a in the current state s; r is the immediate reward, γ is the discount factor, and s′ and a′ respectively represent the next state and the possible actions in that state.

[0044] Preferably, in step S5, when the user exits by pressing the button in the middle of the glasses, the system pops up an exit confirmation interface. The specific steps include:

[0045] After pressing the button, the system detects the pressing signal and displays the exit confirmation interface;

[0046] The user confirms the exit by voice or pressing the button again. After confirmation, the system closes the virtual reality environment.

[0047] Preferably, in step S6, the use of optimized computing and processing technology reduces the latency in scene generation and switching. The specific steps include:

[0048] Parse the user's voice command on the local device to extract the scene generation or switching requirements;

[0049] Transmit the parsed instruction to the edge node, and the edge node executes high-computation tasks such as virtual scene generation, 3D model rendering, or item creation;

[0050] After the edge node completes the calculation, it transmits the generated scene or item data back to the local device to achieve real-time update and fast rendering;

[0051] The system dynamically selects the optimal edge node according to the network latency and the load of the edge node to optimize the data transmission speed and scene generation efficiency.

[0052] Preferably, the optimized computing and processing technology also includes alleviating the slight dizziness of the user when exiting the virtual reality scene through scene transition and frame rate adjustment, which specifically includes the following steps:

[0053] After receiving the user's exit command, the visual elements in the virtual reality scene are gradually reduced through the progressive scene fade-out technology to avoid abrupt scene switching;

[0054] The system uses frame rate adjustment to smoothly reduce the frame rate in the VR environment during the exit process;

[0055] The system dynamically adjusts the scene exit animation based on the user's head movement monitored in real time by the sensor, so that the picture changes during the scene exit process are synchronized with the user's head movement, reducing the dizziness that may occur during the exit process.

[0056] Preferably, the privacy protection of the user data comprises the following specific steps:

[0057] The local device performs preliminary training on the user behavior data, and the model parameters generated by the training are transmitted to the federated learning server for global model update;

[0058] The updated model parameters of the global model are transmitted back to the local device to ensure that the user's behavior data does not leave the local device.

[0059] A virtual scene generation and switching system based on voice drive, comprising:

[0060] The speech recognition module is used to obtain and analyze the voice commands input by the user and identify the needs of scene generation, object creation or scene switching;

[0061] An edge computing module is used to process the instructions parsed by the speech recognition module, perform high-computation tasks such as virtual scene generation, three-dimensional model rendering or object creation, and reduce delays in scene generation and switching;

[0062] Dynamic node selection module, which is used to monitor network latency and edge node load, dynamically select the best edge node, and further optimize data transmission speed and scene generation efficiency;

[0063] A frame rate adjustment module, which is used to smoothly reduce the frame rate of the virtual reality after receiving the user's exit instruction. The frame rate adjustment module combines the real-time data of the motion monitoring and synchronization module to reduce the slight dizziness that may accompany the user when exiting the virtual reality scene;

[0064] The scene fade-out module is used to gradually reduce the visual elements in the virtual reality scene through the progressive scene fade-out technology during the exit process to avoid the abrupt feeling of scene switching;

[0065] The motion monitoring and synchronization module includes sensors for real-time monitoring of the user's head movement and dynamically adjusting the exit animation of the scene according to the movement, so that the scene changes during the exit process are synchronized with the user's head movement, thereby reducing the sense of dizziness.

[0066] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the voice-driven virtual scene generation and switching method, the present invention realizes efficient and convenient interaction of users in the metaverse VR device. Users can directly enter the virtual scene by wearing VR devices and enabling the voice recognition function, greatly improving the operation convenience and experience immersion. By using deep neural network and natural language processing technologies to accurately parse the user's voice commands, it can quickly recognize the user's intention of scene generation, item creation or scene switching, realizing instant and flexible virtual content construction, enabling users to intuitively customize personalized virtual worlds through voice. The present invention has superiority in the real-time response to user voice commands, ensuring no delay in scene switching and updating, improving the fluency of interaction, and enhancing the immersion of users in the virtual world. To increase the operation safety, the system designs multiple exit methods, including voice and manual button pressing exit methods, enabling users to quickly and safely exit the virtual environment according to their needs, improving the autonomy and flexibility of the user experience. By optimizing the computing processing technology, the present invention not only reduces the delay in the process of scene generation and switching, but also effectively protects the privacy and security of user data. At the same time, through visual buffer processing, it reduces the slight sense of dizziness when users exit, thus significantly improving the overall virtual reality experience, making it more comfortable, safe and highly customizable. Brief Description of the Drawings

[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0068] Among them:

[0069] Figure 1 is the schematic flow chart of the method of the embodiment of the present invention;

[0070] Figure 2 is the schematic flow chart of the wearing detection and voice recognition startup process in the embodiment of the present invention;

[0071] Figure 3 is the schematic diagram of virtual scene generation and item creation in the embodiment of the present invention;

[0072] Figure 4 is the schematic flow chart of the scene switching and updating operation in the embodiment of the present invention;

[0073] Figure 5 This is the schematic diagram of the exit process in the embodiment of the present invention;

[0074] Figure 6 This is the schematic diagram of the overall structure in the embodiment of the present invention. Detailed implementation manners

[0075] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the protection scope of the present invention.

[0076] Embodiment 1

[0077] As Figure 1 、 Figure 2 shown, this is an embodiment of the present invention, which provides a method for generating and switching a voice-driven virtual scene, including:

[0078] Step S1: The wearer wears the metaverse VR virtual reality glasses on the head, aligns the eyes directly in front of the glasses, and turns on the voice recognizer;

[0079] Among them, in step S1, when the wearer wears the metaverse VR virtual reality glasses on the head, the device automatically detects the wearing position and strength through the built-in inertial measurement unit IMU and pressure sensor. The sensor judges whether the wearing is correct according to the set threshold, and prompts the wearer to adjust the wearing state through voice or vibration feedback.

[0080] In this embodiment, when the wearer wears the virtual reality (VR) glasses on the head, the inertial measurement unit (IMU) and pressure sensor in the glasses will automatically detect the wearing angle and pressure of the glasses. The inertial measurement unit is a sensor used to detect the movement state of the device, which can sense the head movement of the wearer and the spatial position of the device; the pressure sensor monitors the pressure of the device on the skin to ensure wearing comfort. The system judges whether the wearing state is correct according to the set wearing parameter standards, such as wearing position, pressure, etc. If it detects that the wearing does not meet the requirements, it will prompt the user to adjust through voice or vibration feedback. When the wearing state meets the standard, the system will automatically turn on the voice recognizer and enter the standby state. This design realizes automatic wearing detection, helps to optimize the user experience, and enables the user to turn on the voice-driven function without additional operations.

[0081] Further, as Figure 3As shown, step S2: Obtain the voice command of the wearer, parse the command through voice recognition technology, and recognize commands for scene generation, item creation, or scene switching;

[0082] Specifically, in step S2, the voice recognition technology uses a voice recognition model based on a deep neural network DNN, which specifically includes the following steps:

[0083] Perform multi-layer neural network inference and calculation on the user's voice input through the voice recognition model to generate an output text; The voice recognition model is trained using the following formula:

[0084]

[0085] Among them, P(y|x) represents the probability of outputting the corresponding text y after the given input voice signal x; y t represents the output text or text segment corresponding to time step t; x t represents the input voice signal at time step t; is the product symbol, indicating the cumulative multiplication of the probability P(y t |x t ,h t-1 ) for each time step t; h t-1 is the hidden state of the previous moment; T represents the number of time steps of the voice signal sequence; The model generates an output text based on the hidden state of the previous moment in the feedforward neural network for subsequent scene generation or switching operations.

[0086] Specifically, the parsing of the voice command includes the following steps:

[0087] Convert the user's voice input into text through an end-to-end speech-to-text ASR model;

[0088] Use a natural language processing NLP model to perform intent classification on the generated text, and determine the user's intent through the following formula:

[0089] C(y|x) = softmax(W*h + b)

[0090] Among them, W is the classification weight matrix, h is the feature vector of the text input, b is the bias term, and the output generates a classification result through the softmax function to determine the scene generation, item creation, or scene switching operation corresponding to the text.

[0091] After the wearer puts on the device, the system starts to listen for the user's voice commands. The system uses a deep neural network (DNN) as the speech recognition model and processes the speech signal through multiple layers of neural networks. The deep neural network is a machine learning technique that can handle complex data patterns. In this example, it converts the user's voice commands into text and then classifies the generated text through a natural language processing (NLP) model. The natural language processing model is an artificial intelligence technique specifically used to recognize and understand human language. Here, the NLP model will identify the intent of the command text and classify the command into operation requirements such as "scene generation", "item creation", or "scene switching". This dual parsing method improves the accuracy of speech recognition, effectively reduces the risk of misrecognition, and ensures the accurate execution of the user's commands.

[0092] Furthermore, as Figure 4 shown, step S3: According to the recognized voice command, use virtual reality technology to generate the corresponding virtual scene or virtual item, allowing the user to build a personalized metaverse in the virtual world;

[0093] Among them, in step S3, the virtual reality technology uses the generative adversarial network (GAN) and the neural radiance field (NeRF) technology to generate virtual scenes. The specific steps include:

[0094] Use the GAN composed of the generator G and the discriminator D to perform adversarial generation and optimization on the image elements in the virtual scene. Among them, the generator receives the random noise input z and generates a scene. The calculation formula is:

[0095] G(z) = f(W g *z + b g )

[0096] Among them, G(z) is the output image of the generator; f is the activation function, which is used to introduce non-linear factors to make the generated image more diverse and complex; z is the random noise input, which is used to generate new images; W g is the weight matrix in the generator, which controls the conversion of the input noise by the generator; b g is the bias vector of the generator, which is used to adjust the output result;

[0097] The discriminator judges the authenticity of the generated image through adversarial training. The expression is:

[0098] D(x) = σ(W d *x + b d )

[0099] Among them, D(x) is the output of the discriminator, which represents the authenticity score of the generated image; σ is the activation function; x is the generated image or real image, which is input to the discriminator for judging true or false; W dis the weight matrix of the discriminator, which determines the way the discriminator extracts features from the input image; b d is the bias vector of the discriminator, which is used to adjust the discriminant output;

[0100] Based on NeRF technology, high-precision 3D model rendering is achieved through 3D scene reconstruction and ray tracing. The calculation formula is:

[0101]

[0102] Among them, C(r) represents the color value along ray r, which is the color information of the finally generated 3D image; tn and tf respectively represent the starting and ending positions of the ray, which determine the range of ray tracing; T(t) is the ray transfer function; σ(t) is the volume density function, which is used to describe the density of the object at position t and determines the absorption degree of the ray; c(t) is the color vector, which represents the color information of the object at position t; dt represents the tiny distance between each point and the next point;

[0103] The generation process is used to realize realistic and adjustable scenes and items in a virtual reality environment.

[0104] Specifically, the size, shape, color, and position attributes of the generated virtual scene or virtual item are adjusted in real time through the specific parameters input by the user's voice. The adjustment process includes:

[0105] The generator G is used to dynamically generate items and modify the 3D coordinates (x, y, z) of the items in real time. The expression for controlling the size adjustment is:

[0106] S = α * S0

[0107] Among them, S is the adjusted size of the item, S0 is the initial size, and α is the adjustment ratio parameter input by the user.

[0108] After identifying the specific instructions, the system generates corresponding virtual content according to the user's needs. This process uses the generative adversarial network GAN and the neural radiance field NeRF technology. The generative adversarial network GAN is a machine learning method that includes two sub-networks: a generator and a discriminator. The generator is responsible for creating realistic images, while the discriminator judges the authenticity of the images. GAN is used in the present invention to generate the details and textures of virtual items, making the items present a more realistic appearance. The neural radiance field NeRF technology renders the 3D structure and depth details in the virtual scene through ray tracing, making the scene more realistic and having a sense of depth. Through the combination of GAN and NeRF technology, the system can generate highly realistic virtual scenes and items, providing users with a personalized construction experience in the virtual world.

[0109] Further, step S4: Respond to the user's voice commands in real time to switch or update the virtual scene, or create new virtual items according to the user's needs;

[0110] Among them, in step S4, the switching or updating of the virtual scene is optimized through reinforcement learning technology. The system learns the user's scene switching history and preferences through the Deep Q-Network (DQN) algorithm. The specific steps are as follows:

[0111] The state s during the scene switching process is the user's historical operation record, and the action a is the user's scene switching instruction;

[0112] Use the following Q function to calculate the recommended scene:

[0113] Q(s,a) = r + γ max Q(s′,a′) a a ′ xQ(s′,a′)

[0114] Among them, Q(s,a) is the Q value of taking action a in the current state s; r is the immediate reward, γ is the discount factor, and s′ and a′ represent the next state and the possible actions in that state, respectively.

[0115] After the virtual scene is generated, the system will continue to listen for the user's voice commands and respond in real time. Real-time performance is the key feature of this step: The system uses edge computing technology to complete most of the computing tasks on computing devices near the user (such as VR devices or edge servers), thereby reducing the latency of data transmission. Edge computing is a distributed computing method that can complete complex calculations on devices close to the data source, reducing the latency caused by remote server processing. Therefore, every time the user issues a switching or updating instruction, the system can quickly respond and immediately display the result in the virtual scene, making operations such as scene switching and item updating more fluent. This design significantly enhances the interactivity of virtual reality, enabling users to obtain a seamless immersive experience.

[0116] Further, as Figure 5 shown, step S5: When the user inputs an exit instruction, the system performs an exit operation, or the user manually closes the virtual reality environment by pressing the button in the middle of the glasses;

[0117] Among them, in step S5, when the user exits by pressing the button in the middle of the glasses, the system pops up an exit confirmation interface. The specific steps are as follows:

[0118] After pressing the button, the system detects the pressing signal and displays the exit confirmation interface;

[0119] The user confirms the exit by voice or by pressing the button again. After confirmation, the system closes the virtual reality environment.

[0120] In order to facilitate users to quickly exit the virtual reality environment when needed, the system provides two ways to exit: voice exit and physical button exit. Users can directly input the "exit" command through voice, and the system will immediately respond and end the current virtual scene; at the same time, users can also exit manually by pressing the physical button in the middle of the glasses. After pressing the button, the system will detect the pressing signal and pop up the exit confirmation interface. Users can choose to confirm by voice or press the button again to confirm the exit. This design ensures that users can quickly and conveniently exit the virtual reality environment in an emergency, providing users with better operational safety and convenience.

[0121] Further, step S6: using optimized computing processing technology to reduce delays in scene generation and switching, and ensure privacy protection of user data, while reducing the slight dizziness that accompanies the user during the exit process.

[0122] Among them, in step S6, the optimized computing processing technology is used to reduce the delay in scene generation and switching, and the specific steps include:

[0123] Parse the user's voice commands on the local device to extract scene generation or switching requirements;

[0124] Transmitting the parsed instructions to edge nodes, which perform high-computation tasks such as virtual scene generation, three-dimensional model rendering, or object creation;

[0125] After the edge node completes the calculation, it transmits the generated scene or object data back to the local device to achieve real-time update and fast presentation;

[0126] The system dynamically selects the best edge node based on network latency and edge node load, optimizing data transmission speed and scene generation efficiency.

[0127] Specifically, the optimized computing processing technology also includes reducing the slight dizziness of the user when exiting the virtual reality scene through scene transition and frame rate adjustment, which specifically includes the following steps:

[0128] After receiving the user's exit command, the visual elements in the virtual reality scene are gradually reduced through the progressive scene fade-out technology to avoid abrupt scene switching;

[0129] The system uses frame rate adjustment to smoothly reduce the frame rate in the VR environment during the exit process;

[0130] The system dynamically adjusts the scene exit animation based on the user's head movement monitored in real time by the sensor, so that the picture changes during the scene exit process are synchronized with the user's head movement, reducing the dizziness that may occur during the exit process.

[0131] Specifically, the privacy protection of user data includes the following specific steps:

[0132] The local device performs preliminary training on the user behavior data, and the model parameters generated by the training are transmitted to the federated learning server for global model update;

[0133] The model parameters after the global model update are sent back to the local device to ensure that the user's behavior data does not leave the local device.

[0134] During the process of scene generation and switching, the system combines the optimization strategies of edge computing and distributed computing to improve the processing speed, thus achieving instant response to user instructions. To ensure the privacy and security of user data, the system adopts federated learning technology. Federated learning is a distributed data processing method that allows user data to be computed on local devices and then upload the results to the central server without uploading the original data, thus avoiding privacy leakage. During the exit process, to reduce the possible dizziness of users, the system adopts the progressive scene fade-out technology, that is, the visual content of the virtual scene will gradually weaken, combined with the frame rate adjustment and the synchronous control of the head motion sensor, to keep the visual changes and the body perception of the user consistent, thus effectively alleviating the discomfort caused by scene switching.

[0135] As Figure 6 shown, a virtual scene generation and switching system based on voice driving includes:

[0136] A voice recognition module, used to obtain and parse the voice instructions input by the user, and recognize the requirements for scene generation, item creation or scene switching;

[0137] An edge computing module, used to process the instructions parsed by the voice recognition module, execute high-computation tasks such as virtual scene generation, 3D model rendering or item creation, and reduce the latency during scene generation and switching;

[0138] A dynamic node selection module, used to monitor network latency and edge node load, and dynamically select the optimal edge node to further optimize the data transmission speed and scene generation efficiency;

[0139] A frame rate adjustment module, used to smoothly reduce the frame rate of virtual reality after receiving the user's exit instruction. The frame rate adjustment module combines the real-time data of the motion monitoring and synchronization module to reduce the slight dizziness that the user may experience when exiting the virtual reality scene;

[0140] A scene fade-out module, used to gradually reduce the visual elements in the virtual reality scene through the progressive scene fade-out technology during the exit process, avoiding the abruptness of scene switching;

[0141] The motion monitoring and synchronization module includes sensors that are used to monitor the user's head motion in real time and dynamically adjust the exit animation of the scene according to this motion, so that the scene changes during the exit process are synchronized with the user's head motion, thereby reducing the sense of dizziness.

[0142] Embodiment 2

[0143] The present invention is applicable to multiple virtual reality application fields, including game, education, and medical scenarios. The system realizes real-time response to user instructions through voice recognition technology and edge computing to efficiently generate and switch virtual scenes, and has advantages in data security and privacy protection.

[0144] In game applications, users perform immersive game operations through VR head-mounted devices. The system can generate or switch scenes according to voice instructions, such as "generate a battlefield" or "switch to a forest", and quickly render the required virtual environment through edge computing nodes. At the same time, users can generate specific virtual objects through voice (such as "generate a shield" or "create an obstacle") and dynamically adjust their properties such as size and position, improving the smoothness of the game interaction experience and the smoothness of the scene transition, thereby effectively reducing the sense of dizziness during scene switching.

[0145] In education applications, teachers or students can enter the virtual teaching environment through voice instructions and perform operations such as "generate a chemistry laboratory" or "display the layered structure of the earth". The system quickly generates the required scene according to the voice command, enhancing the interactivity and immersion of teaching. Teachers can further adjust the properties of virtual objects (such as atomic models, celestial bodies) through voice, enabling students to more intuitively observe and understand complex concepts. The edge computing function of this system ensures local processing of interaction data, thereby protecting the privacy of students and meeting the security requirements in the teaching environment.

[0146] In medical applications, the system provides convenience for medical training and rehabilitation therapy. Users generate specific scenes through voice control, such as "generate an operating room" or "switch to an emergency simulation environment", facilitating operation training for doctors or trainees. At the same time, users can generate human organ models or medical tools through voice and adjust the position and size for detailed observation and interaction. In addition, during rehabilitation training, the system generates virtual objects such as stairs and grip trainers according to the patient's instructions for cognitive and motor training, and locally processes and stores the patient's operation data through a privacy protection mechanism.

[0147] In summary, through the voice-driven virtual scene generation and switching method and implementation system of the present invention, the problems of large response delay, inconvenient operation, and weak privacy protection in the existing VR devices during scene switching and generation are successfully solved. Through advanced voice recognition and natural language processing technologies, the present invention can accurately parse the user's scene generation and switching instructions, and uses edge computing to accelerate task processing, reduce latency, and ensure the local privacy and security of user data. At the same time, the present invention adopts frame rate smoothing adjustment and progressive scene fade-out technologies to ensure smooth switching of the screen when the user exits the virtual scene, greatly reducing the sense of dizziness and improving the comfort during the exit process. Generally speaking, the present invention provides a highly intelligent VR interaction solution, which not only improves the fluency of the user experience, but also provides safer and more efficient technical support for the metaverse and virtual reality applications, and has broad application and promotion prospects.

[0148] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0149] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved.

[0150] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A voice-driven virtual scene generation and switching method, applied to a metaverse VR virtual reality glasses device, characterized in that: The following steps are involved: Step S1: The wearer puts the Metaverse VR glasses on his head, faces the front of the glasses with his eyes, and turns on the voice recognizer; Step S2: Acquire the wearer's voice command, parse the command through voice recognition technology, and identify the command of scene generation, object creation or scene switching; Step S3: Generate corresponding virtual scenes or virtual objects using virtual reality technology according to the recognized voice command, allowing users to build a personalized metaverse in the virtual world, wherein the generation process includes constructing target image information based on a generative adversarial network (GAN), and performing three-dimensional reconstruction and rendering through neural radiation field (NeRF) technology to restore image realism and spatial structure, and enhance the flexibility and immersion of virtual scene generation; Step S4: Respond to the user's voice command in real time, switch or update the virtual scene, or create a new virtual item according to the user's needs. The switching or updating process is based on the user's past voice operation records, and the deep Q network DQN algorithm is used to learn and evaluate the state value of the candidate scene to achieve personalized recommendation and optimal scene prediction; Step S5: When the user inputs an exit command, the system executes the exit operation, or the user manually closes the virtual reality environment by pressing a button in the middle of the glasses; Step S6: Use optimized computing processing technology to reduce delays in scene generation and switching, including dynamically selecting edge computing nodes to perform image generation and rendering tasks based on network status and task load, and transmitting the rendering results back to the glasses terminal for display after the calculation is completed, ensuring the privacy of user data, and reducing the slight dizziness that accompanies the user during the exit process.

2. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: In step S1, when the wearer wears the Metaverse VR glasses on his head, the device automatically detects the wearing position and strength through the built-in inertial measurement unit IMU and pressure sensor. The sensor determines whether the wearing is correct based on the set threshold, and prompts the wearer to adjust the wearing status through voice or vibration feedback.

3. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: In step S2, the speech recognition technology adopts a speech recognition model based on a deep neural network DNN, which specifically includes the following steps: Perform multi-layer neural network reasoning and calculation on the user's voice input through the voice recognition model to generate output text; The speech recognition model is trained using the following formula: Where P(y|x) represents the probability of outputting the corresponding text y given the input speech signal x; y t represents the output text or text fragment corresponding to time step t; x t represents the input speech signal at time step t; is a multiplication symbol, indicating the probability P(y t |x t ,h t-1 ) to perform cumulative multiplication; h t-1 is the hidden state at the previous moment; T represents the time step number of the speech signal sequence; the model generates output text according to the hidden state at the previous moment in the feedforward neural network for subsequent scene generation or switching operations.

4. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: The analysis of the voice command comprises the following steps: Convert user voice input into text through an end-to-end speech-to-text ASR model; The generated text is classified using the natural language processing (NLP) model, and the user intent is determined using the following formula: C(y|x)=softmax(W*h+b) Wherein, W is the classification weight matrix, h is the feature vector of the text input, b is the bias term, and the output generates the classification result through the softmax function to determine the scene generation, object creation or scene switching operation corresponding to the text.

5. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: In step S3, the virtual reality technology uses the generative adversarial network GAN and neural radiation field NeRF technology to generate a virtual scene. The specific steps include: Using the GAN composed of the generator G and the discriminator D, the image elements in the virtual scene are adversarially generated and optimized, where the generator receives the random noise input z and generates the scene. The calculation formula is: G(z)=f(W g *z+b g ) Where G(z) is the output image of the generator; f is the activation function, which is used to introduce nonlinear factors to make the generated image more diverse and complex; z is the random noise input used to generate new images; W g is the weight matrix in the generator, which controls the generator's transformation of input noise; b g is the bias vector of the generator, used to adjust the output result; The discriminator judges the authenticity of the generated image through adversarial training, and the expression is: D(x)=σ(W d *x+b d ) Where D(x) is the output of the discriminator, which represents the authenticity score of the generated image; σ is the activation function; x is the generated image or the real image, which is input to the discriminator for distinguishing true from false; W d is the weight matrix of the discriminator, which determines the way the discriminator extracts features of the input image; b d is the bias vector of the discriminator, used to adjust the discriminant output; Three-dimensional scene reconstruction and ray tracing based on NeRF technology achieve high-precision three-dimensional model rendering. The calculation formula is: Among them, C(r) represents the color value along ray r, which is the color information of the final generated three-dimensional image; tn and tf represent the starting and ending positions of the light, respectively, and determine the range of ray tracing; T(t) is the light transmission function; σ(t) is the volume density function, which is used to describe the density of the object at position t and determines the degree of light absorption; c(t) is the color vector, which represents the color information of the object at position t; dt represents the small distance between each point and the next point; The generation process is used to achieve realistic and adjustable scenes and items in a virtual reality environment.

6. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: The size, shape, color and position attributes of the generated corresponding virtual scene or virtual object are adjusted in real time according to the specific parameters input by the user's voice. The adjustment process includes: dynamically generating objects using the generator G and modifying the three-dimensional coordinates (x, y, z) of the objects in real time. The expression for controlling the size adjustment is: S=α*S0 Among them, S is the adjusted item size, S0 is the initial size, and α is the adjustment ratio parameter input by the user.

7. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: In step S4, the switching or updating of the virtual scene is optimized by reinforcement learning technology, and the scene switching history and preferences of the user are learned by the deep Q network DQN algorithm. The specific steps include: The state s in the scene switching process is the user's historical operation record, and the action a is the user's scene switching instruction; The recommended scenario is calculated using the following Q function: Among them, Q(s,a) is the Q value of taking action a in the current state s; r is the immediate reward, γ is the discount factor, s′ and a′ represent the next state and the possible action in this state respectively.

8. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: In step S5, when the user exits by pressing the button in the middle of the glasses, the system pops up an exit confirmation interface, and the specific steps include: After pressing the button, the system detects the pressing signal and displays the exit confirmation interface; The user confirms the exit by voice or by pressing the button again, and after confirmation, the system closes the virtual reality environment.

9. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: In step S6, the optimized computing and processing technology is used to reduce the delay in scene generation and switching, and the specific steps include: Parse the user's voice commands on the local device to extract scene generation or switching requirements; Transmitting the parsed instructions to edge nodes, which perform high-computation tasks such as virtual scene generation, three-dimensional model rendering, or object creation; After the edge node completes the calculation, it transmits the generated scene or object data back to the local device to achieve real-time update and fast presentation; The system dynamically selects the best edge node based on network latency and edge node load, optimizing data transmission speed and scene generation efficiency.

10. The method for generating and switching a virtual scene driven by voice according to claim 9, characterized in that: The optimized computing and processing technology also includes alleviating the slight dizziness of the user when exiting the virtual reality scene through scene transition and frame rate adjustment, which specifically includes the following steps: After receiving the user's exit command, the visual elements in the virtual reality scene are gradually reduced through the progressive scene fade-out technology to avoid abrupt scene switching; The system uses frame rate adjustment to smoothly reduce the frame rate in the VR environment during the exit process; The system dynamically adjusts the scene exit animation based on the user's head movement monitored in real time by the sensor, so that the picture changes during the scene exit process are synchronized with the user's head movement, reducing the dizziness that may occur during the exit process.

11. The method for generating and switching a virtual scene driven by voice according to claim 1, characterized in that: The privacy protection of the user data includes the following specific steps: The local device performs preliminary training on the user behavior data, and the model parameters generated by the training are transmitted to the federated learning server for global model update; the model parameters after the global model update are transmitted back to the local device to ensure that the user's behavior data does not leave the local device.

12. A virtual scene generation and switching system based on voice drive as claimed in claim 1, characterized in that: include: The speech recognition module is used to obtain and analyze the voice commands input by the user and identify the needs of scene generation, object creation or scene switching; An edge computing module is used to process the instructions parsed by the speech recognition module, perform high-computation tasks such as virtual scene generation, three-dimensional model rendering or object creation, and reduce delays in scene generation and switching; Dynamic node selection module, which is used to monitor network latency and edge node load, dynamically select the best edge node, and further optimize data transmission speed and scene generation efficiency; A frame rate adjustment module, which is used to smoothly reduce the frame rate of the virtual reality after receiving the user's exit instruction. The frame rate adjustment module combines the real-time data of the motion monitoring and synchronization module to reduce the slight dizziness that may accompany the user when exiting the virtual reality scene; The scene fade-out module is used to gradually reduce the visual elements in the virtual reality scene through the progressive scene fade-out technology during the exit process to avoid the abrupt feeling of scene switching; The motion monitoring and synchronization module includes a sensor for monitoring the user's head movement in real time and dynamically adjusting the exit animation of the scene according to the movement, so that the scene changes during the exit process are synchronized with the user's head movement, thereby reducing the feeling of dizziness.

Citation Information

Patent Citations

  • Virtual reality glasses

    CN105487230A

  • Virtual scene generation method and device, electronic equipment and storage medium

    CN117745987A

  • Calculation method and system for meta universe

    CN118426965A