Intelligent interaction method and system for meta universe scene

By integrating multiple interactive information analysis technologies in the meta-universe virtual scene, the problems of reduced interaction experience and large differences in personalization in the existing technology are solved, and a more efficient and natural interactive experience between users and the virtual world is achieved.

CN120066271APending Publication Date: 2025-05-30GUANGZHOU KUANHENG INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510148803.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing intelligent interaction technology for meta-universe scenes has limitations in speech recognition and understanding, low accuracy of gesture and action recognition, and difficult to effectively implement emotional and contextual understanding, resulting in a reduced user interaction experience and large differences in personalization.

Method used

By obtaining voice interaction information, action interaction information and facial expression information in the metaverse virtual scene, speech recognition technology, deep learning model and convolutional neural network are used for information analysis, multimodal emotion analysis is performed to obtain the user's emotional intentions, and the interaction log is analyzed through the neural network model to provide interactive recommendation results, and the interaction process is optimized.

Benefits of technology

It improves the real-time and accuracy of meta-universe scene interaction, enhances the user experience, and realizes a more natural, efficient and immersive interaction between users and the virtual world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066271A_ABST
    Figure CN120066271A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent interaction method and system for a meta-universe scene, and relates to the technical field of meta-universe scene interaction. The method comprises the following steps: acquiring scene interaction information and performing information analysis to obtain a voice analysis result, an action analysis result and a face analysis result; performing multi-modal sentiment analysis based on an analysis result to obtain a sentiment intention of the user; interacting with the user according to the emotional intention to obtain first interaction feedback information; acquiring a scene interaction log and performing log analysis to obtain an interaction recommendation result; interacting with the user according to the interaction recommendation result to obtain second interaction feedback information; and comprehensively analyzing the interaction feedback information and updating and optimizing the interaction process. According to the method, from two perspectives of scene interaction information and scene interaction logs, emotional intentions of users are obtained by adopting multiple technologies such as deep learning and neural networks, so that the accuracy of results, the usability and selectivity of interaction and the experience effect of the users are improved to a great extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of metaverse scene interaction, and particularly relates to an intelligent interaction method and system for metaverse scenes. Background Art

[0002] The Metaverse is a virtual and digital world, usually composed of multiple virtual spaces, virtual reality (VR), augmented reality (AR), and blockchain technology, etc., aiming to provide an immersive digital experience. Users can interact, create, and communicate in these virtual spaces through digital avatars (such as virtual characters or avatars). Its core features include: Immersive experience: Through virtual reality or augmented reality devices, users can obtain an immersive experience and feel a more intuitive and interactive feeling than the traditional Internet; Social interaction: Users can interact with others in real time. Whether it is for work, entertainment, or socializing, the Metaverse provides a virtual shared space.

[0003] Intelligent interaction based on metaverse scenes refers to the way in which users interact with the virtual environment, other users, and systems through various intelligent technologies in the virtual world of the metaverse. The core purpose of intelligent interaction in metaverse scenes is to make the interaction between users and the virtual world more natural, efficient, and immersive. However, there are still many problems at present, such as the limitations of speech recognition and understanding leading to a reduced user interaction experience, the low accuracy of gesture and action recognition resulting in lag in recognition, difficulties in effectively understanding the emotions, context, and implicit intentions of users in terms of emotion and context understanding, and the integration of multi-modal interaction, etc. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent interaction method and system for metaverse scenes, which can be achieved through the following technical solutions: In the first aspect, an embodiment of the present application provides an intelligent interaction method for metaverse scenes, including the following steps: Obtain scene interaction information from a pre-established metaverse virtual scene; the scene interaction information includes voice interaction information, action interaction information, and facial expression information; Perform information analysis based on the scene interaction information, including: Analyze the voice interaction information using speech recognition technology to obtain a voice analysis result; Analyze the body postures and body movements in the action interaction information using a deep learning model to obtain an action analysis result; Analyze the facial expression information using a convolutional neural network to obtain a facial analysis result; Perform multimodal sentiment analysis based on the speech analysis result, the action analysis result, and the facial analysis result to obtain the user's emotional intention; Interact with the user according to the emotional intention and obtain the first interaction feedback information; Obtain a scene interaction log from a pre-established metaverse virtual scene; the scene interaction log includes interaction log records and multi-channel log records; Construct a neural network model and perform log analysis based on the scene interaction log to obtain an interaction recommendation result; Interact with the user according to the interaction recommendation result and obtain the second interaction feedback information; Comprehensively analyze the first interaction feedback information and the second interaction feedback information, and continuously update and optimize the interaction process based on the analysis results.

[0005] Preferably, establishing the metaverse virtual scene in advance includes: Establish an initial metaverse scene and generate a corresponding metaverse scene description, including: object geometry, object material properties, light source position, light source intensity, and camera position; For each pixel, randomly select multiple sampling points on the pixel to determine the sampling position of each pixel; Emit multiple rays from the camera position through multiple sampling points and track the paths of the multiple rays in the initial metaverse scene until the rays intersect with the light source or an object, and perform Monte Carlo path tracing to estimate the color of the sampling points; For each ray, use a ray-object intersection test with a ray tracing algorithm. If there is an intersection point, return the intersection point closest to the ray and the corresponding object information; otherwise, return a null value; the object information includes: object geometry and object material properties; For the determined intersection point of the ray and the object, calculate the radiance and the corresponding luminous flux of the intersection point; Judge the object property based on the calculated radiance and luminous flux, trace the ray corresponding to the ray according to the object property, and calculate the attenuation of the ray using an attenuation model; Recursively trace the ray according to the light propagation situation in the initial metaverse scene until the maximum recursion depth is reached or the ray no longer intersects with an object; Mix the colors of all sampling points to obtain the color value of the final pixel, forming the final metaverse virtual scene.

[0006] Preferably, the analysis of the voice interaction information using voice recognition technology includes: Perform noise reduction processing, echo cancellation, and voice enhancement on the voice interaction information to obtain preprocessed voice information; Extract features from the preprocessed speech information, specifically: Perform Fourier transform and Mel filter bank processing on the preprocessed speech information, and extract MFCC features representing speech features; Perform frequency domain transformation on the preprocessed speech information to generate a spectrogram, and further extract audio features from the spectrogram using a filter bank; Fuse the MFCC features and the audio features to form speech information features; Process the speech information features using a speech recognition model, specifically: Use an acoustic model to learn the relationship between the speech information features and speech units, and convert the speech information features into basic units of speech; Use a language model to statistically analyze the language rules and context information within the speech information features; Use a decoder to combine the acoustic model and the language model, and output a word sequence according to probability; Post-process the word sequence and output the speech recognition result.

[0007] Preferably, analyzing the body postures and body movements in the action interaction information using a deep learning model includes: Obtain action video frame images from the action interaction information; Extract features from the action video frame images, including: Skeleton extraction and key point detection: Use a deep learning model to extract the joint point data and skeleton data of the human body in the image, and identify the key body parts; the key body parts include the head, shoulders, elbows, hands, knees, and feet; Body posture features: Obtain the body posture features of dynamic changes between each joint point according to the joint point data and the skeleton data, including angles, speeds, and accelerations; Temporal feature extraction: Based on continuous body movements, use a convolutional neural network and a long short-term memory network to model the time series, and extract the action temporal features of the body movements; Perform dynamic analysis and static analysis based on the body posture features to obtain the posture change situation; Perform dynamic analysis based on the action temporal features to obtain the body movement trajectory; Use the posture change situation and the body movement trajectory as the action analysis result.

[0008] Preferably, analyzing the facial expression information using a convolutional neural network includes: Obtain facial expression images from the facial expression information; Use a face detection algorithm to locate the face region in the facial expression image; Extract local features of the face region through the convolutional layer of the convolutional neural network; Perform deeper feature extraction on the local features through multiple convolutional layers to obtain facial expression features; Reduce the dimension of the facial expression features through the pooling layer and retain the facial feature information; Input the facial expression features into the fully connected layer for feature integration and map the facial expression features to the class labels of facial expressions; Output the probability values of each facial expression category through the softmax layer and determine the facial expression category of the facial expression image; Use the facial expression category as the facial analysis result.

[0009] Preferably, the multi-modal emotion analysis based on the speech analysis result, the action analysis result and the facial analysis result includes: Fuse the speech analysis result, the action analysis result and the facial analysis result to form an analysis result set; Extract multi-modal feature data from the analysis result set, where the multi-modal feature data includes speech feature data, action posture feature data and facial expression feature data; Obtain the background information of the user and preprocess the background information to generate background feature data; Fuse the preprocessed multi-modal feature data and the background feature data to generate input data; Extract speech features, action posture features and facial expression features from the input data through the convolutional neural network; Arrange the speech features, the action posture features and the facial expression features respectively according to the time series order to obtain a speech feature matrix, an action posture feature matrix and a facial expression feature matrix; Construct a multi-layer semantic network structure according to the speech feature matrix, the action posture feature matrix and the facial expression feature matrix; Perform semantic mapping between different modal features through the cross-modal alignment formula to obtain a semantic feature matrix; Calculate the true emotional intention of the user according to the semantic feature matrix and the background feature data.

[0010] Preferably, the extracting multi-modal feature data from the analysis result set includes: Preprocess the analysis result set to obtain initial multi-modal feature data; Perform time series analysis on the initial multi-modal feature data to determine the synchronization and priority among the modal feature data; Determine the weight coefficients among the modal feature data according to the synchronization and priority among the modal feature data; Perform feature fusion on the modal feature data according to the weight coefficients among the modal feature data to obtain multi-modal feature data; Among them, performing time series analysis on the initial multi-modal feature data to determine the synchronization and priority among the modal signals specifically includes: Decompose the initial multi-modal feature data into state sequences used to represent the changes of each modal feature data at different time points; Calculate the joint probability distribution of each modal feature data according to all the state sequences, and the joint probability distribution is used to determine the transition probability between the state sequences and the conditional probability of each modal feature data appearing under each state sequence; Analyze the synchronization among the state sequences of the initial multi-modal feature data according to the joint probability distribution to determine the feature data combination; Determine the priority of each modal feature data according to the feature data combination and in combination with the transition probability between the state sequences and the conditional probability of the initial multi-modal feature data under each state sequence.

[0011] Preferably, the construction of the neural network model and the log analysis according to the scenario interaction log include: Train the neural network model according to the interaction log record and the multi-channel log record to obtain a trained interaction scenario judgment model and an interaction result recommendation model; Receive the voice information generated by the user in the current metaverse virtual scene, and perform speech recognition on the voice information to obtain a speech recognition result; Obtain multi-channel real-time information from the metaverse interaction device, and based on the speech recognition result and the multi-channel real-time information, use the interaction scenario judgment model to obtain an interaction scenario judgment result of the current interaction scenario; Based on the interaction scenario judgment result, judge whether the current interaction scenario in the current metaverse virtual scene is a human-computer interaction scenario. If it is a human-computer interaction scenario, input the speech recognition result and the multi-channel real-time information into the interaction result recommendation model to obtain several interaction recommendation results; otherwise, end the current interaction scenario.

[0012] Preferably, the interaction log record represents the interaction information generated by the user during scene interaction in the metaverse virtual scene within a preset time period; the multi-channel log record represents the log record generated by the user through the deployed multi-channel sensors in the metaverse virtual scene.

[0013] In a second aspect, an intelligent interaction system for a metaverse scenario provided by an embodiment of the present application includes: An information acquisition module: to acquire scenario interaction information from a pre-established virtual metaverse scenario; the scenario interaction information includes voice interaction information, action interaction information, and facial expression information; An information analysis module: to perform information analysis based on the scenario interaction information, including: Using speech recognition technology to analyze the voice interaction information to obtain a voice analysis result; Using a deep learning model to analyze the body postures and body movements in the action interaction information to obtain an action analysis result; Using a convolutional neural network to analyze the facial expression information to obtain a facial analysis result; Performing multi-modal sentiment analysis based on the voice analysis result, the action analysis result, and the facial analysis result to obtain the emotional intention of the user; An information feedback module: to interact with the user according to the emotional intention and obtain first interaction feedback information; A log acquisition module: to acquire scenario interaction logs from a pre-established virtual metaverse scenario; the scenario interaction logs include interaction log records and multi-channel log records; A log analysis module: to build a neural network model and perform log analysis based on the scenario interaction logs to obtain interaction recommendation results; A log feedback module: to interact with the user according to the interaction recommendation results and obtain second interaction feedback information; An intelligent interaction module: to comprehensively analyze the first interaction feedback information and the second interaction feedback information and continuously update and optimize the interaction process based on the analysis results.

[0014] The beneficial effects of the present invention are: (1) The present invention analyzes and interacts from two aspects. The first aspect is to obtain scene interaction information from a pre-established metaverse virtual scene, including voice interaction information, action interaction information, and facial expression information. By analyzing the above scene interaction information, different analysis results are obtained, and further comprehensive multi-modal analysis is performed on multiple analysis results to obtain the user's emotional intention. Then, the emotional intention is used to interact with the user to obtain the first interaction feedback information. The second aspect is to obtain scene interaction logs and analyze the interaction logs using a neural network model to obtain interaction recommendation results for the user to select, thereby strengthening the information interaction with the user. Finally, the first interaction feedback information and the second interaction feedback information obtained from the above two aspects are comprehensively analyzed to continuously optimize and update the interaction process, thereby improving the real-time performance and accuracy of metaverse scene interaction.

[0015] (2) The first aspect of the present invention starts from the perspective of scene interaction information and obtains the user's emotional intention by adopting a variety of different technologies, which can greatly improve the accuracy of the results; the second aspect starts from the perspective of scene interaction logs and improves the usability and selectivity of metaverse interaction through a neural network model. The present application improves the intelligence of metaverse scene interaction and the user experience effect by combining the above two aspects. Brief Description of the Drawings

[0016] For better understanding and implementation, the technical solutions of the present application will be described in detail below with reference to the accompanying drawings.

[0017] Figure 1 It is a flowchart of the steps of an intelligent interaction method for a metaverse scene provided by an embodiment of the present application; Figure 2 It is a flowchart of the steps of facial expression information analysis provided by an embodiment of the present application; Figure 3 It is a flowchart of the steps of log analysis provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of an intelligent interaction system for a metaverse scene provided by an embodiment of the present application. Detailed Embodiment

[0018] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the exemplary embodiments will be described in detail here, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of methods and systems consistent with some aspects of the present application as detailed in the appended claims.

[0019] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any and all possible combinations of one or more of the associated listed items.

[0020] The following describes in detail the specific implementation manners, features, and effects of the present invention in conjunction with the accompanying drawings and preferred embodiments.

[0021] Embodiment 1 Please refer to Figure 1 , this embodiment of the application provides an intelligent interaction method for a metaverse scenario, including the following steps: Obtain scene interaction information from a pre-established metaverse virtual scene; the scene interaction information includes voice interaction information, action interaction information, and facial expression information; Perform information analysis based on the scene interaction information, including: Use speech recognition technology to analyze the voice interaction information to obtain a voice analysis result; Use a deep learning model to analyze the body postures and body movements in the action interaction information to obtain an action analysis result; Use a convolutional neural network to analyze the facial expression information to obtain a facial analysis result; Perform multi-modal emotion analysis based on the voice analysis result, the action analysis result, and the facial analysis result to obtain the user's emotional intention; Interact with the user according to the emotional intention and obtain the first interaction feedback information; Obtain a scene interaction log from a pre-established metaverse virtual scene; the scene interaction log includes interaction log records and multi-channel log records; Build a neural network model and perform log analysis according to the scene interaction log to obtain an interaction recommendation result; Interact with the user according to the interaction recommendation result and obtain the second interaction feedback information; Comprehensively analyze the first interaction feedback information and the second interaction feedback information and continuously update and optimize the interaction process based on the analysis results.

[0022] Specifically, in the existing intelligent interaction based on the metaverse scenario, there are still various problems such as the limitations of speech recognition and understanding, which lead to a reduced user interaction experience, the low accuracy of gesture and motion recognition, resulting in lag in recognition, difficulty in effectively understanding the user's emotions, context, and implicit intentions in terms of emotion and context understanding, and the integration of multi-modal interaction. As a result, there are personalized differences in the intelligent interaction based on the metaverse, making it difficult to make the interaction between users and the virtual world more natural, efficient, and immersive. Therefore, to solve the above problems, this application analyzes and interacts from two aspects. The first aspect is to obtain scene interaction information from a pre-established metaverse virtual scene, including voice interaction information, motion interaction information, and facial expression information. By analyzing the above scene interaction information, different analysis results are obtained, and further comprehensive multi-modal analysis of multiple analysis results is carried out to obtain the user's emotional intention. Then, the emotional intention is used to interact with the user to obtain the first interaction feedback information. The second aspect is to obtain the scene interaction log and use a neural network model to analyze the interaction log to obtain interaction recommendation results for the user to select, thereby strengthening the information interaction with the user. Finally, the first interaction feedback information and the second interaction feedback information obtained from the above two aspects are comprehensively analyzed to continuously optimize and update the interaction process, thereby improving the real-time performance and accuracy of the metaverse scene interaction.

[0023] The first aspect mentioned above starts from the perspective of scene interaction information and obtains the user's emotional intention by adopting a variety of different technologies, which can greatly improve the accuracy of the results; the second aspect starts from the perspective of scene interaction logs and improves the usability and selectivity of the metaverse interaction through a neural network model. This application improves the intelligence of the metaverse scene interaction and the user experience effect by combining the above two aspects.

[0024] In an embodiment provided by this application, the metaverse virtual scene is pre-established, including: Establish an initial metaverse scene and generate a corresponding metaverse scene description, including: object geometry, object material properties, light source position, light source intensity, and camera position; For each pixel, randomly select multiple sampling points on the pixel to determine the sampling position of each pixel; Emit multiple rays from the camera position through multiple sampling points and track the paths of the multiple rays in the initial metaverse scene until the rays intersect with the light source or an object, and perform Monte Carlo path tracing to estimate the color of the sampling points; For each ray, perform a ray-object intersection test using a ray tracing algorithm. If an intersection point exists, return the intersection point closest to the ray and the corresponding object information; otherwise, return null. The object information includes: object geometry and object material properties; For the determined intersection point of the ray and the object, calculate the radiance and corresponding luminous flux of the intersection point; Based on the calculated radiance and luminous flux, determine the object properties, trace the ray corresponding to the ray according to the object properties, and use an attenuation model to calculate the attenuation of the ray; Recursively trace the ray according to the light propagation situation in the initial metaverse scene until the maximum recursion depth is reached or the ray no longer intersects with the object; Mix the colors of all sampling points to obtain the color value of the final pixel, forming the final metaverse virtual scene.

[0025] Specifically, in this embodiment, by integrating advanced graphics technology, natural language processing technology, and ray tracing technology, the generation of a realistic and rich metaverse scene and intelligent user interaction are realized. This embodiment has outstanding advantages in the generation and rendering of the metaverse scene. Through advanced graphics technologies such as Monte Carlo path tracing, a realistic and rich metaverse scene can be generated, including details such as the geometry of objects, material properties, and light source positions. These scenes can not only meet the user's visual needs but also provide a more real and vivid interaction environment for users. In addition, this embodiment also has important improvements in ray tracing and attenuation models. Through the ray tracing algorithm, the system can simulate the propagation and reflection process of light in the metaverse scene, realizing the real simulation and rendering of light; at the same time, by introducing an attenuation model, the system can consider the attenuation effect of light during propagation, making the generated metaverse virtual scene more realistic and true.

[0026] In an embodiment provided by the present application, the analysis of the voice interaction information using voice recognition technology includes: Perform noise reduction processing, echo cancellation, and voice enhancement on the voice interaction information to obtain preprocessed voice information; Extract features from the preprocessed voice information. Specifically: Perform Fourier transform and Mel filter bank processing on the preprocessed voice information to extract MFCC features representing voice features; Perform frequency-domain transformation on the preprocessed voice information to generate a spectrogram, and further extract audio features from the spectrogram using a filter bank; Fuse the MFCC features and the audio features to form voice information features; Process the voice information features using a voice recognition model. Specifically: An acoustic model is used to learn the relationship between the speech information features and speech units, and convert the speech information features into basic units of speech; A language model is used to statistically analyze the language rules and context information within the speech information features; A decoder is used to combine the acoustic model and the language model, and output a word sequence according to probabilities; Post-process the word sequence and output the speech recognition result.

[0027] Specifically, in the above-mentioned feature extraction stage, Fourier transform and Mel filter bank processing are performed on the preprocessed speech information to extract MFCC features representing speech features, which effectively represent the spectral features of the speech information. Frequency domain transformation is performed on the preprocessed speech information to generate a spectrogram, and a filter bank is used to further extract audio features. Commonly used filters include Mel frequency filters, which can simulate the auditory perception of the human ear. Other features include pitch, speech rate, loudness, timbre, etc., which help improve the accuracy of speech recognition in subsequent speech analysis. In the above-mentioned feature processing stage, an acoustic model is used to convert the audio signal into basic units of speech by learning the relationship between audio features and speech units (such as phonemes, syllables); commonly used acoustic models include methods based on deep neural networks (DNN), convolutional neural networks (CNN), and long short-term memory networks (LSTM), etc. A language model is also used to improve the accuracy of speech recognition by statistically analyzing language rules and context information. It predicts the probability of the next word or phrase based on the previous context. Commonly used models include N-gram models, RNN language models, and models based on Transformer. Finally, a decoder combines the acoustic model and the language model, and outputs the most likely word sequence according to probabilities to complete the final speech-to-text conversion. The above post-processing includes spelling correction and error correction, removing redundant or invalid information, etc.

[0028] In an embodiment provided by the present application, the use of a deep learning model to analyze the body postures and body movements in the action interaction information includes: Obtain action video frame images from the action interaction information; Perform feature extraction on the action video frame images, including: Skeleton extraction and key point detection: Use a deep learning model to extract the joint point data and skeleton data of the human body in the image, and identify the key body parts; the key body parts include the head, shoulders, elbows, hands, knees, feet; Body posture features: According to the joint point data and the skeleton data, obtain the body posture features of dynamic changes between each joint point, including angles, speeds, and accelerations; Temporal feature extraction: Based on continuous limb movements, a convolutional neural network and a long short-term memory network are used to model the time series, and the action temporal features of limb movements are extracted; Based on the limb posture features, dynamic analysis and static analysis are performed to obtain the posture change situation; Based on the action temporal features, dynamic analysis is performed to obtain the limb movement trajectory; The posture change situation and the limb movement trajectory are used as the action analysis results.

[0029] Specifically, in this embodiment, a static posture recognition model is used to analyze the posture of an individual at a certain moment, such as standing, sitting, squatting, etc. It also analyzes the characteristics of dynamic actions such as walking, running, jumping, waving, etc., including the start, duration, speed, frequency, movement trajectory, etc. of the actions. It also analyzes the relative positions and angular changes of each joint of the human body to identify subtle changes in complex actions, such as the swing of the arm and the movement trajectory of the leg. By analyzing limb postures and limb movements through a deep learning model, the static postures and dynamic actions of the human body can be accurately recognized, and real-time feedback and data support can be provided for obtaining the true intention of the user subsequently.

[0030] As Figure 2 shown, in an embodiment provided by the present application, the analysis of the facial expression information by using a convolutional neural network includes: Obtain a facial expression image from the facial expression information; Use a face detection algorithm to locate the facial area in the facial expression image; the purpose of face detection is to extract the facial area from the original image and reduce the influence of background noise.

[0031] Extract local features of the facial area through the convolutional layer of the convolutional neural network; Perform deeper feature extraction on the local features through multiple convolutional layers to obtain facial expression features, including detailed features of the face and expression changes; Reduce the dimension of the facial expression features through a pooling layer and retain the facial feature information; common pooling methods include max pooling and average pooling, which help to reduce the dimension and improve the calculation efficiency; Input the facial expression features into a fully connected layer for feature integration and map the facial expression features to the class labels of facial expressions; Output the probability value of each facial expression category through a softmax layer and determine the facial expression category of the facial expression image; Use the facial expression category as the facial analysis result.

[0032] Specifically, in this embodiment, a convolutional neural network is used to analyze facial expressions. It mainly goes through steps such as face detection, feature extraction, and expression classification, and finally obtains analysis results such as the category, intensity, and facial action units of the facial expression. These results can serve as the basis for subsequent analysis of the user's emotional intention. The convolutional neural network extracts the details of the facial image layer by layer from local features to global features through multiple layers of convolution. For example, the low-level convolutional layer may capture edges and textures, while the high-level convolutional layer learns more abstract facial expression patterns, such as smiling or frowning. This hierarchical feature extraction helps to accurately identify complex facial expressions. The convolution operation applies the same convolution kernel to each region in the image, which makes the convolutional neural network have a certain invariance to the position and scale changes of the objects in the image. Even if the facial expression has a position offset or scale change in the image, the convolutional neural network can still effectively identify the expression. The pooling layer (such as max pooling, average pooling) retains important information by dimensionality reduction, while reducing the computational complexity and the risk of overfitting, which enables the model to better generalize and adapt to different facial expression changes. Each filter (convolution kernel) in the convolutional layer slides over the entire image and operates with the same weights, which greatly reduces the number of parameters to be learned and the computational complexity of the model. Therefore, using a convolutional neural network for facial expression analysis makes the obtained analysis results more efficient and accurate.

[0033] In an embodiment provided by the present application, the multi-modal emotion analysis based on the speech analysis result, the action analysis result, and the facial analysis result includes: Fusing the speech analysis result, the action analysis result, and the facial analysis result to form an analysis result set; Extracting multi-modal feature data from the analysis result set, where the multi-modal feature data includes speech feature data, action posture feature data, and facial expression feature data; Obtaining the background information of the user and preprocessing the background information to generate background feature data; Fusing the preprocessed multi-modal feature data with the background feature data to generate input data; Extracting speech features, action posture features, and facial expression features from the input data through a convolutional neural network; Arranging the speech features, the action posture features, and the facial expression features respectively according to the time series order to obtain a speech feature matrix, an action posture feature matrix, and a facial expression feature matrix; Constructing a multi-layer semantic network structure based on the speech feature matrix, the action posture feature matrix, and the facial expression feature matrix; Performing semantic mapping between different modal features through a cross-modal alignment formula to obtain a semantic feature matrix; Calculate the user's true emotional intention based on the semantic feature matrix and the background feature data.

[0034] Specifically, in this embodiment, by generating multi-modal feature data, the interaction effect between the user and the metaverse virtual scene can be significantly improved. The multi-modal feature data mentioned includes speech feature data, action gesture feature data, facial expression feature data, etc. These signals, after being processed by the multi-modal fusion algorithm, can extract more-dimensional information from a single user input, improving the accurate understanding of the user's intention. Combining the extracted relevant background information can further enhance the system's situational awareness ability for the user. Especially in a complex metaverse virtual scene, it can respond to the user's operations in real time and dynamically adjust the interaction content and presentation method in the virtual scene, thereby providing a personalized interaction experience. It should be noted that the above multi-modal feature data is extracted from the speech analysis result, action analysis result, and facial analysis result, so it will also show the advantages of the above three analysis results in the result, thus comprehensively improving the interaction experience between the user and the metaverse virtual scene.

[0035] In an embodiment provided by the present application, extracting multi-modal feature data from the analysis result set includes: Preprocess the analysis result set to obtain initial multi-modal feature data; Perform time series analysis on the initial multi-modal feature data to determine the synchronization and priority between each modal feature data; Determine the weight coefficients between each modal feature data according to the synchronization and priority between each modal feature data; Perform feature fusion on each modal feature data according to the weight coefficients between each modal feature data to obtain multi-modal feature data; Among them, performing time series analysis on the initial multi-modal feature data to determine the synchronization and priority between each modal signal specifically includes: Decompose the initial multi-modal feature data into state sequences representing the changes of each modal feature data at different time points; Calculate the joint probability distribution of each modal feature data according to all the state sequences. The joint probability distribution is used to determine the transition probability between each state sequence and the conditional probability of each modal feature data appearing under each state sequence; Analyze the synchronization between the state sequences of the initial multi-modal feature data according to the joint probability distribution to determine the feature data combination; Determine the priority of each modal feature data according to the feature data combination and combining the transition probability between the state sequences and the conditional probability of the initial multi-modal feature data under each state sequence.

[0036] Specifically, through the above content, this embodiment can accurately identify the true intention of the user and make timely adjustments in the virtual environment, avoiding the problem of inaccurate interaction caused by a single input method commonly found in the prior art. In addition, by real-time monitoring of the user feedback information and making adaptive adjustments, the system can continuously optimize its response ability to ensure that the user's intention is accurately responded to, thus significantly improving the user's immersion and satisfaction in the metaverse. Therefore, it not only has higher semantic understanding ability, but also can provide more accurate and real-time interaction responses in a complex multi-modal input environment, greatly enhancing the practicality of the system and the user experience.

[0037] As Figure 3 shown, in an embodiment provided by the present application, the building of the neural network model and the log analysis based on the scenario interaction log include: Training the neural network model according to the interaction log record and the multi-channel log record to obtain a trained interaction scenario judgment model and an interaction result recommendation model; Receiving voice information generated by the user in the current metaverse virtual scene and performing speech recognition on the voice information to obtain a speech recognition result; Obtaining multi-channel real-time information from the metaverse interaction device, and based on the speech recognition result and the multi-channel real-time information, using the interaction scenario judgment model to obtain an interaction scenario judgment result of the current interaction scenario; Based on the interaction scenario judgment result, judging whether the current interaction scenario in the current metaverse virtual scene is a human-computer interaction scenario. If it is a human-computer interaction scenario, inputting the speech recognition result and the multi-channel real-time information into the interaction result recommendation model to obtain a number of interaction recommendation results; otherwise, ending the current interaction scenario.

[0038] Specifically, in this embodiment, by building a neural network model, training the neural network model to obtain a trained interaction scenario judgment model and an interaction result recommendation model, and using the interaction scenario judgment model and the interaction result recommendation model to combine the speech recognition result and the multi-channel real-time information to automatically identify and obtain interaction recommendation results, the accuracy and usability of the metaverse voice interaction are improved; providing a variety of interaction recommendation results for the user to choose from further enhances the user experience effect of the metaverse voice interaction.

[0039] In an embodiment provided by the present application, the interaction log record represents the interaction information generated by the user during scenario interaction in the metaverse virtual scene within a preset time period; the multi-channel log record represents the log record generated by the user through the deployed multi-channel sensors in the metaverse virtual scene.

[0040] Specifically, it should be noted that the above multi-channel sensors include, but are not limited to: microphones for collecting voice data, tracking sensors for collecting eye movement data, inertial sensors for collecting motion data, heart rate sensors for collecting heart rate data, temperature sensors for collecting body temperature data, etc.

[0041] Embodiment 2 Please refer to Figure 4 , this embodiment of the present application provides an intelligent interaction system for a metaverse scenario, applying an intelligent interaction method for a metaverse scenario as described above, including: Information acquisition module: acquiring scene interaction information from a pre-established metaverse virtual scene; the scene interaction information includes voice interaction information, motion interaction information, and facial expression information; Information analysis module: performing information analysis based on the scene interaction information, including: Analyzing the voice interaction information using speech recognition technology to obtain a voice analysis result; Analyzing the body postures and body movements in the motion interaction information using a deep learning model to obtain a motion analysis result; Analyzing the facial expression information using a convolutional neural network to obtain a facial analysis result; Performing multi-modal sentiment analysis based on the voice analysis result, the motion analysis result, and the facial analysis result to obtain the user's emotional intention; Information feedback module: used to interact with the user according to the emotional intention and obtain the first interaction feedback information; Log acquisition module: acquiring scene interaction logs from a pre-established metaverse virtual scene; the scene interaction logs include interaction log records and multi-channel log records; Log analysis module: used to build a neural network model and perform log analysis based on the scene interaction logs to obtain an interaction recommendation result; Log feedback module: used to interact with the user according to the interaction recommendation result and obtain the second interaction feedback information; Intelligent interaction module: used to comprehensively analyze the first interaction feedback information and the second interaction feedback information and continuously update and optimize the interaction process based on the analysis results.

[0042] In the above embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0043] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0044] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0045] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments of equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. An intelligent interaction method for a metaverse scene, characterized in that: The steps include: Obtain scene interaction information based on the pre-established metaverse virtual scene; The scene interaction information includes voice interaction information, action interaction information, and facial expression information; Performing information analysis based on the scene interaction information includes: Analyzing the voice interaction information using voice recognition technology to obtain a voice analysis result; Using a deep learning model to analyze the body postures and body movements in the action interaction information to obtain action analysis results; Using a convolutional neural network to analyze the facial expression information to obtain a facial analysis result; Performing multimodal emotion analysis based on the speech analysis result, the action analysis result and the facial analysis result to obtain the user's emotional intention; interacting with the user according to the emotional intention and obtaining first interaction feedback information; Obtaining a scene interaction log based on a pre-established metaverse virtual scene; the scene interaction log includes an interaction log record and a multi-channel log record; Constructing a neural network model and performing log analysis based on the scenario interaction log to obtain interactive recommendation results; Interacting with the user according to the interactive recommendation result and obtaining second interactive feedback information; Comprehensively analyze the first interaction feedback information and the second interaction feedback information, and continuously update and optimize the interaction process based on the analysis results.

2. The intelligent interaction method of a metaverse scene according to claim 1, characterized in that: The metaverse virtual scene is established in advance, including: Establish an initial metaverse scene and generate the corresponding metaverse scene description, including: object geometry, object material properties, light source position, light source intensity and camera position; For each pixel, multiple sampling points are randomly selected on the pixel to determine the sampling position of each pixel; Launching a plurality of rays from the camera position through a plurality of sampling points, and tracing the paths of the plurality of rays in the initial metaverse scene until the rays intersect with a light source or an object, and performing Monte Carlo path tracing to estimate the color of the sampling points; For each ray, use the ray tracing algorithm to perform a ray-object intersection test. If there is an intersection, return the intersection closest to the ray and the corresponding object information; otherwise, return a null value; the object information includes: object geometry and object material properties; For the determined intersection point between the ray and the object, the radiance of the intersection point and the corresponding luminous flux are calculated; Determine the properties of the object based on the calculated radiance and luminous flux, trace the light rays corresponding to the object properties, and calculate the attenuation of the light rays using the attenuation model; Recursively trace the ray according to the ray propagation in the initial metaverse scene until the maximum recursion depth is reached or the ray no longer intersects the object; The colors of all sampling points are mixed to obtain the color value of the final pixel, forming the final metaverse virtual scene.

3. The intelligent interaction method of a metaverse scene according to claim 1, characterized in that: The adopting speech recognition technology to analyze the speech interaction information includes: Performing noise reduction, echo cancellation and voice enhancement on the voice interaction information to obtain pre-processed voice information; The pre-processed speech information is subjected to feature extraction, specifically: Performing Fourier transform and Mel filter bank processing on the preprocessed speech information to extract MFCC features representing speech features; Performing frequency domain transformation on the preprocessed speech information to generate a spectrogram, and further extracting audio features from the spectrogram using a filter bank; Fusion of the MFCC features and the audio features to form a speech information feature; The speech information features are processed using a speech recognition model, specifically: Using an acoustic model to learn the relationship between the speech information features and speech units, and converting the speech information features into basic units of speech; Using a language model to count the language regularities and context information in the voice information features; Using a decoder to combine the acoustic model with the language model and output a word sequence according to probability; The word sequence is post-processed and a speech recognition result is output.

4. The intelligent interaction method of a metaverse scene according to claim 1, characterized in that: The adopting of a deep learning model to analyze the limb posture and limb movement in the action interaction information includes: Obtaining an action video frame image from the action interaction information; Extracting features from the action video frame image includes: Skeleton extraction and key point detection: Use deep learning models to extract joint point data and skeleton data of the human body in the image, and identify key limb parts; the key limb parts include the head, shoulders, elbows, hands, knees, and feet; Limb posture features: According to the joint point data and the skeleton data, the limb posture features of dynamic changes between each joint point are obtained, including angle, speed and acceleration; Temporal feature extraction: Based on continuous body movements, convolutional neural networks and long short-term memory networks are used to model the time series and extract the temporal features of body movements; Performing dynamic analysis and static analysis based on the limb posture characteristics to obtain posture change conditions; Perform dynamic analysis based on the action timing characteristics to obtain the limb movement trajectory; The posture change and the limb movement trajectory are used as the action analysis result.

5. The intelligent interaction method of a metaverse scene according to claim 1, characterized in that: The method of analyzing the facial expression information by using a convolutional neural network includes: Acquire a facial expression image from the facial expression information; Locating a facial region in the facial expression image using a facial detection algorithm; Extracting local features of the facial region through a convolutional layer of a convolutional neural network; Performing deeper feature extraction on the local features through multiple convolutional layers to obtain facial expression features; Reducing the dimension of the facial expression features through a pooling layer and retaining facial feature information; Inputting the facial expression features into a fully connected layer for feature integration and mapping the facial expression features to a category label of facial expression; Outputting the probability value of each facial expression category through a softmax layer, and determining the facial expression category of the facial expression image; The facial expression category is used as the facial analysis result.

6. The intelligent interaction method of a metaverse scene according to claim 1, characterized in that: The performing multimodal emotion analysis based on the speech analysis result, the action analysis result and the facial analysis result comprises: Merging the speech analysis result, the action analysis result and the facial analysis result to form an analysis result set; Extracting multimodal feature data from the analysis result set, the multimodal feature data comprising speech feature data, action posture feature data and facial expression feature data; Acquire user background information, and pre-process the background information to generate background feature data; The preprocessed multimodal feature data is fused with the background feature data to generate input data; Extracting speech features, action posture features and facial expression features from the input data through a convolutional neural network; Arranging the speech features, the action posture features and the facial expression features respectively according to the time series order to obtain a speech feature matrix, an action posture feature matrix and a facial expression feature matrix; Constructing a multi-layer semantic network structure according to the speech feature matrix, the action posture feature matrix and the facial expression feature matrix; The semantic mapping between different modal features is performed through the cross-modal alignment formula to obtain the semantic feature matrix; The user's real emotional intention is calculated based on the semantic feature matrix and the background feature data.

7. The intelligent interaction method of a metaverse scene according to claim 6, characterized in that: The extracting multimodal feature data from the analysis result set includes: Preprocessing the analysis result set to obtain initial multimodal feature data; Performing time series analysis on the initial multimodal feature data to determine synchronization and priority between the feature data of each modality; Determine the weight coefficients between the feature data of each modality according to the synchronization and priority between the feature data of each modality; According to the weight coefficients between the feature data of each modality, feature fusion is performed on the feature data of each modality to obtain multi-modal feature data; The time series analysis of the initial multimodal feature data is performed to determine the synchronization and priority between the modal signals, specifically including: Decomposing the initial multimodal feature data into a state sequence for representing changes of each modal feature data at different time points; Calculate the joint probability distribution of each modal feature data according to all state sequences, and the joint probability distribution is used to determine the transition probability between each state sequence and the conditional probability of each modal feature data appearing under each state sequence; Analyzing the synchronization between the state sequences of the initial multimodal feature data according to the joint probability distribution to determine a feature data combination; The priority of each modality feature data is determined according to the feature data combination and in combination with the transition probability between state sequences and the conditional probability of the initial multimodal feature data under each state sequence.

8. The intelligent interaction method of a metaverse scene according to claim 1, characterized in that: The constructing of the neural network model and performing log analysis according to the scene interaction log include: According to the interaction log records and the multi-channel log records, the neural network model is trained to obtain a trained interaction scene judgment model and an interaction result recommendation model; Receive voice information generated by the user in the current metaverse virtual scene, and perform voice recognition on the voice information to obtain a voice recognition result; Acquire multi-channel real-time information from the Metaverse interaction device, and obtain an interaction scene judgment result of the current interaction scene using the interaction scene judgment model based on the speech recognition result and the multi-channel real-time information; Based on the interaction scene judgment result, it is determined whether the current interaction scene in the current metaverse virtual scene is a human-computer interaction scene. If it is a human-computer interaction scene, the speech recognition result and the multi-channel real-time information are input into the interaction result recommendation model to obtain several interaction recommendation results; otherwise, the current interaction scene is terminated.

9. The intelligent interaction method of a metaverse scene according to claim 1, characterized in that: The interaction log record represents the interaction information generated when the user performs scene interaction in the metaverse virtual scene within a preset time period; the multi-channel log record represents the log record generated by the user through the multi-channel sensors deployed in the metaverse virtual scene.

10. An intelligent interaction system for a metaverse scene, applying the intelligent interaction method according to any one of claims 1 to 9, characterized in that: include Information acquisition module: obtain scene interaction information based on the pre-established metaverse virtual scene; The scene interaction information includes voice interaction information, action interaction information, and facial expression information; Information analysis module: performs information analysis based on the scene interaction information, including: Analyzing the voice interaction information using voice recognition technology to obtain a voice analysis result; Using a deep learning model to analyze the body postures and body movements in the action interaction information to obtain action analysis results; Using a convolutional neural network to analyze the facial expression information to obtain a facial analysis result; Performing multimodal emotion analysis based on the speech analysis result, the action analysis result and the facial analysis result to obtain the user's emotional intention; Information feedback module: used for interacting with the user according to the emotional intention and obtaining first interaction feedback information; Log acquisition module: acquires scene interaction logs based on the pre-established metaverse virtual scene; the scene interaction logs include interaction log records and multi-channel log records; Log analysis module: used to build a neural network model and perform log analysis based on the scenario interaction log to obtain interactive recommendation results; Log feedback module: used to interact with the user according to the interactive recommendation result and obtain second interactive feedback information; Intelligent interaction module: used to comprehensively analyze the first interaction feedback information and the second interaction feedback information and continuously update and optimize the interaction process based on the analysis results.

Citation Information

Cited By

  • E-commerce virtual scene generation method and system based on digital twinning and visual interaction

    CN120704539A