Accelerator simulation training space building method based on multi-modal fusion and ray tracing

Through interactive panoramic ray tracing and multimodal fusion model, the problem of inconsistency in virtual space construction is solved, and the authenticity and response accuracy of virtual space are improved.

CN120070708APending Publication Date: 2025-05-30XIDIAN UNIV

Patent Information

Application Number
CN202510111406.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing virtual space construction method is difficult to capture the lighting, geometric and material information of the real environment, resulting in inconsistent final images of the fusion of virtual and real, and insufficient authenticity and response accuracy of virtual space.

Method used

The interactive panoramic ray tracing method combined with the multimodal fusion model is used to obtain virtual space construction resources, build 3D virtual space, and estimate irradiance through ray tracing and interpolation algorithms, consider the environmental information of the real space, and improve the authenticity of the virtual space. At the same time, the user intention is identified based on the multimodal fusion model to improve the response accuracy of the virtual space.

Benefits of technology

Through ray tracing and multimodal fusion models, the authenticity and response accuracy of virtual space are improved, ensuring that the integration of virtual space and the real environment is more consistent and accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070708A_ABST
    Figure CN120070708A_ABST
Patent Text Reader

Abstract

The invention provides an accelerator simulation training space building method based on multi-modal fusion and ray tracing. The method comprises the steps that virtual space building resources are acquired; constructing a 3D virtual space; rendering the 3D virtual space by adopting an interactive panoramic ray tracing method; and identifying the intention of the user to interact with the virtual object in the rendered 3D virtual space based on the multi-modal fusion model. According to the invention, an interactive panoramic ray tracing method is adopted to generate a tracing weight for each pixel, tracking is carried out according to the tracing weight, irradiance is estimated through ray tracing and an interpolation algorithm, illumination, geometry and material information of an environment in a real space are considered to calculate a final image of virtual-real fusion, and a virtual space is rendered. The authenticity of the virtual space is improved; the information of the sensor channel, the scene vision channel and the voice channel is integrated through the multi-modal fusion model to obtain the final user intention, and the response accuracy is improved by improving the intention recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of virtual scene construction, and relates to a method for constructing a virtual space, specifically to a method for constructing an accelerator simulation training space based on multimodal fusion and ray tracing, which can be used in the technical fields of digital twin and virtual scene construction. Background Art

[0002] Mixed reality is a technical category that encompasses virtual reality (VR) and augmented reality (AR). It combines computer-generated virtual content with the physical world of the real world. Mixed reality technology can not only add virtual elements to the user's field of vision, but also enable these elements to interact with the real environment, generating a sense of reality and interactivity.

[0003] Virtual reality technology mainly includes aspects such as simulated environment, perception, natural skills, and sensing devices. The simulated environment is a three-dimensional stereoscopic realistic image generated by a computer in real time. Perception means that an ideal VR should have all the perceptions that a person has. In addition to the visual perception generated by computer graphics technology, there are also perceptions such as hearing, touch, force sense, movement, and even smell and taste. Natural skills refer to the rotation of a person's head, eyes, gestures, or other human behavioral actions. The computer processes the data adapted to the actions of the participant and makes a real-time response to the user's input, and feeds back to the five senses of the user respectively. With the continuous progress of virtual reality technology, people can project any object into the virtual space and perform virtual interactions to meet their needs.

[0004] Virtual space technology constructs a virtual space scene artificially or by computer. After wearing a head-mounted display, the processor transmits the three-dimensional scene information to the display in front of the human eyes. When the user walks in place, the sensor detects the human motion information, and the processor correspondingly changes the three-dimensional scene information to achieve the effect of the user being on the scene.

[0005] For example, in "A 3D Virtual Space Application Method, Device, Equipment and Medium Based on Artificial Intelligence" (Patent Application No.: 202411338372.5, Publication No. CN118860161A) by Yudian (Shanghai) Technology Co., Ltd., a virtual space construction method is disclosed. The method includes: obtaining virtual space construction resources, which include virtual scene construction resources and virtual portrait construction resources; generating a 3D virtual space based on the virtual space construction resources, where the 3D virtual space includes a 3D virtual scene and 3D virtual characters; performing improved visual rendering on the 3D virtual space and obtaining an improved 3D virtual space; the user logs into the improved 3D virtual space through VR technology and controls the 3D virtual user for virtual interaction; identifying the user's emotional tendency during the virtual interaction, and controlling the 3D virtual non-player character to perform virtual interaction with the 3D virtual user according to the user's emotional tendency, making the constructed virtual space closer to the human eye visual effect and realizing providing emotional value for users scientifically and humanely. However, it is difficult for this invention to capture the lighting, geometry, and material information of the real environment, resulting in inconsistent final images of virtual-real fusion, and the authenticity of the virtual space still needs to be improved. At the same time, this invention responds to the user's interaction according to the user's emotional tendency during the virtual interaction, and it is difficult for the virtual space to accurately respond by predicting the user's operations. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and propose a method for constructing an accelerator simulation training space based on multi-modal fusion and ray tracing to improve the authenticity of the virtual space and the accuracy of response.

[0007] To achieve the above purpose, the technical solutions adopted by the present invention include the following steps:

[0008] (1) Obtain virtual space construction resources:

[0009] Obtain virtual space construction resources including virtual scenes and virtual portraits;

[0010] (2) Construct a 3D virtual space:

[0011] Construct a 3D virtual space including virtual objects and virtual portraits based on the virtual space construction resources;

[0012] (3) Render the 3D virtual space using the interactive panoramic ray tracing method:

[0013] Render the 3D virtual space using the interactive panoramic ray tracing method to obtain a rendered 3D virtual space including virtual objects and the real world;

[0014] (4) Identify the intention of the user to interact with the virtual object in the rendered 3D virtual space based on the multi-modal fusion model:

[0015] Based on the intention recognition of the user's interaction with virtual objects in the rendered 3D virtual space by a multi-modal fusion model, a 3D virtual environment is obtained that can respond to the user's operations according to the user's intention after rendering.

[0016] Compared with the prior art, the present invention has the following advantages:

[0017] 1. The present invention uses an interactive panoramic ray tracing method to generate tracing weights for each pixel, tracks based on the tracing weights, estimates irradiance through ray tracing and interpolation algorithms, calculates the final image of virtual-real fusion by considering the lighting, geometry, and material information of the environment in the real space, and renders a hybrid scene of virtual objects and the real world based on the final image, improving the authenticity of the virtual space.

[0018] 2. The present invention identifies the user's intention during the virtual interaction process through a multi-modal fusion model. The multi-modal fusion integrates the entire input information from the sensor channel, scene visual channel, and voice channel to obtain the final user intention, and responds according to the user intention. In the intention fusion process, the improved coefficient of variation avoids unrealistic values caused by the intention probability of the set being too high or too low, and improves the accuracy of the response in the virtual space by improving the accuracy of intention recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart for the implementation of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0020] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0021] Referring to Figure 1 , the present invention includes the following steps:

[0022] Step 1) Obtain virtual space construction resources:

[0023] Obtain virtual space construction resources including a virtual scene and a virtual portrait;

[0024] The virtual scene construction resources include the real panoramic video of the scene to be constructed, and the geometric and material information of the virtual object to be constructed. In this example, taking the construction of an accelerator simulation training space as an example, obtain the panoramic video of the real training space of the accelerator, and the geometric and material information of the accelerator;

[0025] The virtual portrait construction resources are obtained by: performing frame-by-frame processing on the real portrait video, selecting three consecutive frames of images, determining the region of interest in the frame-by-frame images based on the three-frame difference method, and then performing target segmentation on the region of interest in each frame of image to obtain the virtual portrait construction resources including the portrait of interest;

[0026] Step 2) Construct a 3D virtual space:

[0027] Construct a 3D virtual space including virtual objects and virtual human figures based on the resources built in the virtual space;

[0028] The construction of the 3D virtual space is implemented based on a 3D engine platform such as Unity, UE4, or UDK0. In this embodiment, taking the accelerator simulation training space as an example, a 3D virtual space including virtual human figures and digital twin accelerators is constructed;

[0029] Step 3) Render the 3D virtual space using the interactive panoramic ray tracing method:

[0030] Render the 3D virtual space using the interactive panoramic ray tracing method to obtain a 3D virtual space containing virtual objects and the real world after rendering;

[0031] The implementation steps of rendering the 3D virtual space using the interactive panoramic ray tracing method are as follows:

[0032] (3a) Load the constructed 3D virtual space and the panoramic video of the real scene into the user's head-mounted display, and create a 3D virtual space buffer G containing the geometry, material, texture, and light source information of the virtual objects S and a real scene buffer G including the image data, depth map, and normal map in the real scene P ;

[0033] (3b) Generate a tracking weight w j (p) for the p-th pixel in the i-th frame image displayed in the user's head-mounted display, and divide each pixel of the image frame into three tracking levels l j (p) according to w p ;

[0034] (3c) Use the irradiance estimation algorithm to trace rays for each pixel of the image frame in the user's head-mounted display space according to the tracking weight w j (p) to obtain a sparse irradiance map;

[0035] (3d) Interpolate the pixels in the sparse irradiance map according to the tracking level l p to obtain an interpolated irradiance map;

[0036] (3e) Calculate the final rendered image of the mixed scene containing virtual objects and the real world based on the interpolated irradiance map, and generate a 3D virtual space containing virtual objects and the real world after rendering based on the final image using the 3D engine platform, where:

[0037]

[0038] Where M is the fraction of pixels covered by the real-world environment, and k v is the albedo texture of the virtual object, and L cam is the original real-world G-buffer radiance image, and E RV is the irradiance considering both the real scene and the virtual space, and E R is the irradiance considering only the real scene;

[0039] The tracking weight w j (p) of the p-th pixel and the tracking level l p , and the calculation formula is:

[0040]

[0041] Where p represents the p-th pixel in the j-th frame image, and d j (p) represents the binary parameter calculated for the p-th pixel during the irradiance estimation process, and n p , n q respectively represent the normal vectors of the p-th pixel and the q-th adjacent pixel of the p-th pixel in the neighborhood Ω, and pos p , pos q respectively represent the positions of the pixels corresponding to n p , n q , ||·|| represents the modulo operation, and σ p is the normalization parameter of the pixel position difference, and δ 1 and δ 2 are two thresholds. In this embodiment, σ p = 10;

[0042] By using the two thresholds δ 1 and δ 2 the screen pixels are divided into three regions. For different tracking levels 0, 1, and 2, 1, 1 / 4, and 1 / 16 pixel rays will be tracked. In region 2, only the pixels with coordinates that are multiples of 4 are tracked. In region 1, only the pixels with coordinates that are multiples of 2 are tracked. And in region 0, the pixels with coordinates that are multiples of 1 are tracked, that is, each pixel;

[0043] Generally, when tracing rays from the user's viewpoint to the screen, it is necessary to trace one ray for each pixel to obtain the Monte Carlo estimate of the dense image. However, due to the fact that two physically adjacent points on a planar diffuse surface usually receive similar irradiance from the deep panoramic environment, this method sometimes fails. Step (3c) selects whether to perform ray tracing on specific pixels according to the tracking weight, which improves the virtual environment rendering efficiency and the response speed by enhancing the efficiency of ray tracing;

[0044] The steps to implement ray tracing for pixels according to the tracking weight in the user's head-mounted display space using the irradiance estimation algorithm are:

[0045] (3c1) Load candidate ray r and 3D virtual space buffer G in the user's head-mounted display S and real scene buffer G P ;

[0046] (3c2) Determine whether the type of ray r, r.type = RealObject, holds. If so, execute step (3c3); otherwise, execute step (3c4).

[0047] (3c3) Trace ray r in G P using the PRT algorithm to obtain the irradiance E considering only the real scene R , then determine whether r intersects with the virtual object. If so, calculate the randomized possible reflected ray r' = Reflection(r, Bv, G S ) according to the Reflection at the intersection point Bv of r and the virtual object surface, and trace r' in G S to obtain the irradiance E considering both the real scene and the virtual space RV , otherwise set E RV to E R ;

[0048] (3c4) Trace ray r in G S using the PRT algorithm to obtain E RV , determine whether r intersects with the virtual object. If so, calculate the randomized possible reflected ray r' = Reflection(r, Bv, G S ) according to the Reflection at the intersection point of r and the virtual surface, and trace r' in G S to obtain E RV , otherwise trace r in G P to obtain E RV ;

[0049] (3c5) Write E RV and E R to the 3D virtual space buffer G S ;

[0050] (3c6) After completing the tracing of the rays for all pixels, obtain a sparse irradiance map.

[0051] For each ray, first perform tracing in the virtual space to find interactions within the screen. If not, perform tracing in the real image space to find interactions with the surfaces of the real-world environment, physically correctly model the light interactions between the virtual object and the real world, and make the rendering of the virtual space more realistic.

[0052] Step 4) Identify the intention of the user to interact with virtual objects in the rendered 3D virtual space based on the multimodal fusion model:

[0053] Based on the identification of the intention of the user to interact with virtual objects in the rendered 3D virtual space by the multimodal fusion model, a 3D virtual environment that can respond to the user's operations according to the user's intention after rendering is obtained;

[0054] The implementation steps for identifying the intention of the user to interact with virtual objects in the rendered 3D virtual space based on the multimodal fusion model are as follows:

[0055] (4a) Construct a multimodal fusion model that includes a sensor channel, a scene vision channel, and a voice information channel;

[0056] (4b) The user logs into the rendered 3D virtual space that includes virtual objects and the real world through VR technology by wearing a head-mounted display, smart gloves, and a voice module, and controls the 3D virtual user to interact with the virtual objects in the 3D virtual space;

[0057] (4c) Use sensors in the sensor channel of the multimodal fusion model to collect the user's action data, position, and orientation information. Use a convolutional neural network that includes two convolutional network layers and a fully connected layer to automatically extract features from the raw data of each sensor and classify the user's actions and output probabilities to obtain the probability P(action) of the user performing an action action;

[0058] (4d) Use the YOLOv5 model that includes a Backbone backbone network, a Neck neck structure, and a Head head structure in the scene vision channel of the multimodal fusion model to obtain the probability of the user's hand operation for scene vision information analysis, and obtain the object probability P GH (obj i )), where obj i represents the i-th object element in the operation object set OBJ;

[0059] (4e) Parse the user's speech through speech recognition technology in the voice information channel of the multimodal fusion model to obtain the set I of user intention probabilities in the voice channel vc ;

[0060] (4f) Combine the intention probability sets from the continuous information channel and the voice information channel at the decision-making level using information quantity weights to obtain the final user intention.

[0061] Multimodal intent recognition includes three stages: information collection, intent analysis, and intent fusion. The YOLOv5 technology is used to obtain the probability of the user's hand operation, and the probability that the user attempts to operate an object is derived from continuous information channels such as sensors and scene vision. The discrete speech information is processed using speech recognition technology and Chinese vocabulary analysis technology to derive a set of intent probabilities. Finally, the two sets of probabilities are fused at the decision-making layer of GVVS to generate the final user intent, improving the accuracy of inference and the accuracy of response;

[0062] In the information collection stage, features are automatically extracted from the raw data of each sensor through a convolutional neural network to improve the generalization and robustness of detecting the user's operation behavior. Each of the sensors S i has a corresponding two-layer convolutional sub-network. The convolutional sub-network takes as input a segment of the sensor time series matrix V(S i ) obtained from a window sliding in seconds. Each layer of the convolutional sub-network passes through a two-dimensional convolutional kernel of size (1, 3) to learn the data features, and then the feature maps conv i1 and conv i2 are obtained in sequence. The depth feature maps conv i2 of the sensors are merged to obtain a large depth feature map. The large feature map passes through two convolutional network layers to determine the correlation between multiple sensor features, and the Fconv 1 and Fconv 2 are obtained in sequence. In the fully connected layer, a softmax classifier is used to classify the user's actions and the output probability P(action i );

[0063] conv ij =a ij-1 W ij +b ij (4)

[0064] P(action i ) = softmax(Fconv 2 ) (5)

[0065] where a ij-1 represents the output value of the i-th sensor at the (j - 1)-th layer, the convolutional kernel weight is W ij , and the bias is b ij ;

[0066] Scene vision information analysis:

[0067] The implementation steps of scene vision information analysis are as follows:

[0068] (4d1) Use the trained YOLOv5 model to identify the virtual object obj in the user interaction in the image captured by the depth camera on the user's head-mounted display, i and obtain the object element obj in the operation object set OBJ i with depth information dep i and class probability P(obj i ). According to the incremental change Δdep i of the object depth, calculate the weight W(i) corresponding to obj i in the scene visual channel, and multiply W(i) and the class probability P(obj i ) to obtain the probability P i of selecting the object obj G in the scene visual channel by the intelligent glove i :

[0069]

[0070] P G (obj i ) = W(i)P(obj i ) (7)

[0071] where, Δdep i represents the i-th object element in the operation object set OBJ, which is the change amount of the depth information compared with the previous frame. W(i) is the weight corresponding to obj i in the scene visual channel, representing the ratio of the depth increment of the object obj i to the sum of all object depth increments;

[0072] (4d2) Calculate the distances between the bounding box coordinate points x i , y i of all operation objects and the intersection point (X, Y) formed between the head ray R and the plane where the operation object is located and calculate the probability P i of selecting the object obj i by the head-mounted device in the scene visual channel H (obj i ), and obtain the probability P H (obj i ) that the user selects the operation object obj i in the scene visual channel through P GH (obj i ):

[0073]

[0074]

[0075] Wherein, m represents the m-th object in the current operation object set, and n represents the total number of objects in the operation object set;

[0076] Parse the user's speech through speech recognition technology in the speech information channel of the multi-modal fusion model to obtain the set of user intention probabilities I in the speech channel vc , implementation steps:

[0077] Continuously monitor the user's speech and divide it into a group of verbs s v and a group of nouns s n , then perform the Cartesian product operation on s n and s v , generate a set of temporary intentions, and finally determine the set of experimental intention probabilities I of the user in the speech channel by matching the text similarity with the experimental intention database ES vc ;

[0078] Obtain the final user intention, and the implementation steps are as follows:

[0079] (4f.1) Obtain the set of intention probabilities I under the continuous information channel cs ,

[0080] I cs ={P(action)×P(object)|(action + object)∈ES} (10)

[0081] Wherein, (action + object)∈ES means that the intention result of splicing the action and the object is in the current intention library ES;

[0082] (4f.2) Calculate the coefficient of variation of the continuous information channel and the coefficient of variation of the speech channel

[0083]

[0084] Wherein, g(·) represents the Gaussian distribution of the distance between intention probabilities, and R represents the number of intentions in the intention set;

[0085] (4f.3) Perform a normalization operation on the coefficient of variation of the continuous information channel and the coefficient of variation of the speech channel to obtain the dynamic weight information under the continuous information channel and the speech channel and

[0086] (4f.4) Perform intention fusion according to the dynamic weight information under the continuous information channel and the speech channel to obtain the final user intention as:

[0087]

[0088] In the traditional method, the coefficient of variation is calculated by finding the variance and mean of the probability set with R intentions. The greater the difference between the probabilities in the intention probability set, the higher the accuracy of the intention recognition of the channel, and higher weights are required. Step (4f.2) improves the classical coefficient of variation solution, making the coefficient more compatible with the technology of calculating intentions, avoiding unrealistic values caused by overly high or low intention probabilities in the set, and improving the accuracy of the response in the virtual space by enhancing the accuracy of intention recognition.

Claims

1. A method for constructing an accelerator simulation training space based on multimodal fusion and ray tracing, characterized in that: The steps include: (1) Obtaining virtual space construction resources: Acquire virtual space construction resources including virtual scenes and virtual portraits; (2) Constructing 3D virtual space: Constructing a 3D virtual space including virtual objects and virtual portraits based on virtual space building resources; (3) Rendering 3D virtual space using interactive panoramic ray tracing method: An interactive panoramic ray tracing method is used to render the 3D virtual space, and the rendered 3D virtual space includes virtual objects and the real world. (4) Identify the user’s intention to interact with virtual objects in the rendered 3D virtual space based on the multimodal fusion model: Based on the multimodal fusion model, the user's intention to interact with virtual objects in the rendered 3D virtual space is recognized, and the rendered 3D virtual space is obtained which can respond to the user's operation according to the user's intention.

2. The method according to claim 1, characterized in that The virtual scene construction resources described in step (1) include a real panoramic video of the scene to be constructed and geometric and material information of the virtual objects to be constructed.

3. The method according to claim 1, characterized in that The virtual portrait building resource described in step (1) is obtained by: performing frame processing on the real portrait video, and selecting three consecutive frames of images to determine the region of interest in the image after frame division based on the three-frame difference method, and then performing target segmentation on the region of interest in each frame of the image to obtain a virtual portrait building resource including the portrait of interest.

4. The method according to claim 1, characterized in that: The construction of the 3D virtual space in step (2) is implemented based on a 3D engine platform such as Unity, UE4 or UDK0.

5. The method according to claim 1, characterized in that The interactive panoramic ray tracing method described in step (3) is used to render the 3D virtual space, and the implementation steps are as follows: (3a) Load the constructed 3D virtual space and the panoramic video of the real scene into the user's head-mounted display, and create a 3D virtual space buffer G containing the geometry, material, texture and light source information of the virtual object S and a real scene buffer G including image data, depth map and normal map in the real scene P ; (3b) Generate a tracking weight w for the pth pixel in the i-th frame image displayed in the user's head-mounted display j (p), and according to w j (p) Divide each pixel of the image frame into three tracking levels l p ; (3c) An irradiance estimation algorithm is used to calculate the tracking weight w in the user's head mounted display space. j (p) Tracing rays for each pixel of the image frame to obtain a sparse irradiance map; (3d) According to the tracking level l p Interpolate the pixels in the sparse irradiance map to obtain an interpolated irradiance map; (3e) calculating a final rendered image of a mixed scene including virtual objects and the real world based on the interpolated irradiance map, and generating a rendered 3D virtual space including virtual objects and the real world using a 3D engine platform based on the final image, wherein: Where M is the fraction of pixels covered by the real-world environment, k v is the virtual object albedo texture, L cam is the original real-world G-buffer radiance image, E RV It considers the irradiance of both the real scene and the virtual space. R It only considers the irradiance of the real scene.

6. The method according to claim 5, characterized in that The tracking weight w of the pth pixel described in step (3b) j (p) and trace level l p , the calculation formula is: Among them, d j (p) represents the binary parameter calculated during the irradiance estimation process for the p-th pixel in the j-th frame image, n p 、n q denote the normal of the p-th pixel and its q-th neighboring pixel in the neighborhood Ω, pos p ,pos q Respectively represent n p 、n q The corresponding pixel position, |||| represents the modulus operation, σ p is the normalization parameter of the pixel position difference, and δ1 and δ2 are two thresholds.

7. The method according to claim 4, characterized in that The irradiance estimation algorithm described in step (3c) is used to trace rays for pixels in the user head mounted display space according to the tracing weights, and the implementation steps are: (3c1) Load the candidate ray r and the 3D virtual space buffer G in the user's head mounted display S and the real scene buffer G P ; (3c2) Determine whether the type of the ray r r.type = RealObject is true. If so, execute step (3c3); otherwise, execute step (3c4); (3c3) Through the PRT algorithm in G P Tracing ray r in the above example obtains the irradiance E considering only the real scene. R Then determine whether r intersects with the virtual object. If so, calculate the random possible reflection ray r'=Reflection(r,Bv,G S ), in G S Track r' to obtain the irradiance E that takes into account both the real scene and the virtual space RV Otherwise, E RV Set to E R ; (3c4) Through the PRT algorithm in G S Tracing ray r in the above equation gives E RV , determine whether r intersects with the virtual object. If so, calculate the random possible reflection ray r'=Reflection(r,Bv,G S ), in G S Track r' to get E RV , otherwise in G P Track r to get E RV ; (3c5) E RV and E R Write to 3D virtual space buffer G S ; (3c6) After completing the tracing of the rays for all pixels, a sparse irradiance map is obtained.

8. The method according to claim 1, characterized in that The multimodal fusion model described in step (4) identifies the user's intention to interact with the virtual object in the rendered 3D virtual space, and the implementation steps are as follows: (4a) Construct a multimodal fusion model including sensor channel, scene vision channel and voice information channel; (4b) The user logs into the rendered 3D virtual space containing virtual objects and the real world by wearing a head mounted display, smart gloves and a voice module through VR technology, and controls the 3D virtual user to interact with the virtual objects in the 3D virtual space; (4c) Using sensors in the sensor channel of the multimodal fusion model to collect the user's action data, position, and direction information, a convolutional neural network including two convolutional network layers and a fully connected layer is used to automatically extract features from the raw data of each sensor and classify the user's actions and output probabilities to obtain the user's action probability P(action); (4d) In the scene vision channel of the multimodal fusion model, the YOLOv5 model including the Backbone network, the Neck structure and the Head structure is used to obtain the probability of the user's hand operation to analyze the scene vision information and obtain the object probability P GH (obj i ), where obj i Represents the i-th object element in the operation object set OBJ; (4e) In the voice information channel of the multimodal fusion model, the user's voice is analyzed by voice recognition technology to obtain the user intention probability set I in the voice channel. vc ; (4f) Use information weight to combine the intention probability sets from the continuous information channel and the voice information channel at the decision level to obtain the final user intention.

9. The method according to claim 8, characterized in that The scene visual information analysis described in step (4d) is implemented by: (4d1) Use the trained YOLOv5 model to identify the virtual object obj that the user interacts with in the depth camera image on the user's head-mounted display i , get the object element obj in the operation object set OBJ i Depth information dep i And the category probability P(obj i ), according to the incremental change Δdep of the object depth i , calculate the scene visual channel obj i The corresponding weight W(i) is W(i) and the category probability P(obj i ) to perform multiplication operation to obtain the object obj selected by the smart glove in the scene visual channel i The probability P G (obj i ); (4d2) Calculate the bounding box coordinates x of all the objects after the intersection (X, Y) is formed between the head ray R and the plane where the object is located i ,y i The distance from the intersection point (X,Y) And according to distances i Calculate the object obj selected by the head mounted device in the scene visual channel i The probability P H (obj i ), and through P H (obj i ) The user selects the operation object obj in the scene visual channel i The probability P GH (obj i ): Wherein, m represents the mth object in the current operation object set, and n represents the total number of objects in the operation object set.

10. The method according to claim 8, characterized in that The steps for obtaining the end user's intention described in step (4f) are as follows: (4f.1) Obtaining the intention probability set I under the continuous information channel cs , I cs ={P(action)×P(object)|(action+object)∈ES} Among them, (action+object)∈ES, means that the intention result of splicing action and object is in the current intention library ES; (4f.2) Calculate the coefficient of variation of continuous information channels and the coefficient of variation of the speech channel Among them, g() represents the Gaussian distribution of the distance between intent probabilities, and R represents the number of intents in the intent set; (4f.3) Coefficient of variation for continuous information channels and the coefficient of variation of the speech channel Perform normalization operations to obtain dynamic weight information under continuous information channels and voice channels and (4f.4) The final user intention is obtained by performing intention fusion based on the dynamic weight information under the continuous information channel and the voice channel:

Citation Information

Patent Citations

  • 3D virtual space application method and device based on artificial intelligence, equipment and medium

    CN118860161A

Cited By

  • Virtual scene building processing method and system based on image recognition algorithm

    CN121033353A

  • A virtual scene building processing method and system based on an image recognition algorithm

    CN121033353B