Virtual reality scene fusion method and system

Through the collaborative work of the acquisition, processing, output, and parsing modules, the efficient integration of virtual and real scenes in virtual reality devices is achieved, solving the problem of separation between vision and operation in virtual reality devices, enhancing immersion and interactive realism, and adapting to diverse usage scenarios.

CN122066901APending Publication Date: 2026-05-19WUHAN HIPAI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN HIPAI TECH CO LTD
Filing Date
2026-01-26
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing virtual reality headsets on flight simulators suffer from a separation between vision and operation, resulting in insufficient immersion and realism of interaction. Furthermore, the integration of virtual and real scenes is inefficient and relies on slow manual modeling.

Method used

The acquisition module obtains 3D topological structure, ambient light intensity, and dynamic obstacle displacement data of the real scene to generate a basic data matrix; the processing module performs spatial coordinate calibration and data feature matching to generate a virtual-real fusion scene; the output module performs lighting and shadow calculations and image compositing; the parsing module captures user interaction signals and drives virtual element responses; and the control module optimizes resource allocation in real time to ensure system stability.

Benefits of technology

It achieves efficient unification and precise matching of virtual and real scenes, with lighting and shadow effects that closely match reality, rapid response to user interaction, enhanced immersion and adaptability, adaptable to diverse usage scenarios, and provides a natural and smooth virtual reality experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122066901A_ABST
    Figure CN122066901A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual reality scene fusion method and system, and relates to the field of virtual reality fusion, and the system comprises a collection module which is used for collecting a three-dimensional topological structure, environment light intensity and dynamic obstacle displacement data of a target reality scene, and generating a to-be-fused basic data matrix; the uploading module is used for uploading the preset virtual scene data and storing the preset virtual scene data; by accurately capturing the space structure, light intensity and dynamic obstacle information of the real scene, efficient unification of virtual and real scene coordinates and accurate feature matching are realized, the light and shadow effect fits the real environment and is dynamically adjusted along with the light intensity, the detail fidelity is high, the boundary transition is smooth, the limb and voice interaction requirements of the user can be quickly analyzed, and the user experience is improved. And the virtual element response is triggered in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality fusion technology, specifically to a virtual reality scene fusion method and system. Background Technology

[0002] Virtual reality fusion technology captures real-world environmental data through sensors, overlaying virtual models with real scenes in real time to achieve virtual-real interaction and visual fusion. It relies on computer graphics and positioning tracking technology to break down the boundaries between virtual and real, enhancing the immersion of the scene and the realism of the interaction.

[0003] Patent application number 202110573160.5 discloses a method, system, and flight simulator for fusing real-world and virtual-world scenes. This application aims to address the problem that while using virtual reality head-mounted displays (VR headsets) to display virtual-world scenes can greatly enhance the immersive experience, the biggest drawback is that pilots can only see the displayed scene and not their own actions, resulting in a lack of realism during operation. Therefore, the application of VR headsets in flight simulators is currently very limited, mostly restricted to the initial learning and cognitive stages to enable pilots to perceive the environment and perform simulated operations.

[0004] However, for similar virtual-real interaction scenarios such as games, people hope to quickly integrate virtual and real scenes. But this goal cannot be fully or largely achieved by intelligent systems at present. It still requires manual information collection and assisted modeling of real scenes, and the overall efficiency is quite slow.

[0005] To address this, we propose a virtual reality scene fusion method and system. Summary of the Invention

[0006] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a virtual reality scene fusion method and system, which can effectively solve the problems of the existing technology.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions;

[0008] This invention discloses a virtual reality scene fusion system, comprising:

[0009] The system comprises the following modules: an acquisition module for collecting 3D topological structure, ambient light intensity, and dynamic obstacle displacement data of the target real-world scene, generating a basic data matrix to be fused; an upload module for uploading and storing preset virtual scene data; a processing module for receiving the basic data matrix to be fused and the preset virtual scene data, generating a fusion data set of the virtual-real fusion scene through spatial coordinate calibration and data feature matching; an output module for acquiring the fusion data set, performing real-time scene lighting and shadow calculations and image compositing, and outputting continuous fusion scene images to the virtual reality display device; a parsing module for capturing user body movements and voice command signals, parsing the corresponding interaction requirements, and driving virtual elements in the fusion scene to perform matching response actions; and a control module for real-time monitoring of the operating parameters and data transmission rate of each module, adjusting the resource allocation of each module in real time based on scene rendering quality and interaction response latency.

[0010] The preset virtual scene data includes scene topology, rendering parameters, and interactive element attribute data;

[0011] The acquisition module is interconnected with the upload module and the processing module via a wireless network. The processing module is interconnected with the output module via a wireless network. The output module is interconnected with the parsing module and the control module via a wireless network.

[0012] Furthermore, when the acquisition module acquires the 3D topological structure of the target real-world scene, it obtains the surface feature point set and depth information of each object in the target real-world scene through a preset multi-view image acquisition unit and depth sensing unit. Based on feature point association and depth value fitting, it generates 3D topological structure data and simultaneously performs feature extraction of the 3D topological structure data. The feature extraction of the real-time 3D topological structure follows the following rules:

[0013] ;

[0014] In the formula: The key topological feature value of a single feature point in three dimensions; This represents the total number of feature points collected. The weighting coefficients are the spatial location weighting coefficients of feature points and the weighting coefficients for the influence of depth information. Let be the three-dimensional spatial coordinates of the i-th feature point; Let be the 3D spatial depth value of the i-th feature point; The number of neighboring feature points of the i-th feature point; Let be the Euclidean distance between the i-th feature point and the j-th neighboring feature point; This is a preset minimum constant;

[0015] After obtaining the feature threshold, a feature threshold is set. If T is not less than the feature threshold, the feature point is determined to be a key feature point for constructing the three-dimensional topology and is retained; otherwise, it is discarded.

[0016] Furthermore, during the stage of acquiring ambient light intensity data, the acquisition module acquires the light intensity values ​​of multiple sampling points in the target real scene through a full-spectrum light intensity sensor, and removes noise through mean filtering;

[0017] During the stage of acquiring dynamic obstacle displacement data, the acquisition module obtains the obstacle contour based on the inter-frame difference of consecutive frame images, and predicts the obstacle's displacement trajectory using Kalman filtering. The displacement data is obtained through prediction.

[0018] ;

[0019] In the formula: Let k be the displacement value of the dynamic obstacle at time k. For predicting weighting coefficients; This represents the displacement value at time k-1. Let be the instantaneous velocity of the obstacle at time k-1; The time interval between the acquisition of two adjacent image frames; Let be the displacement value at time t; This is the statistical window length for historical displacement data.

[0020] Furthermore, when the processing module performs spatial coordinate calibration, it uses at least three pre-set non-collinear fixed reference points in the target real-world scene as calibration benchmarks to obtain the real-world spatial coordinates and virtual scene coordinates of each reference point, constructs a coordinate transformation matrix, and then uses the transformation matrix to achieve coordinate unification between the basic data matrix to be fused and the pre-set virtual scene data. The coordinate calibration follows the following rules:

[0021] ;

[0022] In the formula: Let X be the calibrated coordinate vector (x', y', z'); A is a 3×3 rotation and scaling matrix; X is the original coordinate vector (x, y, z) before calibration; and B is a 3×1 translation vector.

[0023] Furthermore, when the processing module performs data feature matching, it extracts the real feature dimension from the basic data matrix to be fused and the virtual feature dimension from the preset virtual scene data;

[0024] Among them, the real-world feature dimension includes the node connection density of the three-dimensional topology, the spectral distribution characteristics of ambient light intensity, and the motion attribute characteristics of dynamic obstacles, while the virtual feature dimension includes the node type parameters of the scene topology, the color space parameters of the rendering parameters, and the response triggering characteristics of interactive elements.

[0025] Weights are assigned based on the importance of each feature dimension, and data feature matching is performed by calculating feature similarity. A match is considered successful only when the feature similarity of each dimension is not lower than a preset similarity threshold, triggering the generation of the fused data set.

[0026] Among them, feature similarity ;

[0027] In the formula: This represents the total number of feature dimensions involved in the matching. The importance weight of the i-th feature dimension; Let be the normalized feature value of the i-th real-world feature dimension; Let be the normalized feature value of the i-th virtual feature dimension; The maximum value of all real feature sample values ​​and virtual feature sample values ​​under the i-th feature dimension; It is the minimum value of all real feature sample values ​​and virtual feature sample values ​​under the i-th feature dimension; Let V be the variance of the actual feature sample values ​​under the i-th feature dimension; For preset small constants; This is a preset small constant.

[0028] Furthermore, the output module performs a real-time scene lighting and shadow calculation stage, using the ambient light intensity data in the fused data set as the real lighting and shadow benchmark, and combining the rendering parameters of the virtual scene with the surface reflection coefficient and diffuse reflection coefficient of the objects in the fused scene to calculate the lighting and shadow intensity of each pixel.

[0029] Among them, light and shadow intensity ;

[0030] In the formula: The weighting coefficient is used to determine the contribution of light intensity in the real environment. The reference value of ambient light intensity obtained by the acquisition module; This is the specular reflection coefficient of the object surface corresponding to the target pixel. The angle between the incident direction of ambient light and the normal vector of the object's surface; This is the diffuse reflectance coefficient of the object's surface; The number of virtual light sources affecting this pixel in the virtual scene; Let be the luminous intensity of the nth virtual light source; The illumination weight of the nth virtual light source for that pixel; This represents the virtual spatial distance from the pixel to the nth virtual light source. This is the adjustment constant;

[0031] When the output module performs image compositing, it merges the virtual scene image with the real scene image by layer overlay according to the light and shadow intensity weight of each pixel;

[0032] The fusion weight of virtual scene images and real scene images is dynamically adjusted according to the change of ambient light intensity; that is, the higher the light intensity, the greater the fusion weight of the real scene images.

[0033] Furthermore, when the parsing module captures the user's limb movement signals, it collects the angle change data and acceleration data of each joint of the user through the pre-worn inertial measurement unit, and combines the limb contour data obtained by the visual capture unit to construct a three-dimensional motion model of the user's limbs.

[0034] When the parsing module captures voice command signals, it collects voice data through a preset pickup unit.

[0035] Furthermore, during the real-time analysis of interaction requirements, a mapping library of body movement features and voice command features is established. Based on feature matching degree and association confidence, the user's interaction intent is determined, driving the corresponding virtual elements in the fusion scene to perform response actions.

[0036] Among them, when the association confidence level is not lower than the preset confidence threshold, the interaction intent is determined to be valid.

[0037] Furthermore, the system operating parameters monitored in real time by the control module include the CPU utilization rate, memory utilization rate, and data cache queue length of each module, and the data transmission rate includes the data interaction rate between modules and the data stream output rate to the virtual reality display device;

[0038] Based on the preset scene rendering quality score and interaction response latency threshold, a resource allocation adjustment model is constructed. When the scene rendering quality score is lower than the preset score threshold or the interaction response latency is higher than the preset latency threshold, the CPU usage ratio and memory allocation are adjusted according to module priority.

[0039] On the other hand, a virtual reality scene fusion method includes:

[0040] The system collects 3D topological structure, ambient light intensity, and dynamic obstacle displacement data of the target real-world scene. After feature extraction, noise removal, and trajectory prediction, a basic data matrix to be fused is generated. Preset virtual scene data containing scene topology, rendering parameters, and interactive element attribute data is uploaded and stored. A coordinate transformation matrix is ​​constructed using at least three non-collinear fixed reference points in the real-world scene to achieve coordinate unification. Data matching is completed through multi-dimensional feature similarity calculation to generate a fused data set. The pixel light and shadow intensity is calculated based on the ambient light intensity benchmark and virtual scene parameters. The resulting continuous fused scene image is synthesized by layer overlay according to light and shadow weights and output to the virtual reality display device. User body movements and voice command signals are collected and feature extracted. After parsing the interaction requirements, the corresponding virtual elements in the fused scene are driven to perform matching response actions.

[0041] Compared with the known prior art, the technical solution provided by this invention has the following beneficial effects:

[0042] This invention achieves efficient unification and precise feature matching of virtual and real scene coordinates by accurately capturing information on the spatial structure, light intensity, and dynamic obstacles of real-world scenes. Its lighting and shadow effects closely match the real environment and dynamically adjust with light intensity, resulting in high detail fidelity and smooth boundary transitions. It can quickly analyze user's body and voice interaction needs, promptly trigger virtual element responses, and simultaneously adapt to the system's operating status in real time, optimizing resource allocation to ensure stable frame rates and low interaction latency. This effectively enhances the immersiveness and adaptability of virtual-real fusion. Furthermore, this technical solution is adaptable to diverse usage scenarios, allowing users to obtain a natural, smooth, and believable virtual reality experience. By also considering scene presentation quality and interaction response efficiency, it further improves the intelligence level of virtual reality scene fusion. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0044] Figure 1 This is a schematic diagram of the structure of a virtual reality scene fusion system;

[0045] Figure 2 This is a flowchart illustrating a virtual reality scene fusion method. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0047] The present invention will be further described below with reference to embodiments.

[0048] Example 1:

[0049] This embodiment provides a virtual reality scene fusion system, such as Figure 1 As shown, it includes:

[0050] The acquisition module is used to acquire the three-dimensional topology, ambient light intensity, and dynamic obstacle displacement data of the target real scene, and generate a basic data matrix to be fused.

[0051] When the acquisition module acquires the 3D topological structure of the target real-world scene, it obtains the surface feature point set and depth information of each object in the target real-world scene through a preset multi-view image acquisition unit and depth sensing unit. Based on feature point association and depth value fitting, it generates 3D topological structure data and simultaneously performs feature extraction of the 3D topological structure data. The feature extraction of the real-time 3D topological structure follows the following rules:

[0052] ;

[0053] In the formula: The key topological feature value of a single feature point in three dimensions; This represents the total number of feature points collected. The weighting coefficients are the spatial location weighting coefficients of feature points and the weighting coefficients for the influence of depth information. Let be the three-dimensional spatial coordinates of the i-th feature point; Let be the 3D spatial depth value of the i-th feature point; The number of neighboring feature points of the i-th feature point; Let be the Euclidean distance between the i-th feature point and the j-th neighboring feature point; This is a preset minimum constant;

[0054] The above formula takes into account the spatial location and depth information of feature points, as well as the distribution density of neighboring feature points. By setting weight coefficients to balance the effects of different dimensional parameters, and by using preset minimum constants to avoid the problem of zero denominator in the calculation, the formula accurately selects the core feature points that are crucial to the fusion of virtual and real scenes by obtaining the three-dimensional topological key feature values ​​of a single feature point and comparing them with a threshold. This provides a precise and reliable three-dimensional structural basis for subsequent spatial coordinate calibration and data feature matching.

[0055] It should be noted that in the actual calculation process of the above formula, a scene feature length benchmark based on at least three non-collinear fixed reference points in the target real scene will be introduced first. This benchmark is determined by calculating the maximum Euclidean distance between the above reference points. Then, the three-dimensional spatial coordinates, three-dimensional spatial depth values, Euclidean distance between feature points and preset minimum constants involving length dimensions in the original formula are all normalized by dividing by the scene feature length benchmark and converted into dimensionless values ​​before calculation.

[0056] After obtaining the feature threshold, a feature threshold is set. If T is not less than the feature threshold, the feature point is determined to be a key feature point for constructing the three-dimensional topology and is retained. Otherwise, it is discarded.

[0057] in, The values ​​of all values ​​are within the range of (0,1), and The sum is 1;

[0058] It should be noted that the purpose of feature extraction is to select the key core feature points that are crucial for the fusion of virtual and real scenes, so as to provide an accurate three-dimensional structural benchmark for the spatial coordinate calibration and data feature matching of subsequent processing modules. In addition, the feature types extracted during the feature extraction stage include feature point spatial location features, feature point depth dimension features, and feature point neighborhood distribution features.

[0059] During the stage of acquiring ambient light intensity data, the acquisition module collects light intensity values ​​from multiple sampling points in the target real scene through a full-spectrum light intensity sensor, and removes noise through mean filtering.

[0060] During the acquisition module's phase of collecting dynamic obstacle displacement data, the obstacle contour is obtained based on the inter-frame difference of consecutive frame images. The Kalman filter is then used to predict the obstacle's displacement trajectory; the displacement data is obtained through prediction.

[0061] ;

[0062] In the formula: Let k be the displacement value of the dynamic obstacle at time k. For predicting weighting coefficients; This represents the displacement value at time k-1. Let be the instantaneous velocity of the obstacle at time k-1; The time interval between the acquisition of two adjacent image frames; Let be the displacement value at time t; The statistical window length for historical displacement data;

[0063] The above formula integrates the displacement value, instantaneous velocity and historical displacement of the obstacle at the previous moment, adjusts the influence ratio of real-time motion data and historical statistical data by predicting the weight coefficient, optimizes the prediction logic by combining the acquisition time interval of adjacent frames and the length of historical data statistical window, and with the help of inter-frame difference and Kalman filtering of continuous frame images, it can efficiently capture the displacement trajectory of dynamic obstacles and provide accurate data support for the dynamic adaptation of virtual elements in the fusion scene.

[0064] The upload module is used to upload and store preset virtual scene data.

[0065] The processing module is used to receive the basic data matrix to be fused and the preset virtual scene data, and generate a fused data set of virtual and real fusion scene through spatial coordinate calibration and data feature matching;

[0066] When the processing module performs spatial coordinate calibration, it uses at least three pre-set non-collinear fixed reference points in the target real-world scene as calibration benchmarks to obtain the real-world spatial coordinates and virtual scene coordinates of each reference point, constructs a coordinate transformation matrix, and then uses the transformation matrix to unify the coordinates of the base data matrix to be fused with the pre-set virtual scene data. The coordinate calibration follows the following rules:

[0067] ;

[0068] In the formula: Let X be the calibrated coordinate vector (x', y', z'); A is a 3×3 rotation and scaling matrix; X is the original coordinate vector (x, y, z) before calibration; and B is a 3×1 translation vector.

[0069] When the processing module performs data feature matching, it extracts the real feature dimensions from the basic data matrix to be fused and the virtual feature dimensions from the preset virtual scene data;

[0070] The above formula selects fixed reference points that are not collinear in the target real scene as calibration benchmarks. By constructing a 3×3 rotation scaling matrix A and a 3×1 translation vector B, the corresponding transformation relationship between the original coordinate vector X(x,y,z) before calibration and the coordinate vector X'(x',y',z') of the unified coordinate system after calibration is established. This realizes the coordinate unification between the basic data to be fused and the preset virtual scene data, and builds a spatial consistency foundation for data processing and scene synthesis in subsequent stages.

[0071] Among them, the real-world feature dimension includes the node connection density of the three-dimensional topology, the spectral distribution characteristics of ambient light intensity, and the motion attribute characteristics of dynamic obstacles, while the virtual feature dimension includes the node type parameters of the scene topology, the color space parameters of the rendering parameters, and the response triggering characteristics of interactive elements.

[0072] Weights are assigned based on the importance of each feature dimension, and data feature matching is performed by calculating feature similarity. A match is considered successful only when the feature similarity of each dimension is not lower than a preset similarity threshold, triggering the generation of the fused data set.

[0073] Among them, feature similarity ;

[0074] In the formula: This represents the total number of feature dimensions involved in the matching. The importance weight of the i-th feature dimension; Let be the normalized feature value of the i-th real-world feature dimension; Let be the normalized feature value of the i-th virtual feature dimension; The maximum value of all real feature sample values ​​and virtual feature sample values ​​under the i-th feature dimension; It is the minimum value of all real feature sample values ​​and virtual feature sample values ​​under the i-th feature dimension; Let V be the variance of the actual feature sample values ​​under the i-th feature dimension; For preset small constants; For preset small constants;

[0075] The above formula comprehensively covers the key feature dimensions of both real-world and virtual scenes. Weights are assigned based on the importance of each feature dimension to the fusion effect. Normalized feature values ​​R are obtained by mapping the original real-world feature data and virtual feature data to the [0,1] interval. i V i To eliminate the dimensional differences between different feature dimensions, and to introduce feature sample variance and preset small constants to avoid the denominator being zero during the calculation process, a scientific similarity calculation logic is constructed by combining feature value difference and extreme value range. The matching is only determined when the similarity of all feature dimensions meets the preset requirements, ensuring the comprehensiveness and accuracy of virtual and real data feature matching.

[0076] in, >0, and , ≠ , To avoid The denominator is zero. To avoid , The denominator is zero. This is obtained by mapping the original real-world feature data to the [0,1] interval. This is obtained by mapping the original virtual feature data to the [0,1] interval;

[0077] The output module is used to acquire the fused data set, perform real-time calculation of scene lighting and shadow and image composition, and output continuous fused scene images to the virtual reality display device;

[0078] The output module performs the real-time scene lighting and shadow calculation stage. It uses the ambient light intensity data in the fused data set as the real lighting and shadow benchmark, and combines the rendering parameters of the virtual scene with the surface reflection coefficient and diffuse reflection coefficient of the objects in the fused scene to calculate the lighting and shadow intensity of each pixel.

[0079] Among them, light and shadow intensity ;

[0080] In the formula: The weighting coefficient is used to determine the contribution of light intensity in the real environment. The reference value of ambient light intensity obtained by the acquisition module; This is the specular reflection coefficient of the object surface corresponding to the target pixel. The angle between the incident direction of ambient light and the normal vector of the object's surface; This is the diffuse reflectance coefficient of the object's surface; The number of virtual light sources affecting this pixel in the virtual scene; Let be the luminous intensity of the nth virtual light source; The illumination weight of the nth virtual light source for that pixel; This represents the virtual spatial distance from the pixel to the nth virtual light source. As an adjustment constant, >0;

[0081] The above formula uses the collected real-world ambient light intensity benchmark value as the basic light and shadow reference. It combines the rendering parameters of the virtual scene with the specular reflection coefficient and diffuse reflection coefficient of the object surface. The contribution ratio of real-world light and shadow and virtual light and shadow is balanced by the contribution weight coefficient of real-world ambient light intensity. At the same time, the illumination weight is dynamically adjusted considering the illumination angle, propagation distance and occlusion of the virtual light source. An adjustment constant is introduced to avoid the problem of light and shadow intensity overflow when the distance is too small. This achieves a natural transition and fusion of light and shadow effects in virtual and real scenes, making the light and shadow presentation of the composite image more in line with the human eye's perception of the real environment.

[0082] in, ∈ (0,1], its value increases as the angle between the incident direction of the virtual light source and the normal vector of the object surface corresponding to the target pixel decreases, and increases as the degree of occlusion on the propagation path from the virtual light source to the target pixel decreases, and vice versa. Used to prevent light and shadow intensity from overflowing when the distance is too small;

[0083] When the output module performs image compositing, it blends the virtual scene image with the real scene image by layer overlay according to the light and shadow intensity weight of each pixel;

[0084] Among them, the fusion weight of virtual scene images and real scene images is dynamically adjusted with the change of ambient light intensity, that is, the higher the light intensity, the greater the fusion weight of real scene images.

[0085] The parsing module is used to capture user body movements and voice command signals, parse the corresponding interaction requirements of the signals, and drive virtual elements in the fusion scene to perform matching response actions;

[0086] When the parsing module captures the user's limb movement signals, it collects the angle change data and acceleration data of each joint of the user through the pre-worn inertial measurement unit, and combines it with the limb contour data obtained by the visual capture unit to construct a three-dimensional motion model of the user's limbs.

[0087] When the parsing module captures voice command signals, it collects voice data through a preset pickup unit, and extracts Mel frequency cepstral coefficients and voice rhythm features after noise reduction processing.

[0088] Furthermore, during the real-time analysis of interaction requirements, a mapping library of body movement features and voice command features is established. Based on feature matching degree and association confidence, the user's interaction intent is determined, driving the corresponding virtual elements in the fusion scene to perform response actions.

[0089] Among them, when the association confidence level is not lower than the preset confidence threshold, the interaction intent is determined to be valid;

[0090] The control module is used to monitor the operating parameters and data transmission rate of each module in the system in real time, and adjust the allocation of operating resources of each module in real time according to the scene rendering quality and interaction response latency.

[0091] The system operating parameters monitored in real time by the control module include the CPU utilization rate, memory usage rate, and data cache queue length of each module. The data transmission rate includes the data interaction rate between modules and the data stream output rate to the virtual reality display device.

[0092] Based on the preset scene rendering quality scoring standard and interaction response latency threshold, a resource allocation adjustment model is constructed. When the scene rendering quality score is lower than the preset score threshold or the interaction response latency is higher than the preset latency threshold, the CPU usage ratio and memory allocation are adjusted according to module priority.

[0093] Among them, the output module and the processing module have higher priority than other modules;

[0094] The scene rendering quality score is quantitatively represented by a weighted sum of four dimensions: pixel detail fidelity, lighting and shadow matching, frame rate stability, and smoothness of transition between virtual and real boundaries. Pixel detail fidelity, lighting and shadow matching, frame rate stability, and smoothness of transition between virtual and real boundaries are all normalized and limited to values ​​between 0 and 1 before the weighted summation operation is performed.

[0095] The preset virtual scene data includes scene topology, rendering parameters, and interactive element attribute data;

[0096] The acquisition module is interconnected with the upload module and the processing module via a wireless network. The processing module is interconnected with the output module via a wireless network. The output module is interconnected with the parsing module and the control module via a wireless network.

[0097] In this embodiment, the acquisition module collects the 3D topology, ambient light intensity, and dynamic obstacle displacement data of the target real-world scene, generating a basic data matrix to be fused. The upload module then uploads and stores preset virtual scene data. The processing module further receives the basic data matrix to be fused and the preset virtual scene data, and generates a fusion data set of the virtual-real fusion scene through spatial coordinate calibration and data feature matching. The output module obtains the fusion data set, performs real-time scene lighting and shadow calculation and image synthesis, and outputs continuous fusion scene images to the virtual reality display device. The parsing module captures user body movements and voice command signals, analyzes the corresponding interaction requirements, and drives virtual elements in the fusion scene to perform matching response actions. Finally, the control module monitors the operating parameters and data transmission rate of each module in real time, and adjusts the resource allocation of each module in real time according to the scene rendering quality and interaction response latency.

[0098] In the above embodiments, the system can accurately capture key information and dynamic changes in real-world scenes, achieve natural integration of virtual and real scenes, make lighting and shadow effects fit the real environment, accurately respond to user interaction needs, and dynamically optimize operating resources to ensure smoothness and rendering quality, greatly enhance the immersive experience, adapt to diverse usage scenarios, and make virtual-real interaction more natural, efficient and stable.

[0099] Example 2:

[0100] At the implementation level, based on Example 1, this example refers to... Figure 2 A further detailed description of a virtual reality scene fusion system in Example 1 is provided below:

[0101] A virtual reality scene fusion method, comprising:

[0102] The system collects 3D topological structure, ambient light intensity, and dynamic obstacle displacement data of the target real-world scene. After feature extraction, noise removal, and trajectory prediction, it generates a basic data matrix to be fused.

[0103] Upload and save preset virtual scene data containing scene topology, rendering parameters, and interactive element attribute data;

[0104] Coordinate unification is achieved by constructing a coordinate transformation matrix using at least three non-collinear fixed reference points in a real-world scenario, and data matching is completed through multi-dimensional feature similarity calculation to generate a fused data set.

[0105] The intensity of light and shadow at each pixel is calculated based on the ambient light intensity benchmark and virtual scene parameters. The resulting image is then synthesized into a continuous and blended scene image by layer overlay according to the light and shadow weights and output to the virtual reality display device.

[0106] The system collects user body movements and voice command signals, extracts features, analyzes interaction requirements, and then drives the corresponding virtual elements in the fusion scene to execute matching response actions.

[0107] System Application Example 3:

[0108] A user is conducting a virtual reality adventure experience in their living room. This system is used to integrate the real living room environment with a virtual forest adventure scene.

[0109] After the system starts, the acquisition module uses a multi-view image acquisition unit and a depth sensing unit to acquire surface feature points and depth information of objects such as walls, coffee tables, and sofas in the living room. After calculation and screening, key feature points for constructing a three-dimensional topological structure are selected to form corresponding basic data. At the same time, the light intensity values ​​of multiple sampling points in the living room are acquired through a full-spectrum light intensity sensor. After removing noise through mean filtering, clear ambient light intensity data is obtained. For a moving pet in the living room, the module obtains its outline through inter-frame difference of consecutive frames and calculates the pet's real-time displacement trajectory data by combining Kalman filtering. Finally, the data is integrated to generate a basic data matrix to be fused.

[0110] The upload module completes the uploading and storage of virtual forest scene data in advance. This data includes the forest scene topology, rendering parameters, and attribute information of interactive elements such as virtual animals and guides.

[0111] After receiving the basic data matrix to be fused and the virtual forest scene data, the processing module uses three non-collinear fixed points—the corner of the living room wall, the TV stand, and the corner of the bookshelf—as calibration benchmarks to obtain their coordinates in the real and virtual scenes, and constructs a coordinate transformation matrix to achieve coordinate unification. Subsequently, it extracts features such as the topological node connection density, light intensity spectral distribution, and pet movement attributes of the real scene, as well as dimensions such as the topological node type, color space parameters, and virtual element response triggering features of the virtual scene. After assigning weights to each feature, it calculates the similarity and determines that the similarity of all dimensions meets the standard, thus generating a fused dataset.

[0112] The output module is based on the fusion dataset. Taking the light intensity of the real living room as the benchmark, it combines the rendering parameters of the virtual forest and the surface reflection coefficient of objects in the scene to calculate the light and shadow intensity of each pixel. By overlaying layers, the virtual forest image and the real living room image are fused according to the light and shadow intensity weight. Since the light intensity of the living room is moderate, the fusion weight of the two is kept balanced. The continuous image is output to the virtual reality display device, presenting a fusion effect in which the virtual forest is naturally connected with the living room wall and coffee table.

[0113] The parsing module collects joint angle and acceleration data through the user's wearable inertial measurement unit and constructs a three-dimensional motion model by combining it with the visually captured limb contours. At the same time, it collects the user's voice command to "summon the virtual guide" through the sound pickup unit, extracts features after noise reduction, and determines the validity of the interaction intent by comparing it with the association mapping library, thereby driving the virtual guide to appear in the fusion scene and respond.

[0114] The control module monitors the CPU utilization, memory usage, cache queue length, and data transfer rate of each module in real time. Based on the scene rendering quality score and interaction response latency, it prioritizes the allocation of running resources to the output and processing modules to ensure pixel fidelity, lighting and shadow matching, stable frame rate, smooth transition between virtual and real boundaries, and zero interaction response, thus guaranteeing a smooth overall experience.

[0115] In summary, the systems and methods described above achieve efficient unification and precise feature matching of virtual and real scene coordinates by accurately capturing information on the spatial structure, light intensity, and dynamic obstacles of the real scene. Their lighting effects closely match the real environment and dynamically adjust with light intensity, exhibiting high detail fidelity and smooth boundary transitions. They can quickly analyze user gestures and voice interaction needs, promptly triggering virtual element responses. Simultaneously, they adapt to the system's operating status in real time, optimizing resource allocation to ensure stable frame rates and low interaction latency. This effectively enhances the immersiveness and adaptability of virtual-real fusion. Furthermore, this technical solution is adaptable to diverse usage scenarios, providing users with a natural, smooth, and believable virtual reality experience. By also considering scene presentation quality and interaction response efficiency, it further improves the intelligence level of virtual reality scene fusion.

[0116] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A virtual reality scene fusion system, characterized in that, include: The acquisition module is used to acquire the three-dimensional topology, ambient light intensity, and dynamic obstacle displacement data of the target real scene, and generate a basic data matrix to be fused. The upload module is used to upload and store preset virtual scene data. The processing module is used to receive the basic data matrix to be fused and the preset virtual scene data, and generate a fused data set of virtual and real fusion scene through spatial coordinate calibration and data feature matching; The output module is used to acquire the fused data set, perform real-time calculation of scene lighting and shadow and image compositing, and output continuous fused scene images to the virtual reality display device; The parsing module is used to capture user body movements and voice command signals, parse the corresponding interaction requirements of the signals, and drive virtual elements in the fusion scene to perform matching response actions; The control module is used to monitor the operating parameters and data transmission rate of each module in the system in real time, and adjust the allocation of operating resources of each module in real time according to the scene rendering quality and interaction response latency. The preset virtual scene data includes scene topology, rendering parameters, and interactive element attribute data.

2. The virtual reality scene fusion system according to claim 1, characterized in that, When the acquisition module acquires the 3D topological structure of the target real-world scene, it obtains the surface feature point set and depth information of each object in the target real-world scene through a preset multi-view image acquisition unit and depth sensing unit. Based on the feature point association and depth value fitting, it generates 3D topological structure data and simultaneously performs feature extraction of the 3D topological structure data. The feature extraction of the real-time 3D topological structure follows the following rules: ; In the formula: The key topological feature value of a single feature point in three dimensions; This represents the total number of feature points collected. The weighting coefficients are the spatial location weighting coefficients of feature points and the weighting coefficients for the influence of depth information. Let be the three-dimensional spatial coordinates of the i-th feature point; Let be the 3D spatial depth value of the i-th feature point; The number of neighboring feature points of the i-th feature point; Let be the Euclidean distance between the i-th feature point and the j-th neighboring feature point; This is a preset minimum constant; After obtaining the feature threshold, a feature threshold is set. If T is not less than the feature threshold, the feature point is determined to be a key feature point for constructing the three-dimensional topology and is retained; otherwise, it is discarded.

3. The virtual reality scene fusion system according to claim 1, characterized in that, During the stage of acquiring ambient light intensity data, the acquisition module acquires light intensity values ​​at multiple sampling points in the target real scene through a full-spectrum light intensity sensor, and removes noise through mean filtering. During the stage of acquiring dynamic obstacle displacement data, the acquisition module obtains the obstacle contour based on the inter-frame difference of consecutive frame images, and predicts the obstacle's displacement trajectory using Kalman filtering. The displacement data is obtained through prediction. ; In the formula: Let k be the displacement value of the dynamic obstacle at time k. For predicting weighting coefficients; This represents the displacement value at time k-1. Let be the instantaneous velocity of the obstacle at time k-1; The time interval between the acquisition of two adjacent image frames; Let be the displacement value at time t; This is the statistical window length for historical displacement data.

4. The virtual reality scene fusion system according to claim 1, characterized in that, When the processing module performs spatial coordinate calibration, it uses at least three pre-set non-collinear fixed reference points in the target real scene as calibration benchmarks to obtain the real spatial coordinates and virtual scene coordinates of each reference point, constructs a coordinate transformation matrix, and then uses the transformation matrix to unify the coordinates of the basic data matrix to be fused with the pre-set virtual scene data. The coordinate calibration follows the following rules: ; In the formula: Let X be the calibrated coordinate vector (x', y', z'); A is a 3×3 rotation and scaling matrix; X is the original coordinate vector (x, y, z) before calibration; and B is a 3×1 translation vector.

5. A virtual reality scene fusion system according to claim 1, characterized in that, When the processing module performs data feature matching, it extracts the real feature dimension from the basic data matrix to be fused and the virtual feature dimension from the preset virtual scene data. Among them, the real-world feature dimension includes the node connection density of the three-dimensional topology, the spectral distribution characteristics of ambient light intensity, and the motion attribute characteristics of dynamic obstacles, while the virtual feature dimension includes the node type parameters of the scene topology, the color space parameters of the rendering parameters, and the response triggering characteristics of interactive elements. Weights are assigned based on the importance of each feature dimension, and data feature matching is performed by calculating feature similarity. A match is considered successful only when the feature similarity of each dimension is not lower than a preset similarity threshold, triggering the generation of the fused data set. Among them, feature similarity ; In the formula: This represents the total number of feature dimensions involved in the matching. The importance weight of the i-th feature dimension; Let be the normalized feature value of the i-th real-world feature dimension; Let be the normalized feature value of the i-th virtual feature dimension; The maximum value of all real feature sample values ​​and virtual feature sample values ​​under the i-th feature dimension; It is the minimum value of all real feature sample values ​​and virtual feature sample values ​​under the i-th feature dimension; Let V be the variance of the actual feature sample values ​​under the i-th feature dimension; For preset small constants; This is a preset small constant.

6. A virtual reality scene fusion system according to claim 1, characterized in that, The output module performs a real-time scene lighting and shadow calculation stage. It uses the ambient light intensity data in the fused data set as the real lighting and shadow benchmark, and combines the rendering parameters of the virtual scene with the surface reflection coefficient and diffuse reflection coefficient of the objects in the fused scene to calculate the lighting and shadow intensity of each pixel. Among them, light and shadow intensity ; In the formula: The weighting coefficient is used to determine the contribution of light intensity in the real environment. The reference value of ambient light intensity obtained by the acquisition module; This is the specular reflection coefficient of the object surface corresponding to the target pixel. The angle between the incident direction of ambient light and the normal vector of the object's surface; This is the diffuse reflectance coefficient of the object's surface; The number of virtual light sources affecting this pixel in the virtual scene; Let be the luminous intensity of the nth virtual light source; The illumination weight of the nth virtual light source for that pixel; This represents the virtual spatial distance from the pixel to the nth virtual light source. This is the adjustment constant; When the output module performs image compositing, it merges the virtual scene image with the real scene image by layer overlay according to the light and shadow intensity weight of each pixel; The fusion weight of virtual scene images and real scene images is dynamically adjusted according to the change of ambient light intensity; that is, the higher the light intensity, the greater the fusion weight of the real scene images.

7. A virtual reality scene fusion system according to claim 1, characterized in that, When the parsing module captures the user's limb movement signals, it collects the angle change data and acceleration data of each joint of the user through the pre-worn inertial measurement unit, and combines it with the limb contour data obtained by the visual capture unit to construct a three-dimensional motion model of the user's limbs. When the parsing module captures voice command signals, it collects voice data through a preset pickup unit. Furthermore, during the real-time analysis of interaction requirements, a mapping library of body movement features and voice command features is established. Based on feature matching degree and association confidence, the user's interaction intent is determined, driving the corresponding virtual elements in the fusion scene to perform response actions. Among them, when the association confidence level is not lower than the preset confidence threshold, the interaction intent is determined to be valid.

8. A virtual reality scene fusion system according to claim 1, characterized in that, The system operating parameters monitored in real time by the control module include the CPU utilization rate, memory usage rate, and data cache queue length of each module. The data transmission rate includes the data interaction rate between modules and the data stream output rate to the virtual reality display device. Based on the preset scene rendering quality score and interaction response latency threshold, a resource allocation adjustment model is constructed. When the scene rendering quality score is lower than the preset score threshold or the interaction response latency is higher than the preset latency threshold, the CPU usage ratio and memory allocation are adjusted according to module priority.

9. A virtual reality scene fusion system according to claim 1, characterized in that, The acquisition module is interconnected with the upload module and the processing module via a wireless network. The processing module is interconnected with the output module via a wireless network. The output module is interconnected with the parsing module and the control module via a wireless network.

10. A virtual reality scene fusion method, wherein the method is an implementation method of a virtual reality scene fusion system as described in any one of claims 1-9, characterized in that, include: The system collects 3D topological structure, ambient light intensity, and dynamic obstacle displacement data of the target real-world scene. After feature extraction, noise removal, and trajectory prediction, it generates a basic data matrix to be fused. Upload and save preset virtual scene data containing scene topology, rendering parameters, and interactive element attribute data; Coordinate unification is achieved by constructing a coordinate transformation matrix using at least three non-collinear fixed reference points in a real-world scenario, and data matching is completed through multi-dimensional feature similarity calculation to generate a fused data set. The intensity of light and shadow at each pixel is calculated based on the ambient light intensity benchmark and virtual scene parameters. The resulting image is then synthesized into a continuous and blended scene image by layer overlay according to the light and shadow weights and output to the virtual reality display device. The system collects user body movements and voice command signals, extracts features, analyzes interaction requirements, and then drives the corresponding virtual elements in the fusion scene to execute matching response actions.