Scene analysis method and system based on virtual reality
Through multi-scale hierarchical sampling, autoencoder reconstruction and hierarchical Transformer feature fusion, combined with multi-channel context attention and convolutional neural network, the problem of insufficient expression of complex spatial relationships in point cloud analysis is solved, and higher analysis accuracy and classification accuracy are achieved.
Patent Information
- Application Number
- CN202510702772.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The prior art is difficult to fully express the complex spatial relationships of the scene when analyzing point clouds, lacks hierarchical, multi-scale and dynamic adaptive processing mechanisms, and has limited ability to express multi-category semantic structures.
Multi-scale hierarchical sampling and masking mechanism, autoencoder reconstruction, hierarchical Transformer multi-scale feature fusion and multi-channel context attention-driven global feature generation are adopted, and scene analysis is performed through convolutional neural network combining multi-scale features and global features.
It effectively improves the expression adequacy and intelligent analysis accuracy of point cloud scenes, and improves the accuracy of scene classification.
Smart Images

Figure CN120236145A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual reality analysis, and in particular, to a method and system for scene analysis based on virtual reality. Background Art
[0002] With the rapid development of virtual reality (VR) technology, scene analysis plays an increasingly important role in many fields such as 3D perception, digital twin, and intelligent manufacturing. As the main data carrier of scene space information, 3D point clouds are widely used in the real reconstruction and intelligent analysis of virtual reality scenes due to the maturity of multi-sensor technologies such as lidar, structured light, and depth cameras. In recent years, scene analysis methods based on point clouds have evolved from traditional geometric feature extraction and rule-based methods to the integration of multiple technical means such as deep learning and graph neural networks. Existing mainstream research mainly focuses on tasks such as point cloud segmentation, object detection, scene classification, and relationship parsing. Some methods introduce strategies such as multi-scale sampling, neighborhood feature aggregation, attention mechanism, and global context embedding to address the sparsity and disorder of point clouds. For example, network architectures such as PointNet series, DGCNN, and PointTransformer have promoted the continuous improvement of the accuracy and efficiency of point cloud structure understanding. However, the current related technologies still have deficiencies. Existing technologies are difficult to fully express the complex spatial relationships of scenes when analyzing point clouds, lack hierarchical, multi-scale, and dynamic adaptive processing mechanisms, and have limited expression ability for multi-category semantic structures. Summary of the Invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides a method and system for scene analysis based on virtual reality, which solves the problems that existing technologies are difficult to fully express the complex spatial relationships of scenes when analyzing point clouds, lack hierarchical, multi-scale, and dynamic adaptive processing mechanisms, and have limited expression ability for multi-category semantic structures.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for scene analysis based on virtual reality, which includes using a sensor device to acquire scene 3D point cloud data in real time and perform preprocessing; Perform multi-scale hierarchical sampling on the three-dimensional point cloud data. Set the neighborhood radius for each layer of the sampled point set, and apply a masking operation to each layer to obtain visible points and masked points. Calculate the neighborhood scale based on the neighborhood radius to generate a neighborhood point cloud set. Perform multi-scale neighborhood modeling on the visible points to extract neighborhood feature vectors, and use a convolutional autoencoder to encode the neighborhood feature vectors and then fuse them to obtain the final embedded features. Reconstruct the masked points through the autoencoder combined with the final embedded features of the visible points, extract the reconstructed embedded features of the masked points, and perform multi-level Transformer feature fusion to obtain multi-scale features; Assign point labels to the points in the point cloud to construct semantic context relationships. Perform weighted aggregation through multi-channel context attention weights to output the context-enhanced features of the points, and splice through multiple layers of semantics to obtain the fused high-order features of the points. Perform fusion through a two-layer perceptron and fusion channel weights, and then use a global pooling operation to obtain global features; Perform scene analysis through a convolutional neural network combined with multi-scale features and global features, output the scene category, and visualize the analysis results.
[0006] As a preferred solution of the scene analysis method based on virtual reality according to the present invention, wherein: the reconstruction of the masked points through the autoencoder combined with the final embedded features of the visible points refers to obtaining the set of final embedded features visible in each layer for recording, searching for the nearest visible points among the masked points in each layer and extracting the final embedded features of the visible points , and using the final embedded features of the visible points and the set of final embedded features of all layers below the i-th layer as the input to the decoder, and fitting the spatial coordinates of the masked points through the decoder and using the Chamfer Distance as the loss metric standard.
[0007] As a preferred solution of the scene analysis method based on virtual reality according to the present invention, wherein: the assignment of point labels to the points in the point cloud to construct semantic context relationships and the output of the context-enhanced features of the points through weighted aggregation by multi-channel context attention weights refer to performing K-means clustering on the preprocessed three-dimensional point cloud data, and assigning clustering labels to each point according to the clustering results , embedding the clustering labels into the aggregated features of each point, and calculating the multi-channel context attention weights of each point ; After weighted aggregation through multi-channel context attention weights, obtain the context-enhanced features of each point .
[0008] As a preferred solution of the virtual reality-based scene analysis method of the present invention, wherein: the fused high-order features of points are obtained by splicing through multiple layers of semantics, and after fusion through a two-layer perceptron and fusion channel weights, global pooling operation is used to obtain global features, which means obtaining the context reinforcement features of each point in each layer based on the sampling point set of each layer, and splicing the context reinforcement features of the corresponding points of all layers according to the corresponding relationship of the sampling points to obtain each point 's fused high-order features , calculating the channel mean value z of the point cloud data; Generating fusion channel weights s through a two-layer perceptron; Using the fusion channel weights s to perform weighted output on the fused high-order features of each point ; ; Performing maximum pooling and average pooling operations on the weighted outputs of all point cloud points respectively, and then splicing to obtain the global scene feature vector ; .
[0009] As a preferred solution of the virtual reality-based scene analysis method of the present invention, wherein: the scene analysis output of the scene category by combining multi-scale features and global features through a convolutional neural network means constructing a convolutional neural network, setting the input as the multi-scale features and global scene feature vector of the point cloud data, and the output as the scene category; Training the convolutional neural network with the training data set, defining the loss function and Adam optimizer to iteratively optimize the parameters of the convolutional neural network until the loss converges; Inputting the multi-scale features and global scene feature vector into the trained convolutional neural network to obtain the scene category.
[0010] As a preferred solution of the virtual reality-based scene analysis method of the present invention, wherein: the real-time acquisition and preprocessing of the scene three-dimensional point cloud data by using a sensor device means using a binocular stereo camera and an inertial measurement unit to real-time acquire the three-dimensional point cloud data of the scene and perform denoising and normalization processing.
[0011] As a preferred solution of the virtual reality-based scene analysis method of the present invention, wherein: the visualization display of the analysis result means forming a visualization table of the analyzed scene category and the collected point cloud data for display, and synchronously storing the analysis result and the point cloud data in a database for backup.
[0012] In a second aspect, the present invention provides a virtual reality-based scene analysis system, including, A point cloud acquisition module, configured to collect scene point cloud data in real time through a sensor and perform preprocessing; The multi-scale feature extraction module is used to perform multi-scale hierarchical sampling on the three-dimensional point cloud data, apply a masking operation to each layer to obtain visible points and masked points, perform multi-scale neighborhood modeling on the visible points to extract neighborhood feature vectors, and reconstruct the masked points through an autoencoder, extract the masked point features, and perform multi-level Transformer feature fusion to obtain multi-scale features; The global feature extraction module is used to assign point cloud point labels to construct semantic context relationships, output point features through multi-channel context attention weight assignment, and obtain global features through multi-level semantic fusion; The scene analysis module is used to perform scene analysis through a convolutional neural network in combination with multi-scale features and global features, output the scene category, and visually display the analysis results.
[0013] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the virtual reality-based scene analysis method as described in the first aspect of the present invention is implemented.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the virtual reality-based scene analysis method as described in the first aspect of the present invention is implemented.
[0015] The beneficial effects of the present invention are as follows: By real-time sampling the scene point cloud, adopting multi-scale hierarchical sampling and masking mechanism, autoencoder reconstruction, hierarchical Transformer multi-scale feature fusion, and multi-channel context attention-driven global feature generation, the present invention effectively improves the expression sufficiency and intelligent analysis accuracy of the point cloud scene, and improves the accuracy of scene classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0017] Figure 1 It is a flowchart of the virtual reality-based scene analysis method in Embodiment 1.
[0018] Figure 2 It is a structural diagram of the virtual reality-based scene analysis system in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following detailed description of the specific embodiments of the present invention will be provided in conjunction with the accompanying drawings of the specification.
[0020] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art may make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0021] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that excludes other embodiments.
[0022] Example 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides a method for scene analysis based on virtual reality, including the following steps: S1. Use a sensor device to continuously obtain three-dimensional point cloud data of the scene and perform preprocessing; Specifically, using a sensor device to continuously obtain three-dimensional point cloud data of the scene and perform preprocessing means using a binocular stereo camera and an inertial measurement unit to continuously obtain the three-dimensional point cloud data of the scene and perform denoising and normalization processing.
[0023] S2. Perform multi-scale hierarchical sampling on the three-dimensional point cloud data. Set a neighborhood radius for each layer of sampled point set, and apply a masking operation to each layer to obtain visible points and masked points. Calculate the neighborhood scale in combination with the neighborhood radius to generate a neighborhood point cloud set. Extract neighborhood feature vectors through multi-scale neighborhood modeling based on the visible points, and use a convolutional autoencoder to encode and fuse the neighborhood feature vectors to obtain the final embedded feature. Reconstruct the masked points through the autoencoder in combination with the final embedded feature of the visible points, and extract the reconstructed embedded feature of the masked points for multi-level Transformer feature fusion to obtain multi-scale features; Specifically, multi-scale hierarchical sampling is performed on the three-dimensional point cloud data. For each layer of the sampled point set, a neighborhood radius is set, and a masking operation is applied to each layer to obtain visible points and masked points. The neighborhood scale is generated by combining the neighborhood radius to calculate the neighborhood point cloud set. Based on the visible points, multi-scale neighborhood modeling is performed to extract neighborhood feature vectors, and a convolutional autoencoder is used to encode the neighborhood feature vectors and then fuse them to obtain the final embedded feature. The masked points are reconstructed by the autoencoder in combination with the final embedded feature of the visible points, and the reconstructed embedded feature of the masked points is extracted for multi-level Transformer feature fusion to obtain multi-scale features. According to the preprocessed three-dimensional point cloud data, Farthest Point Sampling is used to obtain the sampled point set of each layer :
[0024] Among them is the point set of the previous layer, initially the preprocessed three-dimensional point cloud data set, is the target number of sampled points in the i-th layer, is to iteratively select mutually farthest representative points in the input point set; For the sampled point set of each layer, a neighborhood radius for each layer is set :
[0025] Among them is the initial neighborhood radius, is the radius increment coefficient; In the sampled point set of each layer, a random masking operation is performed according to the set masking ratio to obtain the visible point subset and the masked point subset :
[0026] Among them is the masking ratio of the i-th layer, is the random masking function; The masking ratio is set according to the formula, expressed as:
[0027] Among them is the lowest layer masking ratio, is the increment amplitude of each layer; According to the neighborhood radius of each layer Combined with the multi-scale coefficient to generate the neighborhood scale of each layer :
[0028] Among them are multi-scale coefficients, such as 1, 2, 4, etc., and m is the scale label; Through the neighborhood scale, for each subset of visible points in each layer For each visible point in Generate a neighborhood point cloud set :
[0029] where is the j-th point in the sampling point set of the i-th layer; According to the relative coordinates between each point in the neighborhood point cloud set and the visible point Construct the neighborhood feature vector for each scale :
[0030] where is the relative coordinate between point pairs; Encode the neighborhood feature vector of each visible point at each scale through a convolutional autoencoder to obtain the depth embedding feature , and use weighted fusion to fuse the depth embedding features of each scale to obtain the final embedding feature of each visible point ; Obtain the set of final embedding features visible for each layer Record, search for the nearest visible point for the masked points in each layer and extract the final embedding feature of the visible point , and use the final embedding feature of the visible point and the set of all final embedding features below the i-th layer Input to the decoder, and fit the spatial coordinates of the masked points through the decoder and use ChamferDistance as the loss metric:
[0031] where is the set of masked points obtained by fitting, a and b are the masked points obtained by fitting and the true masked points respectively, is the loss; Obtain the reconstructed embedding feature based on the reconstructed masked points obtained by fitting and determine the neighborhood, calculate the neighborhood localization attention score for each point in each layer :
[0032] where and are the query vector and the key vector respectively, is the key vector of the p-th point, is the dimension of the projection vector, determined by the embedded features; Use the neighborhood localization attention scores of each point to perform weighted aggregation on each point to obtain the fused feature output of the point :
[0033] where is the value vector; Calculate the Squeeze-Excitation weights of the fused feature output of each layer of points and perform weighted output on the fused feature output to obtain the aggregated feature :
[0034]
[0035] where is the channel-wise multiplication; Concatenate the aggregated features of all points in each layer to obtain the multi-scale features .
[0036] Multi-scale hierarchical sampling and masking operations not only form a hierarchical and progressive spatial perception network, but also provide a strongly robust input for irregular point cloud data in real application scenarios. For example, problems such as occlusion, missing, or density change commonly seen in virtual reality scenarios can be transformed into effective information redundancy and complementation relationships under the mechanism of the present invention, leaving redundant space for multi-level feature expression. At the same time, by setting random masks instead of fixed-point removal, sampling bias can be maximally prevented, making subsequent feature learning more generalizable. Secondly, the combination of multi-scale neighborhood modeling and relative coordinate feature extraction lays a foundation for the efficient representation of local and global geometric relationships of point clouds. Extracting neighborhood structures at different scales can not only capture the microscopic details of point clouds, but also fuse macroscopic structures, realizing hierarchical information coupling, and having inherent advantages in applications such as complex scene restoration and virtual object recognition.
[0037] By coordinating the masking mechanism with the auto-encoding network, feature restoration and spatial completion are achieved in cases of undersampling, occlusion, and data loss. This mechanism not only improves the restoration accuracy of virtual scenes but also significantly enhances the discriminative ability of feature expression through deep learning of embedded features, providing a solid foundation for subsequent high-order tasks (such as scene semantic segmentation, object detection, etc.). The decoder uses the features of the nearest visible points and a multi-layer embedding set, taking into account both local consistency and global context, avoiding the drawback of traditional interpolation being vulnerable to outliers. In addition, based on the features of the reconstructed masked points, a localized attention mechanism is introduced, which can accurately focus on the regions with closely related features in the point cloud, suppress interference, and enhance the discriminative ability, significantly making up for the shortcoming of the dynamic fusion of global and local information in point cloud analysis.
[0038] The Squeeze-Excitation weighting mechanism and multi-scale feature splicing provide a multi-level and full-channel adaptive information aggregation method for the final scene understanding task. This step not only strengthens the expression of information in each scale and semantic channel but also can adaptively allocate feature weights, improving the generalization ability of the model for different task objectives.
[0039] S3. Assign point labels to the point cloud to construct semantic context relationships, perform weighted aggregation through multi-channel context attention weights to output the context-enhanced features of points, and obtain the fused high-order features of points through multi-layer semantic splicing. After fusion through a two-layer perceptron and fused channel weights, a global pooling operation is used to obtain global features; Specifically, assigning point labels to the point cloud to construct semantic context relationships and performing weighted aggregation through multi-channel context attention weights to output the context-enhanced features of points means performing K-means clustering on the preprocessed three-dimensional point cloud data and assigning a clustering label to each point according to the clustering result , embedding the clustering label into the aggregated feature of each point, and calculating the multi-channel context attention weights of each point :
[0040] where is the projection weight, MLP is the multi-layer perceptron, is the aggregated feature after embedding the clustering label of the i-th point, and are the aggregated features after embedding the clustering labels of the j-th and k-th points in the neighborhood of the i-th point, respectively; After weighted aggregation through multi-channel context attention weights, the context-enhanced features of each point are obtained :
[0041] where To improve the feature expression ability of neighborhood points after transformation through MLP.
[0042] K-means clustering and label embedding break the traditional point cloud processing mode that only depends on Euclidean space distribution or single attribute aggregation. By semantically partitioning the point cloud space and embedding high-dimensional labels, when each point updates its features, it can obtain, in addition to its own attributes and position, the distribution in the higher-dimensional semantic space. This processing is particularly crucial for complex situations where there are category transitions, blurred boundaries, or multiple objects nested in the scene. For example, in virtual reality scene analysis, objects that are adjacent but have obvious morphological differences can be distinguished by semantic labels, reducing cross-class mis-clustering. For point clouds in multi-scale, dense, or sparse regions, semantic labels provide a more stable intermediary for the attribution and relationship network of points, effectively improving the robustness of subsequent context awareness and feature transmission. This concept is significantly superior to traditional local methods such as only Jaccard or distance neighborhood.
[0043] By learning the "correlation weights" between points in the neighborhood with the clustering label as the dominant feature, each point finally appears as an information expression body that fuses the information of its strongly correlated points. Different from the current mainstream single modeling methods of spatial attention or channel attention, the multi-channel scheme can parallelly strengthen information interaction in both the feature channel and neighborhood semantics spaces. The artificially designed aggregation weight mode often has performance bottlenecks due to the lack of global context and difficulty in dynamically adapting to scene changes, while the proposed scheme of the present invention can achieve adaptive weighting for semantic differences and spatial structures, fully highlighting the point features with key contributions in the context environment, automatically suppressing redundant or low-correlation information, and significantly reducing the interference of irrelevant noise.
[0044] Furthermore, the fused high-order features of points are obtained by splicing through multiple layers of semantics. After fusion through a two-layer perceptron and fused channel weights, a global pooling operation is used to obtain global features, which means obtaining the context-enhanced features of each point in each layer based on the sampled point set of each layer, and splicing the context-enhanced features of the corresponding points of all layers according to the corresponding relationship of the sampled points to obtain the fused high-order features of each point of each point , calculating the channel mean z of the point cloud data:
[0045] where N is the number of points in the point cloud; Generating fused channel weights s through a two-layer perceptron:
[0046] where and are the weights of the fully connected layers respectively, and are the biases, is the ReLU activation function, is the Sigmoid function; By fusing the channel weight s with the fused high-order features at each point perform weighted output :
[0047] where is channel-wise multiplication; The weighted output of all point cloud points uses max pooling and average pooling operations respectively and then concatenates them to obtain the global scene feature vector .
[0048] By obtaining context-enhanced features on the sampled point sets at each layer, it can be compatible with semantic and structural information at different spatial scales, effectively covering two types of difficulties that may exist in the point cloud: local high complexity and loose global connection. Compared with traditional methods that only use single-scale, single-semantic or basic geometric features, the present invention grafts multi-level expressions onto a single point, significantly improving the semantic integrity and discrimination of point cloud analysis. For example, in large-scale indoor and outdoor scenes, semantic cues from complex object components, overall structures to local details are all captured together, greatly reducing information gaps and spatial ambiguities, and having irreplaceable comprehensiveness in practical applications such as multi-object discrimination and anomaly detection.
[0049] Automatically generate global channel weights through a two-layer perceptron, realizing the adaptive regulation of the importance of feature channel expressions. This mechanism can effectively highlight the feature channels that contribute more to discrimination, while suppressing redundant or noisy features, achieving the best balance between semantic compression and multi-dimensional preservation. The benefit is that the model can automatically complete optimal information extraction according to the actual data distribution in various scene segmentation, object recognition, and high-order spatial reasoning tasks, thereby obtaining higher generalization ability and robustness.
[0050] S4. Perform scene analysis through a convolutional neural network combined with multi-scale features and global features, output the scene category, and visually display the analysis result; Specifically, performing scene analysis through a convolutional neural network combined with multi-scale features and global features and outputting the scene category means constructing a convolutional neural network, setting the input as the multi-scale features and global scene feature vector of the point cloud data, and the output as the scene category; Train the convolutional neural network through a training data set, define a loss function and an Adam optimizer to iteratively optimize the parameters of the convolutional neural network until the loss converges; Input the multi-scale features and global scene feature vector into the trained convolutional neural network to obtain the scene category.
[0051] As a structured deep learning method, the convolutional neural network is different from the point-by-point or neighborhood-by-neighborhood static operations of traditional point cloud analysis. Its multi-core convolutional layer can automatically explore local and global high-order correlations in spatial relationships, and is particularly suitable for feature extraction and discrimination of three-dimensional point cloud data that is highly redundant, disordered, and lacks topological structure. By inputting multi-scale features and global scene feature vectors into the network together, the model can take into account both detail discrimination and global structure understanding in the same inference process, greatly improving the recognition accuracy and generalization ability for complex scene categories.
[0052] Taking the convolutional neural network optimized repeatedly as a scene classifier essentially realizes end-to-end high-order semantic mapping, that is, through the fusion of input multi-scale and global features, it outputs the category discrimination of the world-class complex environment. This is particularly important for fields such as virtual reality, digital twins, and smart city perception.
[0053] Furthermore, visualizing the analysis results means forming a visualization table of the analyzed scene categories and the collected point cloud data for display, and synchronously storing the analysis results and point cloud data in a database for backup.
[0054] This embodiment also provides a scene analysis system based on virtual reality, including: A point cloud acquisition module, which is used to collect scene point cloud data in real time through a sensor and perform preprocessing; A multi-scale feature extraction module, which is used to perform multi-scale hierarchical sampling on three-dimensional point cloud data, apply a masking operation to each layer to obtain visible points and masked points, extract neighborhood feature vectors based on the visible points for multi-scale neighborhood modeling, and perform masked point reconstruction through an autoencoder to extract masked point features for multi-level Transformer feature fusion to obtain multi-scale features; A global feature extraction module, which is used to assign point cloud point labels to build semantic context relationships, output point features through multi-channel context attention weight allocation, and obtain global features through multi-level semantic fusion; A scene analysis module, which is used to perform scene analysis through a convolutional neural network combined with multi-scale features and global features to output scene categories and visualize the analysis results.
[0055] This embodiment also provides a computer device applicable to the case of the scene analysis method based on virtual reality, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the scene analysis method based on virtual reality proposed in the above embodiment.
[0056] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad set on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0057] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for realizing scene analysis based on virtual reality as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk, or optical disc.
[0058] The present invention effectively improves the expression sufficiency and intelligent analysis accuracy of the point cloud scene, and improves the accuracy of scene classification by real-time sampling of the scene point cloud, adopting a multi-scale hierarchical sampling and masking mechanism, auto-encoder reconstruction, hierarchical Transformer multi-scale feature fusion, and multi-channel context attention-driven global feature generation.
[0059] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for scene analysis based on virtual reality, characterized in that: including using a sensor device to obtain 3D point cloud data of a scene in real time and perform preprocessing; performing multi-scale hierarchical sampling on the 3D point cloud data, setting a neighborhood radius for each layer of sampled point sets, applying a masking operation to each layer to obtain visible points and masked points, calculating the neighborhood scale based on the neighborhood radius to generate a neighborhood point cloud set, extracting neighborhood feature vectors through multi-scale neighborhood modeling based on the visible points, encoding the neighborhood feature vectors using a convolutional autoencoder and then fusing them to obtain the final embedded features, reconstructing the masked points through the autoencoder combined with the final embedded features of the visible points, and extracting the reconstructed embedded features of the masked points for multi-level Transformer feature fusion to obtain multi-scale features; assigning point cloud point labels to construct semantic context relationships, performing weighted aggregation through multi-channel context attention weights to output context-enhanced features of the points, and obtaining fused high-order features of the points through multi-layer semantic concatenation, and performing fusion through a two-layer perceptron and fused channel weights and then using a global pooling operation to obtain global features; performing scene analysis through a convolutional neural network combined with multi-scale features and global features to output the scene category and visually display the analysis results.
2. The method for scene analysis based on virtual reality according to claim 1, wherein: The above-mentioned method of reconstructing masked points by combining the final embedding features of visible points using an autoencoder refers to obtaining the set of final embedding features visible in each layer for recording. Search for the nearest visible points among the masked points in each layer and extract the final embedding features of the visible points . Take the final embedding features of the visible points and the set of final embedding features of all layers below the i-th layer as the input to the decoder, and fit the spatial coordinates of the masked points through the decoder and use the Chamfer Distance as the loss metric standard.
3. The method for scene analysis based on virtual reality according to claim 2, wherein: The step of assigning point labels to the point cloud to construct a semantic context relationship and weighted aggregation output of the context-enhanced features of the points through multi-channel context attention weights refers to performing K-means clustering on the preprocessed three-dimensional point cloud data and assigning a clustering label to each point according to the clustering result , embedding the clustering label into the aggregated feature of each point, and calculating the multi-channel context attention weights of each point ; Obtain the context-enhanced features of each point after weighted aggregation through multi-channel context attention weights .
4. The method for scene analysis based on virtual reality according to claim 3, wherein: The fused high-order features of points obtained by splicing through multiple-layer semantics, after being fused through a two-layer perceptron and fused channel weights, are subjected to a global pooling operation to obtain global features, which means obtaining the context-enhanced features of each point in each layer based on the sampled point set of each layer, and splicing the context-enhanced features of the corresponding points of all layers according to the corresponding relationship of the sampled points to obtain each point of the fused high-order features , calculating the channel mean z of the point cloud data; generating fused channel weights s through a two-layer perceptron; The fused high-order features of each point are weighted and output by fusing the channel weight s for weighted output ; Weighted output for all point cloud points After performing max pooling and average pooling operations respectively, they are concatenated to obtain the global scene feature vector .
5. The method for scene analysis based on virtual reality according to claim 4, characterized in that: the performing scene analysis through a convolutional neural network combined with multi-scale features and global features to output the scene category means constructing a convolutional neural network, setting the input as the multi-scale features of the point cloud data and the global scene feature vector, and the output as the scene category; training the convolutional neural network using a training data set, defining a loss function and an Adam optimizer to iteratively optimize the parameters of the convolutional neural network until the loss converges; inputting the multi-scale features and the global scene feature vector into the trained convolutional neural network to obtain the scene category.
6. The method for scene analysis based on virtual reality according to claim 5, wherein: the using a sensor device to obtain 3D point cloud data of a scene in real time and perform preprocessing means using a binocular stereo camera and an inertial measurement unit to obtain the 3D point cloud data of the scene in real time and perform denoising and normalization processing.
7. The method for scene analysis based on virtual reality according to claim 6, wherein: the visually displaying the analysis results means forming a visualization table of the analyzed scene category and the collected point cloud data for display, and synchronously storing the analysis results and the point cloud data in a database for backup.
8. A virtual reality-based scene analysis system, based on the virtual reality-based scene analysis method according to any one of claims 1 to 7, characterized in that: including a point cloud acquisition module for collecting scene point cloud data through a sensor in real time and performing preprocessing; a multi-scale feature extraction module for performing multi-scale hierarchical sampling on the 3D point cloud data, applying a masking operation to each layer to obtain visible points and masked points, performing multi-scale neighborhood modeling based on the visible points to extract neighborhood feature vectors, and reconstructing the masked points through an autoencoder, and extracting the masked point features for multi-level Transformer feature fusion to obtain multi-scale features; a global feature extraction module for assigning point cloud point labels to construct semantic context relationships, distributing output point features through multi-channel context attention weights, and obtaining global features through multi-layer semantic fusion; a scene analysis module for performing scene analysis through a convolutional neural network combined with multi-scale features and global features to output the scene category and visually display the analysis results.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, the steps of the method for scene analysis based on virtual reality according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the method for scene analysis based on virtual reality according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Self-supervised point cloud completion method and device based on scene flow prior guidance
CN117764881A
Multi-view three-dimensional reconstruction system based on feature aggregation Transform
CN117765175A
Point cloud scene segmentation method fusing double neighborhood features and global space perception
CN117934840A
Point cloud classification segmentation method and system of global mask auto-encoder based on fused voxel
CN119762877A
Three-dimensional point cloud semantic segmentation method based on graph convolution and grouping vector attention mechanism
CN119963842A