Virtual environment reconstruction method and device, equipment and medium
By adding feature patterns to the natural environment map and using bag-of-word model and K-dimensional tree structure for image matching, the reconstruction failure and difficulty in positioning caused by few feature points in virtual environment reconstruction are solved, and efficient virtual environment reconstruction is achieved.
Patent Information
- Application Number
- CN202510184521.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-24
AI Technical Summary
In the reconstruction of virtual environments, if there are fewer feature points, the reconstruction failure or environmental positioning is prone to problems such as difficulty.
By adding feature patterns to the natural environment map, the descriptors are classified and clustered using the pre-trained bag-of-word model, and image matching and fusion are combined with the K-dimensional tree structure to generate an overall virtual environment map.
The accuracy and efficiency of virtual environment reconstruction are improved, the reconstruction failure caused by few feature points and the difficulty of environmental positioning is avoided, and the matching relationship between images with many feature points is achieved.
Smart Images

Figure CN120198607A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and in particular, to a method, device, equipment and medium for reconstructing a virtual environment. Background Art
[0002] VR multi-person interaction is a multi-person online interaction system based on virtual reality technology. Through this system, users can interact with other users in a virtual environment generated by a computer.
[0003] When establishing a virtual environment, the current technical solutions usually collect image data of the natural environment, process the image data, extract and match features, and finally obtain an environmental map based on structure from motion recovery.
[0004] When the current technical solutions reconstruct the environment, if there are few feature points in a certain part of the environment, problems such as reconstruction failure or difficult environment positioning are likely to occur. Summary of the Invention
[0005] The present invention provides a method, device, equipment and medium for reconstructing a virtual environment, which can reconstruct the virtual environment based on a synthetic environment map with more feature points, and solve problems such as reconstruction failure or difficult environment positioning caused by few feature points.
[0006] According to one aspect of the present invention, a method for reconstructing a virtual environment is provided, and the method includes:
[0007] Extract feature points and descriptors from each synthetic environment map to obtain the feature points and descriptors corresponding to each synthetic environment map; the synthetic environment map is obtained by adding feature patterns to the natural environment map;
[0008] Classify the descriptors in the synthetic environment map based on a pre-trained bag-of-words model, and determine the bag-of-words model vector corresponding to the synthetic environment map according to the classification result;
[0009] Cluster the bag-of-words model vectors corresponding to the synthetic environment maps, and save the clustered image set in a target data structure; the target data structure storing the image set can reflect the association information between the image sets;
[0010] Match the images located in the target data structure to obtain a matching result, and fuse the synthetic environment maps based on the matching result to obtain an overall virtual environment map.
[0011] According to another aspect of the present invention, a device for reconstructing a virtual environment is provided, including:
[0012] A feature extraction module, configured to extract feature points and descriptors from each synthetic environment image, so as to obtain the feature points and descriptors corresponding to each synthetic environment image; the synthetic environment image is obtained by adding a feature pattern to a natural environment image;
[0013] A bag-of-words model vector determination module, configured to classify the descriptors in the synthetic environment image based on a pre-trained bag-of-words model, and determine the bag-of-words model vector corresponding to the synthetic environment image according to the classification result;
[0014] An image clustering module, configured to cluster the bag-of-words model vectors corresponding to the synthetic environment images, and save the clustered image set in a target data structure; the target data structure storing the image set can reflect the association information between the image sets;
[0015] An overall virtual environment map determination module, configured to match the images located in the target data structure to obtain a matching result, and fuse the synthetic environment images based on the matching result to obtain an overall virtual environment map.
[0016] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the virtual environment reconstruction method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the virtual environment reconstruction method according to any embodiment of the present invention when executed by a processor.
[0021] The technical solution of the embodiment of the present application includes: extracting feature points and descriptors for each synthetic environment map to obtain the feature points and descriptors corresponding to each synthetic environment map; the synthetic environment map is obtained by adding a feature pattern to a natural environment map; classifying the descriptors in the synthetic environment map based on a pre-trained bag-of-words model, and determining the bag-of-words model vector corresponding to the synthetic environment map according to the classification result; clustering the bag-of-words model vectors corresponding to the synthetic environment map, and saving the clustered image set in a target data structure; the target data structure storing the image set can reflect the association information between the image sets; matching the images located in the target data structure to obtain a matching result, and fusing the synthetic environment maps based on the matching result to obtain an overall virtual environment map. Through the reconstruction of the synthetic environment map provided with the feature image, since there are many feature points therein and there are matching relationships between the images, the environment can be accurately located and the environment reconstruction can be realized.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 is a flowchart of a method for reconstructing a virtual environment provided in Embodiment 1 of the present application;
[0025] Figure 2 is a flowchart of a method for reconstructing a virtual environment provided in Embodiment 2 of the present application;
[0026] Figure 3 is a schematic diagram of a K-dimensional tree storing an image set provided in Embodiment 2 of the present application;
[0027] Figure 4 is a schematic structural diagram of a device for reconstructing a virtual environment provided in Embodiment 3 of the present application;
[0028] Figure 5 is a schematic structural diagram of an electronic device for implementing the method for reconstructing a virtual environment of the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first", "second", "target", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0031] Embodiment 1
[0032] Figure 1 The following is a flowchart of a method for reconstructing a virtual environment provided in the first embodiment of the present application. The embodiments of the present application are applicable to the situation of reconstructing a virtual environment. This method can be executed by a virtual environment reconstruction device, which can be implemented in the form of hardware and / or software, and the virtual environment reconstruction device can be configured in an electronic device with data processing capabilities. As Figure 1 shown, the method includes:
[0033] S110, extract feature points and descriptors for each synthesized environment map to obtain the feature points and descriptors corresponding to each synthesized environment map.
[0034] Among them, the synthesized environment map is obtained by adding a feature pattern to a natural environment map. The feature pattern is used to add features at positions with few feature points in the natural environment map. The feature pattern can be an artificial pattern. Exemplarily, if there are parts with few features such as white walls in the natural environment map, a feature pattern can be added to this part to obtain a synthesized environment map.
[0035] Specifically, after obtaining the synthesized environment map, the synthesized environment map can be preprocessed (such as noise reduction, etc.), and based on the preprocessed synthesized environment map, feature points and descriptors corresponding to the feature points are extracted to obtain the feature points and descriptors corresponding to each synthesized environment map.
[0036] Exemplarily, mark the synthetic environment images as: S = {I1, I2,... I N}; where I1 is the synthetic environment image with serial number 1, I2 is the synthetic environment image with serial number 2, and N is the number of synthetic environment images. For all images in S = {I1, I2,... I N}, extract the SIFT feature point positions and feature point descriptors. Denote the set of all image feature points obtained as F, and the descriptor set as W;
[0037] F = { {f1, f2,..., f m1}1, {f1, f2,..., f m2}2,... {f1, f2,..., f mi} N};
[0038] W = { {w1, w2, w3..., w m1}1, {w1, w2,..., w m2}2,... {w1, w2,..., w mi} N};
[0039] where mi is the number of feature points extracted from image I i . The feature point positions and descriptors correspond one by one, and the feature points can be encoded to ensure their uniqueness and recognizability.
[0040] S120. Classify the descriptors in the synthetic environment images based on a pre-trained bag-of-words model, and determine the bag-of-words model vector corresponding to the synthetic environment image according to the classification result.
[0041] Among them, the bag-of-words model is also known as the BOW bag-of-words model, which is usually used for tasks such as feature representation and classification of images.
[0042] Specifically, obtain the visual words of the pre-trained bag-of-words model, classify the descriptors in the synthetic environment images into the visual words to which they belong to obtain the classification result. Since there are usually multiple feature points in each synthetic environment image, and each feature point has a classified visual word, generate a bag-of-words model vector based on the number of visual words to which the feature points on the synthetic environment image are classified. Perform the above processing on each synthetic environment image to obtain the bag-of-words model vector corresponding to each synthetic environment image.
[0043] S130. Cluster the bag-of-words model vectors corresponding to the synthetic environment images, and save the clustered image set in the target data structure.
[0044] Among them, the target data structure storing the image set can reflect the association information between image sets.
[0045] Specifically, cluster the bag-of-words model vectors corresponding to all synthetic environment maps to obtain the clusters after clustering. The BOW representations in each cluster are similar. Then, determine the synthetic environment maps corresponding to the bag-of-words model vectors in each cluster as an image set. In this case, the images in the image set can be understood as similar images or images containing common features. Then, save each image set in the target data structure to perform fusion based on the association information between image sets during subsequent image fusion.
[0046] S140, match the images located in the target data structure to obtain a matching result, and fuse the synthetic environment maps based on the matching result to obtain an overall virtual environment map.
[0047] Specifically, identify the images located in the target data structure, and determine whether there is a matching relationship between the feature points of the images. If there is a matching relationship, determine the matching result according to the image information and the information of the matching feature points. Since the matching result reflects which feature points of which images are corresponding and can also reflect the association information between image sets, the synthetic environment maps can be fused based on the matching result to obtain an overall virtual environment map.
[0048] The technical solution of the embodiment of the present application includes: extracting feature points and descriptors for each synthetic environment map to obtain the feature points and descriptors corresponding to each synthetic environment map; the synthetic environment map is obtained by adding a feature pattern to a natural environment map; classifying the descriptors in the synthetic environment map based on a pre-trained bag-of-words model, and determining the bag-of-words model vector corresponding to the synthetic environment map according to the classification result; clustering the bag-of-words model vectors corresponding to the synthetic environment map, and saving the clustered image sets in the target data structure; the target data structure storing the image set can reflect the association information between image sets; match the images located in the target data structure to obtain a matching result, and fuse the synthetic environment maps based on the matching result to obtain an overall virtual environment map. Through the reconstruction of the synthetic environment map provided with feature images, since there are many feature points and there are matching relationships between images, the environment can be accurately located and the environment reconstruction can be realized.
[0049] Embodiment Two
[0050] Figure 2 It is a flowchart of a method for reconstructing a virtual environment provided by the second embodiment of the present application. The second embodiment of the present application is optimized based on the above embodiment.
[0051] Such as Figure 2As shown in the figure, the method of the embodiment of the present application specifically includes the following steps:
[0052] S210, perform noise reduction processing on the synthesized environment map through a Gaussian function to obtain a noise-reduced synthesized environment map.
[0053] Exemplarily, perform noise reduction on the image through Gaussian blur (σ = 0.5, Gaussian kernel size is 5x5). When performing noise reduction, determine the intensity of the noise-reduced image through the following formula:
[0054]
[0055] where G(x,y,σ) is the Gaussian function, I(x,y) is the image intensity at the image (x,y), the image intensity at (x,y) after Gaussian blur.
[0056] This technical solution removes the noise in the synthesized environment map, making the features in the synthesized environment map more obvious and easier to extract, improving the visual effect and usage value of the image, improving the accuracy and robustness of subsequent feature point and descriptor extraction, and ensuring the smooth completion of subsequent reconstruction.
[0057] S220, extract feature points and descriptors from each noise-reduced synthesized environment map to obtain the feature points and descriptors corresponding to each synthesized environment map.
[0058] In the embodiment of the present application, since some noun explanations and the specific content of the step operations have been elaborated in the above embodiments, they will not be elaborated too much here.
[0059] Exemplarily, when extracting feature points and descriptors (such as SIFT, etc.) from each image, each local feature descriptor can be a vector with a fixed dimension (such as SIFT is 128-dimensional), and the features extracted from one image include several feature points and the descriptors corresponding to the feature points.
[0060] S230, classify the descriptors in the synthesized environment map based on a pre-trained bag-of-words model, and determine the bag-of-words model vector corresponding to the synthesized environment map according to the classification result.
[0061] In the embodiment of the present application, optionally, classifying the descriptors in the synthesized environment map based on a pre-trained bag-of-words model and determining the bag-of-words model vector corresponding to the synthesized environment map includes: obtaining the codebook of the pre-trained bag-of-words model; calculating the distance between the descriptor and the visual words in the codebook, and classifying the descriptor into the visual word in each visual word that is the closest to the descriptor; generating the bag-of-words model vector corresponding to the synthesized environment map based on the quantity information of the visual words to which the descriptors in the synthesized environment map are classified.
[0062] Exemplarily, obtain the codebook of the pre-trained bag-of-words (BOW) model, which contains K visual words, i.e., K clustering centers. For each descriptor, calculate its distances to all the visual words in the BOW model (the distance can be the Euclidean distance). Classify each descriptor into the category of the visual word with the closest distance. For the convenience of subsequent processing, (the descriptors of each image can be quantified into the corresponding visual word index sequence) count the occurrence frequencies of each visual word in each synthetic environment image to obtain the bag-of-words model vector, and normalize the bag-of-words model vector to improve the robustness of clustering. The bag-of-words model vector can be represented in the form of a K-dimensional histogram, etc.
[0063] Exemplarily, if the codebook has 1000 visual words, then each synthetic environment image can be represented as a 1000-dimensional bag-of-words model vector, and each dimension corresponds to the word frequency of a visual word.
[0064] This technical solution represents the synthetic environment image in the form of a vector, which is convenient for subsequent classification or retrieval.
[0065] S240. Cluster the bag-of-words model vector corresponding to the synthetic environment image, and save the clustered image set in the K-dimensional tree.
[0066] Exemplarily, the bag-of-words model vector can be clustered by the K-means algorithm. Specifically, obtain the bag-of-words model vectors corresponding to all synthetic environment images (with a dimension of N×K, where N is the number of synthetic environment images and K is the codebook size). Initialize the K-means algorithm and set the target number of clusters M. Iteratively optimize the clustering centers until convergence or the maximum number of iterations is reached. Obtain M clustering centers and the cluster labels of each image (since the synthetic environment image and the bag-of-words model vector are in one-to-one correspondence, after clustering the bag-of-words model vector, the cluster information after clustering the bag-of-words model vector can be determined as the cluster label of the corresponding synthetic environment image). Allocate the synthetic environment images to the corresponding M sets according to the cluster labels. The synthetic environment images within the same cluster have similar bag-of-words model vectors, so they can be grouped into the same set.
[0067] S250. Perform descriptor matching on the images of the same layer of the K-dimensional tree from the leaf node to the root node direction. If it is determined that the descriptors in the two synthetic environment images currently being matched meet the matching requirements, record the matching relationship of the two synthetic environment images currently being matched.
[0068] Among them, the matching relationship reflects the feature point information corresponding to the matched descriptors and the information of the synthetic environment images where the matched descriptors are located. The K-dimensional tree (KD Tree, K-Dimensional Tree) is a data structure for partitioning data points in a K-dimensional space.
[0069] Figure 3Schematic diagram of a K-dimensional tree storing an image set. It can be seen that in Figure 3 , a total of M image sets are located at leaf nodes. Specifically, starting from the leaf nodes, the descriptors of the images in the K-dimensional tree are matched from bottom to top (i.e., from leaf nodes to the root node) for images in the same layer. Each time a match is performed, if the descriptors in the two synthetic environment images currently being matched meet the matching requirements, the matching relationship between the two synthetic environment images currently being matched is recorded, and this matching relationship is the matching result. Repeat the above operation until the image matching requirements are not met, that is, the matching deviation exceeds the threshold.
[0070] In an embodiment of the present application, optionally, performing descriptor matching on the images at the leaf nodes includes: pairwise matching the synthetic environment images under the same image set based on a brute-force matching method to obtain a matching result.
[0071] Exemplarily, when performing brute-force matching, assume that the two images currently being matched are image A and image B. For each feature point in image A, calculate the distance (e.g., Euclidean distance) between it and the descriptors of all feature points in image B. The point in image B that is closest to the current feature point in image A is used as a candidate matching point. If the closest distance is less than a certain threshold, it is determined that the feature point in image A matches the candidate matching point successfully, and the matching relationship between these two feature points can be recorded.
[0072] Based on the K-dimensional tree for descriptor matching in this solution, when performing matching, since the images within the image set at the leaf nodes are pairwise matched by the brute-force matching method, it avoids unnecessary calculations across sets while ensuring the accuracy of the matching. With the help of the hierarchical structure of the K-dimensional tree, when matching the nodes above the leaf nodes, the search range can be quickly narrowed, and it is not necessary to traverse all image sets. The possible matching sets can be quickly located with the help of the hierarchical structure of the K-dimensional tree.
[0073] In an embodiment of the present application, optionally, when recording the matching relationship, that is, the determination process of the matching relationship includes: determining the image coding value as the sum of the index of the first synthetic environment image and the adjusted value of the second image index; the adjusted value of the second image index is the product of the index of the second synthetic environment image and the first constant, and the first constant is used to make the image coding value a unique value; the first synthetic environment image and the second synthetic environment image are two matching images; determining the feature point coding value as the sum of the index of the first feature point and the adjusted value of the second feature point index; the adjusted value of the second feature point index is the product of the index of the second feature point and the second constant, and the second constant is used to make the feature point coding value a unique value; the first feature point and the second feature point are two matching feature points; determining the image coding value and the feature point coding value as the matching relationship.
[0074] Exemplarily, the matching relationship is recorded in the following form:
[0075]
[0076] where i is the image coding value and j is the feature point coding value;
[0077] Exemplarily, the image coding value and the feature point coding value can be expressed by the following formula:
[0078] i = index(I k ) + 2 16 * index(I q );
[0079] j = index(f l ) + 2 8 * index(f m );
[0080] where index(I k ) is the index of the first synthesized environment map, 2 16 is the first constant, index(I q ) is the index of the second synthesized environment map, index(f l ) is the index of the first feature point, 2 8 is the second constant, and index(f m ) is the index of the second feature point.
[0081] With this setting in this solution, in the obtained matching relationship, both the image coding value and the feature point coding value are unique, that is, the matching relationship of each pair of matching feature points is unique, and there will be no problem of chaotic matching relationship during subsequent image fusion.
[0082] S260. Based on the matching result, fuse the synthesized environment maps to obtain an overall virtual environment map.
[0083] In the embodiments of the present application, optionally, fusing the synthesized environment maps based on the matching result to obtain an overall virtual environment map includes: establishing a visual projection constraint on the synthesized environment maps in the image set based on the matching relationship, and calculating the image pose and the three-dimensional points in space by using the bundle adjustment method to obtain the local virtual environment maps corresponding to each image set; fusing the local virtual environment maps based on the matching relationship to obtain an overall virtual environment map.
[0084] Specifically, since the matching relationship reflects the information of the feature points of the matching synthetic environment maps in the image set, the local virtual environment map can be reconstructed based on this; the matching relationship also reflects the corresponding information of the feature points between the image sets, and the local virtual environment maps can be fused based on this to obtain the overall virtual environment map.
[0085] Exemplarily, for each synthetic environment map in each image set and the matching relationship, a visual projection constraint is established, and the bundle adjustment method is used to calculate the image poses and three-dimensional spatial points, obtaining local virtual environment maps. The number of these local virtual environment maps is the same as that of the image sets, and there are a total of M local virtual environment maps corresponding to the image sets. The local virtual environment maps are fused in combination with the matching relationship, a visual projection constraint is established for multiple small maps, and the bundle adjustment method is used to globally optimize all the image poses and three-dimensional spatial points to obtain the overall virtual environment map.
[0086] In the technical solution of the embodiment of the present application, since the feature pattern is added to the synthetic environment map, it has the advantage of having many feature points, which can avoid problems such as double eyelid failure and difficult environmental positioning caused by few feature points. After clustering the synthetic environment maps through the bag-of-words model vectors, an image set is obtained. The images in the same image set have similar features. Then, the image set is stored in a K-dimensional tree, and the structure of the K-dimensional tree is used to match the feature points, so that when matching between image sets, similar images can be quickly located with the help of the K-dimensional tree, avoiding unnecessary calculations, and the recording method of the matching relationship has the feature of unique coding, and there will be no problem of matching relationship disorder. Furthermore, the virtual environment map can be quickly fused based on the obtained matching results and image sets.
[0087] Embodiment III
[0088] Figure 4 FIG. 3 is a schematic structural diagram of a virtual environment reconstruction device provided in Embodiment III of the present application. The device can execute the virtual environment reconstruction method provided in any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. As Figure 4 shown, the device includes:
[0089] A feature extraction module 310, configured to extract feature points and descriptors for each synthetic environment map, obtaining the feature points and descriptors corresponding to each synthetic environment map; the synthetic environment map is obtained by adding a feature pattern to a natural environment map;
[0090] A bag-of-words model vector determination module 320, configured to classify the descriptors in the synthetic environment map based on a pre-trained bag-of-words model, and determine the bag-of-words model vector corresponding to the synthetic environment map according to the classification result;
[0091] An image clustering module 330 is configured to cluster the bag-of-words model vectors corresponding to the synthesized environment maps, and save the obtained image sets in a target data structure; the target data structure storing the image sets can reflect the association information between the image sets;
[0092] An overall virtual environment map determination module 340 is configured to match the images located in the target data structure to obtain a matching result, and fuse the synthesized environment maps based on the matching result to obtain an overall virtual environment map.
[0093] The technical solution of the embodiment of the present application includes: a feature extraction module 310, configured to extract feature points and descriptors for each synthesized environment map to obtain the feature points and descriptors corresponding to each synthesized environment map; the synthesized environment map is obtained by adding feature patterns to a natural environment map; a bag-of-words model vector determination module 320, configured to classify the descriptors in the synthesized environment map based on a pre-trained bag-of-words model, and determine the bag-of-words model vector corresponding to the synthesized environment map according to the classification result; an image clustering module 330, configured to cluster the bag-of-words model vectors corresponding to the synthesized environment maps, and save the obtained image sets in a target data structure; the target data structure storing the image sets can reflect the association information between the image sets; an overall virtual environment map determination module 340, configured to match the images located in the target data structure to obtain a matching result, and fuse the synthesized environment maps based on the matching result to obtain an overall virtual environment map. Through the reconstruction of the synthesized environment map provided with feature images, since there are many feature points therein and there are matching relationships between images, the environment can be accurately located and the environment reconstruction can be realized.
[0094] Optionally, the target data structure is: a K-dimensional tree;
[0095] Correspondingly, the overall virtual environment map determination module 340 includes:
[0096] A matching relationship determination unit, configured to perform descriptor matching on the images of the same layer of the K-dimensional tree from the leaf node to the root node direction. If it is determined that the descriptors in the two currently matched synthesized environment maps meet the matching requirements, record the matching relationship between the two currently matched synthesized environment maps; the matching relationship reflects the feature point information corresponding to the matched descriptors and the information of the synthesized environment maps where the matched descriptors are located.
[0097] Optionally, the matching relationship determination unit includes:
[0098] A brute-force matching unit, configured to perform pairwise matching on the synthesized environment maps in the same image set based on a brute-force matching method to obtain a matching result.
[0099] Optionally, the device further includes: a matching relationship determination module, including:
[0100] An image encoding value determination unit, configured to determine the sum of the index of the first synthetic environment map and the second image index adjustment value as the image encoding value; the second image index adjustment value is the product of the index of the second synthetic environment map and a first constant, and the first constant is used to make the image encoding value a unique value; the first synthetic environment map and the second synthetic environment map are two matching images;
[0101] A feature point encoding value determination unit, configured to determine the sum of the index of the first feature point and the second feature point index adjustment value as the feature point encoding value; the second feature point index adjustment value is the product of the index of the second feature point and a second constant, and the second constant is used to make the feature point encoding value a unique value; the first feature point and the second feature point are two matching feature points;
[0102] A matching relationship determination unit, configured to determine the image encoding value and the feature point encoding value as the matching relationship.
[0103] Optionally, the overall virtual environment map determination module 340 includes:
[0104] A local virtual environment map determination unit, configured to establish a visual projection constraint on the synthetic environment maps in the image set based on the matching relationship, and calculate the image poses and spatial three-dimensional points by using the bundle adjustment method to obtain the local virtual environment maps corresponding to each image set;
[0105] An overall virtual environment map determination unit, configured to fuse the local virtual environment maps based on the matching relationship to obtain an overall virtual environment map.
[0106] Optionally, the device further includes:
[0107] A synthetic environment map noise reduction module, configured to perform noise reduction processing on the synthetic environment map through a Gaussian function to obtain a noise-reduced synthetic environment map;
[0108] Correspondingly, the feature extraction module 310 includes:
[0109] A feature extraction unit, configured to extract feature points and descriptors from each noise-reduced synthetic environment map to obtain the feature points and descriptors corresponding to each synthetic environment map.
[0110] Optionally, the bag-of-words model vector determination module 320 includes:
[0111] A codebook acquisition unit, configured to acquire the codebook of the pre-trained bag-of-words model;
[0112] A descriptor classification unit for calculating the distance between a descriptor and a visual word in the codebook and classifying the descriptor into the visual word with the closest distance to the descriptor among all visual words;
[0113] A bag-of-words model vector generation unit for generating a bag-of-words model vector corresponding to the synthetic environment map based on the quantity information of the visual words to which the descriptors in the synthetic environment map are classified.
[0114] The reconstruction device of a virtual environment provided by an embodiment of the present application can execute the reconstruction method of a virtual environment provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0115] Embodiment 4
[0116] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0117] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0118] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0119] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for reconstructing a virtual environment.
[0120] In some embodiments, the method for reconstructing a virtual environment can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for reconstructing a virtual environment described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for reconstructing a virtual environment by any other suitable means (e.g., by means of firmware).
[0121] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0122] A computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0123] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0124] In order to provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0125] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0126] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0127] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0128] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for reconstructing a virtual environment, characterized in that: include: Extracting feature points and descriptors from each synthetic environment map to obtain feature points and descriptors corresponding to each synthetic environment map; The synthetic environment map is obtained by adding characteristic patterns to the natural environment map; Classify the descriptors in the synthetic environment map based on the pre-trained bag-of-words model, and determine the bag-of-words model vector corresponding to the synthetic environment map according to the classification result; Clustering the bag-of-words model vectors corresponding to the synthetic environment map, and storing the clustered image set in a target data structure; the target data structure storing the image set can reflect the association information between the image sets; The images in the target data structure are matched to obtain matching results, and the synthetic environment map is fused based on the matching results to obtain an overall virtual environment map.
2. The method according to claim 1, characterized in that The target data structure is: a K-dimensional tree; Accordingly, matching is performed on the image in the target data structure to obtain a matching result, including: Descriptor matching is performed on the same-layer images of the K-dimensional tree from the leaf node to the root node. If it is determined that the descriptors in the two synthetic environment images currently being matched meet the matching requirements, the matching relationship between the two synthetic environment images currently being matched is recorded; the matching relationship reflects the feature point information corresponding to the matched descriptors, and the information of the synthetic environment image where the matched descriptors are located.
3. The method according to claim 2, characterized in that Perform descriptor matching on the leaf node image, including: Based on the brute force matching method, the synthetic environment images in the same image set are matched pairwise to obtain the matching results.
4. The method according to claim 2, characterized in that: The process of determining the matching relationship includes: The sum of the index of the first synthetic environment image and the second image index adjustment value is determined as the image encoding value; the second image index adjustment value is the product of the index of the second synthetic environment image and a first constant, and the first constant is used to make the image encoding value a unique value; the first synthetic environment image and the second synthetic environment image are two matching images; The sum of the index of the first feature point and the index adjustment value of the second feature point is determined as the feature point encoding value; the second feature point index adjustment value is the product of the index of the second feature point and a second constant, and the second constant is used to make the feature point encoding value a unique value; the first feature point and the second feature point are two matching feature points; The image encoding value and the feature point encoding value are determined to be in a matching relationship.
5. The method according to claim 2, characterized in that: The synthetic environment map is fused based on the matching results to obtain an overall virtual environment map, including: Based on the matching relationship, a visual projection constraint is established for the synthetic environment map in the image set, and a bundled optimization method is used to calculate the image pose and spatial three-dimensional points to obtain a local virtual environment map corresponding to each image set; The local virtual environment maps are fused based on the matching relationship to obtain an overall virtual environment map.
6. The method according to claim 1, characterized in that Before extracting feature points and descriptors from each synthetic environment map to obtain feature points and descriptors corresponding to each synthetic environment map, the method further includes: Performing noise reduction processing on the synthetic environment map by using a Gaussian function to obtain a synthetic environment map after noise reduction; Accordingly, feature points and descriptors are extracted from each synthetic environment map to obtain feature points and descriptors corresponding to each synthetic environment map, including: Feature points and descriptors are extracted from each synthetic environment map after denoising to obtain feature points and descriptors corresponding to each synthetic environment map.
7. The method according to claim 1, characterized in that The descriptors in the synthetic environment graph are classified based on the pre-trained bag-of-words model, and the bag-of-words model vector corresponding to the synthetic environment graph is determined according to the classification result, including: Get the codebook of the pre-trained bag-of-words model; Calculating the distance between the descriptor and the visual words in the codebook, and classifying the descriptor as the visual word with the closest distance to the descriptor among the visual words; Based on the quantity information of the visual words to which the descriptors in the synthetic environment graph are classified, a bag-of-words model vector corresponding to the synthetic environment graph is generated.
8. A virtual environment reconstruction device, characterized in that: include: A feature extraction module is used to extract feature points and descriptors from each synthetic environment map to obtain feature points and descriptors corresponding to each synthetic environment map; the synthetic environment map is obtained by adding feature patterns to the natural environment map; A bag-of-words model vector determination module is used to classify the descriptors in the synthetic environment graph based on the pre-trained bag-of-words model, and determine the bag-of-words model vector corresponding to the synthetic environment graph according to the classification result; An image clustering module is used to cluster the bag-of-words model vectors corresponding to the synthetic environment map, and store the clustered image set in a target data structure; the target data structure storing the image set can reflect the association information between the image sets; The overall virtual environment map determination module is used to match the images in the target data structure to obtain matching results, and fuse the synthetic environment map based on the matching results to obtain the overall virtual environment map.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the method for reconstructing a virtual environment according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for reconstructing a virtual environment according to any one of claims 1 to 7 when executed.