A method, device, electronic device and storage medium for generating virtual data

By constructing a virtual building model and performing multi-view sampling and data enhancement, virtual data is generated for neural network training, which solves the problem that real building data is susceptible to external factors and improves the detection accuracy of neural networks.

CN114065928BActive Publication Date: 2025-05-13BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010750439.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-30
Publication Date
2025-05-13
Estimated Expiration
2040-07-30

AI Technical Summary

Technical Problem

Real building data is susceptible to external factors in neural network training, resulting in reduced training accuracy and decreased detection ability.

Method used

By building a virtual building model, performing multi-view sampling and data augmentation, virtual data is generated for neural network training.

Benefits of technology

It improves the quality of neural network training samples, reduces dependence on external factors, and improves the detection accuracy of neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114065928B_ABST
    Figure CN114065928B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a virtual data generation method, device, electronic device and storage medium, the method comprising: selecting a pre-built virtual building model, the virtual building model is composed of multiple building planes, each building plane is provided with a corresponding building label; performing multi-view sampling on the virtual building model to obtain multiple perspective images of the virtual building model, each perspective image includes at least one building perspective plane, each building perspective plane is a view of its corresponding building plane under the perspective corresponding to the perspective image; performing enhancement processing on each building perspective plane in each perspective image and the building label corresponding to the building perspective plane to obtain the training plane and training label corresponding to each building perspective plane; using the training plane and training label as virtual data for neural network training. The method provided by the present disclosure can effectively improve the quality of training data of the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision, and in particular to a method, device, electronic device and storage medium for generating virtual data. Background Art

[0002] In the related technology, in the field of computer vision, the computing device can acquire the ability to detect specific objects through the method of neural network training. The quality of the neural network's detection ability depends on whether the neural network has a stable network structure and accurate network parameters. Therefore, in the actual neural network training process, the selection of training data becomes very critical. Factors such as the amount and quality of the training data, the quality of the data labels, and the cost of data collection largely determine the accuracy of the detection ability after the neural network training is completed.

[0003] Furthermore, when training a neural network for building surface detection, the source of training data is generally real building data. During the collection of real building data, it is greatly affected by external factors, and after the collection of real building data is completed, it is difficult to manually process data labels, which in turn affects the training accuracy of the neural network for building surface detection and reduces the detection ability of the neural network. Summary of the invention

[0004] The present disclosure provides a virtual data generation method, device, electronic device and storage medium to at least solve the problem in the related art that when real building data is used as training data, it is greatly affected by external factors, which affects the training accuracy of the neural network used for building surface detection and reduces the detection ability of the neural network. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a method for generating virtual data, comprising:

[0006] Selecting a pre-built virtual building model, the virtual building model is composed of a plurality of building planes, each of the building planes is provided with a corresponding building label; the building plane is an outer surface area of ​​the virtual building model;

[0007] Performing multi-view sampling on the virtual building model to obtain a plurality of view images of the virtual building model, each of the view images comprising at least one building view plane, and each of the building view planes being a view of a corresponding building plane at a view angle corresponding to the view image;

[0008] Performing enhancement processing on each building viewing plane and the building label corresponding to the building viewing plane in each viewing image to obtain a training plane corresponding to each building viewing plane and the training label corresponding to the training plane;

[0009] Each training plane in each of the viewing angle images and a training label corresponding to each training plane are used as virtual data for neural network training.

[0010] Optionally, the process of constructing the virtual building model includes:

[0011] Obtain a preset geometric body, wherein the geometric body is composed of a plurality of triangular faces;

[0012] Randomly select a triangular face from the multiple triangular faces as a starting triangular face;

[0013] A first building plane corresponding to the starting triangular face is divided in the geometric body, and a corresponding building label is assigned to the first building plane, wherein the first building plane is composed of at least one triangular face, each triangular face constituting the first building plane includes at least the starting triangular face, and the similarity between normals of any two adjacent triangular faces in each triangular face constituting the first building plane is less than a preset similarity threshold;

[0014] A new starting triangular facet is randomly selected from the remaining triangular faces among the multiple triangular faces, and a building plane corresponding to the new starting triangular facet is divided out in the geometric body and a building label is assigned, until all triangular faces in the geometric body are divided into their corresponding building planes, thereby completing the construction process of the virtual building model.

[0015] Optionally, performing multi-view sampling on the virtual building model to obtain multiple view images of the virtual building model includes:

[0016] Using a virtual camera to take photos of the virtual building model from multiple perspectives to obtain multiple building images of the virtual building model;

[0017] For each of the building images, when the building image contains the first viewing plane of the virtual building model, the areas that do not meet the sampling conditions in each of the first viewing planes in the building image are removed to obtain multiple viewing images of the virtual building model.

[0018] Optionally, removing areas that do not meet a sampling condition in each of the first viewing planes in the building image to obtain multiple viewing images of the virtual building model includes:

[0019] In response to determining that there is a first viewing plane including an occlusion map among the first viewing planes in the building image, performing an intersection operation on the first viewing plane including the occlusion map and the occlusion map included therein, removing the occlusion map, and obtaining an intersection plane corresponding to the first viewing plane including the occlusion map;

[0020] Determine each first viewing angle plane that does not contain the occlusion map in the building image and each intersection plane as a target plane;

[0021] A detection frame detection is performed on each of the target planes, and the target planes whose detection frame area is smaller than a preset detection area are eliminated to obtain a perspective image corresponding to the building image, and the remaining target planes whose detection frame area is not smaller than the preset detection area are determined as the building perspective planes in the perspective image.

[0022] Optionally, the process of determining whether there is a first viewing plane including an occlusion map among the first viewing planes in the building image includes:

[0023] At a shooting angle corresponding to the building image, a first image corresponding to the building image is shot; the first image includes contour information of each of the first viewing planes at the shooting angle and a building label of each of the first viewing planes;

[0024] For each first viewing plane, in response to detecting that there is an area without a building label in the contour information of the first viewing plane, it is determined that the first viewing plane includes a corresponding occlusion map.

[0025] Optionally, the process of determining whether there is a first viewing plane including an occlusion map among the first viewing planes in the building image includes:

[0026] Acquire a standard image of a shooting angle corresponding to the building image, wherein the standard image includes standard viewing planes corresponding to each first viewing plane of the building image in a simulated scene without obstacles;

[0027] Each first viewing plane of each of the building images is compared with a standard viewing plane corresponding to the first viewing plane, and in response to detecting that the first viewing plane is inconsistent with the standard viewing plane, the first viewing plane is determined to be a first viewing plane including an occlusion map.

[0028] Optionally, the step of performing enhancement processing on each building viewing plane and a building label corresponding to the building viewing plane in each viewing image to obtain a training plane corresponding to each building viewing plane and a training label corresponding to the training plane includes:

[0029] Enhance the plane features of each building viewing plane and the label features of each building label corresponding to the building viewing plane to obtain an enhanced building viewing plane and an enhanced building label;

[0030] The plane features of each building perspective plane and the plane features of the enhanced building perspective plane corresponding to the building perspective plane are feature stitched to obtain the training plane corresponding to the building perspective plane, and the label features of the building label corresponding to the building plane and the label features of the enhanced building label corresponding to the building label are feature stitched to obtain the training label corresponding to the training plane.

[0031] According to a second aspect of an embodiment of the present disclosure, there is provided a virtual data generating device, comprising:

[0032] A selection unit is configured to select a pre-constructed virtual building model, wherein the virtual building model is composed of a plurality of building planes, each of which is provided with a corresponding building label; the building plane is an outer surface area of ​​the virtual building model;

[0033] A sampling unit is configured to perform multi-view sampling on the virtual building model to obtain a plurality of view images of the virtual building model, each of the view images comprising at least one building view plane, and each of the building view planes is a view of a corresponding building plane at a view corresponding to the view image;

[0034] an enhancement unit, configured to perform enhancement processing on each building viewing plane and the building label corresponding to the building viewing plane in each viewing image, to obtain a training plane corresponding to each building viewing plane and the training label corresponding to the training plane;

[0035] The processing unit is configured to use each training plane in each of the perspective images and the training label corresponding to each training plane as virtual data for neural network training.

[0036] Optionally, the virtual data generating device further comprises a virtual building model building unit; the virtual building model building unit is configured to execute:

[0037] Obtain a preset geometric body, wherein the geometric body is composed of a plurality of triangular faces;

[0038] Randomly select a triangular face from the multiple triangular faces as a starting triangular face;

[0039] A first building plane corresponding to the starting triangular face is divided in the geometric body, and a corresponding building label is assigned to the first building plane, wherein the first building plane is composed of at least one triangular face, each triangular face constituting the first building plane includes at least the starting triangular face, and the similarity between normals of any two adjacent triangular faces in each triangular face constituting the first building plane is less than a preset similarity threshold;

[0040] A new starting triangular facet is randomly selected from the remaining triangular faces among the multiple triangular faces, and a building plane corresponding to the new starting triangular facet is divided out in the geometric body and a building label is assigned, until all triangular faces in the geometric body are divided into their corresponding building planes, thereby completing the construction process of the virtual building model.

[0041] Optionally, the sampling unit includes:

[0042] a photographing subunit, configured to use a virtual camera to take photos of the virtual building model from multiple perspectives to obtain multiple building images of the virtual building model;

[0043] The perspective image processing subunit is configured to perform, for each of the building images, when the building image contains the first perspective plane of the virtual building model, removing areas that do not meet the sampling conditions in each of the first perspective planes in the building image to obtain multiple perspective images of the virtual building model.

[0044] Optionally, the viewing angle image processing subunit includes:

[0045] an occlusion map removal subunit, configured to, in response to determining that there is a first perspective plane including the occlusion map in each first perspective plane in the building image, perform an intersection operation on the first perspective plane including the occlusion map and the occlusion map included therein, remove the occlusion map, and obtain an intersection plane corresponding to the first perspective plane including the occlusion map;

[0046] A first determining subunit is configured to determine each first viewing angle plane that does not include an occlusion map in the building image and each intersection plane as a target plane;

[0047] The plane elimination subunit is configured to perform detection frame detection on each of the target planes, eliminate the target planes whose detection frame area is smaller than a preset detection area, obtain a perspective image corresponding to the building image, and determine the remaining target planes whose detection frame area is not smaller than the preset detection area as the building perspective planes in the perspective image.

[0048] Optionally, the first determining subunit includes:

[0049] The image capturing subunit is configured to capture a first image corresponding to the building image at a capturing angle corresponding to each of the building images; the first image includes contour information of each of the first viewing planes at the capturing angle and a building label of each of the first viewing planes;

[0050] The second determining subunit is configured to execute, for each first viewing plane, in response to detecting that there is an area without a building label in the contour information of the first viewing plane, determining that the first viewing plane includes a corresponding occlusion map.

[0051] Optionally, the first determining subunit includes:

[0052] An acquisition subunit is configured to acquire a standard image corresponding to a shooting angle of each of the building images, wherein the standard image includes standard viewing planes corresponding to each first viewing plane of the building image in a simulated scene without obstacles;

[0053] The third determination subunit is configured to compare each first viewing plane of each of the building images with a standard viewing plane corresponding to the first viewing plane, and in response to detecting that the first viewing plane is inconsistent with the standard viewing plane, determine that the first viewing plane is a first viewing plane including an occlusion map.

[0054] Optionally, the enhancement unit includes:

[0055] The first processing subunit is configured to enhance the plane feature of each building viewing plane and the label feature of each building label corresponding to the building viewing plane to obtain an enhanced building viewing plane and an enhanced building label;

[0056] The second processing sub-unit is configured to perform feature stitching on the plane features of each building perspective plane and the plane features of the enhanced building perspective plane corresponding to the building perspective plane to obtain a training plane corresponding to the building perspective plane, and to perform feature stitching on the label features of the building label corresponding to the building plane and the label features of the enhanced building label corresponding to the building label to obtain a training label corresponding to the training plane.

[0057] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the virtual data generation method as described in the first aspect.

[0058] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the virtual data generating method as described in the first aspect.

[0059] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the virtual data generating method as described in the first aspect.

[0060] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0061] The present disclosure relates to a virtual data generation method, device, electronic device and storage medium, the method comprising: selecting a pre-constructed virtual building model, the virtual building model being composed of a plurality of building planes, each of the building planes being provided with a corresponding building label; performing multi-perspective sampling on the virtual building model to obtain a plurality of perspective images of the virtual building model, each of the perspective images comprising at least one building perspective plane, each of the building perspective planes being a view of its corresponding building plane at a perspective corresponding to the perspective image; performing enhancement processing on each building perspective plane in each of the perspective images and the building label corresponding to the building perspective plane to obtain a training plane corresponding to each building perspective plane and a training label corresponding to the training plane; using each training plane in each of the perspective images and the training label corresponding to each training plane as virtual data for training the neural network. By applying the virtual data generation method provided by the present invention, a virtual building model consistent with the actual building structure is obtained, the virtual building model is sampled, and the sampled data is enhanced to obtain virtual data that is close to the actual sample data. That is, by sampling the virtual building model, data from various perspectives of different buildings can be easily obtained, thereby obtaining training samples for enriching the neural network. During the sampling process, there is no need to go to the actual building site for sampling, so that the sampling process will not be affected by various external interference factors, which can improve the quality of the training samples of the neural network and further improve the detection capability of the neural network.

[0062] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0064] Figure 1 is a flow chart of a method for generating virtual data according to an exemplary embodiment;

[0065] Figure 2 is a flowchart of a process of obtaining multiple perspective images of a virtual building model according to an exemplary embodiment;

[0066] Figure 3 is a flowchart showing another process of obtaining multiple perspective images of a virtual building model according to an exemplary embodiment;

[0067] Figure 4 This is an example diagram showing a virtual camera photographing a virtual building model according to an exemplary embodiment;

[0068] Figure 5 is an example diagram of a building image according to an exemplary embodiment;

[0069] Figure 6 is an example diagram of a sampling effect of a virtual building model according to an exemplary embodiment;

[0070] Figure 7 is an example diagram of a sampling effect of an actual building shown according to an exemplary embodiment;

[0071] Figure 8 is a schematic diagram of a process of obtaining virtual data according to an exemplary embodiment;

[0072] Fig. 9 is a block diagram of a virtual data generating device according to an exemplary embodiment;

[0073] Fig.10 is a block diagram of an electronic device according to an exemplary embodiment;

[0074] Fig.11 is a block diagram of another electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0075] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0076] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0077] Figure 1 is a flowchart of a method for generating virtual data according to an exemplary embodiment. Figure 1 As shown, the virtual data generation method is used in an electronic device, comprising the following steps:

[0078] In step S11, a pre-constructed virtual building model is selected, wherein the virtual building model is composed of a plurality of building planes, each of which is provided with a corresponding building label, and the building plane is an outer surface area of ​​the virtual building model.

[0079] In this method, before executing step S11, multiple virtual building models are pre-constructed, each virtual building model can correspond to a physical building, and the virtual building model is consistent with the corresponding physical building in terms of structure, color, appearance ratio, etc. When executing step S11, the corresponding virtual building model can be selected according to the actual needs of virtual data generation. Specifically, the virtual building model can be a virtual building model specified by the user, and the user can specify the virtual building model by sending a model selection instruction to the electronic device.

[0080] Specifically, the virtual building is composed of a plurality of building planes, each of the building planes is an outer surface area of ​​the virtual building, and the overall outer surface of the virtual building is composed of the respective outer surface areas.

[0081] Among them, each of the building planes is a part of the overall outer surface area of ​​the virtual building. It can be understood that each of the building planes represents the visual presentation of its corresponding outer surface area in the virtual environment, which may include the shape, building components, building materials and colors of the outer surface area. The visual presentation of each building plane in the virtual environment can be a plane or a quasi-plane.

[0082] Optionally, each building plane is provided with a building label, which is used to describe the building plane, such as one or more of the size, color and descriptive information of each object contained on the building plane, and the objects contained on the building plane may be air conditioners, billboards, and the like.

[0083] In step S12, multi-perspective sampling is performed on the virtual building model to obtain multiple perspective images of the virtual building model, each of the perspective images includes at least one building perspective plane, and each of the building perspective planes is a view of its corresponding building plane at the perspective corresponding to the perspective image.

[0084] In the process of sampling the virtual building model, the virtual building may be sampled at various preset viewing angles, thereby obtaining multiple viewing angle images of the virtual building model.

[0085] Specifically, for each virtual building plane of the virtual building model, the view presented by the virtual building plane may be different at different viewing angles, and the view may refer to the displayed portion of the building plane at the viewing angle.

[0086] Optionally, the area of ​​each building viewing plane in the viewing image is not less than a preset threshold value, and the threshold value can be set according to actual needs and is not limited here.

[0087] In step S13, each building viewing plane in each viewing image and the building label corresponding to the building viewing plane are enhanced to obtain a training plane corresponding to each building viewing plane and a training label corresponding to the training plane.

[0088] Specifically, a generative adversarial network or other forms of networks can be used to enhance the building perspective plane and the building label corresponding to the building perspective plane, so as to obtain a training plane corresponding to each building perspective plane and a training label corresponding to the training plane, which can reduce the data feature domain gap between the building perspective plane and the building perspective plane in the real scene, and reduce the data feature domain gap between the building label and the real building label; for example, it can reduce the gap between the color feature of the building perspective plane and the color feature of the real building perspective plane, as well as the feature gap between the color label corresponding to the building perspective plane and the real color label.

[0089] In step S14, each training plane in each of the viewing angle images and a training label corresponding to each training plane are used as virtual data for the neural network training.

[0090] By applying the method provided by the embodiment of the present disclosure, virtual data can be obtained through a virtual building model, so that data of different buildings can be easily obtained, enriching the training samples of the neural network model. In addition, there is no need to manually label the collected data, which can reduce the data labeling cost. The virtual data after enhancement processing can be close to the real data, and in the process of obtaining the virtual data, it will not be interfered by the outside world, which can greatly improve the quality of the training samples, and thus can effectively improve the detection accuracy of the neural network.

[0091] In the method provided in the embodiment of the present disclosure, based on the above implementation process, specifically, the process of constructing a virtual building model includes:

[0092] Obtain a preset geometric body, wherein the geometric body is composed of a plurality of triangular faces;

[0093] Randomly select a triangular face from the multiple triangular faces as a starting triangular face;

[0094] A first building plane corresponding to the starting triangular face is divided in the geometric body, and a corresponding building label is assigned to the first building plane, wherein the first building plane is composed of at least one triangular face, each triangular face constituting the first building plane includes at least the starting triangular face, and the similarity between normals of any two adjacent triangular faces in each triangular face constituting the first building plane is less than a preset similarity threshold;

[0095] A new starting triangular facet is randomly selected from the remaining triangular faces among the multiple triangular faces, and a building plane corresponding to the new starting triangular facet is divided out in the geometric body and a building label is assigned, until all triangular faces in the geometric body are divided into their corresponding building planes, thereby completing the construction process of the virtual building model.

[0096] Among them, a geometric body is constructed by a 3D construction method. The geometric body may be a building geometric body. The structure, material, color and other information of the geometric body may be constructed with reference to an actual building. The geometric body does not carry a building label.

[0097] Specifically, the first building plane is composed of a starting triangle facet or a plurality of sequentially connected triangle faces, and the plurality of sequentially connected triangle faces include at least the starting triangle facet. If the first building plane is composed of a plurality of sequentially connected triangle faces, then the similarity between the normals of any two adjacent triangle faces among the triangle faces constituting the first building plane is less than a similarity threshold.

[0098] Among them, the smaller the similarity between the normals of two adjacent triangular faces, the more it can be explained that the two adjacent triangular faces tend to a plane; the similarity between the normals of any two adjacent triangular faces among the triangular faces constituting the first building plane can be determined by the angle between the normals of the two triangular faces.

[0099] Optionally, when the directions of the normals of two triangular faces are inconsistent, the smaller the angle between the normals of the two triangular faces is, the smaller the similarity between the two triangular faces is.

[0100] In the method provided by the embodiment of the present disclosure, a way to divide the first building plane corresponding to the starting triangle in a geometric body can be: determine the triangle connected to each edge of the starting triangle as the triangle to be compared, and compare the normal of the starting triangle and the normal of each triangle to be compared for similarity, that is, compare the direction of the starting triangle and the direction of the triangle connected to each edge of the starting triangle for similarity, wherein the similarity between the two normals is determined by the size of the angle between the directions of the two normals, and when the similarity between the normal of the triangle to be compared and the normal of the starting triangle is less than a sub-similarity threshold, determine the triangle to be compared as the target triangle; wherein the similarity threshold can be twice the sub-similarity threshold; specifically, if the triangle connected to any edge of the starting triangle is the target triangle If there is a target triangle face, the triangle face connected by the other two edges of the target triangle face can be expanded into a new triangle face to be compared of the starting triangle face, and it is continued to be determined whether the two newly added triangle faces to be compared are the target triangle faces. If so, the expansion is continued, and the triangle face connected by the other two edges of the new target triangle face is expanded into a triangle face to be compared with the starting triangle face. When the newly added triangle face to be compared is not the target triangle face, the expansion of the newly added triangle face to be compared is stopped. When all the newly added triangle faces to be compared cannot be expanded, corresponding plane labels are assigned to the starting triangle face and each target triangle face, so that the starting triangle face and each target triangle face to which plane labels have been assigned form the first building plane, a building label is assigned to the first building plane, and a new starting triangle face is selected from the remaining triangle faces, and it is continuously iterated to complete the division of the geometric body.

[0101] For example, a triangle is randomly selected from multiple triangles as the starting triangle, and the starting triangle is compared with the triangles connected to each edge of the starting triangle in turn for similarity, that is, the similarity is compared with the three triangles, and the triangle with a similarity less than the sub-similarity threshold is used as the target triangle. If all three triangles are target triangles, the triangles connected to the remaining two edges of each target triangle are used as the triangles to be compared, that is, at this time, each target triangle is connected to two new triangles to be compared, for a total of 6 triangles to be compared. If one of the newly added 6 triangles to be compared is If there are 4 target triangles, continue to expand these 4 target triangles to obtain 8 new triangles to be compared, until all the newly added triangles to be compared are not target triangles, then the target triangles and the starting triangle form the first building plane, and select the new starting triangle from the triangles that constitute the geometric body except the triangles that constitute the first building plane. After the geometric body is divided, set the corresponding building label for each first building plane to obtain each building plane, and use the geometric body to which the building plane with the assigned label belongs as a virtual building model.

[0102] By applying the method provided in the embodiment of the present disclosure, multiple triangular faces connected in sequence with small normal similarity are regarded as a building plane, so that the building plane can be visually presented as a plane. In addition, in the process of dividing the building plane, a corresponding building label is assigned to each building plane, which can improve the efficiency of setting the building labels.

[0103] In the method provided in the embodiment of the present disclosure, based on the above implementation process, specifically, the process of performing multi-view sampling on the virtual building model to obtain multiple view images of the virtual building model is as follows: Figure 2 As shown, specifically including:

[0104] In step S21, a virtual camera is used to take photos of the virtual building model from multiple perspectives to obtain multiple building images of the virtual building model.

[0105] Among them, the multiple building images of the virtual building model can be obtained by taking pictures in a simulated real scene, and the simulated real scene can include one or more obstacles such as trees, people, vehicles, clouds, etc., that is, the building image can be the virtual building model in the simulated real scene, and the image is obtained by using a virtual camera to photograph the virtual building model at a corresponding viewing angle.

[0106] Specifically, the virtual camera may be a virtual camera capable of taking color images, and the building image may be a color image.

[0107] In step S22, for each of the building images, when the building image contains the first viewing plane of the virtual building model, the areas that do not meet the sampling conditions in each of the first viewing planes in the building image are removed to obtain multiple viewing images of the virtual building model.

[0108] The first viewing angle plane may be a building plane including the virtual building model and an image simulating a real scene captured by a virtual camera at a viewing angle corresponding to the building image.

[0109] Specifically, the occlusion map area in each first viewing plane and the small plane area in each first viewing plane may be removed, wherein the small plane is a first viewing plane whose area of ​​the partial area not including the occlusion map is smaller than the preset detection area.

[0110] Among them, the building image that does not include the first viewing plane can be deleted.

[0111] By applying the method provided in the embodiment of the present disclosure, the area in the first viewing angle plane that does not meet the sampling conditions can be removed, so that the acquired building viewing angle image can be free from interference, thereby improving the quality of the training data.

[0112] In the method provided by the embodiment of the present disclosure, based on the above implementation process, specifically, the region that does not meet the sampling condition in each of the first viewing planes in the building image is removed to obtain multiple viewing images of the virtual building model, such as Figure 3 As shown, this may include:

[0113] In step S31, in response to determining that there is a first viewing plane including an occlusion map in each first viewing plane in the building image, an intersection operation is performed on the first viewing plane including the occlusion map and the occlusion map it includes, and the occlusion map is removed to obtain an intersection plane corresponding to the first viewing plane including the occlusion map.

[0114] The occlusion map may be an image of obstacles that block the building plane of the virtual building at the viewing angle corresponding to the building image, that is, an image of at least one obstacle between the lens of the virtual camera and the building plane.

[0115] Specifically, in some viewing angles, there may be obstacles between the lens of the virtual camera and the building plane of the virtual building, see Figure 4 , is an example of a virtual building model photographed by a virtual camera provided by the present disclosure, that is, there is a portion of the building plane blocked by a person between the lens of the virtual camera and the building plane of the virtual building, and the building image photographed at this viewing angle is as follows Figure 5As shown, Figure 5 The portrait part in the first viewing plane of building plane 2 is the occlusion map corresponding to the first viewing plane corresponding to building plane 2, and there is no occlusion object in the first viewing plane of building plane 1.

[0116] Specifically, the intersection plane refers to the portion of the occlusion map in the first perspective plane removed, wherein the occlusion map of the first perspective plane can be an image of the contour portion of each occluding object, or can be an image of the circumscribed geometric portion of the contour of each occluding object contained in the first perspective plane, and the circumscribed geometry can be a circumscribed rectangle, circle, or ellipse, etc.

[0117] In step S32, each first viewing angle plane in the building image that does not include the occlusion map and each intersection plane are determined as target planes.

[0118] In step S33, detection frame detection is performed on each of the target planes, and the target planes whose detection frame area is smaller than the preset detection area are eliminated to obtain a perspective image corresponding to the building image, and the remaining target planes whose detection frame area is not smaller than the preset detection area are determined as the building perspective planes in the perspective image.

[0119] Among them, by performing detection frame detection on each target plane, the target plane with a smaller area can be eliminated, simulating the situation of manually eliminating small planes.

[0120] Specifically, a feasible method of performing detection frame detection on each target plane, eliminating the target plane whose detection frame area is smaller than a preset detection area, and obtaining the perspective image corresponding to the building image may include:

[0121] Apply a preset straight line detection algorithm to identify each of the target planes to obtain at least one detection frame marked on each of the target planes; determine the detection frame area of ​​each detection frame on each of the target planes, and determine the target detection frame with the largest detection frame area on each of the target planes; compare the detection frame area of ​​each target detection frame with a preset detection area; if there is a target detection frame with a detection frame area smaller than the preset detection area, then eliminate the target plane to which the target detection frame with a detection frame area smaller than the preset detection area belongs, and obtain a perspective image corresponding to the building image.

[0122] When the target plane is divided into multiple parts by obstacles, a straight line detection algorithm can be used to generate multiple detection frames when identifying the target plane.

[0123] Specifically, for each target plane, if there is one detection frame for the target plane, the detection frame can be determined as the target detection frame; if there are multiple detection frames for the target plane, the detection frame with the largest area among the detection frames can be determined as the target detection frame, that is, the detection frames can be sorted according to the size of the area, and the largest detection frame can be selected as the target detection frame. By eliminating the target planes to which the target detection frames whose detection frame areas are smaller than the preset detection area belong, the situation of manually eliminating small planes can be simulated, thereby improving the generation efficiency of virtual data and improving the quality of virtual data.

[0124] Optionally, if there is no target detection frame with a detection frame area smaller than a preset detection area, the building image with each target detection frame marked can be used as the perspective image, and each target plane can be determined as the building perspective plane in the perspective image.

[0125] In this embodiment, in another feasible way of performing multi-perspective sampling on the virtual building model to obtain multiple perspective images of the virtual building model, there is another feasible step parallel to step S34, which can be: performing detection frame detection on each target plane to obtain a target detection frame for each target plane, and the target detection frame is the detection frame with the largest area among each target plane; determining the ratio of the area of ​​each target detection frame to the area of ​​the building image, eliminating the target planes to which the target detection frames whose ratios are less than a preset ratio threshold belong, obtaining the perspective image corresponding to the building image, and determining the remaining target planes as the building perspective planes in the perspective image.

[0126] By applying the method provided in the embodiment of the present disclosure, it is possible to simulate the removal of occlusion maps and culling of small planes in actual application scenarios, so as to avoid the occlusion maps and small planes affecting the training effect of the neural network and improve the quality of virtual data.

[0127] In the method provided in the embodiment of the present disclosure, based on the above implementation process, specifically, a feasible manner of determining whether there is a first viewing plane including an occlusion map in each first viewing plane in the building image may include:

[0128] At a shooting angle corresponding to the building image, a first image corresponding to the building image is shot; the first image includes contour information of each of the first viewing planes at the shooting angle and a building label of each of the first viewing planes;

[0129] For each first viewing plane, in response to detecting that there is an area without a building label in the contour information of the first viewing plane, it is determined that the first viewing plane includes a corresponding occlusion map.

[0130] If there is no region without a label of a building in the contour information of the first viewing plane, it can be determined that the first viewing plane does not include an occlusion map.

[0131] Specifically, the first image is obtained by photographing the virtual building model at a certain shooting angle by a virtual camera that can photograph labels. The first image may include a black and white outline image of the virtual building model and the building labels of each first viewing plane of the virtual building at the shooting angle; if there is an area without a building label in the first viewing plane, it means that the building label in the area is blocked by an obstruction. Therefore, it can be determined that the first viewing plane contains a corresponding occlusion map, and the occlusion map may be the part of the first viewing plane corresponding to the area.

[0132] In the process of shooting at any shooting angle, the virtual building can be photographed at the shooting angle using a virtual camera that can shoot color images and a virtual camera that can shoot tags, thereby obtaining a color building image and a first image corresponding to the building image, wherein the first image is obtained by shooting the virtual building at the shooting angle using the virtual camera that can shoot tags.

[0133] By applying the method provided in the embodiment of the present disclosure, by detecting the contour information of each first viewing plane, the first viewing plane including the occlusion map can be accurately and quickly determined in each first viewing plane, and then the occlusion map in the first viewing plane can be removed, thereby improving the quality of virtual data.

[0134] In the method provided in the embodiment of the present disclosure, based on the above implementation process, specifically, another feasible manner of the process of determining whether there is a first viewing plane including an occlusion map in each first viewing plane in the building image may include:

[0135] Acquire a standard image of a shooting angle corresponding to the building image, wherein the standard image includes standard viewing planes corresponding to each first viewing plane of the building image in a simulated scene without obstacles;

[0136] Each first viewing plane of each of the building images is compared with a standard viewing plane corresponding to the first viewing plane, and in response to detecting that the first viewing plane is inconsistent with the standard viewing plane, the first viewing plane is determined to be a first viewing plane including an occlusion map.

[0137] Specifically, in the process of actual application, after constructing a virtual building model, in order to make the sampled data closer to the real data, some objects simulating real scenes can be set around the virtual building model, for example, virtual object models such as clouds, trees, people and vehicles can be set around the virtual building model; before arranging the obstacle model scene around the virtual building model, a virtual camera that can take color images can be used to shoot the virtual building model at various shooting angles to obtain standard images at various shooting angles, that is, the standard image can be obtained by taking the virtual camera that can take color images in the absence of an obstacle model scene, and there are no obstacles that block the building planes between the virtual camera and the various building planes of the virtual building model; after arranging the obstacle model scene around the virtual building model, the building model with the obstacle model scene arranged is photographed at various shooting angles to obtain a building image. At this time, each first-perspective plane in the building image may have an image of the obstacle model, that is, an occlusion map.

[0138] Among them, the standard viewing plane is the image of the building plane to which the first viewing plane belongs at the shooting angle. The first viewing planes contained in the building image correspond one to one with the standard viewing planes of the standard image. The corresponding building image and the standard image can be obtained by shooting the virtual building model with the same virtual camera at the same shooting angle in different scenes.

[0139] Specifically, when the first viewing plane and the standard viewing plane corresponding to the first viewing plane are inconsistent, it means that the first viewing plane contains the corresponding occlusion map, that is, the image area where the first viewing plane and the standard viewing plane are different is the occlusion map; if they are consistent, it means that the first viewing plane does not contain the occlusion map.

[0140] By applying the method provided in the embodiment of the present disclosure, by comparing each first viewing plane with the corresponding standard viewing plane, the first viewing plane including the occlusion map can be accurately and quickly determined in each first viewing plane, and then the occlusion map in the first viewing plane can be removed, thereby improving the quality of virtual data.

[0141] In the method provided by the embodiment of the present disclosure, based on the above implementation process, specifically, the enhancing process is performed on each building viewing plane in each viewing image and the building label corresponding to the building viewing plane to obtain a training plane corresponding to each building viewing plane and a training label corresponding to the training plane, including:

[0142] Enhance the plane features of each building viewing plane and the label features of each building label corresponding to the building viewing plane to obtain an enhanced building viewing plane and an enhanced building label;

[0143] The plane features of each building perspective plane and the plane features of the enhanced building perspective plane corresponding to the building perspective plane are feature stitched to obtain the training plane corresponding to the building perspective plane, and the label features of the building label corresponding to the building plane and the label features of the enhanced building label corresponding to the building label are feature stitched to obtain the training label corresponding to the training plane.

[0144] In the method provided by the embodiment of the present disclosure, a preset generative adversarial network can be used to enhance the plane features of each building perspective plane and the label features of the building label in each perspective image. Other forms of data enhancement networks can also be used to enhance the plane features of the building perspective plane and the label features of the building label to improve the visual effect of the building perspective plane and the building label, making the building perspective plane and the building label clearer. In order to avoid the enhanced data from deviating from the real data, the data before and after enhancement are feature stitched to make the plane features of the building perspective plane and the label features of the building label close to the real data.

[0145] By applying the method provided in the embodiment of the present disclosure, by enhancing the building view plane and the building label and performing feature stitching on the data before and after the enhancement, the data feature domain gap between the virtual data and the real data can be reduced, and the features of the virtual data can be evenly distributed.

[0146] The virtual data generation method provided by the embodiment of the present disclosure can, during application, first process the building labels by simulating the method of manually marking buildings, so as to reduce the feature domain gap between the building labels and the real data labels.

[0147] In the process of simulating manual labeling of buildings, the labeling rules of manual data can be standardized to avoid errors caused by differences in different manual labeling methods. The rules of manual labeling can be customized according to actual conditions, but the standard is that the labeling rules of all real data must be consistent.

[0148] For each geometric structure of a virtual building model, an independent building plane label is generated for each building plane of the virtual building model. The label generation method can be: determine the geometry of the virtual building model, which is composed of multiple triangular faces, select any triangular face as the initial triangular face, obtain the normal N_0 of the initial triangular face, and then select any triangular face as the FACE-ID generation circle (initial 3 edges), and mark the FACE-ID as F_0, then check the normals N_n of the faces connected to the three edges in turn, calculate the similarity between N_0 and N_n, if the direction similarity between the face connected to any edge and the initial triangular face is less than a threshold, then assign F_0 to the face connected to the edge, and expand the boundary of F_0 (add 2 new edges), otherwise it will not be expanded. In this way, F_0 will eventually stop expanding. After stopping the expansion, each triangular face marked with F_0 is regarded as a building plane, and a building label is set for the building plane. Then select a new triangular face without FACE-ID mark and repeat the above process. In this way, the virtual building model is obtained as a whole, and the virtual building model includes a plurality of building planes with building labels.

[0149] Among them, the present disclosure reduces the data feature domain gap between the training plane and the real data through an optimal sampling method, see Figure 6 As shown in FIG. 1 , an example diagram of a sampling effect of a virtual building model provided in this embodiment is specifically a data statistics of the normal direction of the building floor of the virtual building model, wherein the black dots are the sampling visualization effects, see Figure 7 , which is an example of a sampling effect of an actual building provided in this embodiment, specifically, the data statistics of the normal direction of the building floor of the actual building, wherein the black dots are visualized effects, compared Figure 6 and Figure 7 It can be seen that the distribution range of the virtual data obtained by sampling covers the distribution of the real data, achieving the effect of "superset".

[0150] See also Figure 8 , which is a schematic diagram of a process for obtaining virtual data provided by an embodiment of the present invention, wherein a building view plane and a building label corresponding to the building view plane are enhanced by a generative adversarial network or other forms of networks to obtain an enhanced building plane and an enhanced building label, the building view plane is spliced ​​with the enhanced building view plane to obtain a training plane; the building label is spliced ​​with the enhanced building label to obtain a training label; and the training plane and the training label are used as virtual data for training a building detection neural network.

[0151] Fig. 9is a block diagram of a virtual data generating device according to an exemplary embodiment. Fig. 9 The device includes a selection unit 901, a sampling unit 902, an enhancement unit 903, and a processing unit 904.

[0152] The selection unit 901 is configured to select a pre-constructed virtual building model, wherein the virtual building model is composed of a plurality of building planes, each of which is provided with a corresponding building label; the building plane is an outer surface area of ​​the virtual building model;

[0153] The sampling unit 902 is configured to perform multi-view sampling on the virtual building model to obtain multiple view images of the virtual building model, each of which includes at least one building view plane, and each of which is a view of the corresponding building plane at the view corresponding to the view image;

[0154] The enhancement unit 903 is configured to perform enhancement processing on each building viewing plane and the building label corresponding to the building viewing plane in each viewing image, so as to obtain a training plane corresponding to each building viewing plane and a training label corresponding to the training plane;

[0155] The processing unit 904 is configured to use each training plane in each of the perspective images and the training label corresponding to each training plane as virtual data for neural network training.

[0156] In another embodiment provided by the present disclosure, based on the above solution, optionally, the virtual data generating device further includes a virtual building model building unit; the virtual building model building unit is configured to execute:

[0157] Obtain a preset geometric body, wherein the geometric body is composed of a plurality of triangular faces;

[0158] Randomly select a triangular face from the multiple triangular faces as a starting triangular face;

[0159] A first building plane corresponding to the starting triangular face is divided in the geometric body, and a corresponding building label is assigned to the first building plane, wherein the first building plane is composed of at least one triangular face, each triangular face constituting the first building plane includes at least the starting triangular face, and the similarity between normals of any two adjacent triangular faces in each triangular face constituting the first building plane is less than a preset similarity threshold;

[0160] A new starting triangular facet is randomly selected from the remaining triangular faces among the multiple triangular faces, and a building plane corresponding to the new starting triangular facet is divided out in the geometric body and a building label is assigned, until all triangular faces in the geometric body are divided into their corresponding building planes, thereby completing the construction process of the virtual building model.

[0161] In another embodiment provided by the present disclosure, based on the above solution, optionally, the sampling unit 902 includes:

[0162] a photographing subunit, configured to use a virtual camera to take photos of the virtual building model from multiple perspectives to obtain multiple building images of the virtual building model;

[0163] The perspective image processing subunit is configured to perform, for each of the building images, when the building image contains the first perspective plane of the virtual building model, removing areas that do not meet the sampling conditions in each of the first perspective planes in the building image to obtain multiple perspective images of the virtual building model.

[0164] In another embodiment provided by the present disclosure, based on the above solution, optionally, the viewing angle image processing subunit includes:

[0165] an occlusion map removal subunit, configured to, in response to determining that there is a first perspective plane including the occlusion map in each first perspective plane in the building image, perform an intersection operation on the first perspective plane including the occlusion map and the occlusion map included therein, remove the occlusion map, and obtain an intersection plane corresponding to the first perspective plane including the occlusion map;

[0166] A first determining subunit is configured to determine each first viewing angle plane that does not include an occlusion map in the building image and each intersection plane as a target plane;

[0167] The plane elimination subunit is configured to perform detection frame detection on each of the target planes, eliminate the target planes whose detection frame area is smaller than a preset detection area, obtain a perspective image corresponding to the building image, and determine the remaining target planes whose detection frame area is not smaller than the preset detection area as the building perspective planes in the perspective image.

[0168] In another embodiment provided by the present disclosure, based on the above solution, optionally, the first determining subunit includes:

[0169] The image capturing subunit is configured to capture a first image corresponding to the building image at a capturing angle corresponding to each of the building images; the first image includes contour information of each of the first viewing planes at the capturing angle and a building label of each of the first viewing planes;

[0170] The second determining subunit is configured to execute, for each first viewing plane, in response to detecting that there is an area without a building label in the contour information of the first viewing plane, determining that the first viewing plane includes a corresponding occlusion map.

[0171] In another embodiment provided by the present disclosure, based on the above solution, optionally, the first determining subunit includes:

[0172] An acquisition subunit is configured to acquire a standard image corresponding to a shooting angle of each of the building images, wherein the standard image includes standard viewing planes corresponding to each first viewing plane of the building image in a simulated scene without obstacles;

[0173] The third determination subunit is configured to compare each first viewing plane of each of the building images with a standard viewing plane corresponding to the first viewing plane, and in response to detecting that the first viewing plane is inconsistent with the standard viewing plane, determine that the first viewing plane is a first viewing plane including an occlusion map.

[0174] In another embodiment provided by the present disclosure, based on the above solution, optionally, the enhancement unit 903 includes:

[0175] The first processing subunit is configured to enhance the plane feature of each building viewing plane and the label feature of each building label corresponding to the building viewing plane to obtain an enhanced building viewing plane and an enhanced building label;

[0176] The second processing sub-unit is configured to perform feature stitching on the plane features of each building perspective plane and the plane features of the enhanced building perspective plane corresponding to the building perspective plane to obtain a training plane corresponding to the building perspective plane, and to perform feature stitching on the label features of the building label corresponding to the building plane and the label features of the enhanced building label corresponding to the building label to obtain a training label corresponding to the training plane.

[0177] Fig.10 1 is a block diagram of an electronic device 1000 according to an exemplary embodiment. For example, the electronic device 1000 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0178] Reference Fig.10 , the electronic device 1000 may include one or more of the following components: a processing component 1002 , a memory 1004 , a power component 1006 , a multimedia component 1008 , an audio component 1010 , an input / output (I / O) interface 1012 , a sensor component 1014 , and a communication component 1016 .

[0179] The processing component 1002 generally controls the overall operation of the electronic device 1000, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 1002 may include one or more processors 1020 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 1002 may include one or more modules to facilitate the interaction between the processing component 1002 and other components. For example, the processing component 1002 may include a multimedia module to facilitate the interaction between the multimedia component 1008 and the processing component 1002.

[0180] The memory 1004 is configured to store various types of data to support operations on the device 1000. Examples of such data include instructions for any application or method operating on the electronic device 1000, contact data, phone book data, messages, pictures, videos, etc. The memory 1004 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0181] The power supply component 1006 provides power to various components of the electronic device 1000. The power supply component 1006 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 1000.

[0182] The multimedia component 1008 includes a screen that provides an output interface between the electronic device 1000 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1008 includes a front camera and / or a rear camera. When the device 1000 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0183] The audio component 1010 is configured to output and / or input audio signals. For example, the audio component 1010 includes a microphone (MIC), and when the electronic device 1000 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 1004 or sent via the communication component 1016. In some embodiments, the audio component 1010 also includes a speaker for outputting audio signals.

[0184] I / O interface 1012 provides an interface between processing component 1002 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0185] The sensor assembly 1014 includes one or more sensors for providing various aspects of status assessment for the electronic device 1000. For example, the sensor assembly 1014 can detect the open / closed state of the device 1000, the relative positioning of components, such as the display and keypad of the electronic device 1000, and the sensor assembly 1014 can also detect the position change of the electronic device 1000 or a component of the electronic device 1000, the presence or absence of user contact with the electronic device 1000, the orientation or acceleration / deceleration of the electronic device 1000, and the temperature change of the electronic device 1000. The sensor assembly 1014 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1014 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1014 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0186] The communication component 1016 is configured to facilitate wired or wireless communication between the electronic device 1000 and other devices. The electronic device 600 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 1016 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1016 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0187] In an exemplary embodiment, the electronic device 1000 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned virtual data generation method.

[0188] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 1004 including instructions, and the instructions can be executed by the processor 1020 of the electronic device 1000 to complete the above virtual data generation method. Optionally, the storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0189] In an exemplary embodiment, a computer program product is also provided, the computer program product includes a readable program code, the readable program code can be executed by the processor 1020 of the electronic device 1000 to complete the virtual data generation method described in any embodiment. Optionally, the program code can be stored in a storage medium of the electronic device 1000, and the storage medium can be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0190] In addition, the electronic device 1000 includes some functional modules not shown, which will not be described in detail here.

[0191] Fig.11 is a schematic diagram of the structure of another electronic device provided by an embodiment of the present disclosure. Fig.11At the hardware level, the electronic device includes a processor, such as a central processing unit (CPU). Optionally, the electronic device also includes an internal bus 1104, a network interface 1102, a memory 1105, and an I / O controller. Among them, the memory may include a memory, such as a high-speed random access memory 1151 (Random-Access Memory, RAM) and a read-only memory 1152 (Read-Only Memory, ROM), and may also include a large-capacity storage device 1106, such as at least one disk storage, etc. Of course, the electronic device may also include hardware required for other services.

[0192] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0193] The memory is used to store processor executable instructions. The processor is configured to execute the instructions stored in the memory, logically forming a virtual data generation device to implement the virtual data generation method provided by any embodiment of the present disclosure and achieve the same technical effect.

[0194] As disclosed above Figure 1The virtual data generation method disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in one or more embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with one or more embodiments of the present disclosure can be directly embodied as a hardware decoding processor for execution, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the field such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0195] Of course, in addition to software implementation, the electronic device disclosed in the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0196] The embodiments of the present disclosure also provide a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the virtual data generation method provided by any embodiment of the present disclosure and achieve the same technical effect.

[0197] The above describes specific embodiments of the present disclosure, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0198] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0199] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for generating virtual data, characterized in that: include: Selecting a pre-built virtual building model, the virtual building model is composed of a plurality of building planes, each of the building planes is provided with a corresponding building label; The building plane is the outer surface area of ​​the virtual building model; A plurality of building images of the virtual building model are obtained; for each of the building images, when the building image contains a first viewing plane of the virtual building model, in response to determining that a first viewing plane including an occlusion map exists in each of the first viewing planes in the building image, an intersection operation is performed on the first viewing plane including the occlusion map and the occlusion map it includes, the occlusion map is removed, and an intersection plane corresponding to the first viewing plane including the occlusion map is obtained; each first viewing plane in the building image that does not contain an occlusion map and each of the intersection planes are determined as target planes; a detection frame detection is performed on each of the target planes, and a target plane whose detection frame area is smaller than a preset detection area is eliminated, and a viewing image corresponding to the building image is obtained, each of the viewing images includes at least one building viewing plane, and each of the building viewing planes is a view of the corresponding building plane at the viewing angle corresponding to the viewing image; Performing enhancement processing on each building viewing plane and the building label corresponding to the building viewing plane in each viewing image to obtain a training plane corresponding to each building viewing plane and the training label corresponding to the training plane; Each training plane in each of the viewing angle images and a training label corresponding to each training plane are used as virtual data for neural network training.

2. The method according to claim 1, characterized in that The construction process of the virtual building model includes: Obtain a preset geometric body, wherein the geometric body is composed of a plurality of triangular faces; Randomly select a triangular face from the multiple triangular faces as a starting triangular face; A first building plane corresponding to the starting triangular face is divided in the geometric body, and a corresponding building label is assigned to the first building plane, wherein the first building plane is composed of at least one triangular face, each triangular face constituting the first building plane includes at least the starting triangular face, and the similarity between normals of any two adjacent triangular faces in each triangular face constituting the first building plane is less than a preset similarity threshold; A new starting triangular facet is randomly selected from the remaining triangular faces among the multiple triangular faces, and a building plane corresponding to the new starting triangular facet is divided out in the geometric body and a building label is assigned, until all triangular faces in the geometric body are divided into their corresponding building planes, thereby completing the construction process of the virtual building model.

3. The method according to claim 1, characterized in that The process of determining whether there is a first viewing plane including an occlusion map among the first viewing planes in the building image comprises: At a shooting angle corresponding to the building image, a first image corresponding to the building image is shot; the first image includes contour information of each of the first viewing planes at the shooting angle and building labels of each of the first viewing planes; For each first viewing plane, in response to detecting that there is an area without a building label in the contour information of the first viewing plane, it is determined that the first viewing plane includes a corresponding occlusion map.

4. The method according to claim 1, characterized in that: The process of determining whether there is a first viewing plane including an occlusion map among the first viewing planes in the building image comprises: Acquire a standard image of a shooting angle corresponding to each of the building images, wherein the standard image includes standard viewing planes corresponding to each first viewing plane of the building image in a simulated scene without obstacles; Each first viewing plane of each of the building images is compared with a standard viewing plane corresponding to the first viewing plane, and in response to detecting that the first viewing plane is inconsistent with the standard viewing plane, the first viewing plane is determined to be a first viewing plane including an occlusion map.

5. The method according to claim 1, characterized in that The step of performing enhancement processing on each building viewing plane and the building label corresponding to the building viewing plane in each viewing image to obtain a training plane corresponding to each building viewing plane and the training label corresponding to the training plane includes: Enhance the plane features of each building viewing plane and the label features of each building label corresponding to the building viewing plane to obtain an enhanced building viewing plane and an enhanced building label; The plane features of each building perspective plane and the plane features of the enhanced building perspective plane corresponding to the building perspective plane are feature stitched to obtain the training plane corresponding to the building perspective plane, and the label features of the building label corresponding to the building plane and the label features of the enhanced building label corresponding to the building label are feature stitched to obtain the training label corresponding to the training plane.

6. A virtual data generating device, characterized in that: include: A selection unit is configured to select a pre-built virtual building model, wherein the virtual building model is composed of a plurality of building planes, and each of the building planes is provided with a corresponding building label; The building plane is the outer surface area of ​​the virtual building model; A sampling unit is configured to execute the acquisition of multiple building images of the virtual building model; for each of the building images, when the building image contains a first viewing plane of the virtual building model, in response to determining that there is a first viewing plane including an occlusion map in each of the first viewing planes in the building image, an intersection operation is performed on the first viewing plane including the occlusion map and the occlusion map it includes, and the occlusion map is removed to obtain an intersection plane corresponding to the first viewing plane including the occlusion map; each first viewing plane in the building image that does not contain the occlusion map and each of the intersection planes are determined as target planes; a detection frame detection is performed on each of the target planes, and a target plane whose detection frame area is smaller than a preset detection area is eliminated to obtain a viewing image corresponding to the building image, each of the viewing images includes at least one building viewing plane, and each of the building viewing planes is a view of the corresponding building plane at the viewing angle corresponding to the viewing image; an enhancement unit, configured to perform enhancement processing on each building viewing plane and the building label corresponding to the building viewing plane in each viewing image, to obtain a training plane corresponding to each building viewing plane and the training label corresponding to the training plane; The processing unit is configured to use each training plane in each of the perspective images and the training label corresponding to each training plane as virtual data for neural network training.

7. The device according to claim 6, characterized in that The virtual building model construction unit is further included; the virtual building model construction unit is configured to execute: Obtain a preset geometric body, wherein the geometric body is composed of a plurality of triangular faces; Randomly select a triangular face from the multiple triangular faces as a starting triangular face; A first building plane corresponding to the starting triangular face is divided in the geometric body, and a corresponding building label is assigned to the first building plane, wherein the first building plane is composed of at least one triangular face, each triangular face constituting the first building plane includes at least the starting triangular face, and the similarity between normals of any two adjacent triangular faces in each triangular face constituting the first building plane is less than a preset similarity threshold; A new starting triangular facet is randomly selected from the remaining triangular faces among the multiple triangular faces, and a building plane corresponding to the new starting triangular facet is divided out in the geometric body and a building label is assigned, until all triangular faces in the geometric body are divided into their corresponding building planes, thereby completing the construction process of the virtual building model.

8. The device according to claim 6, characterized in that The first determining subunit includes: The image capturing subunit is configured to capture a first image corresponding to the building image at a capturing angle corresponding to each of the building images; the first image includes contour information of each of the first viewing planes at the capturing angle and a building label of each of the first viewing planes; The second determining subunit is configured to execute, for each first viewing plane, in response to detecting that there is an area without a building label in the contour information of the first viewing plane, determining that the first viewing plane includes a corresponding occlusion map.

9. The device according to claim 6, characterized in that The first determining subunit includes: An acquisition subunit is configured to acquire a standard image corresponding to a shooting angle of each of the building images, wherein the standard image includes standard viewing planes corresponding to each first viewing plane of the building image in a simulated scene without obstacles; The third determination subunit is configured to compare each first viewing plane of each of the building images with a standard viewing plane corresponding to the first viewing plane, and in response to detecting that the first viewing plane is inconsistent with the standard viewing plane, determine that the first viewing plane is a first viewing plane including an occlusion map.

10. The device according to claim 6, characterized in that The enhancement unit comprises: The first processing subunit is configured to enhance the plane feature of each building viewing plane and the label feature of each building label corresponding to the building viewing plane to obtain an enhanced building viewing plane and an enhanced building label; The second processing sub-unit is configured to perform feature stitching on the plane features of each building perspective plane and the plane features of the enhanced building perspective plane corresponding to the building perspective plane to obtain a training plane corresponding to the building perspective plane, and to perform feature stitching on the label features of the building label corresponding to the building plane and the label features of the enhanced building label corresponding to the building label to obtain a training label corresponding to the training plane.

11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the virtual data generating method according to any one of claims 1 to 5.

12. A storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the virtual data generating method as claimed in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image data labeling method and device

    CN110189406A