A method for virtualizing a scene
By generating model data of virtual entities based on interactive information and video data, the problems of high complexity and time-consuming in three-dimensional scene modeling are solved, and efficient virtual scene creation and realization are achieved.
Patent Information
- Application Number
- CN202210614156.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The generation process of existing three-dimensional scene modeling schemes is complex and time-consuming, resulting in low realism of virtual scenes and lag in operation.
By determining the physical entities within the scene boundary based on interactive information indicating the scene boundary, and using video data to capture the information of the physical entities, generating model data of the virtual entity, and finally creating a virtual scene.
The scene model generation process is simplified, the reality and operation efficiency of the virtual scene are improved, and the calculation complexity and time cost are reduced.
Smart Images

Figure CN114972599B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of virtual reality and digital twins, and more particularly to a method for virtualizing a scene, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Digital Twins (English: Digital Twins) is a process that makes full use of data such as physical models, sensor updates, and operating history, integrates multi-disciplinary, multi-physical quantity, multi-scale, and multi-probability simulation processes, and completes mapping in a virtual space, thereby reflecting the entire life cycle process of the corresponding physical entity. Digital Twins is a concept that transcends reality and can be regarded as a digital mapping system of one or more important and interdependent equipment systems.
[0003] Digital Twin technology can also be combined with Extended Reality (XR) technology. Extended Reality technology specifically includes Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), etc.
[0004] Digital Twin technology has been widely applied in the field of engineering construction, especially in the field of three-dimensional scene modeling. Visual three-dimensional scene applications based on three-dimensional scene models have become widely popular. Currently, there are three-dimensional engines that can assist in the research and development of visual three-dimensional scene applications. In addition, due to the virtualization attributes of three-dimensional scenes, it often involves the simultaneous operation of scene modeling applications and virtual reality applications. However, the model generation process of current three-dimensional scene modeling solutions is not only complex and time-consuming, but also requires a large amount of data to be collected in advance. Therefore, in the actual application process, there are often situations of lag and too low realism of the simulated virtual scene.
[0005] Therefore, the present disclosure proposes a method for virtualizing a scene, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical problems of high computational complexity and long time consumption in the process of scene virtualization. Summary of the Invention
[0006] An embodiment of the present disclosure provides a method for virtualizing a physical scene, including: determining the scene boundary based on interaction information for indicating the scene boundary; determining physical entities within the scene boundary based on the scene boundary, and capturing video data corresponding to the physical entities; determining model data of virtual entities corresponding to the physical entities based on the video data corresponding to the physical entities; and creating a virtual scene corresponding to the physical scene based on the model data of the virtual entities corresponding to the physical entities.
[0007] For example, the video data includes a plurality of video frames, and different video frames among the plurality of video frames correspond to different lighting conditions, shooting positions, or shooting angles.
[0008] For example, the determining of the model data of the virtual entity corresponding to the physical entity based on the video data corresponding to the physical entity further includes: extracting a plurality of discrete points from each video frame in the video data; generating, based on the plurality of discrete points of each video frame, a stereoscopic model data represented by Thiessen polygons as the stereoscopic model data of the video frame; and determining the model data of the virtual entity corresponding to the physical entity based on the stereoscopic model data of each video frame.
[0009] For example, the determining of the model data of the virtual entity corresponding to the physical entity based on the video data corresponding to the physical entity further includes: obtaining one or more of a building information model, global geographical location information, and building positioning space data; and determining the model data of the virtual entity corresponding to the physical entity by using the video data corresponding to the physical entity based on one or more of the building information model, the global geographical location information, and the building positioning space data.
[0010] For example, the determining of the model data of the virtual entity corresponding to the physical entity based on the video data corresponding to the physical entity further includes: obtaining one or more of urban traffic data, urban planning data, and urban municipal data; and determining the model data of the virtual entity corresponding to the physical entity by using the video data corresponding to the physical entity based on one or more of the urban traffic data, the urban planning data, and the urban municipal data.
[0011] For example, the method further includes: displaying relevant information of the virtual scene based on the virtual scene corresponding to the physical scene.
[0012] For example, the displaying of the relevant information of the virtual scene further includes: selecting a plurality of video frames from the video data; performing texture compression and / or texture scaling processing on the plurality of video frames to generate texture mapping data; and rendering the virtual scene corresponding to the physical scene based on the texture mapping data and displaying the rendered virtual scene.
[0013] For example, the step of performing texture compression and / or texture scaling on the multiple video frames to generate texture map data further includes: performing texture compression on the multiple video frames to generate texture-compressed texture map data; determining material resource data and material resource data corresponding to the texture map data based on the texture-compressed texture map data; determining parameters corresponding to the texture scaling process based on the material resource data and material resource data corresponding to the texture map data; and performing texture scaling on the texture-compressed texture map data based on the parameters corresponding to the texture scaling process to generate texture-scaled texture map data.
[0014] Some embodiments of the present disclosure provide an electronic device, including: a processor; a memory storing computer instructions, which when executed by the processor implement the above method.
[0015] Some embodiments of the present disclosure provide a computer-readable storage medium, having stored thereon computer instructions, which when executed by a processor implement the above method.
[0016] Some embodiments of the present disclosure provide a computer program product, which includes computer-readable instructions that, when executed by a processor, cause the processor to execute the above method.
[0017] Thus, in response to the requirements of application service visualization and scenario virtualization, various embodiments of the present disclosure utilize video data to implement scenario virtualization, which helps to solve the technical problems of high complexity and long duration in the process of generating scenario models. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for the description of the embodiments. The drawings in the following description are merely exemplary embodiments of the present disclosure.
[0019] Figure 1 is an exemplary schematic diagram showing an application scenario according to an embodiment of the present disclosure.
[0020] Figure 2 is a flowchart showing an exemplary method for virtualizing a physical scenario according to an embodiment of the present disclosure.
[0021] Figure 3 is a schematic diagram showing a physical scenario, interaction information, and physical entities according to an embodiment of the present disclosure.
[0022] Figure 4 is an exemplary schematic diagram showing the change of an interface when a terminal obtains interaction information according to an embodiment of the present disclosure.
[0023] Figure 5It is a schematic diagram showing the acquisition of interaction information according to an embodiment of the present disclosure.
[0024] Figure 6 It is a schematic diagram showing the processing of video frames according to an embodiment of the present disclosure.
[0025] Figure 7 It is a schematic diagram showing the processing of video frames combined with building information according to an embodiment of the present disclosure.
[0026] Figure 8 It is a schematic diagram showing the processing of video frames combined with geographical information according to an embodiment of the present disclosure.
[0027] Figure 9 It is a schematic diagram showing the architecture of a scene modeling application and / or a virtual reality application according to an embodiment of the present disclosure.
[0028] Figure 10 It is a schematic diagram showing the operation of a rendering engine according to an embodiment of the present disclosure.
[0029] Figure 11 It shows a schematic diagram of an electronic device according to an embodiment of the present disclosure.
[0030] Figure 12 It shows a schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure.
[0031] Figure 13 It shows a schematic diagram of a storage medium according to an embodiment of the present disclosure. Detailed implementation manners
[0032] In order to make the objectives, technical solutions, and advantages of the present disclosure more apparent, exemplary embodiments according to the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0033] In this specification and the accompanying drawings, operations and elements that are substantially the same or similar are denoted by the same or similar reference numerals, and repeated descriptions of these operations and elements will be omitted. Also, in the description of the present disclosure, terms such as "first", "second", etc. are used to distinguish between identical or similar items with substantially the same function and role. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various examples, the first data may be referred to as the second data, and similarly, the second data may be referred to as the first data. Both the first data and the second data can be data, and in some cases, they can be separate and different data. The meaning of the term "at least one" in this application refers to one or more, and the meaning of the term "a plurality of" in this application refers to two or more. For example, a plurality of audio frames means two or more audio frames.
[0034] It should be understood that the terms used in the description of the various examples herein are only for the purpose of describing specific examples and are not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0035] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this application generally represents an "or" relationship between the preceding and following associated objects.
[0036] It should also be understood that in the various embodiments of this application, the magnitude of the serial numbers of the various processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic and should not constitute any limitation on the implementation process of the embodiments of this application. It should also be understood that determining B based on (or on the basis of) A does not mean determining B only based on (or on the basis of) A. B can also be determined based on (or on the basis of) A and / or other information.
[0037] It should also be understood that the term "comprising" (also referred to as "includes", "including", "Comprises" and / or "Comprising") when used in this specification specifies the presence of the stated features, integers, operations, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, operations, operations, elements, components, and / or their groupings.
[0038] It should also be understood that the term "if" can be interpreted to mean "when" ("when" or "upon") or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined..." or "if [the stated condition or event] is detected" can be interpreted to mean "when it is determined..." or "in response to determining..." or "when [the stated condition or event] is detected" or "in response to detecting [the stated condition or event]".
[0039] To facilitate the description of the present disclosure, the following concepts related to the present disclosure are introduced.
[0040] First, with reference to Figure 1 the application scenarios of various aspects of the present disclosure are described. Figure 1 A schematic diagram of application scenario 100 according to an embodiment of the present disclosure is shown, in which a server 110 and a plurality of terminals 120 are schematically shown. The terminals 120 and the server 110 can be directly or indirectly connected by wired or wireless communication means, and the present disclosure does not limit this here.
[0041] As Figure 1 shown, the embodiments of the present disclosure adopt Internet technology, especially Internet of Things technology. The Internet of Things can be an extension of the Internet, which includes the Internet and all resources on the Internet and is compatible with all applications on the Internet. With the application of Internet of Things technology in various fields, various new application fields of intelligent Internet of Things such as smart home, intelligent transportation, and intelligent health have emerged.
[0042] According to some embodiments of the present disclosure, it is used to process scenario data. These scenario data may be data related to Internet of Things technology. The scenario data includes XX. Of course, the present disclosure is not limited thereto.
[0043] For example, the method according to some embodiments of the present disclosure can be wholly or partially implemented on the server 110 to process scene data, such as scene data in the form of pictures. For example, the server 110 will be used to analyze the scene data and determine model data based on the analysis results. Here, the server 110 can be an independent server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), location services, and big data and artificial intelligence platforms. The embodiments of the present disclosure do not make specific limitations in this regard. Hereinafter, the server 110 will also be referred to as the cloud.
[0044] For example, the method according to the embodiments of the present disclosure can also be wholly or partially implemented on the terminal 120 to process scene data. For example, the terminal 120 will be used to collect the above-mentioned scene data in the form of pictures. For another example, the terminal 120 will be used to present the scene data so that the user can interact with the constructed three-dimensional model in the virtual scene. For example, the terminal 120 can be an interactive device that can provide 3D digital virtual objects and includes a display device with a user interface, and can display 3D digital virtual objects through the user interface. The user can interact with the interactive device for information. For another example, the terminal 120 will also be used to analyze the above-mentioned building data. The present disclosure does not make any limitations in this regard.
[0045] For example, each of the multiple terminals 120 can be a fixed terminal such as a desktop computer, or a mobile terminal with network functions such as a smart phone, a tablet computer, a portable computer, a handheld device, a personal digital assistant, a smart wearable device (such as smart glasses), a smart head-mounted device, a camera, a vehicle-mounted terminal, or any combination thereof. The embodiments of the present disclosure do not make specific limitations in this regard. Each of the multiple terminals 120 may further include various sensors or data collection devices, such as Figure 1 the temperature sensor shown in. In some examples, the scene data is related to the lighting conditions, so the terminal can also be a brightness sensor. In still other examples, the terminal 120 can also be a camera (such as an infrared camera) or a distance detector.
[0046] Each of the above-mentioned terminals 120 can incorporate augmented reality (AR) technology and virtual reality (VR) technology. Among them, augmented reality technology is a technology that integrates virtual scene data with the real scene, widely using various technical means such as multimedia, 3D modeling, real-time tracking and registration, intelligent interaction, and sensing. After simulating and emulating virtual information such as text, images, 3D models, music, and videos generated by a computer, it is applied to the real world, and the two types of information complement each other, thereby realizing the "augmentation" of the real world. Virtual reality uses a computer to simulate a virtual world in a three-dimensional space for a real scene, providing users with simulations of senses such as vision, making the user feel as if they are on the scene and can observe things in the three-dimensional space in real time and without restrictions. When the user moves their position, the computer can immediately perform complex calculations and transmit the accurate three-dimensional world image back to generate a sense of presence.
[0047] Take Figure 1 the smart glasses shown in as an example to further illustrate the terminal 120 incorporating augmented reality technology and virtual reality technology. The smart glasses not only include various optical components and support components of conventional glasses, but also include a display component for displaying the above-mentioned augmented reality information and / or virtual reality information. The smart glasses also include corresponding battery components, sensor components, network components, and so on. Among them, the sensor component can include a depth camera (for example, a Kinect depth camera), which captures the depth information in the real scene through the principle of amplitude-modulated continuous wave (AMCW) time-of-flight (TOF) ranging and uses near-infrared light (NIR) to generate a depth map corresponding to the real scene. The sensor component can also include various acceleration sensors, gyroscope sensors, and geomagnetic field sensors, etc., for detecting the user's posture and position information, thereby providing reference information for the processing of scene data. Various eye-tracking accessories may also be integrated on the smart glasses to build a bridge between the real world, the virtual world, and the user through the user's eye movement data, thereby providing a more natural user experience. Those skilled in the art should understand that although the terminal 120 is further illustrated by taking the smart glasses as an example, the present disclosure does not impose any restrictions on the types of terminals.
[0048] It can be understood that the embodiments of the present disclosure may further involve artificial intelligence services to intelligently provide the above-mentioned virtual scenes. The artificial intelligence services can be executed not only on the server 110, but also on the terminal 120, or jointly executed by the terminal and the server. The present disclosure does not limit this. In addition, it can be understood that the device for analyzing and reasoning the scene data using the artificial service of the embodiments of the present disclosure can be either a terminal, a server, or a system composed of a terminal and a server.
[0049] Currently, digital twin technology has been widely applied in the field of engineering construction, especially in the field of three-dimensional scene modeling. Visual three-dimensional scene applications based on three-dimensional scene models have become widely popular. There are currently many three-dimensional engines that can assist in the research and development of visual three-dimensional scene applications. In addition, due to the virtual nature of three-dimensional scenes, it often involves the simultaneous operation of scene modeling applications and virtual reality applications. However, in the current three-dimensional scene modeling solution, the model generation process is not only complex and time-consuming, but also requires a large amount of data to be collected in advance. Therefore, in the actual application process, there are often situations of lag and too low realism of the simulated virtual scene.
[0050] For example, there is currently such a technical solution: taking six pictures of a certain scene from six fixed angles of up, down, left, right, front, and back from a fixed point, and then pasting these six pictures onto a spatial scene model in the form of a cube through a texture mapping scheme.
[0051] Since in the actual display process, it is necessary to stretch and deform the texture data, the virtual three-dimensional scenes generated by such a scheme often have poor realism. In addition, since there are often differences in the shooting times of these six pictures, the six pictures all correspond to different lighting scenes. Therefore, the actually generated virtual scenes often have difficulty simulating real lighting conditions, resulting in distortion of the virtual scenes. Furthermore, since these six pictures are simply pasted onto a spatial scene model in the form of a cube, it often requires a large amount of information collected in advance and a large amount of computing resources to accurately determine the information that meets the requirements of scene modeling applications, resulting in difficulty in simultaneously running scene modeling applications and virtual reality applications.
[0052] Therefore, the embodiments of the present disclosure provide a method for virtualizing a physical scene, including: determining physical entities within the scene boundary based on interaction information indicating the scene boundary, and capturing video data corresponding to the physical entities; determining model data of virtual entities corresponding to the physical entities based on the video data; and creating a virtual scene corresponding to the physical scene based on the model data corresponding to the virtual entities. Thus, in response to the requirements of application business visualization and scene virtualization, the embodiments of the present disclosure utilize video data to achieve scene virtualization, which helps to solve the technical problems of high complexity and long time consumption in the scene model generation process.
[0053] Hereinafter, reference will be made to Figures 2 to 12 for a further description of the embodiments of the present disclosure.
[0054] As an example, Figure 2 is a flowchart showing an example method 20 for virtualizing a physical scene according to an embodiment of the present disclosure. Figure 3 is a schematic diagram showing a physical scene, interaction information, and physical entities according to an embodiment of the present disclosure.
[0055] See Figure 2 , the exemplary method 20 may include one or all of operations S201 - S203, or may include more operations. The present disclosure is not limited thereto. As described above, operations S201 to S203 are executed in real time by the terminal 120 / server 110, or are executed offline by the terminal 120 / server 110. The present disclosure does not limit the execution entity of each operation of the exemplary method 200, as long as it can achieve the purpose of the present disclosure. Each step in the exemplary method may be executed in whole or in part by a virtual reality application and / or a scene modeling application. The virtual reality application and the scene modeling application may be integrated into a large application, or the virtual reality application and the scene modeling application may be two independent applications, but transmit interaction information, video data, model data, etc. through the mutually open interfaces between the two. The present disclosure is not limited thereto.
[0056] For example, in operation S201, based on the interaction information indicating the scene boundary, the scene boundary is determined. In operation S202, based on the scene boundary, physical entities within the scene boundary are determined, and video data corresponding to the physical entities is captured.
[0057] For example, the interaction information may be collected by the terminal 120 in Figure 1 , which indicates which physical entities in the physical scene need to be further virtualized. For example, as shown in Figure 3 , it shows an example of a physical scene, interaction information, and physical entities, and schematically shows an example of a physical scene including physical entities such as a sofa, curtains, a moon, a table lamp, a storage cabinet, and books. For such a physical scene, interaction information shown in a circular frame can be obtained, which indicates that only the physical entities and the physical scene within the circular frame need to be virtualized. That is, in the example of Figure 3 , it can correspondingly determine that the physical entities in the scene only include a table lamp, a storage cabinet, and books. Then, video data corresponding to the table lamp, the storage cabinet, and the books can be captured. Although the scene boundary is shown in the form of a circular frame in Figure 3 , those skilled in the art should understand that the present disclosure is not limited thereto. Specifically, the scene boundary can also be indicated by any connected shape. Various examples of interaction information will be described in detail with reference to Figures 4 to 5 , and the present disclosure will not elaborate herein.
[0058] As an example, the video data corresponding to the physical entity refers to a continuous sequence of images, which is essentially composed of groups of continuous images. Each image in the image sequence is also called a video frame, which is the smallest visual unit that makes up the video. It can be with reference to Figure 1The various terminals 120 described are used to collect the video data. For example, devices such as smart glasses, mobile terminals, and depth cameras can be used to collect the video data. Since the video data captures images (video frames) of a physical entity over a period of time, different video frames in the multiple video frames correspond to different lighting conditions, shooting positions, or shooting angles. Therefore, each video frame in the video data includes various information about the physical entity. According to various experiments adopting the embodiments of the present disclosure, it can be determined that sufficient information capable of characterizing the physical entity can be extracted from the video data including 300 frames, so as to implement the modeling process of a virtual entity with high authenticity.
[0059] In operation S203, based on the video data corresponding to the physical entity, model data of the virtual entity corresponding to the physical entity is determined.
[0060] Optionally, although the video data is collected by the terminal 120, the analysis and processing of the video data can be performed by the server 110. For example, the terminal 120 can transmit the video data to the server by streaming, and then the server 110 can process the video data corresponding to the physical entity (such as image processing, etc.) to obtain the model data of the virtual entity corresponding to the physical entity. In addition, the server 110 can also combine various known information or connect to public or non-public databases through various interfaces to obtain information related to the physical entity as the model data of the virtual entity.
[0061] For example, the model data of the virtual entity indicates any data that can be used to build the virtual entity in the virtual scene. For example, it can be the edge information, position information, depth information, vertex information, height information, width information, length information, etc. of the virtual entity extracted from each video frame of the video data. The model data of the virtual entity can also be the environmental information where the virtual entity is located extracted from each video frame of the video data, such as lighting information, relative position relationship information, etc. Even when the physical entity is an Internet of Things device, the model data of the virtual entity can also include Internet of Things related information, such as network status, registration request information, registered entity information, device operation information, etc. Or, any data related to the physical entity can also be pulled from the Internet / database based on the analysis of the video data. The present disclosure does not limit this. Then, various examples of interaction information will be described in detail with reference to Figure 6 and various examples of interaction information will not be elaborated herein.
[0062] In operation S204, based on the model data corresponding to the virtual entity, a virtual scene corresponding to the physical scene is created.
[0063] Optionally, the virtual scene is a three-dimensional virtual scene, which is a virtualization of a real physical scene. A three-dimensional virtual model corresponding to the virtual entity is placed within the three-dimensional virtual scene. The three-dimensional virtual model is also referred to as a 3D model, which can be created using various 3D software. In combination with the various embodiments of the present disclosure described in detail below, the software for creating the 3D model in the present disclosure is, for example, CAD (CAD - Computer Aided Design) software. In these examples, a 3D model file in STL format can be obtained through this software; then, the STL format file is imported into the slicing processing pipeline of a 3D software capable of slicing to obtain the three-dimensional virtual model. In addition, before creating the three-dimensional virtual model, the model data can be structurally optimized to save computing resources and improve processing efficiency. It should be noted that the present disclosure does not limit the type of 3D software. For example, it can be software for 3D model analysis, 3D software for visual art creation, 3D software for 3D printing, and so on; in addition, a three-dimensional model can also be created and generated through a computer graphics library (i.e., the graphics library used during self-programming); for example, (Open Graphics Library), DirectX (Direct eXtension), and so on.
[0064] Optionally, method 20 may further include operation S205. In operation S205, based on the virtual scene corresponding to the physical scene, relevant information of the virtual scene is displayed. For example, the virtual scene is displayed in a three-dimensional form.
[0065] Optionally, various three-dimensional rendering engines can be used to visualize the virtual scene. The three-dimensional rendering engine can realize generating a displayable two-dimensional image from a digital three-dimensional scene. The generated two-dimensional image can be realistic or non-realistic. The process of three-dimensional rendering needs to rely on a 3D rendering engine to generate. In combination with the various embodiments of the present disclosure described in detail below, the example rendering engine in the present disclosure can use the "ray tracing" technology, which generates an image by tracing the light rays from the camera through the virtual plane of the pixels and simulating the effect of their encounter with objects. The example rendering engine in the present disclosure can also use the "rasterization" technology, which determines the values of each pixel in the two-dimensional image by collecting relevant information of various patches. The present disclosure does not limit the type of 3D rendering engine and the technology adopted.
[0066] Thus, in response to the requirements of application service visualization and scene virtualization, method 20 utilizes video data to achieve scene virtualization, which helps to solve the technical problems of high complexity and long duration in the process of generating a scene model.
[0067] Next, refer to Figure 4 and Figure 5To further describe examples of operations S201 to S202. Among them, Figure 4 is a schematic diagram showing example interface changes when a terminal obtains interaction information according to an embodiment of the present disclosure. Figure 5 is a schematic diagram showing the acquisition of interaction information according to an embodiment of the present disclosure.
[0068] As Figure 4 shown, the terminal 120 may be equipped with a scene modeling application and / or a virtual reality application. In response to the activation of the scene modeling application and / or the virtual reality application, the terminal 120 may trigger a "gesture selection" related function for obtaining interaction information indicating the scene boundary. Specifically, in response to the terminal 120 being a smart glasses or a smart phone, the left figure in Figure 4 may be seen through the smart glasses or using the camera of the smart phone. By triggering a dialog box on the display screen, the smart glasses or the smart phone will capture the user's gesture. For example, the user may gesture an irregular area range in the air in front of the smart glasses. Or, for example, the user may hold the smart phone in one hand and gesture an irregular area range in the area that can be photographed by the camera of the smart phone with the other hand. The smart glasses or the smart phone will identify the gesture to obtain a scene boundary that can be described by a vectorial continuous vector. When it is closed in the head and tail orientation, it can generate a convex polygon closed area as shown in Figure 4 and Figure 5 shown.
[0069] Furthermore, as Figure 5 shown, starting from the imaging component (for example, the camera of the smart glasses or the smart phone), based on the distances from multiple points on the edge of the above-mentioned convex polygon closed area to the vertical plane where the starting point is located. Based on the distances from the multiple points to the vertical plane where the starting point is located, select the shortest distance as the shortest distance corresponding to the convex polygon closed area. Based on the shortest distance corresponding to the convex polygon closed area, determine the first vertical plane. For example, the first vertical plane is perpendicular to the horizontal plane, and the horizontal distance between the first vertical plane and the imaging component is the shortest distance corresponding to the convex polygon closed area. Then, based on the first vertical plane, determine a circular plane area. The circular plane area is used to assist in determining whether a certain physical entity is located within the scene boundary.
[0070] For example, the highest point and the lowest point on the convex polygon closed region can be projected onto the first vertical plane, and a line connecting the projection of the highest point and the projection of the lowest point on the first vertical plane can be used as the diameter. Taking the center of this line as the center of the circle, the circular plane region can be determined. For another example, the leftmost point and the rightmost point on the convex polygon closed region can be projected onto the first vertical plane, and a line connecting the projection of the leftmost point and the projection of the rightmost point on the first vertical plane can be used as the diameter. Taking the center of this line as the center of the circle, the circular plane region can be determined. For yet another example, the longest diagonal of the convex polygon closed region can be projected onto the first vertical plane, and using the projection of the longest diagonal as the diameter and the center of the projection of the longest diagonal as the center of the circle, the circular plane region can be determined. The present disclosure does not further limit the method for determining the circular plane region.
[0071] Similarly, starting from the imaging component, the distances from multiple points on the edge of the physical entity to the vertical plane where the starting point is located are determined. Based on the distances from multiple points on the edge of the physical entity to the vertical plane where the starting point is located, the shortest distance corresponding to the physical entity is selected. Based on the shortest distance corresponding to the physical entity, a second vertical plane is determined. For example, the second vertical plane is perpendicular to the horizontal plane, and the horizontal distance between the second vertical plane and the imaging component is the shortest distance corresponding to the physical entity. Based on the ratio of the shortest distance corresponding to the convex polygon closed region to the shortest distance corresponding to the physical entity, an equally enlarged circular plane region is determined on the second vertical plane. The ratio of the diameter of the circular plane region to the diameter of the equally enlarged circular plane region is equal to the ratio of the shortest distance corresponding to the convex polygon closed region to the shortest distance corresponding to the physical entity, and the center of the circular plane region and the center of the equally enlarged circular plane region are on the same horizontal line.
[0072] If the projection of the physical entity on the equally enlarged circular plane region is entirely within the equally enlarged circular plane region, then it can be determined that the physical entity is inside the scene boundary. As Figure 4 and Figure 5 shown, it can be determined that the gray-marked physical entity is inside the scene boundary, while the white-marked physical entity is outside the scene boundary. Thus, determining the first vertical plane and the second vertical plane based on the shortest horizontal distance corresponding to the convex polygon closed region can achieve smaller errors. Of course, the present disclosure is not limited thereto.
[0073] Figure 4 and Figure 5This is only an example solution for using a gesture tracking solution to obtain interaction information indicating the scene boundary and determining physical entities within the scene boundary. The present disclosure is not limited thereto. For example, in a virtual reality application, it can also first determine multiple physical entities that can be captured by a camera component through an infrared sensing or dynamic image recognition solution, and prompt the user to select from the multiple physical entities through a voice or text dialog box. In such a case, the information selected by the user from the multiple physical entities will be used as the interaction information indicating the scene boundary. For another example, in a virtual reality application, it can also first capture a static image, perform edge extraction on the static image to draw buttons covering the captured physical entities on the static image, and the user triggers the buttons by clicking / touching / gesture indication, etc. to select the physical entities that need to be virtualized from the multiple physical entities. In such a case, the information triggered by the user for the buttons can also be used as the interaction information indicating the scene boundary.
[0074] Next, the camera component will capture video data corresponding to the physical entities within the scene boundary. For example, the camera component can continuously automatically / manually adjust shooting parameters during the shooting period, such as adjusting the focus, focal length, position of the camera component, intermittently turning on the flash, intermittently turning on the high beam, intermittently turning on the low beam, etc. to capture the video data corresponding to the physical entities, so that the video data includes more information. Of course, in some examples, the camera component can also not make any adjustment to the shooting parameters during the shooting period. Since during the operation of the virtual reality application, the ambient light often has changes that can be captured by the device, the captured video data often also includes sufficient information to provide sufficient model data for the virtual entities.
[0075] Thus, through the virtual reality application, various aspects of the present disclosure provide interaction information for indicating the scene boundary by adopting rich human-computer interaction methods, and can conveniently determine the physical entities within the scene boundary, providing sufficient model data for the subsequent creation of the virtual scene.
[0076] Next, refer to Figures 6 to 8 to further describe an example of operation S202. Among them, Figure 6 is a schematic diagram showing the processing of a video frame according to an embodiment of the present disclosure. Figure 7 is a schematic diagram showing the processing of a video frame combined with building information according to an embodiment of the present disclosure. Figure 8 is a schematic diagram showing the processing of a video frame combined with geographical information according to an embodiment of the present disclosure.
[0077] Optionally, operation S202 includes extracting a plurality of discrete points from each video frame in the video data; generating stereoscopic model data represented by Thiessen polygons as the stereoscopic model data of each video frame based on the plurality of discrete points of each video frame; and determining model data of a virtual entity corresponding to the physical entity based on the stereoscopic model data of each video frame.
[0078] Figure 6 An example of a scene modeling application and / or a virtual reality application for a video frame in video data is shown. The video data captures a physical entity shown in the form of a cup. Those skilled in the art should understand that Figure 6 This is only a schematic diagram for illustrating the solution of the present disclosure, and the actual video data may also include more or fewer pixels and information in a single video frame.
[0079] As an example, a scene modeling application and / or a virtual reality application will extract the video frame marked as 601 from the video data. Then, a plurality of discrete points marked with black dots in the image marked as 602 in the video frame marked as 601 can be extracted. Each of the plurality of discrete points indicates information associated with the physical entity. Examples of discrete points can be the vertices, center points, feature points, and points with the most drastic changes in brightness and darkness of the cup. As an example, 20 to 30 discrete points can be extracted in a single video frame. Of course, the embodiments of the present disclosure are not limited thereto.
[0080] The discrete points can be extracted in various ways, and the present disclosure does not limit the way of extracting discrete points. For example, a grayscale image can be generated from the video frame to determine the brightness and darkness changes of each pixel. Then, a heat map is generated based on the brightness and darkness changes of each pixel to obtain the brightness and darkness change distribution of the video frame. Based on the brightness and darkness change distribution, the coordinates of a plurality of discrete points are determined, and these discrete points all indicate the brightness and darkness change information of the video frame.
[0081] For another example, a neural network can be used to intelligently identify multiple discrete points in the video frame, and each discrete point can be a feature point in the video frame. Various neural network models can be used to determine these discrete points. For example, a deep neural network (DNN) model, a factorization machine (FM) model, etc. can be adopted. These neural network models can be implemented as acyclic graphs, where neurons are arranged in different layers. Generally, a neural network model includes an input layer and an output layer, and the input layer and the output layer are separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating an output in the output layer. Network nodes are fully connected to nodes in adjacent layers via edges, and there are no edges between nodes within each layer. The data received at the nodes of the input layer of the neural network is propagated to the nodes of the output layer via any one of a hidden layer, an activation layer, a pooling layer, a convolutional layer, etc. The input and output of the neural network model can take various forms, and the present disclosure places no restrictions thereon.
[0082] Continuing with this example, a solid model data characterized by Delaunay triangles can be generated based on the extracted discrete points. For example, any one of these discrete points can be selected as the first discrete point, and then the point closest to this point is found as the second discrete point, and the first discrete point and the second discrete point are connected as the first baseline. The point closest to the first baseline is found as the third discrete point, and the first discrete point and the third discrete point are connected as the second baseline and the second discrete point and the third discrete point are connected as the third baseline. The first baseline, the second baseline, and the third baseline form the triangle marked in block 603. Then, the discrete points closest to the second baseline and the third baseline are found, and multiple triangles are repeatedly generated until the triangular mesh marked in block 604 is generated. Based on this triangular mesh, a solid model structure is formed by using the method of generating Delaunay triangles. Delaunay triangle generation takes any one discrete point as the center point, and then the center point is respectively connected to multiple surrounding discrete points, and then the perpendicular bisectors of the straight lines are respectively made. The polygon formed by the intersection of these perpendicular bisectors (thus, the area called the vicinity of the center point) is the Delaunay triangle. Thus, for each video frame, a solid model structure characterized by Delaunay triangles can be generated.
[0083] Since it is difficult for the physical structure and physical surface of the same physical entity to change in a short time (for example, within the period when video data is captured), for temporally adjacent or close video frames, the same discrete points in multiple video frames can be determined according to the similarity between the discrete points extracted from the video frames. Combining the principle that objects closer appear larger and those farther appear smaller, the depth information at each discrete point can be calculated. The depth information at each discrete point will be used as an example of the model data of the virtual entity corresponding to the physical entity.
[0084] Such as Figure 7As shown, if a scene modeling application and / or a virtual reality application needs to virtualize a scene including a large building (where the large building will be a physical entity), then the building information model (BIM model) of the large building can be further combined to determine the model data of the virtual entity corresponding to the physical entity. The BIM model, that is, the building information model, whose full English name is Building Information Modeling. In a BIM model, there is not only a three-dimensional model of a building, but also information such as the material properties, color, designer, manufacturer, constructor, inspector, date and time, area, volume, etc. of the building can be set. Each monitored virtual entity can be set as an entity object in the BIM model, which correspondingly includes an object identifier, geometric data of the object, reference geometric data of the object, data collected by the object in real time, and so on. The present disclosure is not limited thereto.
[0085] In addition, the global geographical location information corresponding to the large building can be further combined to determine the model data of the virtual entity corresponding to the physical entity. Among them, the global geographical location information can be the information found in the map database according to some features of the physical entity. For example, the longitude and latitude information corresponding to the physical entity can be found through various navigation map applications as the global geographical location information. For another example, based on the position data of the terminal 120 determined by the positioning module of the terminal 120 (such as a GPS positioning module, a Beidou system positioning module), the position of the physical entity within a certain range from the mobile phone can be further determined. The present disclosure does not further limit the global geographical location information.
[0086] In addition, the building positioning space data corresponding to the large building can be further combined to determine the model data of the virtual entity corresponding to the physical entity. For example, the terminal 120 can pull the building positioning space data of the corresponding building from the building positioning space database, which includes the length, width and height data of the building, wall data, various design data when the building was approved, and so on. The present disclosure does not further limit the building positioning space data.
[0087] For example, the illumination information can be extracted from the above video data, and then the illumination information can be combined with the above building information model to determine the model data of the virtual entity corresponding to the physical entity. For another example, in combination with Figure 6 the method described therein, the stereoscopic model data of each video frame in the video data is generated, and in combination with one or more of the stereoscopic model data, the building information model, the global geographical location information, and the building positioning space data, the model data of the virtual entity corresponding to the physical entity is determined, so as to enable the presentation of virtual scenes under different illumination conditions. The present disclosure does not limit this.
[0088] As Figure 8 shown, if the scene modeling application and / or virtual reality application needs to virtualize a distant view (which includes multiple large buildings, and each large building will be regarded as a physical entity), then the model data of the virtual entity corresponding to the physical entity can be further determined by combining urban traffic data, urban planning data, urban municipal data, etc. The urban traffic data, urban planning data, and urban municipal data can be directly obtained from the web information related to the city or retrieved from the relevant database. The present disclosure does not limit this. The urban traffic data, urban planning data, and urban municipal data are all exemplary geographic information, and the present disclosure will not elaborate on this here.
[0089] Next, reference is made to Figure 9 to further describe an example of operation S203. Among them, Figure 9 is a schematic architecture diagram showing a scene modeling application and / or virtual reality application according to an embodiment of the present disclosure.
[0090] As Figure 9 shown, in the scene modeling application and / or virtual reality application, video data can be obtained from a data acquisition module (such as a camera), and then the video data is preliminarily parsed by the underlying function module. The support components of the data acquisition module can include any hardware device SDK or WebSocket client, and the underlying function module includes: a serialization function module for generating a serialized Xml / Json file summary based on the video data, a listening function module for determining the activity of each program / service, a file format conversion module, and so on.
[0091] According to the above preliminary parsing of the video data, the I / O module can also be used to process the video data into a file that can be transmitted. For example, the I / O module can include multiple service modules, such as a file listening module that provides a file listening service and a file transfer module for FTP transfer of files, and so on.
[0092] Then, the scene modeling application and / or virtual reality application installed on the terminal 120 transmits the video data in file form to the server 110 for further parsing. Specifically, the server 110 also includes a communication module similarly. The support components of this communication module can also include any hardware device SDK or WebSocket client. Even to improve the transmission speed, a pipeline transmission module can be correspondingly included. The server 110 also includes various databases, such as a model database, a material database, and a texture database. The server 110 can use its analysis module, combine the various databases to execute the above operation S202, and then return the model data of the virtual entity to the scene modeling application and / or virtual reality application.
[0093] Subsequently, the scene modeling application and / or virtual reality application will utilize the rule conversion module to convert the rules in the physical world into the rules in the virtual scene (e.g., perform coordinate conversion), and create the virtual scene corresponding to the physical scene in combination with the rules in the virtual scene. It should be noted that the terminal receiving the model data of the virtual entity is not necessarily the terminal sending the video data file. For example, the terminal A can be used to collect video data and send it to the server, and then the server sends the model data to the terminal B, so as to achieve remote multi-location collaborative operation. Provide corresponding dynamic references for users outside the physical scene to assist the user in performing off-site analysis and virtual scene restoration of the virtual scene.
[0094] In addition, the scene modeling application and / or virtual reality application may also include a rendering process and a control process to implement the visualization process of the virtual scene. For example, the rendering process and the control process can communicate with each other to achieve the visualization of the virtual scene. In addition, the rendering process also provides simulation feedback information to the control process to indicate the comparison information between the above virtual scene and the physical scene. Of course, the present disclosure is not limited thereto.
[0095] The various embodiments of the present disclosure have strong scalability. They can not only combine various gesture recognition algorithms for in-depth vertical development to provide model data and auxiliary data to ordinary users of the terminal 120, but also perform horizontal expansion development to provide scene supervision services to regulators in certain special industries, and achieve real-time scene detection through real scene restoration. In addition, the various embodiments of the present disclosure can also be output as JAR packages / dynamic link libraries that can be used by the corresponding platforms for integration by multiple systems.
[0096] Next, refer to Figure 10 to further describe an example of operation S204, where Figure 10 is a schematic diagram showing the operation of the rendering engine according to an embodiment of the present disclosure.
[0097] As an example, operation S204 includes: selecting a plurality of video frames from the video data; performing texture compression and / or texture scaling processing on the plurality of video frames to generate texture map data; and rendering the virtual scene corresponding to the physical scene based on the texture map data, and displaying the rendered virtual scene.
[0098] For example, the interface glCompressedTexImage2D(…, format, …, data) of OpenGL ES can be used to perform texture compression on the multiple video frames. It should be noted that the present disclosure does not limit the format of the texture data, and it can convert the texture data into any format according to the SDK or documentation of the vendor. For example, assume that the display screen of the terminal 120 is adapted with 32 MB of display memory. A single video frame image of 2 MB can be texture-compressed to generate texture map data in the ECT (Ericsson Texture Compression) format to ensure that there are more than 16 texture maps of texture map data.
[0099] In some cases, the texture map data obtained after texture compression may be distorted in proportion. Therefore, texture scaling can be further used in the 3D rendering engine to adjust the texture map data. For example, texture (Texture) resource data can be generated for the texture map data (for example, Figure 10 the parameters of Material A to Material C shown). Based on the texture resource data, the rendering engine will correspondingly generate material (Material) resource data (for example, Figure 10 the parameters such as color, specular, and metal shown). Combining with the model data of the virtual entity corresponding to the physical entity obtained from the video data, based on the texture resource data and the material resource data, the parameters corresponding to the texture scaling process can be determined (for example, the pixel data in some texture maps can be directly characterized by the texture scaling parameters). Based on the parameters corresponding to the texture scaling process, the texture map data can be further subjected to texture scaling processing to further reduce the file size of the texture map data and ensure the running speed of the virtual reality application.
[0100] Thus, in view of the requirements of application service visualization and scene virtualization, various embodiments of the present disclosure utilize video data to implement the virtualization of the scene, which helps to solve the technical problems of high complexity and long time consumption in the process of generating the scene model.
[0101] In addition, according to another aspect of the present disclosure, there is also provided a device for virtualizing a physical scene. The device includes: a first module configured to determine physical entities within the scene boundary based on interaction information for indicating the scene boundary and capture video data corresponding to the physical entities; a second module configured to determine model data of a virtual entity corresponding to the physical entity based on the video data corresponding to the physical entity; and a third module configured to create a virtual scene corresponding to the physical scene based on the model data corresponding to the virtual entity.
[0102] For example, the video data includes multiple video frames, and different video frames in the multiple video frames correspond to different lighting conditions, shooting positions, or shooting angles.
[0103] For example, the second module is further configured to: extract a plurality of discrete points from each video frame in the video data; generate stereoscopic model data characterized by Thiessen polygons as the stereoscopic model data of the video frame based on the plurality of discrete points of each video frame; and determine the model data of the virtual entity corresponding to the physical entity based on the stereoscopic model data of each video frame.
[0104] For example, the second module is further configured to: obtain one or more of a building information model, global geographical location information, and building positioning space data; and determine the model data of the virtual entity corresponding to the physical entity by using the video data corresponding to the physical entity based on one or more of the building information model, the global geographical location information, and the building positioning space data.
[0105] For example, the second module is further configured to: obtain one or more of urban traffic data, urban planning data, and urban municipal data; and determine the model data of the virtual entity corresponding to the physical entity by using the video data corresponding to the physical entity based on one or more of the urban traffic data, the urban planning data, and the urban municipal data.
[0106] For example, the device further includes a fourth module configured to: display relevant information of the virtual scene based on the virtual scene corresponding to the physical scene.
[0107] For example, the displaying of the relevant information of the virtual scene further includes: selecting a plurality of video frames from the video data; performing texture compression and / or texture scaling processing on the plurality of video frames to generate texture mapping data; and rendering the virtual scene corresponding to the physical scene based on the texture mapping data and displaying the rendered virtual scene.
[0108] For example, the performing of the texture compression and / or texture scaling processing on the plurality of video frames to generate texture mapping data further includes: performing texture compression on the plurality of video frames to generate texture-compressed texture mapping data; determining the material resource data and material resource data corresponding to the texture mapping data based on the texture-compressed texture mapping data; determining the parameters corresponding to the texture scaling processing based on the material resource data and material resource data corresponding to the texture mapping data; and performing texture scaling processing on the texture-compressed texture mapping data based on the parameters corresponding to the texture scaling processing to generate texture-scaled texture mapping data.
[0109] In addition, according to another aspect of the present disclosure, an electronic device is further provided for implementing the method according to the embodiments of the present disclosure. Figure 11 A schematic diagram of an electronic device 2000 according to an embodiment of the present disclosure is shown.
[0110] As shown Figure 11 in the figure, the electronic device 2000 may include one or more processors 2010 and one or more memories 2020. Among them, computer-readable code is stored in the memory 2020, and when the computer-readable code is run by the one or more processors 2010, it can execute the search request processing method described above.
[0111] The processor in the embodiments of the present disclosure may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.
[0112] Generally speaking, the various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0113] For example, the method or device according to the embodiments of the present disclosure may also be implemented by means of Figure 12 the architecture of the computing device 3000 shown in the figure. As shown Figure 7 in the figure, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to the network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 7 the architecture shown in the figure is only exemplary, and when implementing different devices, one or more components shown in the Figure 7 computing device may be omitted according to actual needs.
[0114] According to another aspect of the present disclosure, a computer-readable storage medium is also provided. Figure 13 A schematic diagram of the storage medium 4000 according to the present disclosure is shown.
[0115] As Figure 13 shown, computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the methods according to the embodiments of the present disclosure described with reference to the above drawings can be executed. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories. It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories.
[0116] Embodiments of the present disclosure also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods according to the embodiments of the present disclosure.
[0117] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0118] Generally speaking, various example embodiments of the present disclosure can be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or a controller or other computing devices, or some combination thereof.
[0119] The example embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for virtualizing a physical scene, comprising: Determining the scene boundary based on interaction information for indicating the scene boundary; Determining physical entities within the scene boundary based on the scene boundary, and capturing video data corresponding to the physical entities; Determining model data of virtual entities corresponding to the physical entities based on the video data corresponding to the physical entities; And Creating a virtual scene corresponding to the physical scene based on the model data of the virtual entities corresponding to the physical entities; Wherein, the determining the model data of the virtual entities corresponding to the physical entities based on the video data corresponding to the physical entities further comprises: Extracting a plurality of discrete points from each video frame in the video data; Generating stereoscopic model data represented by Thiessen polygons as the stereoscopic model data of each video frame based on the plurality of discrete points of each video frame; Determining the model data of the virtual entities corresponding to the physical entities based on the stereoscopic model data of each video frame.
2. The method according to claim 1, wherein, The video data includes a plurality of video frames, and different video frames in the plurality of video frames correspond to different lighting conditions, shooting positions or shooting angles.
3. The method according to claim 1, wherein The determining the model data of the virtual entities corresponding to the physical entities based on the video data corresponding to the physical entities further comprises: Obtaining one or more of a building information model, global geographical location information, and building positioning space data; Determining the model data of the virtual entities corresponding to the physical entities by using the video data corresponding to the physical entities based on one or more of the building information model, the global geographical location information, and the building positioning space data.
4. The method according to claim 1, wherein The determining the model data of the virtual entities corresponding to the physical entities based on the video data corresponding to the physical entities further comprises: Obtaining one or more of urban traffic data, urban planning data, and urban municipal data; Determining the model data of the virtual entities corresponding to the physical entities by using the video data corresponding to the physical entities based on one or more of the urban traffic data, the urban planning data, and the urban municipal data.
5. The method according to claim 1, further comprising: Displaying relevant information of the virtual scene based on the virtual scene corresponding to the physical scene.
6. The method according to claim 5, wherein, The displaying the relevant information of the virtual scene further comprises: Selecting a plurality of video frames from the video data; Performing texture compression and / or texture scaling processing on the plurality of video frames to generate texture mapping data; Rendering the virtual scene corresponding to the physical scene based on the texture mapping data, Displaying the rendered virtual scene.
7. The method according to claim 6, wherein, The performing texture compression and / or texture scaling processing on the plurality of video frames to generate texture mapping data further comprises: Performing texture compression on the plurality of video frames to generate texture-compressed texture mapping data; Determining material resource data and material resource data corresponding to the texture mapping data based on the texture-compressed texture mapping data; Determining parameters corresponding to the texture scaling processing based on the material resource data and the material resource data corresponding to the texture mapping data; Based on the parameters corresponding to texture scaling processing, perform texture scaling processing on the texture-compressed map data to generate the map data after texture scaling processing.
8. An electronic device, comprising: Processor; A memory storing computer instructions, which when executed by the processor implement the method according to any one of claims 1-7.
9. A computer-readable storage medium having computer instructions stored thereon, which when executed by a processor implement the method according to any one of claims 1-7.
10. A computer program product comprising computer-readable instructions, which when executed by a processor cause the processor to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
VR scene building method and system
CN108074286A