Virtual and real scene fusion method and device, electronic equipment and storage medium

By constructing a virtual scene in a real scene and removing the invisible parts, combining a depth camera or optical perspective device, the consistency problem of virtual view and real view in motion scenes is solved, and a high-precision virtual entity fusion is achieved.

CN120355569AActive Publication Date: 2025-07-22NANJING RUICHEN XINCHUANG NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510804180.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-22
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing virtual view is directly superimposed on the real view, and the dynamic consistency of the virtual view and the real view in the motion scene cannot be guaranteed, affecting the position accuracy and accuracy of the virtual entity in the fused view.

Method used

Build a virtual scene based on real scenes, obtain the position information of real entities and twins in real time, eliminate the invisible parts and twins of virtual entities, and fuse the real and target virtual scenes through a depth camera or optical perspective device to generate a fused view.

Benefits of technology

It effectively ensures the dynamic consistency between virtual and real scenes, and improves the position accuracy and depth information of virtual entities in the scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355569A_ABST
    Figure CN120355569A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual and real scene fusion method and device, electronic equipment and a storage medium, and the method comprises the steps: building a virtual scene based on a real scene, and enabling the virtual scene to comprise virtual entities and twinborn bodies corresponding to the real entities one by one; the pose information of a real entity is obtained in real time, and the pose information of the twinborn body is synchronized; obtaining pose information of a real viewpoint, and obtaining an initial virtual scene based on the pose information of the real viewpoint and the virtual scene; removing an invisible part and a twinborn body of the virtual entity from the initial virtual scene to obtain a target virtual scene; and based on the target virtual scene, obtaining a fused scene. The shielded part of the virtual entity is removed based on the twinborn body of the real entity, and then the virtual entity is fused with the real scene, so that the motion of the virtual entity is limited by the real scene, the dynamic consistency of the virtual scene and the real scene is effectively ensured, and the real scene is more vivid. And the position precision of the virtual entity in the fused scene and the accuracy of the fused depth information are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of virtual and real vision, and particularly relates to a method, device, electronic device, and storage medium for fusing virtual and real vision. Background Art

[0002] The fusion of virtual and real vision refers to the process of fusing real vision and virtual vision to form a fused vision and visualizing the fused vision. Existing virtual and real vision fusion technologies directly superimpose virtual entities onto real vision, and the movement of virtual entities is not restricted by the real scene. During the fusion process of virtual vision and real vision for a moving scene, it is impossible to ensure the dynamic consistency between virtual vision and real vision, which affects the position accuracy of virtual entities in the final fused vision and the accuracy of fused depth information. Summary of the Invention

[0003] The purpose of this application is to provide a method, device, electronic device, and storage medium for fusing virtual and real vision to solve the technical problem in the prior art that when virtual vision is directly superimposed onto real vision, it is impossible to ensure the dynamic consistency between virtual vision and real vision during the process of virtual vision and real vision for a moving scene.

[0004] To achieve the above purpose, the first aspect of this application provides a method for fusing virtual and real vision, including: Based on the real scene, construct a virtual scene, where the real scene includes real entities, and the virtual scene includes virtual entities and twins corresponding one-to-one to the real entities; Obtain the pose information of the real entity in real time and synchronize the pose information of the twin; Obtain the pose information of the real viewpoint, and based on the pose information of the real viewpoint and the virtual scene, obtain an initial virtual vision; In the initial virtual vision, remove the invisible parts of the virtual entities and the twins to obtain a target virtual vision; Based on the target virtual vision, obtain a fused vision.

[0005] In one or more embodiments, the step of obtaining an initial virtual vision based on the pose information of the real viewpoint and the virtual scene includes: Based on the pose information of the real viewpoint, synchronize virtual viewpoints with the same pose information in the virtual scene; Based on the pose information of the virtual viewpoint and the virtual scene, obtain an initial virtual vision.

[0006] In one or more embodiments, the step of removing the invisible parts of the virtual entities and the twins in the initial virtual vision to obtain a target virtual vision includes: Traverse each pixel point of the initial virtual view, and determine whether the current pixel point being traversed has both the virtual entity and the twin; If so, determine whether the depth value of the virtual entity at the current pixel point is greater than the depth value of the twin at the current pixel point; If so, remove the current pixel point of the virtual entity; After the traversal is completed, remove the twin.

[0007] In one or more embodiments, the real viewpoint is a depth camera, and the step of removing the current pixel point of the virtual entity is specifically: set the texture information and depth information of the current pixel point of the virtual entity to not be written, and the step of removing the twin is specifically: set the texture information and depth information of the twin to not be written; or, The real viewpoint is the human eye, and there is an optical perspective device on the observation path of the real viewpoint. The step of removing the current pixel point of the virtual entity is specifically: set the texture information of the current pixel point of the virtual entity to not be written, and the step of removing the twin is specifically: set the texture information of the twin to not be written.

[0008] In one or more embodiments, the real viewpoint is a depth camera; The steps of obtaining the fused view based on the target virtual view include: Obtain the real view based on the pose information of the real viewpoint and the real scene; Fuse the real view and the target virtual view to obtain the fused view.

[0009] In one or more embodiments, the steps of fusing the real view and the target virtual view to obtain the fused view include: Fuse the depth matrix of the real view and the depth matrix of the target virtual view according to the following formula to obtain the fused depth matrix, where the depth matrix is used to describe the depth value of each pixel point, , In the formula, D f represents the fused depth matrix, D r represents the depth matrix of the real view, D v represents the depth matrix of the target virtual view, and i, j represent the coordinates of the pixel point; Fuse the pixel matrix of the real view and the pixel matrix of the target virtual view according to the following formula to obtain the fused pixel matrix, where the pixel matrix is used to describe the color value of each pixel point, , Wherein, P f represents the fused pixel matrix, P r represents the pixel matrix of the real visual scene, P v represents the pixel matrix of the target virtual visual scene, and i, j represent the coordinates of the pixel points; Based on the fused depth matrix and the fused pixel matrix, a fused visual scene is obtained.

[0010] In one or more embodiments, the real view point is the human eye, and an optical perspective device is provided on the observation path of the real view point; The steps of obtaining a fused visual scene based on the target virtual visual scene include: Projecting the target virtual visual scene onto the optical perspective device, so that the target virtual visual scene and the real visual scene are fused and imaged in the human eye to obtain a fused visual scene.

[0011] To achieve the above object, a second aspect of the present application provides a fusion device for virtual and real visual scenes, including: A virtual scene construction module, configured to construct a virtual scene based on a real scene, the real scene includes real entities, and the virtual scene includes virtual entities and twins corresponding to the real entities one by one; A synchronization module, configured to acquire the pose information of the real entity in real time and synchronize the pose information of the twin; A virtual visual scene acquisition module, configured to acquire the pose information of the real view point, and based on the pose information of the real view point and the virtual scene, obtain an initial virtual visual scene; A virtual visual scene processing module, configured to remove the invisible parts of the virtual entities and the twins in the initial virtual visual scene to obtain a target virtual visual scene; A fusion module, configured to obtain a fused visual scene based on the target virtual visual scene.

[0012] To achieve the above object, a third aspect of the present application provides an electronic device, including: At least one processor; and A memory, the memory stores instructions, when the instructions are executed by the at least one processor, the at least one processor is caused to execute the fusion method of virtual and real visual scenes as described in any of the above embodiments.

[0013] To achieve the above object, a fourth aspect of the present application provides a machine-readable storage medium, which stores executable instructions, and when the instructions are executed, the machine is caused to execute the fusion method of virtual and real visual scenes as described in any of the above embodiments.

[0014] Different from the prior art, the beneficial effects of the present application are: This application constructs and synchronizes a virtual scene based on a real scene, obtains a virtual view from the virtual scene, removes the occluded parts of the virtual entities based on the twins of real entities, and then fuses it with the real view. The movement of the virtual entities is restricted by the real scene, effectively ensuring the dynamic consistency between the virtual view and the real view, and ensuring the position accuracy of the virtual entities in the fused view and the accuracy of the fused depth information. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 It is a flowchart of an implementation manner of the method for fusing virtual and real views of the present application; Figure 2 It is a schematic diagram of an implementation manner of the virtual scene of the present application; Figure 3 It is Figure 1 A flowchart of an implementation manner corresponding to S300 in Figure 4 It is a schematic diagram of an implementation manner of the initial virtual view of the present application; Figure 5 It is Figure 1 A flowchart of an implementation manner corresponding to S400 in Figure 6 It is a schematic diagram of an implementation manner of the target virtual view of the present application; Figure 7 It is a schematic diagram of the effect of removal in the present application; Figure 8 It is a schematic diagram of an implementation manner of the fused view of the present application; Figure 9 It is a schematic diagram of an application scenario of the real view, virtual view, and fused view of the present application; Figure 10 It is a schematic diagram of the structure of an implementation manner of the device for fusing virtual and real views of the present application; Figure 11 It is a schematic diagram of the structure of an implementation manner of an electronic device of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0018] Currently, the method for fusing virtual visual scenes adopts the method of directly superimposing virtual entities onto real visual scenes. Among them, the virtual entities themselves are not restricted by the real scenes, resulting in the inability to guarantee the motion consistency between virtual entities and real entities. Especially in the fusion of real and virtual visual scenes during the motion process, there are problems such as position misalignment and depth misalignment between real entities and virtual entities, resulting in a poor fusion effect.

[0019] To solve the above problems, the applicant has developed a new method for fusing real and virtual visual scenes. This method processes the positional relationship and occlusion relationship between virtual entities and real entities in the virtual scene, and realizes the real-virtual fusion based on the processed virtual visual scene, effectively solving the problem of improving the motion consistency between virtual entities and real entities, ensuring the fusion accuracy, and improving the fusion effect.

[0020] Specifically, please refer to Figure 1 , Figure 1 which is a schematic flowchart of an implementation manner of the method for fusing real and virtual visual scenes in this application.

[0021] As Figure 1 shown, this method includes: S100. Construct a virtual scene based on the real scene.

[0022] Among them, the real scene includes real entities, and the virtual scene includes virtual entities and twins corresponding to the real entities one by one.

[0023] The real scene and the virtual scene are based on a unified coordinate system, and all elements of the real scene can be accurately modeled through digital twin technology to ensure their consistency.

[0024] Specifically, any commonly used modeling engine in the art can be used to construct a virtual scene based on the collected real scene information, such as the Unreal Engine, etc., which will not be elaborated here. The user can design virtual entities based on the fusion requirements and add them to the virtual scene, thereby obtaining a virtual scene that includes both twins and virtual entities.

[0025] Exemplarily, please refer to Figure 2 , Figure 2 which is a schematic diagram of an implementation manner of the virtual scene in this application. Figure 2The virtual scene includes twins corresponding one by one to real entities, and also includes virtual entities constructed by users.

[0026] S200. Obtain the pose information of the real entity in real time and synchronize the pose information of the twins.

[0027] After the virtual scene is constructed, the twins in the virtual scene need to be synchronized with the entities in the real scene at all times. The synchronized pose information can include position, orientation, attitude, movement, etc.

[0028] The synchronization of the virtual scene ensures the dynamic consistency between the virtual scene and the real scene, and further ensures the dynamic consistency between the virtual entity and the real entity in subsequent fusion processing.

[0029] S300. Obtain the pose information of the real viewpoint, and based on the pose information of the real viewpoint and the virtual scene, obtain the initial virtual view.

[0030] The real viewpoint refers to the view image acquisition point in the real scene. By obtaining the pose information such as the position, attitude, viewing angle, and focal length of the real viewpoint, the image information that can be collected within its field of view, that is, the view, can be obtained.

[0031] In one implementation, the real viewpoint can be a depth camera. Based on the sensors arranged on the depth camera and the preset parameters of the depth camera, the pose information of the depth camera can be obtained.

[0032] In another implementation, the real viewpoint can also be the human eye. Based on the sensors arranged on the human head, the pose information of the human eye can be obtained.

[0033] Based on the pose information of the real viewpoint, the image information that can be collected within the field of view can be obtained from the virtual scene, that is, the initial virtual view.

[0034] Specifically, please refer to Figure 3 , Figure 3 is Figure 1 a schematic flowchart of an implementation corresponding to S300 in

[0035] As Figure 3 shown, the method for obtaining the initial virtual view includes: S301. Based on the pose information of the real viewpoint, synchronize the virtual viewpoints with the same pose information in the virtual scene.

[0036] S302. Based on the pose information of the virtual viewpoint and the virtual scene, obtain the initial virtual view.

[0037] It can be understood that by constructing virtual viewpoints with the same pose information in the virtual scene, the field of view range of the virtual viewpoint can be obtained, and thus the initial virtual view can be obtained.

[0038] Exemplarily, refer to Figure 4 , Figure 4 , which is a schematic diagram of an embodiment of the initial virtual view of the present application. As Figure 4 shown, a virtual viewpoint is constructed in the virtual scene, and the initial virtual view is obtained by collecting the field of view based on the virtual viewpoint.

[0039] S400. In the initial virtual view, the invisible parts of the virtual entities and the twins are removed to obtain the target virtual view.

[0040] In this embodiment, the positional and depth relationships between the virtual entities and the twins of the real entities are processed in the initial virtual view, thereby ensuring the motion consistency between the virtual entities and the real entities.

[0041] Specifically, refer to Figure 5 , Figure 5 , which is Figure 1 a schematic flowchart of an embodiment corresponding to S400 in

[0042] As Figure 5 shown, the method for generating the target virtual view includes: S401. Traverse each pixel point of the initial virtual view, and determine whether both a virtual entity and a twin exist at the currently traversed pixel point.

[0043] If so, then: S402. Determine whether the depth value of the virtual entity at the current pixel point is greater than the depth value of the twin at the current pixel point.

[0044] If so, then: S403. Remove the current pixel point of the virtual entity.

[0045] S404. After the traversal is completed, remove the twin.

[0046] Based on the above steps, the occluded virtual entities can be removed from the virtual view, and then the twins are removed to obtain the target virtual view for fusing with the real view.

[0047] Exemplarily, refer to Figure 6 , Figure 6 , which is a schematic diagram of an embodiment of the target virtual view of the present application. As Figure 6 shown, the occluded virtual entities and the twins in the initial virtual view are removed to obtain the target virtual view.

[0048] Specifically, the above elimination method can be achieved by adjusting the parameter drawing settings of the simulation engine for objects. Specifically, the simulation engine of the virtual scene supports the drawing of object textures and depths. If it is set that a certain type of object does not write texture information during rendering, in the final image, this type of object will appear as black without texture on the final image. If it is set that the object does not write texture information and depth information during rendering, the object will not appear on the final image.

[0049] Please refer to Figure 7 , Figure 7 which is a schematic diagram of the elimination effect in this application. Among them, a is the original image. In b, the texture information of the grass is not written, and the grass appears black. In c, neither the texture information nor the depth information of the grass is written, and the grass does not appear.

[0050] Based on the above principle, the elimination of objects in the virtual scene can be achieved through rendering control. Specifically, in one embodiment, when the real viewpoint is a depth camera, at this time, the finally fused visual scene image needs to be transmitted to the display screen. Therefore, the texture information and depth information of the objects to be eliminated can be set not to be written. At this time, the objects to be eliminated in the target virtual visual scene do not appear. When finally fused, the eliminated part will be replaced by the corresponding part of the real visual scene, achieving the effect of virtual-real fusion.

[0051] In another embodiment, when the real viewpoint is a human eye, at this time, the virtual visual scene needs to be projected onto an optical lens device such as a semi-transparent lens in front of the human eye. Therefore, the texture information of the objects to be eliminated can be set not to be written. At this time, the objects to be eliminated in the target virtual visual scene appear as black. When projected onto the optical lens device, the corresponding position of this part does not emit light, and the finally observed black part by the human eye presents the semi-transparent state of the semi-transparent lens itself. Furthermore, this part can be replaced by the real visual scene, achieving the effect of virtual-real fusion.

[0052] S500. Obtain a fused visual scene based on the target virtual visual scene.

[0053] Since the target virtual visual scene includes visible virtual entities, a fused visual scene can be obtained based on the target virtual visual scene.

[0054] In one embodiment, when the real viewpoint is a depth camera, the fused visual scene needs to be output to a display device to achieve fused visualization.

[0055] At this time, the real visual scene can be obtained based on the pose information of the real viewpoint and the real scene first; then the real visual scene and the target virtual visual scene are fused to obtain a fused visual scene.

[0056] Among them, the method of fusing the real visual scene and the target virtual visual scene can include: The depth matrix of the real view and the depth matrix of the target virtual view are fused according to the following formula to obtain a fused depth matrix, where the depth matrix is used to describe the depth value of each pixel point. , where D f represents the fused depth matrix, D r represents the depth matrix of the real view, and D v represents the depth matrix of the target virtual view. i and j represent the coordinates of the pixel point. The pixel matrix of the real view and the pixel matrix of the target virtual view are fused according to the following formula to obtain a fused pixel matrix, where the pixel matrix is used to describe the color value of each pixel point. , where P f represents the fused pixel matrix, P r represents the pixel matrix of the real view, and P v represents the pixel matrix of the target virtual view. i and j represent the coordinates of the pixel point. Based on the fused depth matrix and the fused pixel matrix, a fused view is obtained.

[0057] Exemplarily, please refer to Figure 8 , Figure 8 which is a schematic diagram of an embodiment of the fused view of the present application. As Figure 8 shown, the real view and the target virtual view are fused to obtain a fused view, and then the fused view is visually presented.

[0058] In another embodiment, when the real viewpoint is the human eye, an optical see-through device is provided on the observation path of the real viewpoint. The user observes the real scene through this optical see-through device. Therefore, the target virtual view can be directly projected onto the optical see-through device, so that the real view and the target virtual view can be fused and imaged in the human eye to obtain a fused view when the user observes.

[0059] Among them, the optical see-through device can be any device such as a semi-transparent lens that can display images and through which the human eye can observe the external environment.

[0060] Based on the fusion methods of the above embodiments, a virtual scene is constructed and synchronized based on the real scene, a virtual view is obtained from the virtual scene, and the occluded part of the virtual entity is removed based on the twin of the real entity, and then it is fused with the real view. The movement of the virtual entity is restricted by the real scene, effectively ensuring the dynamic consistency between the virtual view and the real view, and ensuring the position accuracy of the virtual entity and the accuracy of the fused depth information in the fused view.

[0061] The fusion method of virtual and real visual scenes in the above embodiments can be applied to simulation training scenarios. For example, the real training ground can be surveyed by drone oblique photography, and the site modeling can be completed based on the survey data to construct the virtual scene. During the simulation training process, the participating entities collect data such as position, attitude, and movement in real time through devices such as wearable positioning sensors, attitude sensors, and motion capture sensors, which are used to synchronize the twins in the virtual scene in real time.

[0062] The participating entities obtain real visual scene data by wearing cameras. The virtual visual scene acquisition module synchronizes the virtual camera according to the position, attitude, field of view angle, and focal length in the real visual scene data, and obtains virtual visual scene data.

[0063] The virtual and real visual scene fusion module fuses the images in the real visual scene and the virtual visual scene to obtain fused visual scene data, and outputs the fused image to the head-mounted display terminal to realize the visualization of the fused visual scene.

[0064] Exemplarily, please refer to Figure 9 , Figure 9 which is a schematic diagram of an application scenario of the real visual scene, virtual visual scene, and fused visual scene of this application. As Figure 9 shown, by fusing the real visual scene and the virtual visual scene, a fused visual scene for simulation training can be obtained.

[0065] This application also provides a virtual and real visual scene fusion device. Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of an embodiment of the virtual and real visual scene fusion device of this application. As Figure 10 shown, the device includes a virtual scene construction module 21, a synchronization module 22, a virtual visual scene acquisition module 23, a virtual visual scene processing module 24, and a fusion module 25.

[0066] Among them, the virtual scene construction module 21 is used to construct a virtual scene based on the real scene. The real scene includes real entities, and the virtual scene includes virtual entities and twins corresponding to the real entities one by one; The synchronization module 22 is used to obtain the pose information of the real entity in real time and synchronize the pose information of the twin; The virtual visual scene acquisition module 23 is used to obtain the pose information of the real viewpoint, and based on the pose information of the real viewpoint and the virtual scene, obtain the initial virtual visual scene; The virtual visual scene processing module 24 is used to remove the invisible parts of the virtual entities and the twins in the initial virtual visual scene to obtain the target virtual visual scene; The fusion module 25 is used to obtain the fused visual scene based on the target virtual visual scene.

[0067] As described above with reference to Figures 1 to 9, a method for fusing virtual and real visual scenes according to the embodiments of the present specification is described. The details mentioned in the above description of the method embodiments are equally applicable to the virtual and real visual scene fusion device of the embodiments of the present specification. The above virtual and real visual scene fusion device can be implemented by hardware, or can be implemented by software or a combination of hardware and software.

[0068] The present application also provides an electronic device. Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of an embodiment of the electronic device of the present application. As Figure 11 shown, the electronic device 30 may include at least one processor 31, a memory 32 (such as a non-volatile memory), a memory 33, and a communication interface 34, and at least one processor 31, the memory 32, the memory 33, and the communication interface 34 are connected together via a bus 35. At least one processor 31 executes at least one computer-readable instruction stored or encoded in the memory 32.

[0069] It should be understood that the computer-executable instructions stored in the memory 32, when executed, cause at least one processor 31 to perform the various operations and functions described above in the respective embodiments of the present specification in combination with Figures 1 - 9 the description.

[0070] In the embodiments of the present specification, the electronic device 30 may include, but is not limited to: a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile electronic device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable electronic device, a consumer electronic device, and the like.

[0071] According to one embodiment, a program product such as a machine-readable medium is provided. The machine-readable medium may have instructions (i.e., the above elements implemented in software form), which when executed by the machine, cause the machine to perform the various operations and functions described above in the respective embodiments of the present specification in combination with Figures 1 - 9 the description. Specifically, a system or device equipped with a readable storage medium may be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device is caused to read and execute the instructions stored in the readable storage medium.

[0072] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present specification.

[0073] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, program code can be downloaded from a server computer or the cloud via a communication network.

[0074] Those skilled in the art should understand that various modifications and variations can be made to the above-disclosed embodiments without departing from the essence of the invention. Therefore, the scope of protection of this specification should be defined by the appended claims.

[0075] It should be noted that not all steps and units in the above-mentioned processes and system structure diagrams are necessary, and some steps or units can be ignored according to actual needs. The execution order of each step is not fixed and can be determined according to requirements. The device structures described in the above embodiments can be physical structures or logical structures. That is, some units may be implemented by the same physical entity, or some units may be implemented separately by multiple physical entities, or some components in multiple independent devices may be jointly implemented.

[0076] In the above embodiments, the hardware units or modules can be implemented mechanically or electrically. For example, a hardware unit, module, or processor can include permanent dedicated circuits or logic (such as a dedicated processor, FPGA, or ASIC) to perform corresponding operations. The hardware unit or processor can also include programmable logic or circuits (such as a general-purpose processor or other programmable processors), which can be temporarily set by software to perform corresponding operations. The specific implementation method (mechanical method, or dedicated permanent circuit, or temporarily set circuit) can be determined based on cost and time considerations.

[0077] The specific embodiments described above in conjunction with the accompanying drawings describe exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the scope of protection of the claims. The term "exemplary" used throughout this specification means "serving as an example, instance, or illustration", and does not mean "preferred" or "advantageous" compared to other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described embodiments.

[0078] The foregoing description of the disclosure is provided to enable any person of ordinary skill in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those of ordinary skill in the art, and the generic principles herein may be applied to other variations without departing from the scope of the disclosure. Thus, the disclosure is not limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for fusing virtual and real visual scenes, characterized in that Including: Based on a real scenario, a virtual scenario is constructed. The real scenario includes real entities, and the virtual scenario includes virtual entities and twins corresponding one-to-one to the real entities; The pose information of the real entities is obtained in real time, and the pose information of the twins is synchronized; The pose information of the real viewpoint is obtained, and an initial virtual view is obtained based on the pose information of the real viewpoint and the virtual scenario; In the initial virtual view, the invisible parts of the virtual entities and the twins are removed to obtain a target virtual view; Based on the target virtual view, a fused view is obtained.

2. The fusion method according to claim 1, characterized in that, The step of obtaining an initial virtual view based on the pose information of the real viewpoint and the virtual scenario includes: Based on the pose information of the real viewpoint, virtual viewpoints with the same pose information are synchronized in the virtual scenario; Based on the pose information of the virtual viewpoints and the virtual scenario, an initial virtual view is obtained.

3. The fusion method according to claim 1, wherein The step of removing the invisible parts of the virtual entities and the twins in the initial virtual view to obtain a target virtual view includes: Traverse each pixel point of the initial virtual view, and determine whether the current pixel point being traversed has both the virtual entity and the twin; If so, determine whether the depth value of the virtual entity at the current pixel point is greater than the depth value of the twin at the current pixel point; If so, remove the current pixel point of the virtual entity; After the traversal is completed, the twins are removed.

4. The fusion method according to claim 3, characterized in that The real viewpoint is a depth camera. The step of removing the current pixel point of the virtual entity is specifically: setting the texture information and depth information of the current pixel point of the virtual entity not to be written. The step of removing the twins is specifically: setting the texture information and depth information of the twins not to be written; or, The real viewpoint is a human eye, and an optical see-through device is provided on the observation path of the real viewpoint. The step of removing the current pixel point of the virtual entity is specifically: setting the texture information of the current pixel point of the virtual entity not to be written. The step of removing the twins is specifically: setting the texture information of the twins not to be written.

5. The fusion method according to claim 1, wherein The real viewpoint is a depth camera; The step of obtaining a fused view based on the target virtual view includes: Based on the pose information of the real viewpoint and the real scenario, a real view is obtained; The real view and the target virtual view are fused to obtain a fused view.

6. The fusion method according to claim 5, characterized in that The step of fusing the real view and the target virtual view to obtain a fused view includes: The depth matrices of the real view and the target virtual view are fused according to the following formula to obtain a fused depth matrix. The depth matrix is used to describe the depth value of each pixel point, , Where, D f represents the fused depth matrix, D r represents the depth matrix of the real visual scene, D v represents the depth matrix of the target virtual visual scene, and i, j represent the coordinates of the pixel points; The pixel matrices of the real view and the target virtual view are fused according to the following formula to obtain a fused pixel matrix. The pixel matrix is used to describe the color value of each pixel point, , Wherein, P f represents the fused pixel matrix, P r represents the pixel matrix of the real visual scene, P v represents the pixel matrix of the target virtual visual scene, and i, j represent the coordinates of the pixel points; Based on the fused depth matrix and the fused pixel matrix, a fused view is obtained.

7. The fusion method according to claim 1, wherein The real viewpoint is a human eye, and an optical see-through device is provided on the observation path of the real viewpoint; The steps of obtaining a fused view based on the target virtual view include: Projecting the target virtual view onto the optical see-through device so that the target virtual view and the real view are fused and imaged in the human eye to obtain a fused view.

8. A fusion device for virtual and real visual scenes, characterized in that, Including: A virtual scene construction module for constructing a virtual scene based on a real scene, the real scene including real entities, and the virtual scene including virtual entities and twins corresponding one-to-one to the real entities; A synchronization module for acquiring the pose information of the real entities in real time and synchronizing the pose information of the twins; A virtual view acquisition module for acquiring the pose information of a real viewpoint and obtaining an initial virtual view based on the pose information of the real viewpoint and the virtual scene; A virtual view processing module for removing the invisible parts of the virtual entities and the twins from the initial virtual view to obtain a target virtual view; A fusion module for obtaining a fused view based on the target virtual view.

9. An electronic device, characterized in that, Including: At least one processor; And A memory storing instructions that, when executed by the at least one processor, cause the at least one processor to execute the method for fusing virtual and real views according to any one of claims 1 to 7.

10. A machine-readable storage medium, characterized in that, Stored with executable instructions that, when executed, cause the machine to execute the method for fusing virtual and real views according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Virtual and real object synthesis method and device

    CN108182730A

  • Virtual-real fusion implementation method and system and electronic equipment

    CN112346572A

  • Augmented reality equipment and virtual and real object shielding display method

    CN113066189A

  • Virtual-real occlusion relation processing method and device and electronic equipment

    CN118505949A

  • Realistic occlusion for a head mounted augmented reality display

    WO2013155217A1