Virtual-real image fusion method, virtual-real image fusion system and non-transient computer-readable medium
By using a camera in a head-mounted display to obtain three-dimensional spatial pictures and calculate spatial correction parameters, the problem of users in the prior art requiring themselves to obtain correction parameters is solved, and automated virtual and real image fusion is realized, improving efficiency and accuracy.
Patent Information
- Application Number
- CN202110011498.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-14
- Filing Date
- 2021-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-01-06
AI Technical Summary
In the fusion of virtual and real images, existing head-mounted displays require users to operate and obtain spatial correction parameters by themselves. The process is time-consuming and affected by the user's operation method.
The first camera obtains the picture in the three-dimensional space, including the screen picture and the label picture of the entity label. The processor calculates the spatial correction parameters based on the corresponding point data, and automatically adjusts the screen picture to realize the overlapping effect of virtual and real images.
An automated spatial correction process is realized, reducing the time and human error of user operations, and improving the efficiency and accuracy of virtual and real image fusion.
Smart Images

Figure CN114359516B_ABST
Abstract
Description
Technical Field
[0001] The present case relates to a virtual-real image fusion method, a virtual-real image fusion system and a non-transitory computer-readable medium, and in particular to a virtual-real image fusion method, a virtual-real image fusion system and a non-transitory computer-readable medium applied to an optical see-through head-mounted display. Background Art
[0002] When the head-mounted display leaves the factory, due to the different positions / orientations of the screen relative to the camera (external parameters), the size / angle / position of the display screen projected into the three-dimensional space will also vary with the pupil and focal length (internal parameters) of the wearer. If you want to produce the effect of superimposing the virtual image with the actual scene, virtual space correction is required. The current common practice is that after the head-mounted display leaves the factory, the user operates the mixed reality virtual-reality superposition application software to obtain personalized space correction parameters. However, the above correction method will be affected by the user's operation method and is time-consuming. Summary of the invention
[0003] One aspect of the present invention is to provide a virtual-real image fusion method. The method comprises the following steps: obtaining a picture in a three-dimensional space by a first camera, wherein the picture comprises a screen picture and a label picture of a physical label, wherein the screen picture is projected onto the physical label; obtaining corresponding point data of the physical label on the screen picture by a processor according to the picture; obtaining a spatial correction parameter by the processor according to the corresponding point data; and displaying an image on the screen picture by the processor according to the spatial correction parameter.
[0004] In some embodiments, the method further includes: performing an anti-distortion conversion process on the image.
[0005] In some embodiments, the method further includes: obtaining a position of the screen frame in the frame to perform perspective conversion processing on the frame to generate a screen space frame.
[0006] In some embodiments, the method further includes: obtaining a position of the tag frame of the physical tag in the screen space frame to obtain the corresponding point data.
[0007] In some embodiments, the method further includes: displaying a background color or a border on the screen image so that the processor can identify the position of the screen image.
[0008] In some embodiments, the method further includes: obtaining the spatial correction parameter by performing a singular value decomposition (SVD) operation according to the corresponding point data.
[0009] In some embodiments, it also includes: obtaining a position of the physical tag relative to a second camera in the three-dimensional space, wherein the screen image is displayed on a display, and the second camera and the display are located in a head-mounted display device; and obtaining the corresponding point data based on the position of the physical tag relative to the second camera in the three-dimensional space and the image captured by the first camera.
[0010] In some embodiments, the method further includes: obtaining an image of the physical tag from the second camera; estimating a parameter of the second camera; and obtaining the position of the physical tag relative to the second camera in the three-dimensional space based on the image of the physical tag and the parameter of the second camera.
[0011] Another aspect of the present invention is to provide a virtual-real image fusion system, characterized by comprising a display, a first camera, and a processor. The display is used to display a screen image. The first camera is used to obtain an image in a three-dimensional space. The image includes a screen image and a label image of a physical tag, wherein the screen image is projected onto the physical tag. The processor is used to obtain corresponding point data of the physical tag on the screen image according to the image, obtain a spatial correction parameter according to the corresponding point data, and display an image on the screen image according to the spatial correction parameter.
[0012] In some embodiments, the processor is further configured to perform a dewarping conversion process on the frame.
[0013] In some embodiments, the processor is further configured to obtain a position of the screen frame in the frame to perform perspective conversion processing on the frame to generate a screen space frame.
[0014] In some embodiments, the processor is further configured to obtain a position of the tag frame of the physical tag in the screen space frame to obtain the corresponding point data.
[0015] In some embodiments, the processor is further configured to display a background color or a border on the screen image so that the processor can identify the position of the screen image.
[0016] In some embodiments, the processor is further configured to obtain the spatial calibration parameter by performing a singular value decomposition (SVD) operation according to the corresponding point data.
[0017] In some embodiments, the processor is further used to obtain a position of the physical tag relative to a second camera in the three-dimensional space, wherein the second camera corresponds to the screen image, and obtain the corresponding point data based on the position of the physical tag relative to the second camera in the three-dimensional space and the image captured by the first camera.
[0018] In some embodiments, it also includes: a second camera for obtaining an image of the physical tag; wherein the processor is also used to estimate a parameter of the second camera, and obtain the position of the physical tag relative to the second camera in the three-dimensional space based on the image of the physical tag and the parameter of the second camera.
[0019] In some embodiments, the display and the second camera are located in a head mounted display device and are communicatively connected to the processor and the first camera.
[0020] Another aspect of the present invention is to provide a non-transitory computer-readable medium, characterized in that it includes at least one program instruction for executing a virtual-real image fusion method. The virtual-real image fusion method includes the following steps: a processor obtains corresponding point data of a physical tag on a screen according to a picture, wherein the picture includes a screen and a label picture of the physical tag, wherein the screen is projected on the physical tag. A spatial correction parameter is obtained according to the corresponding point data. An image is displayed on the screen according to the spatial correction parameter.
[0021] In some embodiments, the method further includes: performing an anti-distortion conversion process on the image.
[0022] In some embodiments, the method further includes: obtaining a position of the screen image in the image to perform perspective transformation processing on the image to generate a screen space image; and obtaining a position of the tag image of the physical tag in the screen space image to obtain the corresponding point data. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to make the above and other objects, features, advantages and embodiments of the present disclosure more clearly understood, the accompanying drawings are described as follows:
[0024] Figure 1 It is a schematic diagram of the traditional calibration of a head mounted display device;
[0025] Figure 2 is a schematic diagram of a virtual-real image fusion system according to some embodiments of the present invention;
[0026] Figure 3 is a schematic diagram of a virtual-real image fusion method according to some embodiments of the present invention;
[0027] Figure 4 is a schematic diagram of a screen depicted according to some embodiments of the present invention;
[0028] Figure 5 is a schematic diagram of a screen space picture according to some embodiments of the present invention; and
[0029] Figure 6is a schematic diagram of a screen image according to some embodiments of the present invention.
[0030]
Explanation of symbols
[0031] 1A, 1B, 1C: Screen
[0032] 100: Head-mounted display device
[0033] 110: Camera
[0034] 130: Display
[0035] 135: Screen
[0036] 150: Entity tag
[0037] 200: Virtual and real image fusion system
[0038] 210: Head-mounted display device
[0039] 212: Display
[0040] 214: Second Camera
[0041] 230: First Camera
[0042] 250: Processor
[0043] 300: Virtual and Real Image Fusion Method
[0044] S310 to S370: Steps
[0045] 400:Screen
[0046] 410: Screen
[0047] 430: Entity Tag
[0048] 500: Screen space image
[0049] 610: Cuboid DETAILED DESCRIPTION
[0050] The following disclosure provides many different embodiments or examples for implementing different features of the present invention. The components and configurations in the specific examples are used to simplify the present case in the following discussion. Any examples discussed are only used for illustrative purposes and do not limit the scope and significance of the present invention or its examples in any way.
[0051] See also Figure 1 . Figure 1It is a schematic diagram of a conventional calibration of a head-mounted display device 100. As shown in screen 1A, the head-mounted display device 100 includes a camera 110 and a display 130. The display 130 can be a single-screen display or a dual-screen display. Screen 1B is the content drawn on the screen image 135 of the display 130. Screen 1C is the virtual-real superimposed image seen by the user in screen 1A through the display 130. As shown in screen 1C, after the screen image 135 in the image seen by the user is projected onto the physical tag 150, the screen image 135 and the physical tag 150 seen by the user overlap with each other. In the conventional calibration method, the user adjusts the position and angle by himself, aims at and aligns the physical target, and then takes pictures to collect data for calibration. In a preferred embodiment, the head-mounted display device 100 refers to an optical see-through head-mounted display (Optical See-through HMD).
[0052] See also Figure 2 . Figure 2 FIG. 2 is a schematic diagram of a virtual-real image fusion system 200 according to some embodiments of the present invention. Figure 2 As shown, the virtual-real image fusion system 200 includes a head-mounted display device 210, a first camera 230, and a processor 250. The head-mounted display device 210 includes a display 212 and a second camera 214, and the display 212 is used to display the screen image 135. In terms of connection relationship, the head-mounted display device 210 is connected to the first camera 230 for communication, the first camera 230 is connected to the processor 250 for communication, and the processor 250 is connected to the head-mounted display device 210 for communication. Figure 2 The virtual-real image fusion system 200 is only used for illustration and is not limited to this embodiment. Figure 3 Provide explanation.
[0053] See also Figure 3 . Figure 3 FIG. 3 is a schematic diagram of a virtual-real image fusion method 300 according to some embodiments of the present invention. The implementation of the present invention is not limited thereto.
[0054] It should be noted that this control method can be applied to Figure 2 The virtual-real image fusion system 200 in FIG. 1 is a system having the same or similar structure as the virtual-real image fusion system 200 in FIG. 1 . Figure 2 The operation method is described by taking the example of Figure 2 The application is limited.
[0055] It should be noted that, in some embodiments, the virtual-real image fusion method 300 can also be implemented as a computer program and stored in a non-transitory computer-readable medium, so that a computer, an electronic device, or the like can be used. Figure 2 The processor 250 in the virtual-real image fusion system 200 reads the recording medium and executes the operation method. The processor may be composed of one or more chips. The non-transitory computer-readable recording medium may be a read-only memory, a flash memory, a floppy disk, a hard disk, an optical disk, a portable disk, a magnetic tape, a database accessible by a network, or a non-transitory computer-readable recording medium having the same function that can be easily conceived by a person familiar with the art.
[0056] In addition, it should be understood that the operations of the blind scanning method mentioned in this embodiment, except for those whose order is specifically stated, can be adjusted in order according to actual needs, and can even be executed simultaneously or partially simultaneously.
[0057] Furthermore, in different embodiments, these operations may be adaptively increased, replaced, and / or omitted.
[0058] See also Figure 3 The virtual-real image fusion method 300 comprises the following steps.
[0059] In step S310: the camera obtains a picture in the three-dimensional space. The picture includes a screen picture and a label picture of the physical label. The screen picture is projected onto the physical label. For the method of projecting the screen picture onto the physical label, please refer to Figure 1 As shown, various known projection methods are also applicable. Please also refer to Figure 2 In some embodiments, step S310 may be performed as follows Figure 2 The first camera 230 in FIG. Figure 2 The first camera 230 is used to simulate the user's eyes. Figure 4 , Figure 4 is a schematic diagram of a screen 400 according to some embodiments of the present invention.
[0060] like Figure 4 drawn, Figure 4 The first camera 230 is a picture 400 in a three-dimensional space captured by the display 212 . The picture 400 includes a screen picture 410 and a label picture of a physical label 430 . The screen picture 410 is a display picture of the display 212 .
[0061] In step S330, the processor obtains corresponding point data of the physical tag on the screen according to the picture. Figure 2 In some embodiments, step S330 may be performed as follows: Figure 2 Executed by processor 250 in.
[0062] Please also read Figure 4 .when Figure 2 After the first camera 230 in the image 400 is acquired, the first camera 230 transmits the image 400 to the processor 250. In step S330, after the processor 250 receives the image 400, the processor 250 performs an anti-distortion conversion process. Due to the difference in the field of view and focal length of the first camera 230, the image 400 captured by the first camera 230 will be deformed, and the anti-distortion conversion process is used to eliminate the influence of the deformation on the image 400.
[0063] In addition, in step S330, the processor 250 also performs perspective conversion processing on the image 400 to generate the following Figure 5 The screen space image 500 is shown. In some embodiments, the processor 250 obtains the position of the screen image 410 in the image 400 to perform perspective conversion processing. For example, when the first camera 230 obtains the image 400, the processor 250 can control the screen image 410 to display a single background color or a frame, so that the processor 250 can identify the position of the screen image 410 in the image 400. In some embodiments, the position of the screen image 410 obtained by the processor 250 is the coordinates of the four corners of the screen image 410, but the implementation method of the present case is not limited to this.
[0064] Please also read Figure 5 . Figure 5 is a schematic diagram of a screen space picture 500 according to some embodiments of the present invention. Figure 5 The range of the dotted line frame is the screen space frame 500 . After the processor 250 obtains the position of the screen frame 410 , the processor 250 generates the screen space frame 500 according to the screen frame 410 .
[0065] In some embodiments, when performing perspective conversion, since the screen image 410 is a rectangle, the screen image 410 obtained in the image 400 becomes a parallelogram due to the distance between the first camera 230 and the display 212. After the perspective conversion, the processor 250 converts the screen image 410 back to a rectangular screen space image 500.
[0066] Next, the processor 250 obtains the position of the label screen of the physical tag 430 in the screen space frame 500. In some embodiments, the position of the label screen of the physical tag 430 is the coordinates of the label screen of the physical tag 430 in the screen space frame 500. For example, the processor 250 obtains the coordinates of the four corners of the label screen of the physical tag 430 in the screen space frame 500 to obtain the position of the label screen of the physical tag 430.
[0067] Next, the processor 250 obtains corresponding point data according to the position of the tag image of the physical tag 430 in the screen space image 500. The corresponding point data is the corresponding value of the position of the physical tag in the three-dimensional space in the screen space image 500 in the two-dimensional space.
[0068] Please refer back to Figure 2 In some embodiments, in step S330, the second camera 214 captures an image of the physical tag 430 and transmits the acquired image of the physical tag 430 to the processor 250. After the processor 250 estimates the parameters of the second camera 214, the processor 250 obtains the position of the physical tag 430 relative to the second camera 214 in the three-dimensional space based on the image of the physical tag 430 and the parameters of the second camera 214. Since the second camera 214 and the display 212 are connected to each other and are close to each other, the position of the physical tag 430 relative to the second camera 214 in the three-dimensional space is similar to the position of the physical tag 430 relative to the display 212 in the three-dimensional space. When the processor 250 obtains the corresponding point data of the physical tag on the screen in step S330, the processor 250 generates the corresponding point data based on the position of the physical tag 430 relative to the second camera 214 in the three-dimensional space.
[0069] In step S350, the processor obtains the spatial calibration parameters according to the corresponding point data. In some embodiments, step S350 is performed as follows: Figure 2 The processor 250 is shown as executing the above-described process. In some embodiments, the processor 250 obtains the spatial calibration parameters by performing a singular value decomposition (SVD) operation based on the corresponding point data obtained in step S330. In some embodiments, the singular value decomposition operation adopts a known singular value operation method.
[0070] In step S370, the processor displays the image on the screen according to the spatial calibration parameters. In some embodiments, step S370 is performed by Figure 2 For example, please refer to Figure 6 . Figure 6 is a schematic diagram of a screen image 410 according to some embodiments of the present invention. Figure 6 As shown, the processor 250 performs the spatial correction based on the spatial correction parameters. Figure 2 The cuboid 610 is displayed on the screen 410 of the display 212 in the image. According to the space correction parameter, the processor 250 displays the cuboid 610 in the portion where it is expected to overlap with the image in the real space. When the user uses the head mounted display device 210, the image seen by the user through the display 212 is as shown in FIG. Figure 6 shown.
[0071] In some embodiments, the processor 250 may be located in a mobile phone, a server, or other devices. In some embodiments, the processor 250 may be a server, circuit, central processor unit (CPU), microprocessor (MCU) or other devices with equivalent functions, such as storage, calculation, data reading, receiving signals or messages, and sending signals or messages. In some embodiments, the camera 110, the first camera 230, and the second camera 214 may be circuits or other devices with equivalent functions, such as image capture and photo taking. The first camera 230 includes a position and angle adjuster to simulate the position and angle of the human eye. In some embodiments, the display 130 may be a circuit or other device with equivalent functions, such as image display.
[0072] like Figure 2 The virtual-real image fusion system 200 is only used for illustration and the present invention is not limited thereto. For example, in some embodiments, the processor 250 may be located in the head mounted display device 210 .
[0073] From the above-mentioned implementation methods of the present case, it can be seen that the embodiments of the present case provide a virtual-real image fusion method, a virtual-real image fusion system, and a non-transient computer-readable medium, using a camera to simulate human eyes, and adjusting different focal lengths through the device, without the need for manual adjustment, and generating a large amount of data in the form of automated equipment. In addition, the image obtained by the camera simulating the human eye is analyzed to find the position of the entity label in the image. In conjunction with the conventionally known camera estimation algorithm, the position of the entity label of the augmented reality in the three-dimensional space can be obtained to establish the corresponding data between the two-dimensional screen space image and the three-dimensional space, and then the space correction parameters are obtained by known methods. After knowing the camera position and the position of the object in the three-dimensional space, the space correction parameters can draw the corresponding virtual image on the screen to achieve the effect of virtual-real superposition.
[0074] In addition, the above examples include sequential exemplary steps, but the steps do not have to be performed in the order shown. Performing the steps in different orders is within the scope of the present disclosure. Within the spirit and scope of the embodiments of the present disclosure, the steps may be added, replaced, changed in order, and / or omitted as appropriate.
[0075] Although the present invention has been disclosed in the above embodiments, it is not intended to limit the present invention. Anyone familiar with the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the scope defined in the attached claims.
Claims
1. A virtual and real image fusion method, characterized in that: Include: Acquire a picture in a three-dimensional space by a first camera, wherein the picture includes a screen picture and a label picture of a physical label, wherein the screen picture is projected onto the physical label; A processor obtains corresponding point data of the physical tag on the screen according to the screen; The processor obtains a position of the screen frame in the frame to perform perspective conversion processing on the frame to generate a screen space frame; The processor obtains a spatial correction parameter according to the corresponding point data; and The processor displays an image on the screen according to the spatial calibration parameter.
2. The virtual-real image fusion method according to claim 1, characterized in that: Also includes: The image is subjected to an anti-warping conversion process.
3. The virtual-real image fusion method according to claim 1, characterized in that: Also includes: A position of the tag frame of the physical tag is obtained in the screen space frame to obtain the corresponding point data.
4. The virtual-real image fusion method according to claim 1, characterized in that: Also includes: A background color or a frame is displayed on the screen so that the processor can identify the position of the screen.
5. The virtual-real image fusion method according to claim 1, characterized in that: Also includes: The spatial calibration parameters are obtained by performing a singular value decomposition (SVD) operation according to the corresponding point data.
6. The virtual-real image fusion method according to claim 1, characterized in that: Also includes: Obtaining a position of the physical tag relative to a second camera in the three-dimensional space, wherein the screen image is displayed on a display, and the second camera and the display are located in a head-mounted display device; as well as The corresponding point data is obtained according to the position of the physical tag relative to the second camera in the three-dimensional space and the picture captured by the first camera.
7. The virtual-real image fusion method according to claim 6, characterized in that: Also includes: Acquire an image of the physical tag by the second camera; estimating a parameter of the second camera; and The position of the physical tag relative to the second camera in the three-dimensional space is obtained according to the image of the physical tag and the parameter of the second camera.
8. A virtual and real image fusion system, characterized in that: Include: A display for displaying a screen image; a first camera for acquiring a picture in a three-dimensional space, wherein the picture includes a screen picture and a label picture of a physical label, wherein the screen picture is projected onto the physical label; and A processor is used to obtain corresponding point data of the physical tag on the screen image based on the image, obtain a position of the screen image in the image to perform perspective transformation on the image to generate a screen space image, obtain a space correction parameter based on the corresponding point data, and display an image on the screen image based on the space correction parameter.
9. The virtual-real image fusion system according to claim 8, characterized in that: The processor is also used to perform anti-distortion conversion processing on the picture.
10. The virtual-real image fusion system according to claim 8, characterized in that: The processor is further used for obtaining a position of the tag frame of the physical tag in the screen space frame to obtain the corresponding point data.
11. The virtual-real image fusion system according to claim 8, characterized in that: The processor is further used to display a background color or a frame on the screen image so that the processor can identify the position of the screen image.
12. The virtual-real image fusion system according to claim 8, characterized in that: The processor is further used to obtain the space correction parameter by performing a singular value decomposition (SVD) operation according to the corresponding point data.
13. The virtual-real image fusion system according to claim 8, characterized in that: The processor is further used to obtain a position of the physical tag relative to a second camera in the three-dimensional space, wherein the second camera corresponds to the screen image, and obtain the corresponding point data based on the position of the physical tag relative to the second camera in the three-dimensional space and the image captured by the first camera.
14. The virtual-real image fusion system according to claim 13, characterized in that: Also includes: a second camera, for acquiring an image of the physical tag; The processor is further used for estimating a parameter of the second camera, and obtaining the position of the physical tag relative to the second camera in the three-dimensional space according to the image of the physical tag and the parameter of the second camera.
15. The virtual-real image fusion system according to claim 14, characterized in that: The display and the second camera are located in a head mounted display device and are communicatively connected with the processor and the first camera.
16. A non-transitory computer-readable medium, characterized in that The invention comprises at least one program instruction for executing a virtual-real image fusion method, wherein the virtual-real image fusion method comprises the following steps: A processor obtains corresponding point data of a physical tag on a screen image according to a picture, wherein the picture includes the screen image and a tag image of the physical tag, wherein the screen image is projected onto the physical tag; The processor obtains a position of the screen frame in the frame to perform perspective conversion processing on the frame to generate a screen space frame; Obtaining a spatial correction parameter according to the corresponding point data; and An image is displayed on the screen according to the spatial calibration parameter.
17. The non-transitory computer readable medium of claim 16, wherein: The virtual-real image fusion method also includes: The image is subjected to an anti-warping conversion process.
18. The non-transitory computer readable medium of claim 16, wherein: The virtual-real image fusion method also includes: A position of the tag frame of the physical tag is obtained in the screen space frame to obtain the corresponding point data.
Citation Information
Patent Citations
Calibrating real and virtual views
US20060152434A1
Information processing apparatus, information processing system, and information processing method
US20150070389A1