Method and device for converting 2D video into 3D video, head-mounted display equipment and storage medium

By calculating the maximum non-perceptible viewing angle of the human eye, adjusting the image to shift the background and magnify the foreground objects based on the parallax, the problem of time-consuming and expensive hole filling in XR glasses is solved, achieving a natural and efficient 2D to 3D video conversion, improving the stereoscopic effect and user experience.

CN122093544APending Publication Date: 2026-05-26DALIAN SITUNE TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN SITUNE TECHNOLOGY CO LTD
Filing Date
2026-04-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing 2D to 3D video conversion technology in XR glasses suffers from problems such as time-consuming and expensive hole filling and unnatural results, especially when displayed near the eye, resulting in poor stereoscopic effect and potentially causing visual discomfort.

Method used

By calculating the maximum unperceived viewing angle of the human eye, adjusting the image's anti-parallax translation of the background and magnifying the foreground object reduces the need for hole filling. Using pixel anti-parallax translation technology and magnifying the foreground object reduces or even eliminates holes, avoiding the use of duplicated pixels or AI-generated backgrounds.

Benefits of technology

It effectively reduces computational costs, improves the naturalness of 3D effects, avoids visual discomfort, and reduces the time and cost of filling in voids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093544A_ABST
    Figure CN122093544A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for converting a 2D video into a 3D video, head-mounted display equipment and a storage medium, and belongs to the technical field of converting the 2D video into the 3D video through XR glasses, the maximum angle at which pixel anti-parallax translation cannot cause discomfort of human eyes is obtained through an experiment, the maximum angle is set as the maximum non-perceptual visual angle of the human eyes, a movable pixel M of the maximum anti-parallax is calculated based on the maximum non-perceptual visual angle of the human eyes, and the 3D video is converted into the 2D video through the head-mounted display equipment. And obtaining a new picture of the adjusted left and right eyes of the XR glasses based on the movable pixel M, pushing the background in the picture to infinity, and then amplifying the foreground object to fill the holes, thereby reducing and even eliminating the holes which need to be filled due to the movement of the foreground object. According to the invention, the method does not need to copy pixels or generate a background through AI to fill up a hole generated by the movement of a foreground object, can effectively reduce the calculation cost, is more natural in 3D effect, and does not cause discomfort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of 2D to 3D video conversion for XR glasses, specifically relating to a method, apparatus, head-mounted display device, and storage medium for 2D to 3D video conversion. Background Technology

[0002] Shooting 3D movies or videos entirely with 3D dual-lens cameras is prohibitively expensive. Currently, most 3D movies or videos are shot in 2D initially and then converted to 3D in post-production using computer graphics processing and artificial intelligence generation methods. Since the XR glasses for watching movies include two screens for the left and right eyes, when the viewer needs to watch 3D videos, the existing 2D to 3D conversion steps are: (1) Generate depth map: The artificial intelligence model generates the depth map of each frame of the image; (2) Differentiate adjustment layers: Identify the background that does not need adjustment and the foreground objects that need adjustment based on the depth map; (3) Calculate adjustment value: Calculate the negative disparity pixels that need to be adjusted for the objects based on the depth map of the foreground objects that need adjustment and the exaggeration desired by the user, that is, obtain the adjustment value of the foreground objects; (4) Adjust the foreground objects: Obtain the pixel displacement value based on the depth map and the adjustment value, and shift each pixel of the foreground objects in each frame of the right eye video to the left, while the left eye screen displays the original image without adjustment; and (5) Fill the holes: Since the foreground objects move pixels to the left, the left-shifted foreground objects will block part of the background pixels on their left side, and holes will appear on the right side of the left-shifted foreground objects. For example, if the foreground objects move 20 pixels to the left, then 20 pixels of blank space will be generated on the right side of the foreground objects. The common approach is to copy the pixels from the right edge of the object to the right to fill the hole, or copy the background pixels to the left to fill the hole, or use the average of the pixels on both sides to fill the hole. The filled area looks like a foggy border. There are also methods to solve this problem by using AI to generate the background filling, but this approach is very time-consuming and expensive. Summary of the Invention

[0003] The purpose of this invention is to provide a method, apparatus, head-mounted display device, and storage medium for converting 2D to 3D video, which calculates the parallax limit by using the maximum non-perceptible viewing angle of the human eye, thereby reducing or even eliminating gaps that need to be filled due to the movement of foreground objects.

[0004] This invention discloses a method for converting 2D to 3D video, applicable to XR glasses, comprising the following steps: Step S1. By shifting the background away through image parallax, the negative parallax of foreground objects that needs adjustment is reduced: The experiment determined the maximum angle at which pixel parallax shift would not cause discomfort to the human eye, and set this as the maximum imperceptible viewing angle for the human eye. Calculate the maximum unperceived visual angle of this person's eye. The corresponding movable pixel M deletes the left M columns of pixels from the left screen image of the XR glasses, and at the same time deletes the right M columns of pixels from the right screen. The number of rows of pixels N that can be deleted on the Y-axis is calculated according to the resolution or aspect ratio of the original image, and N rows of pixels are deleted on the Y-axis of the image to generate new images for the left and right eyes. Step S2. Fill the hole created by the negative parallax adjustment that makes the object appear closer by enlarging the foreground object: Set the adjustment ratio f according to the actual effect, and calculate the negative parallax adjustment pixel P: P=M*f (5); When adjusting the negative parallax, increase the foreground object by P pixels on the X-axis and Y-axis respectively, so that the foreground object becomes larger and the object becomes closer; In the right screen, the right side of the foreground object is kept in the original position and enlarged to the left to become a new foreground object, and in the left screen, the left side of the foreground object is kept in the original position and enlarged to the right to become a new foreground object; If there are still a few holes after the foreground object is enlarged to become a new foreground object, fill them in using the traditional method.

[0005] The calculation of the maximum imperceptible field of view of the human eye The corresponding movable pixel M is as follows: The distance from the human eye to the XR glasses screen includes the thickness of the optical lenses. The total number of pixels (T) and size (S) along the horizontal X-axis of the image displayed on the XR glasses screen, and the maximum perceptible viewing angle for the human eye. First, calculate the maximum unperceptible field of view of the human eye. The corresponding displacement value relative to the real world Then, the displacement value relative to the real world is calculated based on the screen size. The corresponding movable pixel M: (1); (2).

[0006] The calculation of the maximum imperceptible field of view of the human eye The corresponding movable pixel M is as follows: Given the total number of pixels T along the horizontal X-axis of the image displayed on the XR glasses screen, the field of view (FOV) of the XR glasses optical design, and the maximum angle imperceptible to the human eye. Calculate the maximum angle imperceptible to the human eye The corresponding movable pixel M: (3).

[0007] The number of rows of pixels N that can be deleted along the Y-axis is calculated based on the resolution of the original image, specifically as follows: Given the resolution or resolution of the original image displayed on the XR glasses screen, i.e. * After deleting the pixels in column M, the X-axis pixels are: The formula for calculating N rows of pixels that can be deleted along the Y-axis is: ; ; (4); in, The number of columns of pixels on the Y-axis after deleting N rows of pixels from the original image.

[0008] In step S1, N / 2 rows of pixels are deleted above and below the Y-axis of the image. In step S2, P / 2 pixels are added to the top and bottom edges of the image along the Y-axis of the foreground object.

[0009] This invention discloses a device for converting 2D to 3D video, suitable for XR glasses, including a background processing module and a foreground object processing module: The background processing module reduces the negative parallax of foreground objects by shifting the background further away through image anti-parallax translation. The maximum angle at which pixel anti-parallax translation does not cause discomfort to the human eye, determined experimentally, is set as the maximum imperceptible viewing angle for the human eye. Calculate the maximum unperceived visual angle of this person's eye. The corresponding movable pixel M deletes the left M columns of pixels from the left screen image of the XR glasses, and at the same time deletes the right M columns of pixels from the right screen. The number of rows of pixels N that can be deleted on the Y-axis is calculated according to the resolution or aspect ratio of the original image, and N rows of pixels are deleted on the Y-axis of the image to generate new images for the left and right eyes. The foreground object processing module fills the gaps caused by the negative parallax adjustment that makes the object appear closer by enlarging the foreground object: the adjustment ratio f is set according to the actual effect, and the negative parallax adjustment pixel P is calculated: P=M*f (5); during negative parallax adjustment, the foreground object is increased by P pixels on the X-axis and Y-axis respectively, so that the foreground object is enlarged and the object appears closer; in the right screen, the right side of the foreground object is kept in the original position of the image and enlarged to the left to become a new foreground object, and in the left screen, the left side of the foreground object is kept in the original position of the image and enlarged to the right to become a new foreground object; if there are still a few gaps after the foreground object is enlarged to become a new foreground object, they are filled by the traditional method.

[0010] The background processing module deletes N / 2 rows of pixels above and below the Y-axis of the image; The foreground object processing module adds P / 2 pixels to the top and bottom edges of the image along the Y-axis.

[0011] This invention discloses a head-mounted display device, which includes at least two cameras for capturing target images of a target area; the head-mounted display device also includes a memory and a processor, the memory for storing computer programs; the processor for executing the computer programs to implement any of the above-described 2D to 3D video conversion methods.

[0012] This invention discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-described methods for converting 2D to 3D video.

[0013] After adopting the technical solution of the present invention, the movable pixel M with the maximum parallax is calculated based on the maximum non-perceptible viewing angle of the human eye. Based on the movable pixel M, the new images of the left and right eyes of the XR glasses are obtained after adjustment. The background in the image is pushed to infinity, and then the foreground object is magnified to fill the gaps. This reduces or even eliminates the gaps that need to be filled due to the movement of the foreground object. The present invention does not require the method of copying pixels or generating backgrounds with AI to fill the gaps caused by the movement of the foreground object. It can effectively reduce the computational cost, and the 3D effect is more natural and will not cause discomfort. Attached Figure Description

[0014] Figure 1 This is a schematic diagram illustrating the calculation of the movable pixel M with maximum parallax according to the present invention. Figure 2 This is a schematic diagram of the 2D-to-3D video conversion of the present invention; Figure 3 This is a functional structural block diagram of a head-mounted display device according to the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0017] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or solution described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. Explanation of the principle of this invention: (1) Binocular parallax and parallax shift: Human depth perception largely depends on the distance between the two eyes (interpupillary distance, IPD). Each eye sees the world from a slightly different angle. The brain fuses the images seen by the left and right eyes, calculating the difference between them (parallax) to determine depth. In XR glasses, when a 2D image is used to generate a 3D side-by-side image... By When viewed by the left and right eyes (SBS), this parallax is artificially reconstructed. Pixel movement in the generation of 3D images is controlled by parallax rules: Zero Parallax (reference plane): If an object is in exactly the same position in the left and right images, then it appears to be located on the depth plane of the screen. This is the reference plane for judging "in" and "out" of the screen. The object appears to be attached to the screen and is usually the focal point of the movie image. Negative Parallax (foreground object): If an object moves to the right in the left-eye image and to the left in the right-eye image (the two are closer together), the lines of sight of the two eyes will slightly cross when viewing, and the object will appear to "jump out of the screen". Positive Parallax (background): If an object moves to the left in the left-eye image and to the right in the right-eye image (the two are more separated), the lines of sight of both eyes will tend to be parallel. This makes the background appear to extend deeper into the screen, as if the object is "pushed" into the screen, creating a sense of depth. This is how the background is presented in most 3D movies.

[0018] (2) Depth-Image-Based Rendering (DIBR): The DIBR algorithm includes the following process: S1: Depth Map Generation The system analyzes 2D images (using methods such as AI, monocular depth cues, or motion structure recovery) and generates a grayscale depth map, typically with white representing foreground objects and black representing the background. S2: 3D Image Warping (also known as "pixel movement") Using depth map information, foreground object pixels are horizontally moved to generate left-eye and right-eye views. The horizontal displacement (x_shift) of the pixels is inversely proportional to their depth value, and background pixels are systematically stretched further apart to simulate distance.

[0019] (3) Existing AI-based algorithms: This approach combines monocular depth estimation (such as MiDaS or DepthAnything) with novel view synthesis (such as NeRF or 3D GaussianSplatting). Unlike simply "pulling back" the background (leaving holes where the foreground is occluded), the focus is on "image inpainting": as the background is pulled back to create depth, the AI ​​intelligently generates ("imagines") invisible pixel regions in the original 2D image that are occluded by foreground objects, thus filling in these missing holes.

[0020] (4) The problem of filling the void: However, "illusion" pixel generation is generally inaccurate and time-consuming and expensive, especially since the void area is small. A cheap and quick approach is to draw a line along the outside of the foreground object. Therefore, some 2D / 3D conversion software directly copies the edge pixels of the foreground object and the background image, merging them to fill the void. This creates a hazy border around the foreground object. Because XR glasses are very close to the user's eyes, the stereoscopic effect needs to enlarge this void, making the hazy border particularly noticeable. Advanced illusion filling can more effectively present a more realistic background; however, this approach is unreliable and very time-consuming. This is because AI-generated continuous images often have flaws. Assuming a movie frame rate of 90 frames per second, a 90-minute movie would require generating 486,000 frames, and ensuring each frame correctly fills the background void is very time-consuming. In contrast, movie theater screens are far away; generally, adjusting only 20 pixels of parallax pixels in 1080p is sufficient to achieve a noticeable stereoscopic effect. Since XR glasses are near-eye displays, it is difficult to perceive a stereoscopic 3D visual effect with a parallax adjustment of 20 pixels. Usually, more than 60 pixels need to be moved to achieve a good 3D effect, so filling the 60-pixel gap becomes particularly noticeable.

[0021] There are generally two ways to remedy this: (1) Adjust the parallax pixels for both eyes: If the foreground object in the left eye moves to the right, then the 60-pixel parallax becomes adjusting 30 pixels for the left eye and 30 pixels for the right eye, reducing the range of the blurred blank area. However, 30 pixels is still very wide, and it becomes a blurred area in both eyes, which is easy for the human eye to see. (2) In addition to adjusting the parallax for both eyes, the color of the filled pixels is also darkened to resemble the feeling of a shadow. Although shadows can help objects feel more three-dimensional, the shadows generated by filling the image will conflict with the shadows generated by the light source in the actual image. If the parallax is adjusted for both eyes, the shadows generated on both sides will conflict with each other.

[0022] (5) Stereo Window Violation Issue: There's an important psychological rule in 3D viewing called the stereoscopic window. The physical bezel of the monitor or VR lenses acts like a "window frame" through which the viewer sees the 3D world. A window violation occurs when an object appears closer to the viewer than the screen, seemingly "jumping out" of the screen, yet is simultaneously cut off by the left or right edge of the screen. This creates a logical conflict in the brain: "If this object is closer to me than the screen, why is it blocked by the screen bezel?" This contradiction disrupts the 3D illusion and can easily lead to headaches.

[0023] To address this issue, the professional 3D film and stereoscopic imaging industry employs a technique called stereoscopic imaging theory. This technique, known as Horizontal Image Translation (HIT), is used to adjust the Zero Parallax Setting (ZPS) and manage the Stereo Window.

[0024] The theory behind horizontal image translation (HIT): When AI generates 3D images, it creates parallax (the difference between the left and right eye images) based on the depth map. However, AI typically doesn't know the physical screen's position relative to the 3D scene. By deleting and translating edge pixels, it's essentially changing the image's viewfinder position. That is, the content in the left-eye image shifts to the left (closer to the new edge), and the content in the right-eye image shifts to the right. The result is that the distance between the left and right views of every object in the scene is increased overall. In stereoscopic vision theory, increasing the horizontal distance between the left and right views creates positive parallax, thus pushing the entire 3D world deeper into the screen.

[0025] However, we cannot infinitely increase parallax. While the theory of horizontal image shifting (HIT) can stretch the background and increase perceived depth, research in visual perception and ergonomics emphasizes that this technique has strict limitations; otherwise, it can cause visual discomfort or so-called "3D motion sickness." Specifically, these limitations are as follows: 1. Divergence Limit: The human eye is accustomed to viewing near objects by converging (crossing the eyes) and viewing distant objects (such as the horizon) by looking at them with parallel lines of sight. If the background pixels are separated by a distance greater than the viewer's physical interpupillary distance, the viewer's eyes must diverge to fuse the images. This immediately leads to eye strain, headaches, and disrupts the 3D effect.

[0026] 2. "Cardboard Effect": If the background is pushed back as a whole, and there is no smooth depth transition between the foreground object and the background, the foreground object will look like layers of flat cardboard instead of a three-dimensional object with real volume.

[0027] Example 1 Embodiment 1 of the present invention discloses a method for converting 2D to 3D video, applicable to XR glasses, comprising the following steps: Step S1. Push the background away and reduce the negative parallax of foreground objects by using image anti-parallax shifting: This invention posits that what the eye truly perceives is "angular disparity," not the physical distance between the viewer's pupils (IPD). Through repeated experimental testing, this invention has repeatedly verified that the human eye perceives a visual effect of parallax less than 1.5°. Therefore, as long as the parallax ultimately presented on the screen is controlled within 1.5°, which is comfortable for the human eye to blend—meaning there is no discomfort as long as the outward divergence angle of the eyes is controlled within 1.5°—the distance between the "two virtual cameras" during rendering can be set greater than the user's physical interpupillary distance (IPD). This allows the background world to be pushed to infinity, so that only a very small angle falls into the eye. While still acceptable to the brain, the depth effect perceived by the viewer will feel exaggerated yet natural. Simply put, the world spatial baseline is not equal to the screen spatial parallax. Interpupillary distance (IPD) is the distance between the two eyes in the real world (average 63 mm). What the human eye truly perceives is the "relative displacement" of the left and right eye images on the screen.

[0028] To reduce the holes caused by adjusting negative parallax for foreground objects, the original image is shifted inversely to make the background appear infinitely distant. However, the human eye cannot diverge outwards. This invention defines the maximum angle at which pixel-level parallax shifting does not cause discomfort to the human eye as the maximum imperceptible viewing angle. Therefore, this invention proposes the theory that the world spatial baseline is not equal to the screen spatial parallax and uses the maximum imperceptible viewing angle θ to adjust the left and right images displayed on XR glasses. Figure 1 As shown, assuming This is a preset maximum unperceptible field of view for the human eye (which can be obtained through experiments). It is necessary to calculate this maximum unperceptible field of view for the human eye. The corresponding movable pixel M is calculated using two methods, as described below: Method 1: The distance from the human eye to the XR glasses screen is known (including the thickness of the optical lenses). The total number of pixels (T) and size (S) along the horizontal X-axis of the image displayed on the XR glasses screen, and the maximum perceptible viewing angle for the human eye. First, calculate the maximum unperceptible field of view of the human eye. The corresponding displacement value relative to the real world Then, the displacement value relative to the real world is calculated based on the screen size. The corresponding movable pixel M: (1); (2); For example: θ = 1.5º, D view = 24mm, then Δx = D view *tan(θ) = 0.6285 mm; Given T = 1920 mm and S = 18.034 mm, then M = T * (Δx / S) = 67 pixels. Method 2: Given the total number of pixels T along the horizontal X-axis of the image displayed on the XR glasses screen, the field of view (FOV) of the XR glasses optical design, and the maximum imperceptible angle for the human eye. Calculate the maximum angle imperceptible to the human eye The corresponding movable pixel M: (3); For example: θ = 1.5º, T = 1920, FOV = 43º, then M = θ * (T / FOV) = 67 pixels.

[0029] Both methods above yielded M=67 pixels, proving their feasibility.

[0030] By deleting M columns of pixels from the left side of the XR glasses' left screen image and M columns of pixels from the right side of the right screen, the viewing angle for each eye is reduced by 1.5° without causing discomfort to the viewer. To ensure that the overall XY pixel ratio of the image matches the original image, the pixels along the Y-axis also need to be adjusted.

[0031] Assuming the original image displayed on the XR glasses screen is 1920x1080 (commonly known as 1080p), or a 16:9 aspect ratio, when M=67 columns of pixels are removed from the X-axis, only 1920-67=1853 pixels remain. Calculating with the same aspect ratio, the Y-axis pixels should become (1853 / 1920)*1080=1042 pixels, meaning N=38 rows of pixels must be removed from the Y-axis. Should the top or bottom rows be removed? To minimize perceptual change, N / 2 pixels should be removed from both the top and bottom, meaning 19 rows of pixels should be removed from both the top and bottom. The final image will be 1853x1042 resolution, still a 16:9 aspect ratio, consistent with the original image.

[0032] Given the resolution or resolution of the original image displayed on the XR glasses screen, i.e. * After deleting the pixels in column M, the X-axis pixels are: The formula for calculating N rows of pixels that can be deleted along the Y-axis is: ; ; (4); in, The number of columns of pixels on the Y-axis after deleting N rows of pixels from the original image; Corresponding to the maximum imperceptible field of view of the human eye After deleting M columns of pixels, the number of rows of pixels N that can be deleted along the Y-axis is calculated according to the resolution or aspect ratio of the original image. Then, N / 2 rows of pixels are deleted above and below the Y-axis of the image to generate new images for the left and right eyes.

[0033] You can also delete N rows of images with different proportions at the top and bottom depending on the effect. For example, if the original image at the bottom of the screen has subtitles and is not an important viewing area, you can directly delete the bottom N pixels without deleting the pixels at the top of the screen.

[0034] Step S2. Fill the hole created by the negative parallax adjustment that makes the object appear closer by enlarging the foreground object: Traditional 2D-to-3D video conversion methods use negative parallax to make foreground objects appear closer. Taking 1080p as an example, in step S1, based on the maximum imperceptible field of view of the human eye, the negative parallax adjustment to make objects appear closer in XR glasses requires a movement of more than 67 pixels to achieve the viewer's perception of closerness. When the foreground object on the right eye's screen moves 67 pixels to the left, a 67-pixel hole is simultaneously created on the right side of the object. This large hole exceeds 3% of the total X-axis pixels, making it a very noticeable crack. Since step S1 has effectively pushed the background to infinity, the foreground object only needs a slight adjustment in negative parallax to appear very three-dimensional. For example, adjusting the ratio M / 10 or f=1 / 10 results in a significant change. Of course, if the adjustment ratio is increased to f=1 / 5, the effect will be even more dramatic; this adjustment ratio f can be set according to the actual effect.

[0035] Calculate the negative disparity adjustment pixel P: P=M*f(5) During negative parallax adjustment, the foreground object is increased by P pixels on both the X-axis (width) and Y-axis (height) (i.e., new pixels on the X-axis of the foreground object = old pixels on the X-axis of the foreground object + P pixels; new pixels on the Y-axis of the foreground object = old pixels on the Y-axis of the foreground object + P pixels), making the foreground object appear larger and closer. On the right screen, the right side of the foreground object remains in its original position and is enlarged to the left to become a new foreground object; on the left screen, the left side of the foreground object remains in its original position and is enlarged to the right to become a new foreground object. The foreground object can be enlarged by P / 2 pixels on both the top and bottom sides along the Y-axis; alternatively, it can be increased by a total of P pixels on both the top and bottom sides at different ratios depending on the effect. If there are still a few holes after the foreground object is enlarged vertically to become a new foreground object, they can be filled using traditional methods.

[0036] like Figure 2As shown, the top left corner is the original 2D image containing the foreground object (fish); the top right corner is the depth map; the bottom left corner is the display screen of the left screen of the XR glasses after 2D-to-3D conversion, which is the result of deleting M columns on the left and N / 2 rows above and below the original 2D image, and enlarging the foreground object (fish) by M / 2 pixels to the right; the bottom right corner is the display screen of the right screen of the XR glasses after 2D-to-3D conversion, which is the result of deleting M columns on the right and N / 2 rows above and below the original 2D image, and enlarging the foreground object (fish) by M / 2 pixels to the left.

[0037] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0038] Specifically, the steps of the method embodiments in this application can be implemented by integrated logic circuits in the processor hardware and / or instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information in the memory and implements the steps in the above method embodiments in combination with its hardware.

[0039] Example 2 Embodiment 2 of the present invention discloses a device for converting 2D to 3D video, suitable for XR glasses, comprising a background processing module and a foreground object processing module, wherein: The background processing module reduces the negative parallax of foreground objects by shifting the background further away through image anti-parallax translation. The maximum angle at which pixel anti-parallax translation does not cause discomfort to the human eye, determined experimentally, is set as the maximum imperceptible viewing angle for the human eye. Calculate the maximum unperceived visual angle of this person's eye. The corresponding movable pixel M deletes the left M columns of pixels from the left screen image of the XR glasses, and at the same time deletes the right M columns of pixels from the right screen. The number of rows of pixels N that can be deleted on the Y-axis is calculated according to the resolution or aspect ratio of the original image. N / 2 rows of pixels are deleted above and below the Y-axis of the image to generate new images for the left and right eyes. The calculation of the maximum imperceptible field of view of the human eye The corresponding movable pixel M: The distance from the human eye to the XR glasses screen (including the thickness of the optical lenses) is known. The total number of pixels (T) and size (S) along the horizontal X-axis of the image displayed on the XR glasses screen, and the maximum perceptible viewing angle for the human eye. First, calculate the maximum unperceptible field of view of the human eye. The corresponding displacement value relative to the real world Then, the displacement value relative to the real world is calculated based on the screen size. The corresponding movable pixel M: (1); (2); The calculation of the maximum imperceptible field of view of the human eye The corresponding movable pixel M: Given the total number of pixels T along the horizontal X-axis of the image displayed on the XR glasses screen, the field of view (FOV) of the XR glasses optical design, and the maximum angle imperceptible to the human eye. Calculate the maximum angle imperceptible to the human eye The corresponding movable pixel M: (3); The number of rows of pixels N that can be deleted along the Y-axis is calculated based on the resolution of the original image. Given the resolution or resolution of the original image displayed on the XR glasses screen, i.e. * After deleting the pixels in column M, the X-axis pixels are: The formula for calculating N rows of pixels that can be deleted along the Y-axis is: ; ; (4); in, The number of columns of pixels on the Y-axis after deleting N rows of pixels from the original image; The foreground object processing module fills the gaps caused by the negative parallax adjustment that makes the object appear closer by enlarging the foreground object: the adjustment ratio f is set according to the actual effect, and the negative parallax adjustment pixel P is calculated: P=M*f (5); during negative parallax adjustment, the foreground object is increased by P pixels on the X-axis and Y-axis respectively (i.e., the new pixel of the foreground object on the X-axis = the old pixel of the foreground object on the X-axis + P pixels; the new pixel of the foreground object on the Y-axis = the old pixel of the foreground object on the Y-axis + P pixels), so that the foreground object is enlarged to produce the effect of the object appearing closer; in the right screen, the right side of the foreground object is kept in the original position of the image and enlarged to the left to become a new foreground object, and in the left screen, the left side of the foreground object is kept in the original position of the image and enlarged to the right to become a new foreground object; the foreground object can be enlarged by P / 2 pixels on the top and bottom sides in the Y-axis direction; or it can be increased by P pixels in total according to different ratios on the top and bottom sides according to the effect; if there are still a few gaps after the foreground object is enlarged to the top and bottom to become a new foreground object, they are filled by the traditional method.

[0040] Example 3 Embodiment 3 of the present invention provides a head-mounted display device, such as... Figure 3 As shown, the head-mounted display device 700 may include a memory 710 and a processor 720. The memory 710 stores a computer program and transmits the program code to the processor 720. In other words, the processor 720 can call and run the computer program from the memory 710 to implement the method in Embodiment 1 of this application. For example, the processor 720 can be used to execute the method described in Embodiment 1 according to the instructions in the computer program. In some embodiments of this application, the processor 720 may include, but is not limited to: General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0041] In some embodiments of this application, the memory 710 includes, but is not limited to, volatile memory and / or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).

[0042] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 710 and executed by the processor 720 to complete the method of Embodiment 1 provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the head-mounted display device 700.

[0043] like Figure 3 As shown, the head-mounted display device may further include a transceiver 730, which can be connected to the processor 720 or the memory 710. The processor 720 can control the transceiver 730 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 730 may be at least two cameras used to capture target images of a target area.

[0044] It should be understood that the various components in the head-mounted display device 700 are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.

[0045] Example 4 Embodiment 4 of the present invention also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the method described in Embodiment 1 above.

[0046] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for converting 2D to 3D video, suitable for XR glasses, characterized in that... Includes the following steps: Step S1. By shifting the background away through image parallax, the negative parallax of foreground objects that needs adjustment is reduced: The experiment determined the maximum angle at which pixel parallax shift would not cause discomfort to the human eye, and set this as the maximum imperceptible viewing angle for the human eye. Calculate the maximum unperceived visual angle of this person's eye. The corresponding movable pixel M deletes the left M columns of pixels from the left screen image of the XR glasses, and at the same time deletes the right M columns of pixels from the right screen. The number of rows of pixels N that can be deleted on the Y-axis is calculated according to the resolution or aspect ratio of the original image, and N rows of pixels are deleted on the Y-axis of the image to generate new images for the left and right eyes. Step S2. Fill the hole created by the negative parallax adjustment that makes the object appear closer by enlarging the foreground object: Set the adjustment ratio f according to the actual effect, and calculate the negative parallax adjustment pixel P: P=M*f (5); When adjusting the negative parallax, increase the foreground object by P pixels on the X-axis and Y-axis respectively, so that the foreground object becomes larger and the object becomes closer. In the right screen, the right side of the foreground object remains in the original position of the image and is magnified to the left to become a new foreground object; in the left screen, the left side of the foreground object remains in the original position of the image and is magnified to the right to become a new foreground object. If a small number of holes still exist after the foreground object is magnified into a new foreground object, they can be filled using traditional methods.

2. The method for converting 2D to 3D video according to claim 1, characterized in that, The calculation of the maximum imperceptible field of view of the human eye The corresponding movable pixel M is as follows: The distance from the human eye to the XR glasses screen includes the thickness of the optical lenses. The total number of pixels (T) and size (S) along the horizontal X-axis of the image displayed on the XR glasses screen, and the maximum perceptible viewing angle for the human eye. First, calculate the maximum unperceptible field of view of the human eye. The corresponding displacement value relative to the real world Then, the displacement value relative to the real world is calculated based on the screen size. The corresponding movable pixel M: (1); (2)。 3. The method for converting 2D to 3D video according to claim 1, characterized in that, The calculation of the maximum imperceptible field of view of the human eye The corresponding movable pixel M is as follows: Given the total number of pixels T along the horizontal X-axis of the image displayed on the XR glasses screen, the field of view (FOV) of the XR glasses optical design, and the maximum angle imperceptible to the human eye. Calculate the maximum angle imperceptible to the human eye The corresponding movable pixel M: (3)。 4. The method for converting 2D to 3D video according to claim 1, characterized in that, The number of rows of pixels N that can be deleted along the Y-axis is calculated based on the resolution of the original image, specifically as follows: Given the resolution or resolution of the original image displayed on the XR glasses screen, i.e. * After deleting the pixels in column M, the X-axis pixels are: The formula for calculating N rows of pixels that can be deleted along the Y-axis is: ; ; (4); in, The number of columns of pixels on the Y-axis after deleting N rows of pixels from the original image.

5. The method for converting 2D to 3D video according to claim 1, characterized in that, In step S1, N / 2 rows of pixels are deleted above and below the Y-axis of the image. In step S2, P / 2 pixels are added to the top and bottom edges of the image along the Y-axis of the foreground object.

6. A device for converting 2D to 3D video, suitable for XR glasses, comprising a background processing module and a foreground object processing module, characterized in that: The background processing module reduces the negative parallax of foreground objects by shifting the background further away through image anti-parallax translation. The maximum angle at which pixel anti-parallax translation does not cause discomfort to the human eye, determined experimentally, is set as the maximum imperceptible viewing angle for the human eye. Calculate the maximum unperceived visual angle of this person's eye. The corresponding movable pixel M deletes the left M columns of pixels from the left screen image of the XR glasses, and at the same time deletes the right M columns of pixels from the right screen. The number of rows of pixels N that can be deleted on the Y-axis is calculated according to the resolution or aspect ratio of the original image, and N rows of pixels are deleted on the Y-axis of the image to generate new images for the left and right eyes. The foreground object processing module fills the gap caused by the negative parallax adjustment that makes the object appear closer by enlarging the foreground object: the adjustment ratio f is set according to the actual effect, and the negative parallax adjustment pixel P is calculated: P=M*f (5); during negative parallax adjustment, the foreground object is increased by P pixels on the X-axis and Y-axis respectively, so that the foreground object is enlarged and the object appears closer. In the right screen, the right side of the foreground object remains in the original position of the image and is magnified to the left to become a new foreground object; in the left screen, the left side of the foreground object remains in the original position of the image and is magnified to the right to become a new foreground object. If a small number of holes still exist after the foreground object is magnified into a new foreground object, they can be filled using traditional methods.

7. The apparatus for converting 2D to 3D video according to claim 6, characterized in that: The background processing module deletes N / 2 rows of pixels above and below the Y-axis of the image; The foreground object processing module adds P / 2 pixels to the top and bottom edges of the image along the Y-axis.

8. A head-mounted display device, characterized in that, The head-mounted display device includes at least two cameras for capturing target images of a target area; the head-mounted display device also includes a memory and a processor, the memory for storing a computer program; the processor is used to execute the computer program to implement the 2D to 3D video conversion method of any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements any one of the 2D to 3D video conversion methods of claims 1 to 5.