Image perspective conversion display apparatus and method, and storage medium
Through the combined hardware module of a digital signal processor and a network processor with a dedicated integrated circuit, high-precision image internal compensation and energy consumption reduction are achieved, solving the problems of hollows and edge distortion in the viewing angle conversion image, and adapting to new extended reality display devices with high resolution and high frame rates.
Patent Information
- Application Number
- PCT/CN2025/071242
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-01
- Filing Date
- 2025-01-08
- Publication Date
- 2025-08-07
AI Technical Summary
The prior art has problems of hollowness and edge distortion in view angle conversion images, and it is difficult to adapt to new extended reality display devices with high resolution and high frame rate, with high computing power requirements and high energy consumption.
The digital signal processor carries out curling of video images and depth data, the network processor carries out object detection and internal complementation operations, and the application-specific integrated circuit performs layer aliasing. Combined with the data processing characteristics of the hardware module, high-precision image internal complementation and energy consumption reduction are achieved.
Improve the aliasing efficiency of image complement and video images, reduce energy consumption, and adapt to new extended reality display devices with high resolution and high frame rates.
Smart Images

Figure CN2025071242_07082025_PF_FP_ABST
Abstract
Description
Image viewing angle conversion display device, method and storage medium
[0001] This application claims priority to a patent application with a filing date of February 1, 2024, Chinese application number 202410140967.3, entitled “An image perspective conversion display device, method and storage medium”, and requests that all the contents disclosed therein be incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of image processing technology, and in particular to an image viewing angle conversion display device, an image viewing angle conversion display method, and a computer-readable storage medium. Background Art
[0003] When dealing with perspective conversion problems, existing technologies mainly use the geometric methods of traditional three-dimensional visual computing, or optimize the local modules of traditional geometric methods through deep learning algorithm models to simplify traditional geometric methods into several steps such as stereo correction, left and right disparity / depth estimation, image warping processing, and post-processing.
[0004] However, the perspective conversion images obtained by traditional 3D visual computational geometry methods generally contain a large number of holes and outliers, and there is edge distortion, which makes the visual experience poor when there are objects nearby. Although the geometric method optimized by the deep learning algorithm model can overcome the defects of holes and outliers to a certain extent, it generally requires a lot of computing power and still produces certain edge distortions and temporal jitter problems. In addition, due to the high computing power requirements, this technology can only be implemented on lower-resolution images, and it is difficult to adapt to the new extended reality (XR) display devices with high resolution and high frame rate.
[0005] In order to overcome the above-mentioned defects of the existing technology, this field urgently needs an image perspective conversion display technology. By combining the data processing characteristics of the hardware module, high-precision image infill based on depth and target detection can be achieved, thereby improving the aliasing efficiency of the infilled image and the video image, and reducing the energy consumption of image aliasing. Summary of the Invention
[0006] The following is a brief summary of one or more aspects to provide a basic understanding of these aspects. This summary is not an exhaustive overview of all conceivable aspects and is neither intended to identify key or critical elements of all aspects nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that will be provided later.
[0007] In order to overcome the above-mentioned defects of the prior art, the present invention provides an image perspective conversion display device, an image perspective conversion display method, and a computer-readable storage medium, which can realize high-precision image interpolation based on depth and target detection by combining the data processing characteristics of hardware modules, thereby improving the aliasing efficiency of the interpolated image and the video image, and reducing the energy consumption of image aliasing.
[0008] Specifically, the above-mentioned image perspective conversion display device provided according to the first aspect of the present invention includes: a digital signal processor, which is used to warp the first video image from the camera perspective and its corresponding depth data to obtain a second video image from the user's human eye perspective; the network processor is used to perform target detection on the first video image to determine a first image mask indicating an occluded area in the second video image, and perform an internal complement operation based on the second video image and the first image mask to obtain an internal complement image that fills the value of at least one pixel in the occluded area; and a dedicated integrated circuit, which is used to perform correction and layer aliasing on the first video image and the internal complement image to obtain a third video image to be displayed.
[0009] Preferably, in one embodiment of the present invention, the image perspective conversion display device also includes a scene camera and a depth sensor, and the digital signal processor is configured to: obtain the first video image via the scene camera; obtain depth data corresponding to the first video image via the depth sensor; and warp the first video image and the depth data to generate the second video image.
[0010] Preferably, in one embodiment of the present invention, a depth post-processing unit is configured in the digital signal processor, and the step of obtaining depth data corresponding to the first video image via the depth sensor includes: obtaining first depth data from the sensor perspective via the depth sensor; and warping the first depth data via the depth post-processing unit to obtain second depth data from the camera perspective.
[0011] Preferably, in one embodiment of the present invention, the dedicated integrated circuit includes a first dedistortion unit, and the first dedistortion unit is configured to: obtain a high-resolution fourth video image via the scene camera; obtain a camera distortion grid indicating camera distortion of the scene camera, and / or a preset target resolution; and process the fourth video image according to the camera distortion grid and / or the target resolution to obtain a low-resolution first video image with the camera distortion removed.
[0012] Preferably, in one embodiment of the present invention, the first dedistortion unit is connected to the network processor, and the step of processing the fourth video image according to the camera distortion grid and / or the target resolution to obtain a low-resolution first video image with the camera distortion removed includes: acquiring a second image mask via the network processor to determine a first area corresponding to the occluded area and a remaining second area in the fourth video image; assigning a default value to at least one first pixel in the first area; performing dedistortion processing on at least one second pixel in the second area according to the camera distortion grid; and compressing the first area and the second area according to the target resolution to obtain the first video image.
[0013] Preferably, in one embodiment of the present invention, the step of warping the first video image and the depth data to generate the second video image includes: determining first 3D video data under the camera perspective based on the first video image and its corresponding depth data; obtaining a perspective conversion grid indicating the conversion relationship between the camera perspective and the human eye perspective; determining second 3D video data under the human eye perspective based on the first 3D video data and the perspective conversion grid; and determining the second video image based on the second 3D video data.
[0014] Preferably, in one embodiment of the present invention, a target detection unit is configured in the network processor, and the target detection unit is configured to: obtain the first video image; identify and separate the foreground target image and the remaining background image from the first video image through a pre-trained target detection network, wherein the target detection network is trained based on sample images of a mixed reality (MR) display scene; determine a second image mask based on the positions of the foreground target image and the background image in the first video image; and determine the first image mask based on the positions of the foreground target image and the background image in the second video image.
[0015] Preferably, in one embodiment of the present invention, the step of identifying and separating the foreground target image and the remaining background image from the first video image through a pre-trained target detection network includes: obtaining depth data corresponding to the first video image; dividing the first video image into multiple regions according to the gradient of the depth data; and determining the image in the third region with a smaller depth as the foreground target image, and determining the image in the fourth region with a larger depth as the background image.
[0016] Preferably, in one embodiment of the present invention, the network processor is further configured with an internal complement operation unit, and the internal complement operation unit is configured to: obtain the second video image, the first image mask and the second image mask; input the second video image, the first image mask and the second image mask into a pre-configured internal complement algorithm to determine the value of at least one pixel point of the occluded area in the background image; and fill each pixel point of the occluded area according to the value of each pixel point of the occluded area to obtain the internal complement image.
[0017] Preferably, in one embodiment of the present invention, the ASIC is configured with a second dedistortion unit and a third dedistortion unit. The second dedistortion unit is connected to the digital signal processor and configured to: obtain a camera distortion grid indicating camera distortion, an optomechanical distortion grid indicating optomechanical distortion, and / or a perspective conversion grid indicating the conversion relationship between the camera perspective and the human eye perspective to determine a corresponding video distortion grid; and perform correction on the first video image based on the video distortion grid to obtain a fifth video image from the human eye perspective. The third dedistortion unit is connected to the network processor and configured to: obtain the optomechanical distortion grid to determine a corresponding interpolation distortion grid; and correct the first interpolation image obtained via the network processor based on the interpolation distortion grid to obtain a second interpolation image from which the optomechanical distortion is removed. The ASIC performs layer blending on the fifth video image from the human eye perspective and the second interpolation image to obtain the third video image.
[0018] Preferably, in one embodiment of the present invention, the second dedistortion unit is connected to the scene camera and the network processor. The step of correcting the first video image according to the video distortion grid to obtain the fifth video image with a high resolution from the human eye's perspective includes: obtaining the first image mask via the network processor to determine video layer data located in the occluded area in the fourth video image with a high resolution obtained by the scene camera; and warping the video layer data according to the video distortion grid to obtain the fifth video image with a high resolution from the human eye's perspective.
[0019] Preferably, in one embodiment of the present invention, the step of correcting the first interpolated image obtained via the network processor according to the interpolated distortion grid to obtain a second interpolated image with the optomechanical distortion removed includes: determining, according to the first image mask, interpolated layer data located in the occluded area in the first interpolated image; and warping the interpolated layer data according to the interpolated distortion grid to obtain a high-resolution second interpolated image from the perspective of the human eye.
[0020] Preferably, in one embodiment of the present invention, the ASIC is configured with a plurality of layer blending units. A first layer blending unit is configured to perform layer blending on the fifth video image and the second internal complement image to obtain an intermediate image, and a second layer blending unit is configured to perform layer blending on the intermediate image and at least one virtual image to obtain the third video image.
[0021] Preferably, in one embodiment of the present invention, the ASIC is further configured with a first color processing unit and / or a second color processing unit. The first color processing unit is located before the first layer blending unit and is configured to perform a preset first color processing on the video image from the human eye perspective. The second color processing unit is located before the first layer blending unit and is configured to perform a preset second color processing on the interpolated image.
[0022] Preferably, in one embodiment of the present invention, the ASIC is further configured with a first format processing unit and / or a second format processing unit. The first format processing unit is located before the first layer blending unit and is configured to perform a preset first format processing on the data format of the video image from the human eye perspective; the second format processing unit is located before the first layer blending unit and is configured to perform a preset second format processing on the data format of the interpolated image.
[0023] Preferably, in one embodiment of the present invention, the device further comprises: a display screen; and a display light engine. The display light engine is located in front of the display screen and is used to increase the field of view of the third video image and display the third video image with the increased field of view on the display screen.
[0024] In addition, the above-mentioned image perspective conversion and display method provided according to the second aspect of the present invention includes the following steps: obtaining a first video image from the camera perspective and its corresponding depth data; warping the first video image and the depth data via a digital signal processor to obtain a second video image from the user's human eye perspective; performing target detection on the first video image via a network processor to determine an image mask indicating an occluded area in the second video image, and performing an internal complement operation based on the second video image and the image mask to obtain an internal complement image that fills at least one pixel value in the occluded area; and performing layer aliasing on the video image from the human eye perspective and the internal complement image via a dedicated integrated circuit to obtain a third video image to be displayed.
[0025] Preferably, in one embodiment of the present invention, the step of acquiring the first video image includes: acquiring a high-resolution fourth video image via a scene camera; acquiring a camera distortion grid indicating camera distortion of the scene camera, and / or a preset target resolution; and processing the fourth video image according to the camera distortion grid and / or the target resolution to obtain a low-resolution first video image with the camera distortion removed.
[0026] Preferably, in one embodiment of the present invention, the step of processing the fourth video image according to the camera distortion grid and / or the target resolution to obtain a low-resolution first video image with the camera distortion removed includes: obtaining an optomechanical distortion grid indicating optomechanical distortion and / or a perspective conversion grid indicating a conversion relationship between the camera perspective and the human eye perspective; determining a corresponding video distortion grid according to the camera distortion grid, the optomechanical distortion grid and / or the perspective conversion grid; and correcting the fourth video image according to the video distortion grid to obtain a high-resolution fifth video image from the human eye perspective.
[0027] Preferably, in an embodiment of the present invention, the step of acquiring the depth data includes: acquiring first depth data from a sensor perspective via a depth sensor; and performing warping processing on the first depth data to obtain second depth data from the camera perspective.
[0028] Preferably, in one embodiment of the present invention, the step of performing target detection on the first video image to determine an image mask indicating the occluded area in the second video image includes: identifying and separating the foreground target image and the remaining background image from the first video image via a pre-trained target detection network, wherein the target detection network is obtained by training based on sample images of MR scenes; and determining the image mask based on the positions of the foreground target image and the background image in the second video image.
[0029] Preferably, in one embodiment of the present invention, the step of performing an internal complement operation based on the second video image and the image mask to obtain an internal complement image that fills the value of at least one pixel in the occluded area includes: inputting the second video image and the image mask into a pre-configured internal complement algorithm to determine the value of at least one pixel point in the occluded area in the background image; and filling each pixel point in the occluded area according to the value of each pixel point in the occluded area to obtain the internal complement image.
[0030] Furthermore, the computer-readable storage medium provided in accordance with the third aspect of the present invention stores computer instructions, which, when executed by a processor, implement the image viewing angle conversion and display method provided in any one of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above features and advantages of the present invention will be better understood after reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings. In the drawings, the components are not necessarily drawn to scale, and components with similar related properties or characteristics may have the same or similar reference numerals.
[0032] FIG1 shows a schematic diagram of the architecture of an image viewing angle conversion display device provided according to some embodiments of the present invention;
[0033] FIG2 is a schematic diagram showing a flow chart of an image viewing angle conversion and display method according to some embodiments of the present invention;
[0034] 3A to 3D are schematic diagrams showing images at various stages of an internal complementation algorithm according to some embodiments of the present invention; and
[0035] FIG. 4 shows a schematic diagram of layer aliasing according to some embodiments of the present invention. DETAILED DESCRIPTION
[0036] The following specific embodiments illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Although the description of the present invention will be introduced in conjunction with the preferred embodiment, this does not mean that the features of this invention are limited to this embodiment. On the contrary, the purpose of introducing the invention in conjunction with the embodiment is to cover other options or modifications that may be extended based on the claims of the present invention. In order to provide a deep understanding of the present invention, the following description will include many specific details. The present invention can also be implemented without using these details. In addition, in order to avoid confusion or blurring the focus of the present invention, some specific details will be omitted in the description.
[0037] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0038] Furthermore, the terms "upper," "lower," "left," "right," "top," "bottom," "horizontal," and "vertical" used in the following description should be understood to refer to the orientations depicted in that section and the accompanying drawings. These relative terms are used solely for convenience of description and do not necessarily imply that the devices described herein must be manufactured or operated in a specific orientation. Therefore, they should not be construed as limiting the present invention.
[0039] It will be understood that although the terms "first," "second," "third," etc. may be used herein to describe various components, regions, layers, and / or portions, these components, regions, layers, and / or portions should not be limited by these terms, and these terms are merely used to distinguish different components, regions, layers, and / or portions. Thus, a first component, region, layer, and / or portion discussed below may be referred to as a second component, region, layer, and / or portion without departing from some embodiments of the present invention.
[0040] As mentioned above, the perspective conversion images obtained by traditional 3D visual computational geometry methods generally have a large number of holes and outliers, and there is edge distortion, which makes the visual experience poor when there are objects nearby. Although the geometric method optimized by the deep learning algorithm model can overcome the defects of holes and outliers to a certain extent, it generally requires a large amount of computing power and still produces certain edge distortion and time domain jitter problems. In addition, due to the high computing power requirements, this technology can only be implemented on lower resolution images, but it is difficult to adapt to new XR display devices with high resolution (for example: 4K resolution) and high frame rate (for example: 120Hz).
[0041] In order to overcome the above-mentioned defects of the prior art, the present invention provides an image perspective conversion display device, an image perspective conversion display method, and a computer-readable storage medium, which can realize high-precision image interpolation based on depth and target detection by combining the data processing characteristics of hardware modules, thereby improving the aliasing efficiency of the interpolated image and the video image, and reducing the energy consumption of image aliasing.
[0042] In some non-limiting embodiments, the image viewing angle conversion display method provided in the second aspect of the present invention may be implemented via the image viewing angle conversion display device provided in the first aspect of the present invention.
[0043] Please refer to FIG. 1 for details. FIG. 1 shows a schematic diagram of the architecture of an image viewing angle conversion display device provided according to some embodiments of the present invention.
[0044] In the embodiment shown in FIG1 , the image perspective conversion and display device provided in the first aspect of the present invention is configured with a memory 10 and a processor 20. The memory 10 includes, but is not limited to, the computer-readable storage medium provided in the third aspect of the present invention, and stores computer instructions thereon. The processor 20 is connected to the memory 10 and is configured to execute the computer instructions stored in the memory 10 to implement the image perspective conversion and display method provided in the second aspect of the present invention.
[0045] Furthermore, the processor 20 may be implemented by a digital signal processor (DSP) 21, a network processing unit (NPU) 22, and an application specific integrated circuit (ASIC) 23 in cooperation with each other.
[0046] Specifically, the digital signal processor (DSP) 21 can be used to warp the first video image from the camera's perspective and its corresponding depth data to obtain a second video image from the user's human eye perspective. The network processor (NPU) 22 can be used to perform target detection on the first video image to determine a first image mask indicating an occluded area in the second video image, and perform an internal complement operation based on the second video image and the first image mask to obtain an internal complement image that fills at least one pixel value in the occluded area. The application-specific integrated circuit (ASIC) 23 can be used to perform correction and layer aliasing on the first video image and the internal complement image to obtain a third video image to be displayed.
[0047] The following will describe the working principles of the above-mentioned image perspective conversion display device and its processing units such as the digital signal processor (DSP) 21, network processor (NPU) 22, and application-specific integrated circuit (ASIC) 23 in conjunction with some embodiments of the image perspective conversion display method. Those skilled in the art will understand that the embodiments of these image perspective conversion display methods are only some non-restrictive implementation methods provided by the present invention, which are intended to clearly demonstrate the main concept of the present invention and provide some specific solutions that are convenient for the public to implement, rather than to limit all functions or all working modes of the image perspective conversion display device. Similarly, the image perspective conversion display device, digital signal processor (DSP) 21, network processor (NPU) 22, and application-specific integrated circuit (ASIC) 23 are also only a non-restrictive implementation method provided by the present invention, and do not constitute a limitation on the execution subject and execution order of each step in these image perspective conversion display methods.
[0048] Please refer to FIG. 1 and FIG. 2 . FIG. 2 shows a flow chart of an image viewing angle conversion and display method according to some embodiments of the present invention.
[0049] In the embodiments shown in Figures 1 and 2, the image perspective conversion display device may further include a scene camera 31 and a depth sensor 32. During the image perspective conversion display process, the digital signal processor (DSP) 21 may first obtain a first video image from the camera perspective via the scene camera 31, and obtain depth data corresponding to the first video image via the depth sensor 32, and then warp the first video image and the depth data to generate a second video image from the user's human eye perspective.
[0050] Furthermore, in some embodiments, the digital signal processor (DSP) 21 may be configured with a depth post-processing unit 211 and a forward warping unit 212. In response to obtaining first depth data from the depth sensor's perspective through the depth sensor 32, the digital signal processor (DSP) 21 may first perform warping processing on the first depth data via the depth post-processing unit 211 to obtain second depth data from the camera's perspective, and then perform forward warping processing on the first video image and the second depth data from the camera's perspective via the forward warping unit 212 to obtain a second video image from the user's human eye perspective.
[0051] Specifically, during the forward warping process of the first video image and the second depth data from the camera's perspective, the forward warping unit 212 first determines the first 3D video data from the camera's perspective based on the first video image and its corresponding second depth data, and obtains a perspective conversion grid 243 indicating the conversion relationship between the camera's perspective and the human eye's perspective. The forward warping unit 212 then determines the second 3D video data from the human eye's perspective based on the first 3D video data and the perspective conversion grid 243, and then determines the second video image from the user's perspective based on the second 3D video data.
[0052] Furthermore, in some embodiments of the present invention, the application-specific integrated circuit (ASIC) 23 may be configured with a first dedistortion unit 231. Here, the first dedistortion unit 231 stores a camera distortion grid 241 indicating the camera distortion of the scene camera 31. In response to obtaining a high-resolution (e.g., >8M pixel resolution) fourth video image from the camera perspective from the scene camera 31, the application-specific integrated circuit (ASIC) 23 may first remove the camera distortion from the fourth video image using the first dedistortion unit 231 based on the camera distortion grid 241 stored therein, and simultaneously compress the image based on a preset target resolution to obtain a low-resolution (e.g., <1M pixel resolution) first video image with the camera distortion removed. The low-resolution first video image is then transmitted to the forward warping unit 212 of the digital signal processor (DSP) 21 for the warping process described above, thereby reducing the data processing load of the warping process. Here, the target resolution may be consistent with the resolution of the depth map data collected by the depth sensor 32.
[0053] Furthermore, in some embodiments, the first dedistortion unit 231 of the application-specific integrated circuit (ASIC) 23 may be connected to the network processor (NPU) 22 to obtain the second image mask 2221 generated by the NPU 22 via an internal complement algorithm, and determine, based on the Boolean value of the second image mask 222, a first region in the fourth video image corresponding to the occluded region in the second video image and a second region corresponding to the remaining unoccluded region. Subsequently, during the process of removing camera distortion, the first dedistortion unit 231 may skip the dedistortion process and assign a default value to at least one first pixel in the first region, performing dedistortion processing only on at least one second pixel in the remaining second region. The first and second regions are then compressed based on the target resolution to obtain the first video image, thereby further reducing the data processing load of the dedistortion process and improving its real-time performance.
[0054] In addition, in some embodiments of the present invention, the network processor (NPU) 22 may be configured with a target detection unit 221. In response to obtaining a first video image from the camera perspective from the scene camera 31 or the first dedistortion unit 231, the target detection unit 221 may identify and separate the foreground target image and the remaining background image from the first video image via a pre-trained target detection network, and then determine the first image mask 2221 based on the positions of the foreground target image and the background image in the second video image, and determine the second image mask 2222 based on the positions of the foreground target image and the background image in the first video image. Here, the target detection network may use the YOLOV8 network, and be trained using a large number of images of objects of interest to users, such as people, hands, and pets in a mixed reality (MR) display scene, as foreground target image samples. The specific training process does not involve the technical improvement of this application and will not be described in detail here.
[0055] Specifically, in the process of identifying and separating the foreground target image and the remaining background image from the first video image, the object detection unit 221 may first obtain depth data corresponding to the first video image and then divide the first video image into multiple regions based on the gradient of the depth data. The object detection unit 221 may then determine the image in the third region with a smaller depth as the foreground target image and the image in the fourth region with a larger depth as the background image.
[0056] Furthermore, in some embodiments, the network processor (NPU) 22 may also be configured with an internal complement operation unit 223. Referring to FIG2 and FIG3A to FIG3D , FIG3A to FIG3D illustrate schematic diagrams of images at various stages of an internal complement algorithm according to some embodiments of the present invention.
[0057] As shown in FIG2 , during the interpolation operation, the interpolation operation unit 223 may first obtain the first video image shown in FIG3A and the second video image shown in FIG3B from the forward warping unit 212 of the digital signal processor (DSP) 21, and obtain the first image mask 2221 and the second image mask 2222 generated above. Subsequently, the interpolation operation unit 223 may determine the foreground object image 31, the mid-ground image 32, and the distant image 33 in the first video image based on the second image mask 2222, and determine the mid-ground image 32 and / or the distant image 33 as the background image of the first video image. Furthermore, the interpolation operation unit 223 may determine the positions of the foreground object image 31, the mid-ground image 32, and the distant image 33 in the second video image based on the first image mask 2221, and thereby determine at least one occluded region 34 generated after the first video image is converted into the second video image from the user's perspective. Next, the inpainting unit 223 executes a preconfigured inpainting algorithm, using a neural network or interpolation of surrounding pixels from a mask to determine the value of at least one pixel in the occluded area 34 of the background image. Based on the values of each pixel in the occluded area 34, the inpainting unit 223 fills in the pixels in the occluded area 34 to obtain the inpainted image 35 shown in FIG3C . In this way, the network processor (NPU) 22 can implement an inpainting algorithm based on depth, object detection, or a combination of both, using a flexible software architecture to generate the inpainted image 35 of the occluded area 34.
[0058] Furthermore, for foreground target images 31 that can move quickly, such as "hands" and "human bodies", the internal complement operation unit 223 may not perform internal complementation on the blocked areas 34 at their edges, so as to avoid visual effects such as abnormal values and edge distortion caused by the internal complemented images.
[0059] Furthermore, in response to the second image mask 2222 indicating that the content of the foreground target image 31 is a specific object that can move quickly, such as a "hand" or a "human body", the internal complement operation unit 223 can also preferably use an internal complement algorithm dedicated to these specific objects to further enhance the internal complement effect of the occluded area 34 at its edge.
[0060] Those skilled in the art will appreciate that the above-described internal complement algorithm is merely a non-limiting embodiment of the present invention. Its specific process does not involve technical improvements of the present invention and therefore does not require further elaboration. This embodiment is provided solely to clearly illustrate the main concepts of the present invention and to provide some specific solutions that facilitate implementation by the public, and is not intended to limit the scope of protection of the present invention.
[0061] Optionally, in other embodiments, the infill operation unit 223 may also be configured with a novel view synthesis (Novel View Synthesis) task network, and the value of at least one pixel point in the occluded area 44 may be determined through the novel view synthesis task network to thereby generate the effect of the infill image 35.
[0062] Afterwards, the internal complement operation unit 223 may input the internal complement image 35 into the ASIC 23, which performs layer blending on the internal complement image 35 and the second video image from the human eye perspective to obtain the third video image to be displayed as shown in FIG3D .
[0063] Please refer to FIG. 2 and FIG. 4 for details. FIG. 4 shows a schematic diagram of layer aliasing according to some embodiments of the present invention.
[0064] As shown in FIG2 , in some embodiments, the application-specific integrated circuit (ASIC) 23 may be configured with a second dedistortion unit 232 and / or a third dedistortion unit 233. The second dedistortion unit 232 is connected to the digital signal processor (DSP) 22 and is configured to obtain a camera distortion grid 241 indicating camera distortion, an optomechanical distortion grid 242 indicating display optomechanical distortion, and / or a perspective conversion grid 243 indicating the conversion relationship between the camera perspective and the human eye perspective, and perform a matrix-based fusion operation on these grids to determine a corresponding video distortion grid 244. It is understood that when the scene camera directly captures video images from the user's perspective, or first performs perspective conversion on the captured video images before outputting video images from the user's perspective, the perspective conversion grid 243 may not be required in the second dedistortion unit 232.
[0065] In addition, for the embodiment configured with the above-mentioned first dedistortion unit 231, since the acquired second video image has corrected the camera distortion, the third dedistortion unit 233 can only obtain the optomechanical distortion grid 242 indicating the display optomechanical distortion, and through backward warping, find the position of the equidistant grid point position in the output image corresponding to the position of the input image to determine the corresponding internal compensation distortion grid 245.
[0066] Afterwards, as shown in FIG4 , during the layer blending process, the second dedistortion unit 232 can directly obtain a fourth video image with a high resolution (e.g., >8M pixel resolution) from the scene camera 31, and correct the fourth video image according to the video distortion grid 244 to obtain a fifth video image with a high resolution (e.g., >8M pixel resolution) from the human eye perspective. Simultaneously, the third dedistortion unit 233 can obtain a first interpolated image with a low resolution (e.g., <1M pixel resolution) from the network processor (NPU) 22, and correct the first interpolated image according to the interpolated distortion grid 245 to obtain a second interpolated image with the optomechanical distortion removed.
[0067] Specifically, during correction of the fourth video image, the second dedistortion unit 232 may preferably obtain the first image mask 2221 from the network processor (NPU) 22, and perform image operations such as dilation and erosion on the first image mask 2221 to obtain a third image mask 25. The third image mask 25 is then used to determine the video layer data located in the occluded area of the fourth video image. Subsequently, the second dedistortion unit 232 may warp the video layer data based on the video distortion grid 244 to obtain a fifth video image with a high resolution (e.g., >8M pixel resolution) from the user's human eye perspective, while avoiding image problems such as holes in the occluded area.
[0068] In addition, during the process of correcting the first interpolated image, the third dedistortion unit 233 may also obtain the first image mask 2221 from the network processor (NPU) 22 to determine the interpolated layer data located in the occluded area of the first interpolated image, and then perform warping processing on the interpolated layer data according to the interpolated distortion grid 245 to pre-correct the display optical and mechanical distortion that will be generated in the first interpolated image, and interpolate and amplify the first interpolated image to obtain a second interpolated image with high resolution (e.g., >8M pixel resolution) from the perspective of the human eye.
[0069] In addition, referring to FIG4 , the ASIC 23 may also be configured with a first layer blending unit 2341. In response to obtaining a high-resolution fifth video image from the user's human eye perspective from the second dedistortion unit 232 and a high-resolution second internal complement image from the user's human eye perspective from the third dedistortion unit 233, the ASIC 23 may perform layer blending on the color-processed fifth video image and the second internal complement image via the first layer blending unit 2341 to obtain a third video image to be displayed, thereby fully utilizing the high speed and low energy consumption characteristics of the ASIC 23 to complete the blending and fusion of the low-resolution internal complement image and the high-resolution video image.
[0070] Thus, the present invention utilizes, on the one hand, a video distortion grid 244 that accounts for camera distortion, optomechanical distortion, and / or the conversion relationship between camera and eye perspective, and an interpolation distortion grid 245 that accounts only for optomechanical distortion, to differentially correct the video image and the interpolation image, thereby ensuring consistent visual effects for the video image and the interpolation image, consistent with the preceding interpolation algorithm, and thus improving displayed image quality. Furthermore, by first downsampling the high-resolution fourth video image to accommodate low-resolution depth data for target detection and interpolation, and then upsampling the low-resolution first interpolation image for layer blending with the high-resolution fifth video image, the present invention effectively reduces the data processing load for target detection and interpolation without compromising the visual quality of the image, thereby improving the real-time performance of image perspective conversion and display.
[0071] Those skilled in the art will understand that the above-mentioned embodiment of first converting the low-resolution first internal complement image into the high-resolution second internal complement image and then performing layer blending on the low-resolution fifth video image from the user's human eye perspective is only a preferred implementation method provided by the present invention, which is intended to clearly demonstrate the main concept of the present invention and provide some specific solutions that are convenient for the public to implement, and is not intended to limit the scope of protection of the present invention.
[0072] Optionally, in other embodiments, the first layer blending unit 2341 may also directly obtain the low-resolution fifth video image from the user's human eye perspective from the second de-distortion unit 232, and perform layer blending with the low-resolution first internal complement image from the user's human eye perspective obtained from the third de-distortion unit 233, so as to also obtain the third video image to be displayed and basically meet the resolution consistency requirements of layer blending.
[0073] Furthermore, for an image perspective conversion display device involving at least one other layer of virtual layer data, the application-specific integrated circuit (ASIC) 23 may also preferably be configured with at least one second layer blending unit 2342. As shown in FIG4 , the second layer blending unit 2342 may be provided at the back end of the first layer blending unit 2341 to further blend the intermediate image output by the first layer blending unit 2341 with the virtual image of the at least one virtual layer to obtain a third video image to be displayed.
[0074] Furthermore, in some embodiments, the application-specific integrated circuit (ASIC) 23 may also be preferably configured with a first color processing unit 2351 and / or a second color processing unit 2352. The first color processing unit 2351 and / or the second color processing unit 2352 are located before the first layer blending unit 2341 and are configured to perform a preset first color processing on the fifth video image from the human eye's perspective and / or perform a preset second color processing on the acquired second internal complement image to enhance the visual effect of the blending of the two.
[0075] Furthermore, in some embodiments, the application-specific integrated circuit (ASIC) 23 may further include a first format processing unit (not shown) and / or a second format processing unit (not shown). The first format processing unit and / or the second format processing unit are also located before the first layer blending unit 2341 and are configured to perform a preset first format processing on the data format of the fifth video image from the human eye's perspective and / or perform a preset second format processing on the data format of the acquired second internal complement image, thereby achieving format uniformity between the two and facilitating layer blending between the two.
[0076] In addition, in the embodiment shown in Figure 4, the above-mentioned application-specific integrated circuit (ASIC) 23 can also be configured with various other processing units 2361~2363 such as de-distortion, optical-mechanical pre-distortion, and image scaling, which are used to process the above-mentioned fifth video image, the second internal complement image, the intermediate image output by the first layer blending unit 2341 and / or the virtual image of the at least one virtual layer according to the actual display requirements of the image perspective conversion display device, so as to further improve the display effect of the final third video image.
[0077] In addition, in the embodiment shown in FIG4 , the image perspective conversion display device provided by the present invention may further include a display light engine 237 and a display screen. The display light engine 237 is located in front of the display screen and is used to increase the field of view of the third video image and convert the perspective of the third video image with the increased field of view to display on the display screen, thereby further enhancing the user's visual experience.
[0078] In summary, by selecting a digital signal processor (DSP) 21 to perform warping processing of video images and depth data, selecting a network processor (NPU) 22 to perform target detection, internal complement operations and other neural network-based reasoning operations, and selecting an application-specific integrated circuit (ASIC) 23 to perform layer blending, dedistortion, color processing, format processing, uniformity compensation and other image processing, the above-mentioned image perspective conversion display device, image perspective conversion display method, and a computer-readable storage medium provided by the present invention can all combine the data processing characteristics of each hardware module to achieve the combination of ASIC hardware and AI computing power, as well as high-precision image internal complement based on depth and target detection, thereby improving the aliasing efficiency of the internal complement image and the video image, and reducing the energy consumption of image aliasing.
[0079] Although the above methods are illustrated and described as a series of acts for simplicity of explanation, it is to be understood and appreciated that these methods are not limited by the order of the acts, as some acts may occur in a different order and / or concurrently with other acts from those illustrated and described herein or not illustrated and described herein but understandable to those skilled in the art according to one or more embodiments.
[0080] Those skilled in the art will appreciate that information, signals, and data may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips cited throughout the foregoing description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0081] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A skilled person may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as resulting in a departure from the scope of the present invention.
[0082] The various illustrative logic modules and circuits described in conjunction with the embodiments disclosed herein may be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0083] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image viewing angle conversion display device, characterized in that: include: a digital signal processor, configured to warp the first video image from the camera's perspective and its corresponding depth data to obtain a second video image from the user's human eye's perspective; the network processor being configured to perform target detection on the first video image to determine a first image mask indicating an occluded area in the second video image, and to perform an interpolation operation based on the second video image and the first image mask to obtain an interpolated image that fills in at least one pixel value in the occluded area; as well as A dedicated integrated circuit is used to perform correction and layer aliasing on the first video image and the internal complement image to obtain a third video image to be displayed.
2. The image viewing angle conversion display device according to claim 1, wherein: The image perspective conversion display device further includes a scene camera and a depth sensor, and the digital signal processor is configured to: acquiring the first video image via the scene camera; acquiring, via the depth sensor, depth data corresponding to the first video image; as well as Warping the first video image and the depth data to generate the second video image.
3. The image viewing angle conversion display device according to claim 2, wherein: The digital signal processor is provided with a depth post-processing unit, and the step of acquiring depth data corresponding to the first video image via the depth sensor includes: Acquiring first depth data from a sensor perspective via the depth sensor; and The first depth data is warped by the depth post-processing unit to obtain second depth data of the camera perspective.
4. The image viewing angle conversion display device according to claim 2, wherein: The ASIC includes a first de-distortion unit, and the first de-distortion unit is configured to: acquiring a fourth video image with high resolution via the scene camera; Obtaining a camera distortion grid indicating camera distortion of the scene camera, and / or a preset target resolution; as well as The fourth video image is processed according to the camera distortion grid and / or the target resolution to obtain a low-resolution first video image with the camera distortion removed.
5. The image viewing angle conversion display device according to claim 4, wherein: The first dedistortion unit is connected to the network processor, and the step of processing the fourth video image according to the camera distortion grid and / or the target resolution to obtain the low-resolution first video image with the camera distortion removed comprises: acquiring, via the network processor, a second image mask to determine, in the fourth video image, a first region corresponding to the blocked region and a remaining second region; assigning a default value to at least one first pixel in the first area; performing dedistortion processing on at least one second pixel in the second area according to the camera distortion grid; and The first region and the second region are compressed according to the target resolution to obtain the first video image.
6. The image viewing angle conversion display device according to claim 2, wherein: The step of performing warping on the first video image and the depth data to generate the second video image includes: Determining first 3D video data from the camera perspective based on the first video image and its corresponding depth data; Acquire a perspective conversion grid indicating a conversion relationship between the camera perspective and the human eye perspective; determining second 3D video data under the human eye perspective according to the first 3D video data and the perspective conversion grid; and The second video image is determined according to the second 3D video data.
7. The image viewing angle conversion display device according to claim 1 or 5, wherein: The network processor is provided with a target detection unit, and the target detection unit is configured to: Acquire the first video image; Identifying and separating a foreground target image and a remaining background image from the first video image using a pre-trained target detection network, wherein the target detection network is trained based on sample images of MR scenes; determining a second image mask according to positions of the foreground target image and the background image in the first video image; and The first image mask is determined according to positions of the foreground target image and the background image in the second video image.
8. The image viewing angle conversion display device according to claim 7, wherein: The step of identifying and separating the foreground target image and the remaining background image from the first video image via the pre-trained target detection network includes: Acquiring depth data corresponding to the first video image; Dividing the first video image into a plurality of regions according to a gradient of the depth data; and The image in the third region with the smaller depth is determined as the foreground object image, and the image in the fourth region with the larger depth is determined as the background image.
9. The image viewing angle conversion display device according to claim 7, wherein: The network processor is further configured with an internal complement operation unit, and the internal complement operation unit is configured to: Acquire the second video image, the first image mask, and the second image mask; Inputting the second video image, the first image mask, and the second image mask into a preconfigured interpolation algorithm to determine a value of at least one pixel of the blocked area in the background image; as well as Fill the pixel points of the occluded area according to the values of the pixel points of the occluded area to obtain the internal complement image.
10. The image viewing angle conversion display device according to claim 1 or 4, wherein: The ASIC is provided with a second de-distortion unit and a third de-distortion unit, wherein: The second dedistortion unit is connected to the digital signal processor and is configured to: obtain a camera distortion grid indicating camera distortion, an optomechanical distortion grid indicating optomechanical distortion, and / or a perspective conversion grid indicating a conversion relationship between the camera perspective and the human eye perspective to determine a corresponding video distortion grid; and perform correction on the first video image according to the video distortion grid to obtain a fifth video image from the human eye perspective. The third dedistortion unit is connected to the network processor and is configured to: obtain the optomechanical distortion grid to determine a corresponding interpolation distortion grid; and correct the first interpolation image obtained by the network processor according to the interpolation distortion grid to obtain a second interpolation image with the optomechanical distortion removed. The dedicated integrated circuit performs layer blending on the fifth video image from the human eye's perspective and the second internal complement image to obtain the third video image.
11. The image viewing angle conversion display device according to claim 10, wherein: The second dedistortion unit is connected to the scene camera and the network processor, and the step of correcting the first video image according to the video distortion grid to obtain the fifth video image from the human eye perspective includes: Obtaining the first image mask via the network processor to determine video layer data located in the obscured area in a high-resolution fourth video image acquired via the scene camera; and The video layer data is warped according to the video distortion grid to obtain a high-resolution fifth video image from the human eye perspective.
12. The image viewing angle conversion display device according to claim 11, wherein: The step of correcting the first interpolated image obtained by the network processor according to the interpolated distortion grid to obtain a second interpolated image without the optomechanical distortion comprises: determining, according to the first image mask, interpolating layer data located in the blocked area in the first interpolated image; and The interpolated layer data is warped according to the interpolated distortion grid to obtain a high-resolution second interpolated image at the human eye's viewing angle.
13. The image viewing angle conversion display device according to claim 10, wherein: The ASIC is provided with a plurality of layer aliasing units, wherein: The first layer aliasing unit is used to perform layer aliasing on the fifth video image and the second internal complement image to obtain an intermediate image. The second layer blending unit is configured to perform layer blending on the intermediate image and at least one virtual image to obtain the third video image.
14. The image viewing angle conversion display device according to claim 13, wherein: The ASIC is further configured with a first color processing unit and / or a second color processing unit, wherein: The first color processing unit is located before the first layer aliasing unit and is used to perform a preset first color processing on the video image from the human eye perspective. The second color processing unit is located before the first layer blending unit and is used to perform a preset second color processing on the internal complement image.
15. The image viewing angle conversion display device according to claim 13, wherein: The ASIC is further configured with a first format processing unit and / or a second format processing unit, wherein: The first format processing unit is located before the first layer aliasing unit and is used to perform a preset first format processing on the data format of the video image from the human eye perspective. The second format processing unit is located before the first layer aliasing unit, and is used to perform a preset second format processing on the data format of the inner complement image.
16. The image viewing angle conversion display device according to claim 1, wherein: Also includes: Display screen; as well as A display light engine, wherein the display light engine is located in front of the display screen, and is used to increase the field of view of the third video image and display the third video image with the increased field of view on the display screen.
17. A method for displaying an image by converting an image viewing angle, characterized in that: The following steps are involved: Obtaining a first video image from a camera perspective and its corresponding depth data; warping the first video image and the depth data via a digital signal processor to obtain a second video image from a user's eye perspective; performing object detection on the first video image via a network processor to determine an image mask indicating an occluded area in the second video image, and performing an interpolation operation based on the second video image and the image mask to obtain an interpolated image that fills in at least one pixel value in the occluded area; as well as The video image from the human eye's perspective and the internal complement image are layer-mixed via a dedicated integrated circuit to obtain a third video image to be displayed.
18. The image viewing angle conversion and display method according to claim 17, wherein: The step of acquiring the first video image includes: acquiring a high-resolution fourth video image via a scene camera; Obtaining a camera distortion grid indicating camera distortion of the scene camera, and / or a preset target resolution; and The fourth video image is processed according to the camera distortion grid and / or the target resolution to obtain a low-resolution first video image with the camera distortion removed.
19. The image viewing angle conversion and display method according to claim 18, wherein: The step of processing the fourth video image according to the camera distortion grid and / or the target resolution to obtain the low-resolution first video image with the camera distortion removed comprises: Obtaining an optomechanical distortion grid indicating optomechanical distortion and / or a perspective conversion grid indicating a conversion relationship between the camera perspective and the human eye perspective; Determining a corresponding video distortion grid according to the camera distortion grid, the optomechanical distortion grid, and / or the perspective conversion grid; and The fourth video image is corrected according to the video distortion grid to obtain a fifth video image with a high resolution from the human eye perspective.
20. The image viewing angle conversion and display method according to claim 17, wherein: The step of obtaining the depth data includes: Acquiring first depth data from a sensor perspective via a depth sensor; and The first depth data is warped to obtain second depth data of the camera perspective.
21. The image viewing angle conversion and display method according to claim 17, wherein: The step of performing target detection on the first video image to determine an image mask indicating an obscured area in the second video image comprises: Identifying and separating a foreground target image and a remaining background image from the first video image via a pre-trained target detection network, wherein the target detection network is trained based on sample images of MR scenes; and The image mask is determined according to positions of the foreground target image and the background image in the second video image.
22. The image viewing angle conversion and display method according to claim 17, wherein: The step of performing an interpolation operation based on the second video image and the image mask to obtain an interpolated image that fills at least one pixel value in the blocked area includes: Inputting the second video image and the image mask into a pre-configured interpolation algorithm to determine a value of at least one pixel of the blocked area in the background image; and Fill the pixel points of the occluded area according to the values of the pixel points of the occluded area to obtain the internal complement image.
23. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the image viewing angle conversion and display method according to any one of claims 17 to 22 is implemented.
Citation Information
Patent Citations
Three-dimensional target detection method and system based on image restoration
CN111079545A
Image synthesis method and device, electronic equipment and storage medium
CN113870165A
Planar video light field conversion method and system based on time sequence multi-depth plane field
CN116778072A
Display system with machine learning (ML) based stereoscopic view synthesis over a wide field of view
WO2023146882A1
Image processing method and apparatus, nonvolatile readable storage medium and electronic device
WO2023217046A1
Cited By
YOLO-based sorting video identification processing method and system
CN120913131A