Image generation method and apparatus therefor
By acquiring view images with a monocular camera and performing parallax image transformation, the problems of high cost and poor generalization of binocular image acquisition are solved, and efficient image data generation and improved calculation accuracy are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2022-04-19
- Publication Date
- 2026-04-21
AI Technical Summary
Existing binocular image acquisition schemes are costly and have poor generalization ability, resulting in low computational accuracy.
The first view image is acquired by the target camera, the parallax image is determined, and an affine transformation is performed based on the parallax image to generate the second view image. The acquisition of binocular images is simulated by using a monocular camera.
It reduces the cost of binocular image acquisition, improves the generalization and computational accuracy of image data, and avoids complex binocular data acquisition processes and expensive equipment requirements.
Smart Images

Figure CN114881841B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of information technology, and specifically relates to an image generation method and apparatus. Background Technology
[0002] With the development of binocular imaging technology, binocular images have received widespread attention in recent years.
[0003] Currently, the mainstream binocular image acquisition schemes mainly include the following two types: (1) using 3D rendering software to generate fake image data. Due to the large difference between this type of image data and the actual scene, the generalization of the image data is insufficient, and the model trained using the fake image data generally performs poorly on the real dataset. Moreover, this method has a high software usage threshold and time cost, requiring a lot of time to learn how to use 3D rendering software and create image data for various scenes; (2) using RGB cameras and depth cameras to collaboratively acquire binocular images. The annotation accuracy of this type of image data is poor, and the image data acquisition is relatively expensive and laborious. The acquisition scene is also relatively limited, and it is usually only suitable for indoor scenes.
[0004] Therefore, the two existing methods for acquiring binocular images are costly and difficult, and the acquired binocular images have poor generalization ability, resulting in low computational accuracy. How to acquire binocular images has become an urgent problem to be solved. Summary of the Invention
[0005] The purpose of this application is to provide a method for generating stereo data, which can solve the problems of high cost and difficulty in data collection, poor generalization of the collected data, and thus low computational accuracy.
[0006] In a first aspect, embodiments of this application provide an image generation method, which includes: acquiring a first view image through a target camera; determining a first parallax image corresponding to the first view image; performing an affine transformation on the first view image based on the first parallax image to obtain a second view image; wherein the first view image and the second view image are view images from different perspectives of the same shooting scene.
[0007] Secondly, embodiments of this application provide an image generation apparatus, which includes: a shooting module and an execution module, wherein: the shooting module is used to capture a first view image through a target camera; the execution module is used to determine a first parallax image corresponding to the first view image captured by the shooting module; the execution module is further used to perform an affine transformation on the first view image based on the first parallax image to obtain a second view image; the first view image and the second view image are view images from different perspectives of the same shooting scene.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implementing the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0012] In this embodiment, the image generation device acquires a first view image using a target camera, determines a first disparity image corresponding to the first view image, and then performs an affine transformation on the first view image based on the first disparity image to obtain a second view image. The first view image and the second view image are view images from different perspectives of the same shooting scene. Through this method, the image generation device can acquire a first view image using a target camera and perform an affine transformation on the first view image based on its disparity map to obtain a second view image that matches the first view image. Thus, view images can be acquired using only one camera, eliminating the need for complex binocular data acquisition processes and expensive binocular data acquisition equipment, and improving the generalization of the generated second view image. Attached Figure Description
[0013] Figure 1 This is a schematic flowchart of the image generation method provided in the embodiments of this application;
[0014] Figure 2 This is a schematic diagram of image affine transformation processing of a first view image provided in an embodiment of this application;
[0015] Figure 3 This is a schematic diagram of parallax prediction of a fourth view image provided in an embodiment of this application;
[0016] Figure 4 This is a schematic diagram of the structure of an image generation device provided in an embodiment of this application;
[0017] Figure 5 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;
[0018] Figure 6 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0020] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0021] The image generation method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0022] This application provides a shooting method. Figure 1 A flowchart illustrating the image generation method provided in an embodiment of this application is shown. Figure 1 As shown, the image generation method provided in this application embodiment may include the following steps 201 to 203:
[0023] Step 201: The image generation device acquires a first-view image through the target camera.
[0024] In this embodiment of the application, the target camera can be a monocular camera.
[0025] Optionally, in this embodiment, the monocular camera is a camera on an electronic device. For example, the monocular camera is a single camera on an electronic device. This could be a front / rear camera of an electronic device, a camera lens, etc.
[0026] In the embodiments of this application, the first view image can be a left-eye image, i.e., a left view, or the first view image can be a right-eye image, i.e., a right view.
[0027] Optionally, in this embodiment of the application, the first view image described above is an image that includes multiple colors.
[0028] Optionally, in embodiments of this application, the first view image may include multiple view images. For example, the first view image may include 100 left views, or the first view image may include 100 right views.
[0029] Optionally, in this embodiment, the first view image can be multiple images captured by a monocular camera in multiple scenes. Optionally, the scenes can be landscapes, buildings, portraits, and still life, etc. For example, the scenes can be the sea, forest, houses, indoor / outdoor portraits, and food, etc. The above are just some common shooting scenarios; the scenarios can be selected according to actual needs, and this embodiment does not impose any limitations on them.
[0030] For example, the number of the first view images described above can be greater than or equal to 100 and less than or equal to 1000.
[0031] For example, the image generation device can use 100 left views captured by a monocular camera, and these 1000 left views include different types of images captured in different shooting scenarios, such as landscape images, portraits, and building images.
[0032] For example, the image generation device can capture 550 left views using a monocular camera, and these 1000 left views include different types of images captured in different shooting scenarios, such as landscape images, portraits, and building images.
[0033] For example, an image generation device can capture 1,000 view images using a monocular camera, and these 1,000 images include different types of images captured in different shooting scenarios, such as landscape images, portraits, and building images.
[0034] In this way, the image generation device can acquire monocular images containing more scenes through a monocular camera, which is low-cost and effectively improves the generalization ability of monocular image data.
[0035] Step 202: The image generating device determines the first parallax image corresponding to the first view image.
[0036] In this embodiment of the application, the first disparity image includes the disparity value corresponding to each pixel in the first view image.
[0037] Optionally, in this embodiment, the disparity value includes the offset of each pixel in the first view image. For example, the offset can be the offset of the pixel coordinates of each pixel. Further, the offset can be the horizontal offset of the pixel coordinates of each pixel.
[0038] Optionally, the disparity value corresponding to each pixel in the first view image can be the same or different. For example, the disparity value of each pixel in the first view image is 3px; as another example, the disparity value of pixel 1 in the first view image is 3px, and the disparity value of pixel 2 in the first view image is 4px.
[0039] Optionally, in this embodiment of the application, the image generation device can obtain a first disparity image corresponding to the first view image based on a deep learning network (such as a neural network). For example, the deep learning network can be a binocular stereo matching network. The image generation device can input the first view image into the binocular stereo matching network and use the disparity map output by the binocular stereo matching network as the first disparity image corresponding to the first view image.
[0040] Optionally, the image generation device may use the first view image as an input sample set to perform first disparity image prediction using a deep learning network.
[0041] For example, assuming the input sample set includes 1000 image samples, these samples can be divided into 10 batches for prediction, with each batch containing 100 image samples.
[0042] Optionally, in this embodiment of the application, the image generating apparatus may obtain a first disparity image corresponding to the first view image based on the depth information of the first view image. For example, the image generating apparatus may calculate the disparity information of the first view image based on the depth information of the first view image, and obtain the first disparity image corresponding to the first view image based on the disparity information.
[0043] Step 203: The image generation device performs an affine transformation on the first view image based on the first parallax image to obtain the second view image.
[0044] The first view image and the second view image mentioned above are view images from different perspectives of the same shooting scene.
[0045] Optionally, in this embodiment of the application, the image generation device may perform an affine transformation on the first view image based on the disparity value corresponding to each pixel in the first view image to obtain a second view image corresponding to the first view image.
[0046] Optionally, in this embodiment, the image generation apparatus may perform affine transformation processing on a first image region of the first view image based on a first parallax image. For example, the first image region may be the foreground or background of the first view image, or it may be the image region corresponding to a first object in the first view image. For instance, the first image region may be the image region where a person is located in the first view image.
[0047] For example, the first view image is taken as the left-eye image (i.e., the left view). Figure 2 This is a schematic diagram of performing an affine transformation on the first view image. After obtaining the disparity image corresponding to the left view image 21, as shown... Figure 2 As shown, the image generation device performs an affine transformation on the left eye image 21 based on the disparity value of each pixel in the left eye image 21 contained in the disparity image, to obtain the corresponding right eye image 22. Figure 2 As can be seen, the obtained right eye image 22 and left eye image 21 are in the same coordinate system, and each pixel in the right eye image 22 is shifted to the left.
[0048] The left eye image 21 is a color image, and the size of the left eye image 21 is the same as that of its corresponding disparity value image.
[0049] It should be noted that image affine transformation maps each pixel in an image to a new position according to certain rules. Essentially, it is the process of solving for the new horizontal and vertical coordinates of the pixels in the image, that is, mapping M×N pixels of the original image to M×N new positions in the target image.
[0050] It should be noted that image affine transformation (warp) is also called image distortion.
[0051] In the image generation method provided in this application embodiment, the image generation device acquires a first view image through a target camera, determines a first disparity image corresponding to the first view image, and then performs an affine transformation on the first view image based on the first disparity image to obtain a second view image; wherein the first view image and the second view image are view images from different perspectives of the same shooting scene. Through this method, the image generation device can acquire a first view image through a target camera and perform an affine transformation on the first view image according to its disparity map to obtain a second view image that matches the first view image. Thus, view images can be acquired with a single camera, eliminating the need for complex binocular data acquisition processes and expensive binocular data acquisition equipment, and improving the generalization of the generated second view image.
[0052] Optionally, in this embodiment of the application, step 203 may include steps 203a1 and 203a2:
[0053] Step 203a1: The image generation device performs an affine transformation on the first position information of each pixel in the first view image based on the disparity value corresponding to each pixel in the first view image to obtain the second position information.
[0054] Step 203a2: The image generating device generates a second view image based on the second position information and pixel information of each pixel.
[0055] Optionally, the first position information of the aforementioned pixel can be represented by the pixel's coordinates. For example, the pixel's coordinates can be two-dimensional coordinates (x, y). For instance, the position of pixel i can be represented by coordinates (xi, yi).
[0056] Optionally, the second position information can be the position information of each pixel in the first view image within the second view image. Alternatively, the second position information of the pixel can be represented by the coordinates of the pixel. For example, if the position of pixel i in the first view image can be represented by coordinates (xi, yi), then the position of that pixel in the second view image can be represented by coordinates (xi', yi').
[0057] Optionally, the pixel information for each pixel can be the pixel value for each pixel. For example, the pixel value can be a grayscale value.
[0058] For example, taking the first view image as the left-eye image and the second view image as the right-eye image. Assuming that the coordinate position of pixel i in the left-eye image is (xi, yi) and the disparity value corresponding to pixel i is dsi, then the coordinate position of pixel i in the right-eye image is (xi-dsi, yi).
[0059] For example, taking the first view image as the right-eye image and the second view image as the left-eye image. Assuming that the coordinate position of pixel j in the right-eye image is (xj, yj) and the disparity value corresponding to pixel j is dsj, then the coordinate position of pixel j in the right-eye image is (xj+dsj, yj).
[0060] It should be noted that when using the first view image as the left-eye image to determine the right-eye image, each pixel i in the left-eye image needs to be shifted to the left by Di pixels, i.e., subtract Di from the x-coordinate of each pixel i. Conversely, when using the first view image as the right-eye image to determine the left-eye image, each pixel j in the right-eye image needs to be shifted to the right by Dj pixels, i.e., add Dj to the x-coordinate of each pixel j. Here, Di represents the disparity value corresponding to each pixel i in the left-eye image, and Dj represents the disparity value corresponding to each pixel j in the left-eye image.
[0061] Optionally, the image generating apparatus may determine the coordinates of each pixel in the first view image in the second view image, and assign a pixel value to each pixel at each coordinate to generate the second view image.
[0062] Thus, the image generation device can perform an affine transformation on the coordinates of each pixel in the first view image based on the disparity value corresponding to each pixel in the first view image, so as to obtain the coordinates of each pixel in the second view image, thereby obtaining a second view image with a high degree of matching with the first view image.
[0063] Optionally, in this embodiment of the application, step 202 may include steps 202a and 202b:
[0064] Step 202a: The image generation device generates a second parallax image corresponding to the first view image based on the depth information of the first view image.
[0065] Step 202b: The image generating device deletes pixels in invalid value regions of the second parallax image to obtain the first parallax image.
[0066] Optionally, the image generation device can generate a second parallax image corresponding to the first view image based on the depth information of the first view image and the baseline and focal length information of the binocular imaging system.
[0067] For example, the image generation apparatus can use a monocular depth network model to predict the depth of a first view image to obtain the depth information of the first view image. Optionally, the monocular depth network model can be a MiDas network or other networks, and the embodiments of this application do not limit each other.
[0068] For example, the image generation apparatus can acquire disparity information of a first view image and generate a second disparity image based on the disparity information. For example, the disparity information may include disparity values. For instance, the image generation apparatus uses each disparity value of the first view image as pixel information of the second disparity image to generate the second disparity image.
[0069] For example, taking image A as an example, the image generation device inputs image A into the monocular depth model m, performs prediction on image A, and obtains the monocular depth estimation result z of image A, i.e., z = m(a).
[0070] For example, the image generation apparatus may determine the parallax information of the first view image based on the depth information of the first view image.
[0071] For example, the image generation device acquires the baseline b and focal length information f of the binocular imaging system, and calculates the disparity information d of the image A using the disparity formula based on the baseline b, focal length information f and the predicted depth estimation result z.
[0072] It should be noted that the above parallax formula is formula (1) in the following text. For the specific calculation method of obtaining the parallax information d based on the baseline b, focal length information f and the predicted depth estimation result z using this parallax formula, please refer to the following text, which will not be repeated here.
[0073] Optionally, the invalid value region in the second disparity image includes at least one of the following:
[0074] Image regions including scattered noise;
[0075] Image regions including invalid pixel values.
[0076] It should be noted that the geometric constraints of the depth estimation result z obtained by using a monocular depth network are usually weak. Therefore, if the disparity map obtained by using the depth estimation result z is directly used for image affine transformation processing, i.e. warping, the warped result will have some isolated noise points, resulting in poor quality of the generated second view image.
[0077] Optionally, the image generation device can use the Sobel operator to extract regions in the second disparity image with gradients greater than g (for example, the value of g can be 3) to obtain an edge map e, and then delete the edge map e from the second disparity image, thereby deleting image regions in the second disparity image that include scattered noise.
[0078] Optionally, after deleting pixels in invalid value regions of the second parallax image, the image generating apparatus may perform image filling on the second parallax image with the invalid value regions removed.
[0079] For example, the image generation apparatus can perform interpolation processing on a second disparity image with invalid value regions removed to fill the hole regions formed by the removal of invalid value regions. For instance, after removing image regions e that include scattered noise from the disparity map d, a disparity map d' with hole regions is obtained. Then, multivariate interpolation is used to interpolate the hole regions to obtain a filled disparity map ds.
[0080] In this way, the image generation device can remove image regions containing scattered noise from the second parallax image, thereby effectively improving the quality of the generated second view image.
[0081] Optionally, the image generating apparatus may determine the image region where invalid pixels are located in the second parallax image and delete the image region.
[0082] For example, the aforementioned invalid pixels include at least one of the following:
[0083] Pixels whose pixel values are greater than the x-coordinate of their corresponding pixel in the first view image are those that are not within the range represented by the second view image, calculated based on parallax.
[0084] Multiple pixels at the same pixel location in the second view image, that is, multiple pixels in the first view image, belong to the same location in the right view image based on parallax calculation.
[0085] It should be noted that the pixel value of each pixel in the second parallax image is essentially the parallax value of each pixel in the first view image, and each pixel in the second parallax image corresponds one-to-one with each pixel in the first view image.
[0086] It's important to note that, typically, some pixels near the left edge of the left-view image will not be reflected in the right-view image. In other words, some image content in the left-view image is invisible in the right-view image, and similarly, some pixels near the right edge of the right-view image will not be reflected in the left-view image. Therefore, when generating a right-view image from a left-view image, pixels near the left edge of the left-view image can be considered invalid pixels, and vice versa.
[0087] For example, the image generation device can delete pixels in the second parallax image whose pixel values are greater than the horizontal coordinate of their corresponding pixel in the first view image, so that the generated second view image does not include image content near the left or right edge of the first view image, thereby obtaining a first view image and a second view image that strictly match the binocular image acquired by the binocular device.
[0088] For example, taking the first view image as the left-eye image. Assume that the position of pixel i in the left-eye image is (xi, yi), and the disparity of this pixel is dsi. Then the coordinates of this pixel in the right-eye image are (xi-dsi, yi). We can count whether (xi-dsi, yi) is within the size of the right-eye image (i.e., whether xi-dsi is less than 0) and whether there are multiple pixels in the right-eye image that correspond to it. If so, we delete pixel i and obtain the disparity map dt with invalid values.
[0089] Thus, by deleting invalid values in the second parallax image, the generated second view image does not include image content near the left or right edge of the first view image, thereby obtaining a first view image and a second view image that strictly match the binocular image acquired by the binocular device, improving the quality of the generated second view image.
[0090] Further, optionally, step 202a above may include steps A1 and A2:
[0091] Step A1: Calculate the disparity value corresponding to each pixel in the first view image based on the depth information and the first parameter of the first view image.
[0092] Step A2: Obtain the second disparity image based on the disparity value corresponding to each pixel.
[0093] The first parameter mentioned above includes: first distance and focal length information.
[0094] Optionally, the aforementioned first distance is used to characterize the distance between different viewpoints of the same shooting scene.
[0095] For example, the aforementioned first distance can be the baseline of a binocular imaging system.
[0096] It should be noted that a binocular imaging system is typically a system that uses binocular imaging devices (such as binocular cameras) to perform three-dimensional stereo imaging. The baseline of a binocular imaging system is the distance between the two cameras of the binocular imaging device, and the baseline may be different for different binocular imaging devices.
[0097] For example, in a binocular camera system, both cameras can be considered as two horizontally placed pinhole cameras, with the aperture centers of both cameras located on the x-axis. The distance between the two cameras is called the binocular camera's baseline (denoted as b). The longer the baseline, the greater the maximum distance that the binocular camera can measure; conversely, the smaller the baseline, the smaller the maximum distance that the binocular camera can measure.
[0098] Optionally, the depth information of the first view image is: the distance (depth) value from the monocular camera that acquires the first view image to each point in the scene.
[0099] Optionally, the image generation apparatus can use a monocular depth network model to predict the depth of the first view image to obtain the depth information of the first view image. Optionally, the monocular depth network model can be a MiDas network or other networks, and the embodiments of this application do not limit each other.
[0100] For example, taking a first view image that includes 100 left views as an example. The 100 left views are used as a batch of monocular datasets a and input into the monocular depth network model m for depth prediction. The monocular depth estimation result z of the 100 left views is output. The relationship between the monocular dataset a and the monocular depth estimation result z is z = m(a).
[0101] For example, the image generating apparatus can generate a first depth image corresponding to the first view image based on the depth information of the first view image. For example, the pixel value of each pixel in the first depth image is the depth value of the first view image.
[0102] Example 1: When predicting image depth information for the first view image, the monocular image is input into the monocular depth network for depth prediction, and the depth image corresponding to the monocular image is output.
[0103] It should be noted that the monocular image input to the monocular depth network is a color image, and the output depth image is the depth map of that monocular image.
[0104] Optionally, the image generating apparatus may use the disparity value corresponding to each pixel in the first view image as a pixel value to generate a second disparity image corresponding to the first view image.
[0105] For example, the image generation device can calculate the disparity value d based on the baseline b and focal length information f of the binocular imaging system and the first depth map z corresponding to the first view image, as shown in formula (1).
[0106]
[0107] Example 2, in conjunction with Example 1 above, when performing parallax transformation on the first view image, the image generation device performs parallax transformation on the depth image corresponding to the monocular image to obtain the parallax image corresponding to the monocular image.
[0108] It should be noted that for details on parallax transformation processing of depth images, please refer to relevant technologies, which will not be elaborated here.
[0109] In this way, the baseline and focal length can be flexibly selected according to actual needs, and various parallax values of various baselines and focal lengths can be easily generated, and second view images of various baselines and focal lengths can be obtained in the future.
[0110] Further, optionally, in this embodiment of the application, step 203 may include steps 203b1 and 203b2:
[0111] Step 203b1: The image generation device performs an affine transformation on the first view image based on the first parallax image to obtain the third view image.
[0112] Step 203b2: The image generating device fills the target image region of the third view image to obtain the second view image.
[0113] The target image region mentioned above is the image region in the third view image that corresponds to the invalid value region in the second parallax image mentioned above.
[0114] It is understandable that, since some invalid pixels were deleted from the first disparity image, meaning that some pixels in the first view image do not have corresponding disparity values, the third view image obtained by warping the first view image based on the first disparity image contains some missing pixels, that is, some pixels are invalid.
[0115] Optionally, the image generation apparatus can fill the target image region of the third-view image based on the acquired background image. For example, the image generation apparatus can acquire background images in multiple scenes to obtain a background image containing multiple scenes.
[0116] For example, taking the first view image as the left view and the third view image as the right view, the process begins by first collecting multiple background images as a batch of background datasets. Then, the disparity map dt, after deleting pixels from invalid value regions, is used to warp the left view a, resulting in the right view b. At this point, the right view b contains some invalid value regions. Finally, random sampling is performed on the background dataset, and the sampled pixel values are used as the pixel values for the invalid value regions, thus filling in the invalid value regions in the right view b, resulting in the filled right view b.
[0117] Optionally, the image generating apparatus can fill the target image region in the third view image with pixels from the second image region in the third view image. For example, the second image region can be the image region surrounding the target image region.
[0118] For example, taking the first view image as the left view and the third view image as the right view, the left view a is first warped using the disparity map dt after deleting pixels in the invalid value region, resulting in the right view b. At this point, the right view b contains some invalid value regions. Then, using template matching, the N closest matching points to the neighborhood of the pixel to be filled are searched. Finally, the N matching points are sorted by similarity, and the optimal matching point is selected to fill the invalid value region, resulting in the complete right eye image.
[0119] In this way, the image generation device can fill the missing areas of the third-view image with textures from images randomly selected in the training set, or fill the missing areas with textures around the missing areas in the third-view image, thereby improving the realism and accuracy of the final generated second-view image.
[0120] Optionally, in this embodiment of the application, the image generation method provided in this embodiment of the application further includes the following steps 204 to 206:
[0121] Step 204: The image generation device uses the first view image and the second view image to train the binocular stereo matching network to obtain the trained binocular stereo matching network.
[0122] Step 205: The image generation device inputs the fourth view image into the trained binocular stereo matching network and outputs the target disparity image.
[0123] Step 206: The image generation device performs an affine transformation on the fourth view image based on the above target parallax image to obtain the target view image.
[0124] The aforementioned fourth view image and the aforementioned target view image are view images from different perspectives of the same shooting scene.
[0125] Optionally, the aforementioned fourth view image can be a view image captured by a monocular camera.
[0126] For example, the image generation device can use the first view image as the left eye image, the second view image as the right eye image, and the first disparity image as the output ground truth to train the binocular stereo matching network and obtain the trained binocular stereo matching network p.
[0127] For example, the stereo matching network described above can be a RAFT-Stereo network.
[0128] It should be noted that currently, more and more camera algorithms are beginning to use AI intelligent algorithms to replace traditional algorithms. As an important branch of computer vision tasks, binocular stereo matching algorithm has a very wide range of application prospects in fields such as robotics and autonomous driving.
[0129] In related technologies, a large amount of training data is required when calculating the disparity of corresponding pixels in binocular images using AI-based stereo matching algorithms.
[0130] In this embodiment, the image generation device can generate a second view image based on a first view image, and use the first view image and the second view image as a pair of stereo training images as input to train the stereo matching network. Thus, when stereo training data is needed to train the model, only monocular data needs to be collected to automatically generate stereo training data that matches the monocular data, eliminating the need for complex stereo data acquisition processes and expensive stereo data acquisition equipment. This reduces the data acquisition cost for model training and effectively improves the model's accuracy, thereby solving the problems of small training dataset size and poor generalization, and ultimately improving the training effect and robustness of the stereo matching network.
[0131] Optionally, after training the binocular stereo matching network using the first view image and the second view image generated from the first view image to obtain the trained binocular stereo matching network, the trained binocular stereo matching network can be used to estimate the disparity of the input fourth view image to obtain the target disparity image corresponding to the fourth view image.
[0132] Optionally, the fourth view image may be the same as, at least partially the same as, or both different from the first view image.
[0133] In one example, the image generation device can use a trained binocular stereo matching network to re-predict the first view image and obtain the predicted disparity value dp, thereby further optimizing the first disparity image obtained above.
[0134] Example 3, Figure 3 This is a schematic diagram of disparity prediction for a fourth-view image. For example... Figure 3 As shown, the left-eye image 31 (i.e., the left view) is input into the trained binocular stereo matching network 32 for disparity estimation, and the output disparity image 33 corresponding to the monocular image is the output. Since the trained binocular stereo matching network is obtained by optimizing the first view image and the second view image generated based on the first view image, that is, the prediction accuracy of the optimized binocular stereo matching network is higher. Thus, by using the optimized binocular stereo matching network to predict the fourth view image, a more accurate disparity image corresponding to the fourth view image can be obtained.
[0135] It should be noted that the depth information of the scene corresponding to an image is generally described by a grayscale image of the same size. The grayscale value of each pixel in the grayscale image describes the depth value of the scene corresponding to that point. This grayscale image is also called a depth image.
[0136] It should be noted that in practical applications, image datasets formed by multiple images are usually processed. For example, multiple images are input into the network as a batch of image datasets for prediction to obtain the disparity image corresponding to each image. The embodiments of this application illustrate the processing flow of one image among multiple images, and the same processing flow applies to multiple images.
[0137] It should be noted that the stereo matching network trained on the first and second view images has been trained on a large dataset, so the prediction results are better than the disparity value dt obtained based on depth information. Therefore, reusing the trained stereo matching network to predict the first view image will result in a more accurate disparity estimate dp.
[0138] In another example, the image generation device can use a trained binocular stereo matching network to continue predicting new fourth view images, thereby improving the accuracy of the parallax images corresponding to the subsequently generated view images.
[0139] Alternatively, after obtaining the trained binocular stereo matching network, the image generation device can use the binocular stereo matching network to re-predict the disparity value corresponding to the first view image, and determine the depth information of the first view image based on the disparity value and the selected baseline b and focal length f, thereby obtaining a depth map corresponding to the first view image, thus obtaining more accurate depth information.
[0140] Optionally, the image generation device can train the monocular depth network model using the first view image and the corresponding depth map, update the weights of the monocular depth network model, and obtain the trained monocular depth network model. This can optimize the performance of the monocular depth network model, effectively improve its prediction accuracy, and thus enhance its robustness.
[0141] Optionally, after obtaining the optimized monocular depth network model, the depth information of the first view image can be re-predicted using the optimized monocular depth network model. Further, the image generation device performs depth transformation on the re-predicted depth information of the first view image to obtain a disparity image corresponding to the first view image, thereby obtaining a more accurate disparity image, and subsequently a more accurate second view image.
[0142] For example, a monocular image is input into an optimized monocular depth network model for depth prediction, and the output is a depth image corresponding to the monocular image. In this way, the optimized monocular depth network model is used to re-predict the depth information of the first view image, thereby improving the accuracy of the obtained depth information.
[0143] The image generation method provided in this application can be executed by an image generation device. This application uses an image generation device executing the image generation method as an example to illustrate the image generation device provided in this application.
[0144] This application provides an image generation apparatus 400, such as... Figure 4 As shown, the image generation device includes: a shooting module 401 and an execution module 402, wherein: the shooting module 401 is used to capture a first view image through a target camera; the execution module 402 is used to determine a first parallax image corresponding to the first view image captured by the shooting module; the execution module 402 is also used to perform an affine transformation on the first view image based on the first parallax image to obtain a second view image; the first view image and the second view image are view images from different perspectives of the same shooting scene.
[0145] Optionally, in this embodiment of the application, the first parallax image includes: the parallax value corresponding to each pixel in the first view image;
[0146] Optionally, in this embodiment of the application, the first parallax image includes: the parallax value corresponding to each pixel in the first view image;
[0147] The aforementioned execution module 402 is specifically used to perform affine transformation on the first position information of each pixel in the first view image based on the disparity value corresponding to each pixel in the first view image, to obtain the second position information;
[0148] The aforementioned execution module 402 is specifically used to generate a second view image based on the second position information and pixel information of each pixel.
[0149] Optionally, in the embodiments of this application,
[0150] The aforementioned execution module 402 is specifically used to generate a second disparity image corresponding to the first view image based on the depth information of the first view image;
[0151] The aforementioned execution module 402 is specifically used to delete pixels in invalid value regions of the second disparity image to obtain the first disparity image.
[0152] Optionally, in the embodiments of this application,
[0153] The aforementioned execution module 402 is specifically used to perform an affine transformation on the first view image based on the first disparity image to obtain a third view image;
[0154] The aforementioned execution module 402 is specifically used to fill the target image region of the third view image to obtain the second view image;
[0155] The target image region mentioned above is the image region in the third view image that corresponds to the invalid value region in the second parallax image.
[0156] Optionally, in this embodiment of the application, the above-mentioned device further includes: a training module 403 and a processing module 404, wherein,
[0157] The training module 403 is used to train the binocular stereo matching network using the first view image captured by the shooting module 401 and the second view image obtained by the execution module 402, so as to obtain the trained binocular stereo matching network.
[0158] The aforementioned processing module 404 is used to input the fourth view image into the trained binocular stereo matching network and output the target disparity image;
[0159] The aforementioned execution module 402 is further configured to perform an affine transformation on the fourth view image based on the aforementioned target disparity image to obtain the target view image;
[0160] Among them, the fourth view image and the target view image are view images from different perspectives of the same shooting scene.
[0161] In the image generation apparatus provided in this application embodiment, the image generation apparatus acquires a first view image through a target camera, determines a first disparity image corresponding to the first view image, and then performs an affine transformation on the first view image based on the first disparity image to obtain a second view image; wherein the first view image and the second view image are view images from different perspectives of the same shooting scene. Through this method, the image generation apparatus can acquire a first view image through a target camera and perform an affine transformation on the first view image according to its disparity map to obtain a second view image that matches the first view image. Thus, view images can be acquired with a single camera, eliminating the need for complex binocular data acquisition processes and expensive binocular data acquisition equipment, and improving the generalization of the generated second view image.
[0162] The image generation device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0163] The image generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0164] The image generation apparatus provided in this application embodiment can achieve... Figures 1 to 3 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0165] Optionally, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described image generation method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0166] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0167] Figure 6 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0168] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.
[0169] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0170] In this embodiment, the input unit 104 is a target camera used to acquire a first view image; the processor 110 is used to determine a first parallax image corresponding to the first view image acquired by the target camera; the processor 110 is also used to perform an affine transformation on the first view image based on the first parallax image to obtain a second view image; the first view image and the second view image are view images from different perspectives of the same shooting scene.
[0171] Optionally, in this embodiment of the application, the first parallax image includes: the parallax value corresponding to each pixel in the first view image;
[0172] The processor 110 is specifically used to perform affine transformation on the first position information of each pixel in the first view image based on the disparity value corresponding to each pixel in the first view image to obtain the second position information.
[0173] The aforementioned processor 110 is specifically used to generate a second view image based on the second position information and pixel information of each pixel.
[0174] Optionally, in the embodiments of this application,
[0175] The processor 110 is specifically used to generate a second parallax image corresponding to the first view image based on the depth information of the first view image;
[0176] The processor 110 is specifically used to delete pixels in invalid value regions of the second disparity image to obtain the first disparity image.
[0177] Optionally, in the embodiments of this application,
[0178] The processor 110 described above is specifically used to perform an affine transformation on the first view image based on the first parallax image to obtain a third view image;
[0179] The processor 110 described above is specifically used to fill the target image region of the third view image to obtain the second view image;
[0180] The target image region mentioned above is the image region in the third view image that corresponds to the invalid value region in the second parallax image.
[0181] Optionally, in this embodiment of the application, the processor 110 is used to train the binocular stereo matching network using the first view image captured by the monocular camera and the obtained second view image, so as to obtain the trained binocular stereo matching network.
[0182] The processor 110 described above is used to input the fourth view image into the trained binocular stereo matching network and output the target disparity image;
[0183] The processor 110 is further configured to perform an affine transformation on the fourth view image based on the target disparity image to obtain the target view image;
[0184] Among them, the fourth view image and the target view image are view images from different perspectives of the same shooting scene.
[0185] In the electronic device provided in this application embodiment, the electronic device acquires a first view image through a target camera, determines a first disparity image corresponding to the first view image, and then performs an affine transformation on the first view image based on the first disparity image to obtain a second view image; wherein the first view image and the second view image are view images from different perspectives of the same shooting scene. Through this method, the electronic device can acquire a first view image through a target camera and perform an affine transformation on the first view image according to its disparity map to obtain a second view image that matches the first view image. Thus, view images can be acquired through a single camera of the electronic device without the need for complex binocular data acquisition processes and expensive binocular data acquisition equipment, and the generalization of the generated second view image is improved.
[0186] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.
[0187] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or it may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0188] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.
[0189] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image generation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0190] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0191] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image generation method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0192] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0193] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the image generation method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0194] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0195] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0196] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An image generation method, characterized in that, The method includes: First-view images are captured using the target camera; Determine the first parallax image corresponding to the first view image; Based on the first parallax image, an affine transformation is performed on the first view image to obtain a second view image; the first view image and the second view image are view images from different perspectives of the same shooting scene. Determining the first disparity image corresponding to the first view image includes: Based on the depth information of the first view image, a second disparity image corresponding to the first view image is generated; Delete the pixels in the invalid value region of the second disparity image to obtain the first disparity image; The step of generating a second disparity image corresponding to the first view image based on the depth information of the first view image includes: Where d is the disparity value corresponding to each pixel in the first view image, f is the focal length information of the binocular imaging system, b is the baseline of the binocular imaging system, and z is the depth information of the first view image. The second disparity image is generated by using the disparity value corresponding to each pixel in the first view image as the pixel value.
2. The method according to claim 1, characterized in that, The first disparity image includes: the disparity value corresponding to each pixel in the first view image; Based on the first disparity image, an affine transformation is performed on the first view image to obtain a second view image, including: Based on the disparity value corresponding to each pixel in the first view image, an affine transformation is performed on the first position information of each pixel in the first view image to obtain the second position information. A second view image is generated based on the second position information and pixel information of each pixel.
3. The method according to claim 1, characterized in that, The step of performing an affine transformation on the first view image based on the first disparity image to obtain a second view image includes: Based on the first disparity image, an affine transformation is performed on the first view image to obtain the third view image; The target image region of the third view image is filled to obtain the second view image; The target image region is the image region in the third view image that corresponds to the invalid value region in the second parallax image.
4. The method according to claim 1, characterized in that, The method further includes: The binocular stereo matching network is trained using the first view image and the second view image to obtain the trained binocular stereo matching network. The fourth view image is input into the trained binocular stereo matching network, which outputs a target disparity image. Based on the target disparity image, an affine transformation is performed on the fourth view image to obtain the target view image; The fourth view image and the target view image are view images from different perspectives of the same shooting scene.
5. An image generation apparatus, characterized in that, The device includes: a shooting module and an execution module, wherein: The shooting module is used to capture a first view image through the target camera; The execution module is used to determine the first parallax image corresponding to the first view image acquired by the shooting module; The execution module is further configured to perform an affine transformation on the first view image based on the first parallax image to obtain a second view image; the first view image and the second view image are view images from different perspectives of the same shooting scene; The execution module is specifically used to generate a second disparity image corresponding to the first view image based on the depth information of the first view image; The execution module is specifically used to delete pixels in the invalid value region of the second disparity image to obtain the first disparity image; The execution module is specifically used for: Where d is the disparity value corresponding to each pixel in the first view image, f is the focal length information of the binocular imaging system, b is the baseline of the binocular imaging system, and z is the first disparity image. The second disparity image is generated by taking the disparity value corresponding to each pixel in the first view image as the pixel value.
6. The apparatus according to claim 5, characterized in that, The first disparity image includes: the disparity value corresponding to each pixel in the first view image; The execution module is specifically used to perform affine transformation on the first position information of each pixel in the first view image based on the disparity value corresponding to each pixel in the first view image, so as to obtain the second position information. The execution module is specifically used to generate a second view image based on the second position information and pixel information of each pixel.
7. The apparatus according to claim 5, characterized in that, The execution module is specifically used to perform an affine transformation on the first view image based on the first disparity image to obtain a third view image; The execution module is specifically used to fill the target image region of the third view image to obtain the second view image; The target image region is the image region in the third view image that corresponds to the invalid value region in the second parallax image.
8. The apparatus according to claim 5, characterized in that, The device further includes: a training module and a processing module. The training module is used to train the binocular stereo matching network using the first view image captured by the shooting module and the second view image obtained by the execution module, so as to obtain the trained binocular stereo matching network. The processing module is used to input the fourth view image into the trained binocular stereo matching network and output the target disparity image; The execution module is further configured to perform an affine transformation on the fourth view image based on the target disparity image to obtain the target view image; The fourth view image and the target view image are view images from different perspectives of the same shooting scene.
Citation Information
Patent Citations
Method and device for converting single view image into a plurality of view images
CN108124148A