A method and system for quickly processing naked-eye 3D character photos
By generating three-dimensional point clouds and rotating interpolated viewpoint images, combining stereo matching algorithms and cylindrical grating plates, the problems of time-consuming and low quality in the existing technology are solved, and fast and high-quality naked-eye 3D photo generation is achieved.
Patent Information
- Application Number
- CN202510510897.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing naked-eye 3D photo generation methods take a long time and rely on manual operations, resulting in large angle errors in view angle transformation and relying on manual experience to generate depth information, which is difficult to meet the needs of high-quality 3D photos. In addition, the stitching step size calculation relies on manual experience, which is easy to produce image ghosting.
By extracting the depth information of at least two viewpoint images, the interpolated viewpoints are generated, and the interpolated viewpoint images are generated. The pixel extraction and splicing are performed using a stereo matching algorithm and an interpolation method, and the printer resolution and the cylindrical grating plate are combined to generate naked-eye 3D photos.
It realizes the rapid generation of high-quality naked-eye 3D photos, simulating the parallax movement characteristics of human eyes, and the generated photos have a strong sense of three-dimensionality, meeting the needs of the tourism industry to take photos immediately.
Smart Images

Figure CN120047591B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and particularly relates to a method and system for quickly processing naked-eye 3D human photos. Background Art
[0002] In the tourism industry, naked-eye 3D photos usually adopt the method of covering multiplexed images with lenticular gratings. To generate naked-eye 3D photos, first, it is necessary to generate quantitative interpolation viewpoint images as a sequence of images, and then mix the sequence of images into a multiplexed image. Mixing the sequence of images is a process of extracting appropriate pixels from the quantitative interpolation viewpoint images, using the width of each pixel as a step size, and then splicing them into a new image in a certain order. The multiplexed image contains multi-viewpoint information of the original image.
[0003] Currently, in the methods for generating naked-eye 3D photos used in the tourism industry, the most common method is to generate N interpolation viewpoint images based on 2D photos by manual rotation, assign depth information to the spliced image according to personal experience, and then manually select the sequence of images and use and other tools to manually generate the spliced image. This method relies on manually generating viewpoint images, which takes up to 2 - 5 hours, and the angle error rate of the viewing angle transformation between adjacent interpolation viewpoint images is relatively large. In addition, the generation of depth information relies on manual experience annotation, resulting in a distorted sense of three-dimensionality and making it difficult to meet the requirements of high-quality 3D photos. At the same time, the calculation of the splicing step size also relies on manual experience, which is prone to image ghosting and further reduces the visual effect of naked-eye 3D photos. These defects severely limit the application of traditional methods in the tourism industry, and there is an urgent need for an efficient and accurate automated solution.
[0004] In summary, the existing systems and methods for generating naked-eye 3D photos used in the tourism industry have disadvantages such as long processing time and low photo quality. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for generating naked-eye 3D human photos with fast processing speed, high photo quality, and suitable for tourist photography scenarios.
[0006] In the first aspect, the present invention provides a method for quickly processing naked-eye 3D human photos, and the process is as follows:
[0007] Extract the depth information of at least two viewpoint images and generate a three-dimensional point cloud.
[0008] Establish an observation arc line between the viewpoints of the initial at least two viewpoint images, rotate and interpolate viewpoints on the observation arc line, and generate interpolated viewpoint images corresponding to each interpolated viewpoint.
[0009] Pixel extraction and stitching are performed on each viewpoint image to form a stereoscopic image corresponding to the naked-eye 3D portrait photo.
[0010] Preferably, the process of generating the three-dimensional point cloud is as follows:
[0011] Pixel-level matching is performed on two viewpoint images through a stereo matching algorithm to generate a disparity map.
[0012] Color mapping is performed on the disparity map to generate a color disparity image; the color disparity image is converted into a three-dimensional point cloud; the three-dimensional depth information of each point in the three-dimensional point cloud is extracted.
[0013] Preferably, the process of interpolating the rotational viewpoint is as follows: A rotational reference point is selected in the three-dimensional point cloud, and an observation arc line is established; the observation arc line takes the rotational reference point as the center and passes through the viewpoints of the initial two viewpoint images. One or more viewpoints are interpolated on the observation arc line.
[0014] Preferably, the process of generating the interpolated viewpoint image is as follows: According to the position of the interpolated viewpoint on the observation arc line, an interpolation quaternion q corresponding to the interpolated viewpoint is generated t ; According to the interpolation quaternion q t The coordinates p of each point in the three-dimensional point cloud n are rotated to obtain the coordinates p' of each point in the three-dimensional point cloud from the perspective of the interpolated viewpoint n . The three-dimensional point cloud from the perspective of the interpolated viewpoint is projected onto a pixel grid and adjusted to obtain a primary interpolated viewpoint image.
[0015] Preferably, for the initial viewpoint image and the interpolated viewpoint images interpolated and generated on the observation arc line, one or more secondary interpolated viewpoint images are generated between the viewpoint images corresponding to two adjacent viewpoints through a large motion frame interpolation method.
[0016] Preferably, the process of pixel extraction and stitching for the viewpoint image is as follows: The resolution parameter of the printer is divided by the number of grating lines of the lenticular grating plate to obtain an optimal step size ; According to the arrangement order of each viewpoint image, different pixel column intervals are respectively extracted on different viewpoint images; the pixel width of the pixel column interval is equal to the optimal step size . The pixel column intervals extracted on each viewpoint image are stitched in sequence to form a stereoscopic image
[0017] Preferably, grating alignment lines are generated on both sides of the stereoscopic image; the distance between the grating alignment line and the edge of the stereoscopic image is an integer multiple of the optimal step size . According to the position of the grating alignment line, the lenticular grating plate is covered and fixed on the stereoscopic image to obtain the naked-eye 3D portrait photo
[0018] Preferably, the rotation reference point is the position of the center point of the human face in the viewpoint image in the three-dimensional point cloud.
[0019] Preferably, the viewpoint image is preprocessed. The preprocessing includes binarization, dilation, and Gaussian blur processing.
[0020] Preferably, the initial at least two viewpoint images are obtained by shooting with a binocular camera.
[0021] Preferably, the rotation reference point is the position of the focus of the initial viewpoint image in the three-dimensional point cloud.
[0022] In a second aspect, the present invention provides a fast processing system for naked-eye 3D human photos, which is used to execute the foregoing method; the system includes a binocular camera, a point cloud generation module, an interpolation module, a stitching module, and a printing and output module. The binocular camera is used to shoot the initial viewpoint image; the point cloud generation module is used to extract the depth information of the viewpoint image and generate a three-dimensional point cloud. The interpolation module is used to generate an interpolated viewpoint image; the stitching module is used to perform pixel extraction and stitching on each viewpoint image; the printing and output module is used to print the split-view image.
[0023] In a third aspect, the present invention provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and is characterized in that: the memory stores the computer program; the processor executes the foregoing method for fast processing of naked-eye 3D human photos.
[0024] The present invention has the following beneficial effects.
[0025] 1. The present invention automatically calculates and generates an interpolated viewpoint image between the left and right two viewpoint images, and uses multiple viewpoint images to interleave and stitch to form a split-view image, and uses the printer resolution parameter and the grating line number of the lenticular grating plate to align the grating on both sides of the split-view image, so that it is convenient for the staff to quickly cover the lenticular grating plate on the automatically generated split-view image to obtain a 3D naked-eye human photo. The time for the present invention to generate a complete split-view image is about 60 seconds, which meets the on-the-spot photo-taking needs of the tourism industry.
[0026] 2. The present invention extracts depth information based on the left and right two viewpoint images, generates a three-dimensional point cloud, and then obtains the observation arc line for generating the interpolated image; the viewpoint images generated along the observation arc line can simulate the binocular parallax motion characteristics of the human eye, and can make the split-view image synthesized by multiple viewpoint images more three-dimensional.
[0027] 3. The 3D naked-eye photos generated by the present invention have the advantages of strong 3D effect and fast processing speed, which meet the high-quality requirements of tourists for photos and the need to quickly obtain photos. Description of the Drawings
[0028] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. In the accompanying drawings:
[0029] Figure 1 is the overall flowchart of the embodiment of the present invention;
[0030] Figure 2 is the original photo obtained in step S1 of the embodiment of the present invention;
[0031] Figure 3 is the partial interpolated view point image obtained in step S4 of the embodiment of the present invention;
[0032] Figure 4 is the enlarged view of the perspective difference between different interpolated view point images in the embodiment of the present invention;
[0033] Figure 5 is the split view image obtained in step S63 of the embodiment of the present invention. Specific Embodiments
[0034] In order to better understand the principle and steps of the present invention, the specific embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings.
[0035] The embodiment is as Figure 1 shown. A method for generating a naked-eye 3D portrait photo includes the following steps:
[0036] S1. Use a binocular camera to collect images of the same object to obtain an original photo as Figure 2 shown; the original photo includes a left-eye portrait photo and a right-eye portrait photo taken by the binocular camera, and the two are used as the leftmost view point image and the rightmost view point image respectively.
[0037] S2. Extract the 2D coordinates of the focus when the binocular camera takes pictures, which is convenient for finding the rotation reference point of the rotation interpolation frame (combining the quaternion spherical linear interpolation algorithm and the FILM frame filling model). In this embodiment, the three-dimensional coordinates of the focus are used as the rotation reference point.
[0038] S3. Use a binocular stereo matching model to extract the depth information of the leftmost view point image and the rightmost view point image respectively and generate a three-dimensional point cloud, specifically as follows:
[0039] S31. Based on the stereo matching algorithm, perform pixel-level matching on the leftmost view point image and the rightmost view point image to generate a disparity map. In this embodiment, the stereo matching algorithm uses the SGM (semi-global matching) algorithm.
[0040] S32. Perform color mapping on the disparity map, perform pseudo-color coding on the disparity data based on a preset color space conversion model, generate a color disparity image with a visual enhancement effect, convert the color-mapped disparity map into a three-dimensional point cloud, and obtain the three-dimensional depth information of each pixel point in the three-dimensional point cloud;
[0041] In this embodiment, the ratio of the pixel distance x to the actual distance X between any two points in the three-dimensional point cloud is 1:16; the expression of the Euclidean distance d between any two points in the three-dimensional point cloud is as follows:
[0042]
[0043] Where are the horizontal coordinates of the two points respectively; are the vertical coordinates of the two points respectively; are the vertical coordinates of the two points respectively.
[0044] S4. Generate N interpolated view images based on the rotation interpolation method, and use the interpolated view images as a sequence of images, specifically as follows:
[0045] S41. According to the two focal point 2d coordinates obtained in step S2 and the three-dimensional depth information of the point cloud obtained in step S3, obtain the three-dimensional coordinates of the focal points in the point cloud as the rotation reference points.
[0046] S42. Generate interpolated view images based on an arc path. Determine one or more interpolated view coordinates in the point cloud space. The interpolated view coordinates are on the observation arc line. The observation arc line takes the rotation reference point obtained in step S41 as the center of the circle and passes through the positions of the two lenses of the binocular camera. For each interpolated view coordinate, generate a first-level interpolated view image; the first-level interpolated view image is a photo taken with the interpolated view coordinate as the shooting point.
[0047] In this embodiment, determine an interpolated view coordinate, denoted as the middle view coordinate; the middle view coordinate is located at the center point of the two lens positions (i.e., the left view point and the right view point) on the observation arc line.
[0048] The process of generating the first-level interpolated view image for the middle view coordinate is as follows:
[0049] (1) Determine the interpolation ratio according to the set position of the interpolated view coordinate between the two rear view points ; the interpolation ratio represents the proportion of the viewing angle difference from the interpolated view point to the left view point in the viewing angle difference from the left view point to the right view point; in this embodiment, since the interpolated view point is located at the center point of the two lens positions on the observation arc line, so ; in some other embodiments, the interpolation ratio One or more other values can also be taken, such as 1 / 4, 1 / 3, 2 / 3, 3 / 4 or other reasonable numerical values.
[0050] (2) Establish the direction vector of the rotation axis in three-dimensional space ; The rotation axis is a straight line passing through the rotation reference point and perpendicular to the plane where the observed arc line is located. In this embodiment, the focus is used as the rotation reference point and the three-dimensional coordinate origin, so the rotation axis is the Z-axis. .
[0051] (3) Establish the spherical linear interpolation quaternion q of the interpolation viewpoint t as follows:
[0052]
[0053] Among them, and are the viewpoint quaternions corresponding to the leftmost viewpoint image and the rightmost viewpoint image respectively, representing the initial rotation states of the leftmost viewpoint image and the rightmost viewpoint image. The quaternion is calculated by the following formula: ; ; is the viewing angle of the viewpoint; satisfying ; is the viewing angle difference between the leftmost viewpoint image and the rightmost viewpoint image. In this embodiment, it is stipulated that the viewing angle of the leftmost viewpoint image is 0, and the viewing angle of the rightmost viewpoint image is .
[0054] (4) Use the spherical linear interpolation quaternion q t to perform position rotation adjustment on each point in the original three-dimensional point cloud respectively; The specific rotation method is: represent the coordinates of each point in the original three-dimensional point cloud as , where n = 1, 2,..., N, and N is the total number of points in the three-dimensional point cloud;
[0055] Represent the quaternion q t as ; Among them, w t is the real part, x t , y t , z t are the imaginary parts, and i, j, k are three imaginary units; Convert each point in the three-dimensional point cloud into the form of a pure imaginary quaternion ; Calculate the pure imaginary quaternion corresponding to the coordinates of each point in the three-dimensional point cloud at the viewing angle of the interpolation viewpoint according to the following formula
[0056]
[0057] Among them, , Extract the pure imaginary quaternion and obtain the coordinates of each point in the three-dimensional point cloud from the imaginary part under the new perspective. .
[0058] (5) Project the three-dimensional point cloud obtained in step (4) under the new perspective onto the pixel grid corresponding to the interpolation view point, use the bicubic interpolation method to supplement the sawteeth or holes, and use the depth buffer (Z-Buffer) method to solve the visibility conflict when multiple three-dimensional points are projected onto the same pixel. Generate the first-level interpolation view point image.
[0059] In this step, since the established view point change curve is circular arc-shaped, it can simulate the binocular parallax motion characteristics of the human eye, and can make the anaglyph image synthesized from multiple view point images more three-dimensional.
[0060] S43. After step S41, more than three view point images are generated. Further generate the second-level interpolation view point images between adjacent two view point images.
[0061] In this embodiment, through the large motion frame interpolation method (FILM, Frame Interpolation for LargeMotion), one or more second-level interpolation view point images are generated between adjacent two view point images. This step can further encrypt the view point images and reduce the view angle difference between adjacent view point images, so that the transition of different positions of the finally generated anaglyph image is smoother.
[0062] In this step, by generating the interpolation view point images, the large motion occlusion area can be repaired, the observer's field of view can be expanded, and the realism and clarity of the frame can be improved.
[0063] The efficiency of generating view point images by the large motion frame interpolation method is relatively high; and, since the view angle difference between adjacent view point images has been reduced after generating the first-level interpolation view point images through step S42, the view point images directly generated by the large motion frame interpolation method are still close to the observed circular arc line, and can still provide good three-dimensional sense for the anaglyph image.
[0064] S44. Arrange all view point images in a sequence diagram. The number N of images in the sequence diagram is determined according to the smoothness requirement of the anaglyph image; the larger the number N of images, the higher the smoothness of the anaglyph image, but the higher the required image processing computing power, and the longer the time required to generate the anaglyph image.
[0065] Some of the interpolation view point images shown in step S4 are as Figure 3 and 4 shown; as can be seen from Figure 4 , the human figure occlusion relationships of different interpolation view point images are different, which helps to improve the three-dimensional sense of the final 3D photo.
[0066] S5. Optimize the sequence diagram obtained in step S44 through binarization, dilation, and Gaussian blur processing. The specific process is as follows:
[0067] S51. Perform binarization processing on the sequence diagram to convert the image into a black-and-white binary image.
[0068] S52. Perform dilation processing on the sequence diagram. Use a 3×3 rectangular kernel to perform morphological dilation on the binary image to enhance the edge continuity.
[0069] S53. Perform Gaussian blur processing on the sequence diagram. Apply a Gaussian filter to smooth the noise and retain the main body contour, ultimately enhancing the image quality and three-dimensional sense.
[0070] S6. Calculate the optimal step size for stitching to obtain the autostereoscopic image. Stitch the sequence diagram processed in step S5 to obtain the autostereoscopic image, and generate raster alignment lines, then upload them to the printer. The specific process is as follows:
[0071] S61. Automatically optimize and adjust the size (including length and width) of the sequence diagram processed in step S5 to the target size according to the target size of the printed photo.
[0072] S62. Calculate the optimal step size for stitching according to the printer resolution parameter to ensure that the visual effect is natural after the stitched image covers the lenticular grating plate;
[0073] Optimal step size Is the number of horizontal pixels for each step of stitching, and its expression is as follows:
[0074]
[0075] Where, Is the printer resolution parameter (specifically the number of physical ink dots that can be printed per unit length); Is the grating line number of the lenticular grating plate (specifically the number of grating columns per unit length).
[0076] S63. According to the optimal step size , stitch the sequence diagram to obtain the autostereoscopic image, specifically as follows:
[0077] Extract the reserved area on each viewpoint image as follows:
[0078] S i =[ k · N·γ + ( i-1)· γ + 1, k · N·γ + i·γ ]
[0079] where k = 0, 1, 2,... ; i = 1, 2,... N; P is the total number of pixel columns of the target photo; N is the number of viewpoint images; is the serial number of the viewpoint image; is the set of column serial numbers of the reserved area of the viewpoint image i.
[0080] The reserved areas of each viewpoint image are spliced in sequence to obtain a multi-view image. The multi-view image is as Figure 5 shown. The 3D effect of the multi-view image cannot be directly observed, and its 3D effect appears after covering with a lenticular grating sheet. The multi-view image occupies a large amount of memory, and the memory occupied by the multi-view image required for printing a six-inch photo is between 120 Mb and 160 Mb.
[0081] After completing the step S6, the multi-view image is printed, covered with a lenticular grating sheet and framed, and thus the complete finished product of the present invention is obtained.
[0082] S64. Generate grating alignment lines on both sides of the multi-view image; the grating alignment lines are at a distance of .
[0083] S65. Use a printer to print out the multi-view image.
[0084] S66. Stick a lenticular grating sheet on the multi-view image, a naked-eye 3D portrait photo.
[0085] In some embodiments, instead of using the focus as the rotation reference point, the center point of one of the faces in the leftmost viewpoint image and the rightmost viewpoint image is used as the rotation reference point. The position of the face in the image is obtained by identifying through an application network.
Claims
1. A method for quickly processing a naked-eye 3D portrait photo, characterized in that: The method is as follows: Extract the depth information of at least two viewpoint images and generate a three-dimensional point cloud; Establish an observation arc line between the viewpoints of the initial at least two viewpoint images, rotate and interpolate viewpoints on the observation arc line, and generate interpolated viewpoint images corresponding to the interpolated viewpoints; The process of the rotation interpolation viewpoint is as follows: Select a rotation reference point in the three-dimensional point cloud and establish an observation arc line; The observation arc line takes the rotation reference point as the center and passes through the viewpoints of the initial two viewpoint images; Interpolate one or more viewpoints on the observation arc line; The process of generating an interpolated view point image is as follows: according to the position of the interpolated view point on the observation arc line, an interpolation quaternion corresponding to the interpolated view point is generated. q t ; According to the interpolation quaternion q t the coordinates of each point in the three-dimensional point cloud p n are rotated to obtain the coordinates of each point in the three-dimensional point cloud under the perspective of the interpolated view point p ' n ; The three-dimensional point cloud under the perspective of the interpolated view point is projected onto the pixel grid and adjusted to obtain a primary interpolated view point image. Extract and splice pixels of each viewpoint image to form a parallax image corresponding to the naked-eye 3D portrait photo; Generate raster alignment lines on both sides of the parallax image; The distance between the raster alignment line and the edge of the parallax image is an integer multiple of the optimal step γ; According to the position of the raster alignment line, cover and fix a lenticular grating plate on the parallax image to obtain a naked-eye 3D portrait photo.
2. A method for quickly processing a naked-eye 3D portrait photo according to claim 1, characterized in that: The process of generating a three-dimensional point cloud is as follows: Perform pixel-level matching on two viewpoint images through a stereo matching algorithm to generate a disparity map; Perform color mapping on the disparity map to generate a color disparity image; Convert the color disparity image into a three-dimensional point cloud; Extract the three-dimensional depth information of each point in the three-dimensional point cloud.
3. A fast processing method for a naked-eye 3D portrait photo according to claim 1, characterized in that: For the initial viewpoint image and the first-level interpolated viewpoint images generated by interpolation on the observation arc line, generate one or more second-level interpolated viewpoint images between the viewpoint images corresponding to adjacent viewpoints through a large motion frame interpolation method.
4. A method for quickly processing a naked-eye 3D portrait photo according to claim 1, characterized in that: The process of pixel extraction and splicing of the viewpoint image is as follows: Divide the resolution parameter of the printer by the number of grating lines of the lenticular grating plate to obtain the optimal step γ; According to the arrangement order of each viewpoint image, extract different pixel column intervals on different viewpoint images; The pixel width of the pixel column interval is equal to the optimal step γ; Splice the pixel column intervals extracted on each viewpoint image in order to form a parallax image.
5. A method for quickly processing a naked-eye 3D portrait photo according to claim 1, characterized in that: The rotation reference point takes the position of the center point of the face in the viewpoint image in the three-dimensional point cloud.
6. A rapid processing system for naked-eye 3D character photos, characterized in that: For executing the method according to claim 1; The system includes a binocular camera, a point cloud generation module, an interpolation module, a splicing module, and a printing output module; The binocular camera is used to capture the initial viewpoint image; The point cloud generation module is used to extract the depth information of the viewpoint image and generate a three-dimensional point cloud; The interpolation module is used to generate interpolated viewpoint images; The splicing module is used to extract and splice pixels of each viewpoint image; The printing output module is used to print the parallax image.
7. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: The memory stores a computer program; The processor executes a method for quickly processing a naked-eye 3D portrait photo according to any one of claims 1-5.
Citation Information
Patent Citations
Optical grating three-dimensional printing image synthetic method based on binocular camera
CN103702103A
KR20240158102A