Methods, systems, devices, and storage media for generating and producing 3D lenticular images
The integration of 2D-3D image transformation and dynamic viewpoint interpolation techniques with machine learning models automates and enhances the production of high-quality 3D lenticular images, overcoming scalability and accuracy challenges.
Patent Information
- Application Number
- PCT/US2024/037113
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2026-01-15
AI Technical Summary
The production of 3D lenticular images is hindered by manual labor-intensive processes, lack of scalability, and insufficient accuracy in texture information and viewpoint transitions, making large-scale production inefficient and of poor quality.
A method combining 2D-3D image transformation algorithms with dynamic viewpoint interpolation techniques, utilizing machine learning models for depth information estimation and content-adaptive lenticular printing, to automate and enhance the generation of high-quality 3D lenticular images.
Enables efficient, automated production of high-quality 3D lenticular images with improved texture and smooth viewpoint transitions, addressing scalability and accuracy issues.
Smart Images

Figure US2024037113_15012026_PF_FP_ABST
Abstract
Description
METHODS, SYSTEMS, DEVICES, AND STORAGE MEDIA FOR GENERATING ANDPRODUCING 3D LENTICULAR IMAGESTECHNICAL FIELD
[0001] The present disclosure relates to the technical field of 3D lenticular images, and in particular, to a method, a system, a device, and a storage medium for generating and producing a 3D lenticular image.BACKGROUND
[0002] A 3D lenticular image utilizes the principles of binocular disparity and optical refraction to create a flat image that offers a three-dimensional visual effect. The current market demand for 3D lenticular images is steadily increasing, with rising expectations for the visual quality of these images. However, the production process of 3D lenticular images requires a significant amount of manual work and lacks flexibility and scalability, making it difficult to achieve efficient large-scale production. Existing production algorithms for 3D lenticular images suffer from insufficient accuracy, resulting in issues such as inadequate texture information and discontinuous viewpoint transitions in the generated images.
[0003] Accordingly, it is desired to provide a method, a system, a device, and a storage medium for generating and producing a 3D lenticular image, capable of automatically and efficiently generating a 3D lenticular image and improving its quality.SUMMARY
[0004] One or more embodiments of the present disclosure provide a method for generating and producing a 3D lenticular image, comprising: determining depth information of a 2D target image of a target object; generating a multi -viewpoint image corresponding to the 2D target image using a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique based on the depth information, and the multi-viewpoint image including a plurality of image frames corresponding to different shooting angles; and generating a 3D lenticular image of the target object based on the multi -viewpoint image.
[0005] In some embodiments, determining the depth information of a 2D target image of a target object includes obtaining an algorithm selection model, which is a trained machine learning model; utilizing the algorithm selection model to process the 2D target image and select a target depth information estimation algorithm from candidate depth information estimation algorithms; and determining the depth information of the 2D target image by processing the 2D target image using the target depth information estimation algorithm.
[0006] In some embodiments, the algorithm selection model is an adaptive depth perception network.
[0007] In some embodiments, a first training input and a first training label of the algorithm selection model are determined by: obtaining an image generation record, the image generation record includes a historical 2D image, a historical depth information estimation algorithm, and a historical 3D lenticular image, the historical 3D lenticular image is generated based on the historical 2D image and the historical depth information estimation algorithm; and determining whether the historical 3D lenticular image in the image generation record satisfies a preset condition; and in response to determining that the historical 3D lenticular image satisfies the preset condition, designating the historical 2D image in the image generation record as the first training input and the historical depth information estimation algorithm in the image generation record as the first training label.
[0008] In some embodiments, wherein the candidate depth information estimation algorithms include a depth information estimation model. A second training input and a second training label of the depth information estimation model are determined by: obtaining an image generation record, the image generation record includes a historical 2D image, historical depth information of the historical 2D image, and a historical 3D lenticular image corresponding to the historical 2D image, the historical 3D lenticular image is generated based on the historical depth information; and determining whether the historical 3D lenticular image in the image generation record satisfies a preset condition; and in response to determining that the historical 3D lenticular image satisfies the preset condition, designating the historical 2D image in the image generation record as the second training input and the historical depth information in the image generation record as the second training label.
[0009] In some embodiments, wherein the generating a multi-viewpoint image corresponding to the 2D target image using a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique based on the depth information includes generating an initial multiviewpoint image using the 2D-3D image transformation algorithm based on the depth information and preset shooting parameters, the initial multi-viewpoint image including a plurality of initial image frames corresponding to different initial shooting angles, and the preset shooting parameters including at least a preset shooting angle range and a preset shooting angle step; and processing at least a portion of the plurality of initial image frames using the dynamic viewpoint interpolation technique to generate the plurality of image frames of the multiviewpoint image, wherein shooting angles of the plurality of image frames are consistent with the preset shooting angle range and the preset shooting angel step.
[0010] In some embodiments, wherein the generating a 3D lenticular image of the target objectbased on the multi -viewpoint image includes determining arrangement information of each pixel of the plurality of image frames of the 3D lenticular image using a content-adaptive lenticular prints algorithm based on the multi-viewpoint image; and generating the 3D lenticular image based on the arrangement information.
[0011] In some embodiments, wherein the determining arrangement information of each pixel of the plurality of image frames of the 3D lenticular image using a content-adaptive lenticular prints algorithm based on the multi-viewpoint image includes: obtaining a grating device parameter; and determining the arrangement information of each pixel of the plurality of image frames of the 3D lenticular image using the content-adaptive lenticular prints algorithm based on the grating device parameter and the multi-viewpoint image.
[0012] In some embodiments, wherein the obtaining a grating device parameter includes: obtaining the grating device parameter by processing category information of the target object and a target off-screen distance using a grating parameter determination model, the grating parameter determination model being a trained machine learning model.
[0013] In some embodiments, wherein the generating a 3D lenticular image of the target object based on the multi -viewpoint image includes generating the 3D lenticular image by processing the multi-viewpoint image using an image simulation model, wherein the image simulation model is a trained machine learning model, the image simulation model includes a feature extraction module and an image simulation module, the feature extraction module is configured to extract feature information of the multi-viewpoint image, and the image simulation module is configured to generate the 3D lenticular image based on the feature information.
[0014] In some embodiments, wherein an output of the image simulation model is designated as a candidate 3D lenticular image, and the method further comprises determining whether the candidate 3D lenticular image satisfies a preset condition based on a target off-screen distance and an off-screen distance of the candidate 3D lenticular image; in response to determining that the candidate 3D lenticular image satisfies the preset condition, using the candidate 3D lenticular picture as the 3D lenticular image; and in response to determining that the candidate 3D lenticular image does not satisfy the preset condition, processing the candidate 3D lenticular image to generate the 3D lenticular image.
[0015] In some embodiments, wherein the processing the candidate 3D lenticular image to generate the 3D lenticular image includes generating the 3D lenticular image by processing the candidate 3D lenticular image, the target off-screen distance, and the off-screen distance of the candidate 3D lenticular image using an image transformation model, the image transformation model being a trained machine learning model.
[0016] In some embodiments, the method further comprises controlling a printing device toprint the 3D lenticular image; obtaining a first real-time image of a printing result of the printing device, the first real-time image being captured by a first image-capturing device during a printing process; and processing the first real-time image using a quality problem monitoring model to monitor a quality problem during the printing process, the quality problem monitoring model being a trained machine learning model.
[0017] In some embodiments, wherein the printing device includes a grating, and the processing the first real-time image using a quality problem monitoring model to monitor a quality problem during the printing process includes obtaining a second real-time image of the grating, the second real-time image being captured by a second image-capturing device during the printing process; and processing the first real-time image and the second real-time image using the quality problem monitoring model to monitor the quality problem during the printing process.
[0018] One or more embodiments of the present disclosure provide a system for generating and producing a 3D lenticular image, comprising a determination module, configured to determine depth information of a 2D target image of a target object; a first generation module, configured to generate a multi-viewpoint image corresponding to the 2D target image using a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique based on the depth information, and the multi -viewpoint image including a plurality of image frames corresponding to different shooting angles; and a second generation module, configured to generate a 3D lenticular image of the target object based on the multi -viewpoint image.
[0019] One or more embodiments of the present disclosure provide a device for generating and producing a 3D lenticular image, comprising at least one processor and at least one storage device, wherein the at least one storage device is configured to store computer instructions; and the at least one processor is configured to execute at least a portion of the computer instructions to implement a method for generating and producing a 3D lenticular image.
[0020] One or more embodiments of the present disclosure provide a computer-readable storage medium, wherein the storage medium stores computer instructions, and when a computer executes the computer instructions in the storage medium, the computer performs a method for generating and producing a 3D lenticular image.
[0021] According to the above scheme, by combining the 2D-3D image transformation algorithm and the dynamic viewpoint interpolation technique to process the depth information of the 2D target image, the multi-viewpoint image and the 3D lenticular image can be generated, thus achieving automatic and efficient production and improving the quality of the 3D lenticular image.BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The present disclosure is further described in terms of exemplary embodiments. These exemplary embodiments are described in detail with reference to the drawings. The drawings are not to scale. These embodiments are non-limiting exemplary embodiments, in which like reference numerals represent similar structures throughout the several views of the drawings, and wherein:
[0023] FIG. 1 is a schematic diagram of illustrating an exemplary system for generating and producing 3D lenticular images according to some embodiments of the present disclosure;
[0024] FIG. 2 is a block diagram illustrating an exemplary system for generating and producing 3D lenticular images according to some embodiments of the present disclosure;
[0025] FIG. 3 is a flowchart illustrating an exemplary method for generating and producing a 3D lenticular image according to some embodiments of the present disclosure;
[0026] FIG. 4 is a schematic diagram illustrating an exemplary process for determining depth information of a 2D target image according to some embodiments of the present disclosure;
[0027] FIG. 5 is a schematic diagram illustrating an exemplary process for generating a 3D lenticular image based on a multi-viewpoint image according to some embodiments of the present disclosure; and
[0028] FIG. 6 is a schematic diagram illustrating an exemplary process for monitoring a quality problem during a printing process according to some embodiments of the present disclosure.DETAILED DESCRIPTION
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be used in the description of the embodiments will be briefly described below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present disclosure, and it is possible for a person of ordinary skill in the art to apply the present disclosure to other similar scenarios in accordance with these drawings without creative labor. Unless obviously obtained from the context or the context illustrates otherwise, the same numeral in the drawings refers to the same structure or operation.
[0030] It should be understood that as used herein, the terms “system”, “device”, “unit” and / or “module” are used herein as a way to distinguish between different components, elements, parts, sections, or assemblies at different levels. However, the words may be replaced by other expressions if other words accomplish the same purpose.
[0031] As shown in the present disclosure and the claims, unless the context clearly suggests an exception, the words “one,” “a”, “an”, “one kind”, and / or “the” dose not refer specifically to thesingular, but may also include the plural. Generally, the terms “including” and “comprising” suggest only the inclusion of clearly identified steps and elements that do not constitute an exclusive list, and the method or apparatus may also include other steps or elements. |
[0032] Flowcharts are used in the present disclosure to illustrate operations performed by a system in accordance with embodiments of the present disclosure. It should be appreciated that the preceding or following operations are not necessarily performed in an exact sequence. Instead, steps can be processed in reverse order or simultaneously. Also, it is possible to add other operations to these processes or to remove a step or steps from these processes.
[0033] FIG.l is a schematic diagram illustrating an exemplary system for generating and producing 3D lenticular images according to some embodiments of the present disclosure. As shown in FIG. 1, the system 100 may include an image-capturing device 110, a processor 120, a network 130, a storage device 140, a terminal 150, and a printing device 160.
[0034] The image-capturing device 110 may be configured to capture an image (e.g., an RGB image, a point cloud image, etc.). For example, the image-capturing device 110 may take a picture of a printing result of the printing device 160 to obtain a first real-time image. As another example, the image-capturing device 110 may take a picture of a grating of the printing device 160 to obtain a second real-time image. The image-capturing device 110 may include a camera, a webcam, etc. The image-capturing device 110 may include a plurality of devices for taking pictures of different targets (e.g., a printing result of a printing device, a grating, etc.).
[0035] The processor 120 may be configured to process data and / or information obtained from the image-capturing device 110, the storage device 140, and / or the terminal 150. For example, the processor 120 may determine depth information of a 2D target image of a target object. As another example, the processor 120 may generate a multi-viewpoint image corresponding to the 2D target image by utilizing a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique based on the depth information. Further example, the processor 120 may generate a 3D lenticular image of the target object based on the multi -viewpoint image. In some embodiments, the processor 120 may send a processing result to the terminal 150. For example, the processor 120 may send the 3D lenticular image of the target object to the terminal 150 and display the 3D lenticular image on one or more display devices in the terminal 150.
[0036] In some embodiments, the processor 120 may be a single server or a group of servers. The group of servers may be centralized or distributed. In some embodiments, the processor 120 may be local or remote. In some embodiments, the processor 120 may be connected to the image-capturing device 110, the storage device 140, the terminal 150, and / or the printing device 160 via the network 130 or directly to access information and / or data stored thereon. In some embodiments, the processor 120 may be implemented on a cloud platform. By way of exampleonly, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an on-premises cloud, a multi-cloud, etc. or any combination thereof.
[0037] The network 130 may include any suitable network that may facilitate data exchange and communication connections. For example, the processor 120 may obtain the 2D target image of the target object from the image-capturing device 110 via the network 130. In some embodiments, the network 130 may be any one or more of a wired network or a wireless network.
[0038] The storage device 140 may store data, instructions, and / or any other information. In some embodiments, the storage device 140 may store data obtained from the image-capturing device 110, the processor 120, the terminal 150, and / or the printing device 160. In some embodiments, the storage device 140 may store data and / or instructions used by the processor 120 to perform or use exemplary methods that have been accomplished as described in the present disclosure. In some embodiments, the storage device 140 may include mass memory, removable memory, volatile read-write memory, read-only memory (ROM), or the like, or any combination thereof. In some embodiments, the storage device 140 may be implemented on the cloud platform. In some embodiments, the storage device 140 may be part of the processor 120.
[0039] The terminal 150 may enable user interactions. In some embodiments, the terminal 150 may include one of a mobile device 150-1, a tablet 150-2, a laptop 150-3, etc., or other device with input and / or output capabilities or any combination thereof. In some embodiments, the terminal 150 may receive data (e.g., the 3D lenticular image) from the processor 120 and display the data. In some embodiments, the terminal 150 may be omitted.
[0040] The printing device 160 may be used to print the 3D lenticular image. The printing device 160 may include printing elements such as a grating, a printing head, an ink cartridge, or the like. In some embodiments, the processor 120 may send the 3D lenticular image over the network 130 to the printing device 160 for printing.
[0041] It should be noted that the system 100 is provided for illustrative purposes only and is not intended to limit the scope of the present disclosure. For a person of ordinary skill in the art, a variety of modifications or variations may be made in accordance with the description of the present disclosure. For example, the system 100 may be implemented on other devices to achieve similar or different functionality. However, changes and modifications will not depart from the scope of the present disclosure.
[0042] FIG. 2 is a schematic diagram illustrating exemplary modules of a system 200 for generating and producing 3D lenticular images (hereinafter referred to as a system 200)according to some embodiments of the present disclosure. In some embodiments, the system 200 may include a determination module 210, a first generation module 220, and a second generation module 230. In some embodiments, the system 200 may be implemented by the processor 120 shown in FIG. 1.
[0043] The determination module 210 is configured to determine depth information of a 2D target image of a target object. Detail descriptions of determining the depth information can be found in step 310.
[0044] The first generation module 220 is configured to generate a multi-viewpoint image corresponding to the 2D target image based on the depth information using a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique. The multiviewpoint image includes a plurality of image frames corresponding to different shooting angles. Detail descriptions of generating the multi-viewpoint image can be found in step 320.
[0045] The second generation module 230 is configured to generate a 3D lenticular image of the target object based on the multi -viewpoint image. Detail descriptions of generating the 3D lenticular image can be found in step 330.
[0046] In some embodiments, the system 200 further includes a training module 240 and a monitoring module 250.
[0047] The training module 240 is configured to generate machine learning models such as a grating parameter determination model, an image simulation model, or the like. For example, the training module 240 may generate a training input and a training label for training the grating parameter determination model, and determine the grating parameter determination model based on the training input and the training label. Detailed descriptions of the grating parameter determination model, the image simulation model, an image transformation model, a quality problem monitoring model, and a depth information estimation model can be found in FIG. 3 and its related descriptions. Detailed descriptions of the algorithm selection model and the depth information estimation model can be found in FIG. 4 and its related descriptions.
[0048] The monitoring module 250 is configured to control a printing device to print the 3D lenticular image and to monitor a quality problem during a printing process. Detailed descriptions of printing the 3D lenticular image and monitoring the quality problem can be found in steps 340 and 350.
[0049] It should be noted that the above description of the system 200 and its modules is provided only for descriptive convenience, and does not limit the present disclosure to the scope of the cited embodiments. It is to be understood that for a person skilled in the art, after understanding the principle of the system, it may be possible to arbitrarily combine the modules or form a sub-system to be connected to the other modules without departing from the principle.In some embodiments, the determination module 210, the first generation module 220, the second generation module 230, the training module 240, and the monitoring module 250 disclosed in FIG. 2 may be different modules in a system, or different modules in different systems. The system 200 may include one or more additional components and / or one or more components of the system 200 described above may be omitted. Additionally or alternatively, two or more components of the system 200 may be integrated into a single component. A component of the system 200 may be implemented on two or more sub-components.
[0050] FIG. 3 is a flowchart illustrating an exemplary method for generating and producing a 3D lenticular image according to some embodiments of the present disclosure. In some embodiments, the process 300 may be executed by the processor 120 or the system 200. For example, the process 300 may be stored in the storage device 140 in the form of a program or instructions, and the process 300 may be implemented when the processor 120 or system 200 executes the instructions. The process 300 presented below is illustrative. In some embodiments, the process 300 may include one or more additional operations not described, and / or one or more operations discussed below may be omitted. Additionally, the order of operations of the process 300 illustrated in FIG. 3 and described below is not limiting.
[0051] Step 310, depth information of a 2D target image of a target object may be determined.
[0052] The 2D target image is a 2D image to be processed. The target object is an object (including a character, an object, a scene, etc.) in the 2D image to be processed. The 2D target image may be a real image captured using an image-capturing device (e.g., a camera) or a computer-generated virtual image.
[0053] The depth information is related to a distance between each point in the 2D target image and a reference point. The reference point is a point selected as a reference. For example, the reference point may be a location where a lens of the image-capturing device is located. As another example, when the 2D target image is a virtual image, the reference point may be a location where a virtual lens corresponding to the virtual image is located.
[0054] In some embodiments, the processor may utilize a depth information estimation algorithm to process the 2D target image to determine the depth information of the 2D target image. The depth information estimation algorithm may include a LeRes algorithm, a MiDaS algorithm, a Structure From Motion (SfM) algorithm, a ResNet algorithm, a ZoeDepth algorithm, a DPT (Dense Prediction Transformer) algorithm, a Depth Anything algorithm, a BTS (BTS-Net) algorithm, and a Monodepth2 algorithm, among others. In some embodiments, the processor may also process the 2D target image to determine the depth information of the 2D target image using a depth information determination model. The depth information determination model is a machine learning model, e.g., a Deep Neural Network (DNN) model.Detailed descriptions of the depth information determination model can be found in FIG. 4.
[0055] In some embodiments, the processor may utilize an algorithm selection model to process the 2D target image to select a target depth information estimation algorithm. The algorithm selection model may evaluate various factors such as image characteristics, computational efficiency, and accuracy requirements to determine the most suitable depth information estimation algorithm. Once the target depth information estimation algorithm is selected, the processor processes the 2D target image using the chosen algorithm to determine the depth information of the 2D target image. Detailed descriptions of the foregoing embodiments can be found in FIG. 4 and its related description.
[0056] Step 320, a multi-viewpoint image corresponding to the 2D target image may be generated using a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique based on the depth information. This step enables the creation of a more immersive and dynamic viewing experience by transforming 2D images into 3D images and generating intermediate viewpoints.
[0057] The 2D-3D image transformation algorithm is an algorithm for transforming a 2D image into a 3D image. For example, the 2D-3D image transformation algorithm includes a 3D-photo-inpainting algorithm, a stable video diffusion (SVD) algorithm, a Lucid Dreamer algorithm, a Depth-warping algorithm, a Facebook 3D photo algorithm, etc. In some embodiments, the 2D-3D image transformation algorithm is the 3D-photo-inpainting algorithm.
[0058] The dynamic viewpoint interpolation technique (DVIT) is an image interpolation algorithm that may be used to generate an image of a target viewpoint by interpolating images of other viewpoints.
[0059] The multi-viewpoint image is a sequence of frame images that includes a plurality of image frames corresponding to different shooting angles. In some embodiments, the shooting angles of the plurality of image frames are consistent with a preset shooting angle range and a preset shooting angle step. For example, if the preset shooting angle range is -15 to 15 degrees and the preset shooting angle step is 1.5 degrees, the corresponding shooting angles of the plurality of image frames are -15 degrees, -13.5 degrees, -12 degrees, , 12 degrees, 13.5 degrees, and 15 degrees.
[0060] In some embodiments, the processor may utilize the 2D-3D image transformation algorithm to generate an initial multi-viewpoint image based on the depth information and preset shooting parameters; and process at least a portion of a plurality of initial image frames using the dynamic viewpoint interpolation technique to generate the plurality of image frames of the multiviewpoint image.
[0061] The preset shooting parameters are desired shooting parameters of the image frames ofthe multi -viewpoint image. For example, the preset shooting parameters may include a preset shooting distance, a preset shooting angle range, a preset shooting angle step, or the like. The shooting distance is a distance between a lens corresponding to the image frames in the multiview point image and the target object (e.g., 2 meters). The shooting angle range is a range of the shooting angles of the image frames in the multi-viewpoint image (e.g., -15 degrees to 15 degrees, the span of the shooting angle range is 30 degrees). The shooting angle step is a change value between shooting angles corresponding to two neighboring image frames (e.g., 1 degree, 1.5 degrees, 2 degrees, etc.) in the multi-viewpoint image.
[0062] In some embodiments, the preset shooting parameters include at least the preset shooting angle range and the preset shooting angle step. In some embodiments, the preset shooting parameters may be a default value, a preset value, a user-input value, or the like.
[0063] The initial multi-viewpoint image is an initially determined multi-viewpoint image. The initial multi-viewpoint image includes a plurality of initial image frames corresponding to different initial shooting angles. In some embodiments, the processor may utilize the 3D- photo-inpainting algorithm or other 2D-3D image transformation algorithms to process the depth information and the preset shooting parameters to generate the initial multi-viewpoint image.
[0064] Due to certain errors in data simulation, initial shooting angles of some initial image frames obtained by the 2D-3D image transformation algorithm do not necessarily conform to the preset shooting parameters, so the processor may process the initial image frames using the dynamic viewpoint interpolation technique to obtain that image frames that conform to the preset shooting parameters. For example, based on the preset shooting parameters, the shooting angles of the image frames of the multi -viewpoint image may be 0 degrees, 1.5 degrees, and 3 degrees. If the initial shooting angles of the initial image frames obtained by the 2D-3D image transformation algorithm are 0 degrees, 1.4 degrees, and 3 degrees, the initial image frame with the initial shooting angle of 1.4 degrees does not conform to the preset shooting parameters.The processor may generate an image frame with a shooting angle of 1.5 degrees using the dynamic viewpoint interpolation technique based on the initial image frame with the initial shooting angle of 1.4 degrees and other initial image frames of its vicinity (e.g., initial image frames with initial shooting angles of 0 degrees and 3 degrees).
[0065] In some embodiments of the present disclosure, by combining the dynamic viewpoint interpolation technique with the 3D-photo-inpainting algorithm, the shooting angles of the image frames in the multi-viewpoint image are more consistent with the preset shooting parameters, so that the image frames may realize a smooth viewpoint transition, making the change in viewpoints more smooth and improving the viewing experience of a subsequently generated 3D lenticular image.
[0066] Step 330, a 3D lenticular image of the target object may be generated based on the multi-viewpoint image.
[0067] The 3D lenticular image is a flat image that produces a sense of three-dimensionality. The 3D lenticular image may be subsequently printed into an actual image, which can show different contents at different angles and can even achieve the effect of movement and visual deception.
[0068] In some embodiments, the processor may generate the 3D lenticular image of the target object based on the multi -viewpoint image using a 3D reconstruction algorithm.
[0069] In some embodiments, the processor may utilize a content-adaptive lenticular prints (CALP) algorithm to determine arrangement information of each pixel of the plurality of image frames in the 3D lenticular image based on the multi-viewpoint image; and generate the 3D lenticular image based on the arrangement information.
[0070] The arrangement information includes a location and a width of each pixel of each image frame of the multi-viewpoint image in the 3D lenticular image. For example, the multiviewpoint image includes 20 image frames, and the arrangement information includes a location and a width of each pixel of each of the 20 image frames in the 3D lenticular image. After determining the arrangement information, the processor may obtain the 3D lenticular image by arranging each pixel of the plurality of image frames based on arrangement information of each pixel.
[0071] The content-adaptive lenticular prints algorithm may determine arrangement information of an optimal lens based on an input light field. In some embodiments, the processor may determine the input light field based on pixel information of each pixel in the multi-viewpoint image and a desired visual effect; utilize the content-adaptive lenticular prints algorithm to process the input light field to determine the arrangement information of the optimal lens, then the arrangement information of the optimal lens is converted into the arrangement information of the pixels. Using the content-adaptive lenticular prints algorithm can obtain more accurate arrangement information and improve the accuracy of the generated 3D lenticular image.
[0072] In some embodiments, the processor may determine the arrangement information of pixels based on the arrangement information of the optimal lens through a formula N = denotes a serial number of a pixel in the multi -viewpoint image; kdenotes a lateral intercept of the pixel; I denotes a longitudinal intercept of the pixel; a denotes an angle of inclination of a lens with respect to a vertical axis; and X denotes a total number of pixels within grating lines of the lens; Ntotdenotes a total number of viewpoints in the multiviewpoint image, k and I are related to the arrangement information of pixels, and a and X arerelated to the arrangement information of the optimal lens.
[0073] In some embodiments, the processor may also utilize other algorithms (e.g., a content- adaptive parallax barriers (CAPB) algorithm) to process the multi-viewpoint image to determine the arrangement information of each pixel of the plurality of image frames of the 3D lenticular image.
[0074] In some embodiments, the processor may obtain a grating device parameter; and utilize the content-adaptive lenticular prints algorithm to determine the arrangement information of each pixel of the plurality of image frames of the 3D lenticular image based on the grating device parameter and the multi-viewpoint image.
[0075] The grating device parameter is a parameter related to a dimension of the grating. For example, the grating device parameter includes the grating's intercept, radian, thickness, etc. By taking the grating device parameter into account when determining the arrangement information, the determined arrangement information is more consistent with a grating device used in subsequent printing.
[0076] In some embodiments, if a grating to be used by a printing device has been determined, the grating may be measured by an optical device (e.g., a microscope device) to determine the grating device parameter. In some embodiments, if the grating to be used by the printing device has not been determined, the processor may utilize a grating parameter determining model to process category information of the target object and a target off-screen distance so as to obtain a recommended grating device parameter. Then, a grating with the recommended grating device parameter may be selected as the grating of the printing device for print the 3D lenticular image.
[0077] The grating parameter determination model is a machine learning model used to determine a grating device parameter. For example, the grating parameter determination model may include a random forest model.
[0078] In some embodiments, an input to the grating parameter determination model includes the category information of the target object and the target off-screen distance and an output includes the recommended grating device parameter. The category information is a type of the target object. For example, the category information may include characters, animals, cars, etc. The target off-screen distance is a desired off-screen distance of the 3D lenticular image. The off-screen distance measures the visual effect of the target object in the 3D lenticular image (e.g., an out-of-screen effect or an in-screen effect). For example, the target off-screen distance is that the target object in the 3D lenticular image presents an out-of-screen effect of 3 centimeters.
[0079] In some embodiments, the grating parameter determination model may be trained with a plurality of third training samples. Each third training sample includes a third training inputand a third training label. The third training input includes sample category information of a sample object and a sample target off-screen distance, and the third training label includes a sample grating device parameter. The third training input may be obtained based on historical data, and the third training label may be obtained by manual labeling. For example, the third training inputs are input into an initial grating parameter determination model, a value of a loss function may be determined based on the third training labels and an output of the initial grating parameter determination model, then parameters of the grating parameter determination model are iteratively updated based on the value of the loss function. When an iteration preset condition is satisfied, the model training is completed, and a trained grating parameter determination model is obtained. The iteration preset condition may be that the loss function converges, a number of iterations reach a threshold, etc.
[0080] In some embodiments of the present disclosure, determining a recommended grating device parameter based on the category information of the target object and the target off-screen distance can make a determined grating device parameter more suitable for the target object and achieve the desired visual effect (i.e., the target off-screen distance), and improve the accuracy in selecting a grating. In addition, the grating parameter determination model may excavate a deep relationship between various factors, improving the accuracy in determining a grating device parameter and selecting a grating.
[0081] In some embodiments, the processor may utilize an image simulation model to process the multi-viewpoint image to generate the 3D lenticular image.
[0082] The image simulation model is a machine learning model used to fuse a plurality of image frames of the multi -viewpoint image into the 3D lenticular image. For example, the image simulation model may include a Graph Neural Network (GNN) model.
[0083] In some embodiments, the image simulation model includes a feature extraction module and an image simulation module. The feature extraction module is configured to extract feature information of the multi-viewpoint image, for example, the feature extraction module may include a Convolutional Neural Network (CNN). An input of the feature extraction module includes the multi-viewpoint image and an output includes the feature information. The image simulation module is used to generate the 3D lenticular image based on the feature information. For example, the image simulation module may include a generator of a Generative Adversarial Network (GAN). An input of the image simulation module includes the feature information and an output includes the 3D lenticular image.
[0084] In some embodiments, the feature extraction module and the image simulation module may be obtained by training an initial image simulation model using fourth training samples. Each fourth training sample includes a fourth training input and a fourth training label, the fourthtraining input including a sample multi-viewpoint image and the fourth training label including a sample 3D lenticular image. The fourth training sample may be obtained from historical data. For example, a historical multi-viewpoint image in an image generation record is designated as the fourth training input, and a historical 3D lenticular image in the image generation record is designated as the fourth training label. In some embodiments, the historical 3D lenticular image designated as the fourth training label needs to satisfy a preset condition. Detail descriptions of the image generation record and the preset condition can be found in FIG. 4 and its related descriptions.
[0085] Merely by way of example, the initial image simulation model includes an initial feature extraction module, an initial image simulation module, and an initial discrimination module.The fourth training input may be input into the initial feature extraction module to obtain sample feature information; the initial image simulation module processes the sample feature information to obtain a simulated 3D lenticular image; and the initial discrimination module discriminates the authenticity of the simulated 3D lenticular image and the authenticity of the fourth training label. The initial image simulation model is iteratively updated based on a discrimination result and a difference between the simulated 3D lenticular image and the fourth training label. During a training process, the model parameters are continuously adjusted to make the simulated 3D lenticular image generated by the initial discrimination module as close as possible to the fourth training label, so as to confuse the initial discrimination module. At the end of the training, the feature extraction module and the image simulation module are used as the image simulation model.
[0086] In some embodiments of the present disclosure, machine learning techniques are utilized to learn a generation mechanism for the 3D lenticular image, which makes the determination of the 3D lenticular image more efficient and accurate. Meanwhile, using the image simulation model can improve the generation efficiency of the 3D lenticular image.
[0087] In some embodiments, the 3D lenticular image may be generated by performing a process 500 shown in FIG. 5.
[0088] Step 510, a candidate 3D lenticular image is generated based on the multi -viewpoint image.
[0089] Specifically, the processor may use an image determined based on the 3D reconstruction algorithm, the content-adaptive lenticular prints algorithm, or the image simulation model described above as the candidate 3D lenticular image. Merely by way of example, an output of the image simulation model is designated as the candidate 3D lenticular image.
[0090] Step 520, whether the candidate 3D lenticular image satisfies a preset condition is determined.
[0091] The preset condition is a condition related to the quality of a lenticular image. In some embodiments, the processor may determine whether the candidate 3D lenticular image satisfies the preset condition based on a target off-screen distance and an off-screen distance of the candidate 3D lenticular image. For example, the processor may determine a difference between the off-screen distance of the candidate 3D lenticular image and the target off-screen distance, and the candidate 3D lenticular image satisfies the preset condition if a determined difference is less than a first difference threshold. The first difference threshold may be a default value, a preset value, or the like. More information about the off-screen distance can be found in step 330 and its relevant description.
[0092] In response to determining that the candidate 3D lenticular image satisfies the preset condition, the processor may perform step 530 to use the candidate 3D lenticular image as the 3D lenticular image.
[0093] In response to determining that the candidate 3D lenticular image does not satisfy the preset condition, the processor may re-execute the step 510 to generate a new candidate 3D lenticular image. For example, the processor may adopt another generation algorithm to regenerate the candidate 3D lenticular image. Or, in response to determining that the candidate 3D lenticular image does not satisfy the preset condition, the processor may perform step 540 to generate the 3D lenticular image by processing the candidate 3D lenticular image.
[0094] In some embodiments, the processor may utilize an image transformation model to process the candidate 3D lenticular image, the target off-screen distance, and the off-screen distance of the candidate 3D lenticular image to generate the 3D lenticular image. The image transformation model is a machine learning model trained to transform a 3D lenticular image at a certain off-screen distance to a 3D lenticular image at another off-screen distance. For example, the image transformation model includes a Deep Neural Network (DNN) model. The input of the image transformation model includes the candidate 3D lenticular image, the target off-screen distance, and the off-screen distance of the candidate 3D lenticular image, and an output includes a 3D lenticular image corresponding to the target off-screen distance.
[0095] In some embodiments, the image transformation model may be trained by a plurality of fifth training samples. Each fifth training sample includes a fifth training input and a fifth training label, the fifth training input may include a first 3D lenticular image, a first off-screen distance corresponding to the first 3D lenticular image, and a second off-screen distance corresponding to a second 3D lenticular image distance; the fifth training label may include a second 3D lenticular image. The first 3D lenticular image and the second 3D lenticular image are lenticular images of a same sample object and correspond to different off-screen distances. The fifth training sample may be obtained based on historical data. The image transformationmodel is trained similarly as the grating parameter determination model described above.Using the image transformation model can quickly and accurately generate a 3D lenticular image corresponding to the target off-screen distance, improving the quality of the generated image.
[0096] In some embodiments of the present disclosure, the quality of the 3D lenticular image can be improved by evaluating the quality of the candidate 3D lenticular image.
[0097] In some embodiments of the present disclosure, the 2D-3D image transformation algorithm and the dynamic viewpoint interpolation technique are combined to process the depth information of the 2D target image for generating the multi-viewpoint image and the 3D lenticular image. Therefore, automatic and efficient production of the 3D lenticular image can be realized, and the quality of the 3D lenticular image can be improved.
[0098] In some embodiments, the process 300 further includes steps 340 and 350.
[0099] Step 340, a printing device is controlled to print the 3D lenticular image.
[0100] Step 350, a quality problem is monitored during a printing process.
[0101] As shown in FIG. 6, during the printing process, the processor may obtain a first realtime image 610 of a printing result of the printing device; the first real-time image 610 is processed utilizing a quality problem monitoring model 630 to monitor the quality problem during the printing process and obtain a quality monitoring result 640.
[0102] The first real-time image 610 is an image obtained by taking a picture of the printing result, which can reflect a real-time state of the printing result. The first real-time image 610 may include images corresponding to various time points during the printing process. The first image-capturing device may be mounted around the printing device and directed to the printing result.
[0103] The quality problem monitoring model 630 refers to a machine learning model used to monitor the quality problem during the printing process based on input data. For example, the quality problem monitoring model is an improved Recurrent Neural Network (RNN) combined with a Long Short-Term Memory (LSTM) model.
[0104] In some embodiments, an input of the quality problem monitoring model 630 includes the first real-time image 610 and an output includes the quality monitoring result 640. For example, the first real-time image 610 may be input to the quality problem monitoring model 630 in the form of time-series data. The quality monitoring result 640 reflects the quality of the printing result. For example, the quality monitoring result 640 indicates whether there is a quality problem with the printing result and a type of the quality problem (e.g., uneven inking, deviating colors, and other printing defects).
[0105] In some embodiments, the quality problem monitoring model 630 may be trained using a plurality of sixth training samples. Each sixth training sample includes a sixth training inputand a sixth training label, the sixth training input includes a sample image of a sample printing result, and the sixth training label includes a sample quality monitoring result of the sample image. The sixth training sample may be obtained based on historical data. The quality problem monitoring model 630 is trained in a similar manner to the grating parameter determination model described above.
[0106] In some embodiments of the present disclosure, the first real-time image is processed by the quality problem monitoring model to achieve real-time monitoring of a quality problem, which is conducive to improving printing quality.
[0107] In some embodiments, as shown in FIG. 6, the processor may further obtain a second real-time image 620 of the grating of the printing device, which is captured by second imagecapturing device; utilize the quality problem monitoring model 630 to process the first real-time image 610 and the second real-time image 620 for more comprehensive monitoring of quality problems during the printing process.
[0108] The second real-time image 620 is an image obtained by taking a picture of the grating of the printing device, which reflects a real-time state of the grating. The second real-time image 620 may include images corresponding to various time points during the printing process. The second image-capturing device may be mounted around the grating and directed to the grating. In some embodiments, the first image-capturing device and the second imagecapturing device may be the same type or different types of devices.
[0109] When an input into the quality problem monitoring model 630 includes the second realtime image, the quality monitoring result output can simultaneously reflect an operation status of the grating. For example, the quality monitoring result 640 includes whether there is a quality defect in the grating and a type of the quality defect (e.g. scratches, bubbles, uneven spacing, etc.).
[0110] In some embodiments of the present disclosure, by analyzing the first real-time image and / or the second real-time image, various quality problems during the printing process can be monitored in real time, and adjustments can be made during the printing process in a timely manner, so that the printing quality can be improved.[oni] It should be noted that the foregoing description of the process 300 is intended to be exemplary and illustrative only and does not limit the scope of application of the present disclosure. For a person skilled in the art, various corrections and changes can be made to the process 300 under the guidance of the present disclosure. However, these amendments and changes remain within the scope of the present disclosure.
[0112] FIG. 4 is a schematic diagram illustrating an exemplary process for determining depth information of a 2D target image according to some embodiments of the present disclosure.
[0113] In some embodiments, a processor may obtain an algorithm selection model 420; process a 2D target image 410 utilizing the algorithm selection model 420 to select a target depth information estimation algorithm 430 from candidate depth information estimation algorithms; determine depth information 440 of the 2D target image 410 using the target depth information estimation algorithm 430.
[0114] The candidate depth information estimation algorithms include various available candidate depth information estimation algorithms. For example, the candidate depth information estimation algorithms may include a LeRes algorithm, a Midas algorithm, an SFM algorithm, a ResNet algorithm, a Midas algorithm, a ZoeDepth algorithm, the depth information estimation model, or the like. The target depth information estimation algorithm 430 is selected from the candidate depth information estimation algorithms and is a depth information estimation algorithm suitable for processing the 2D target image 410.
[0115] The algorithm selection model 420 is a machine learning model used to select an appropriate depth information estimation algorithm for an input image based on a feature of the input image. For example, the algorithm selection model may include a Deep Neural Network (DNN) model. In some embodiments, the algorithm selection model is an adaptive depth perception network (ADPN).
[0116] In some embodiments, the algorithm selection model 420 may be trained with a plurality of first training samples. Each first training sample includes a first training input and a first training label. The first training input includes a sample 2D target image of a sample object, and the first training label includes a sample target depth information estimation algorithm. The first training input may be determined based on historical data. The first training label may be determined by manual labeling or automatically through data analysis. The algorithm selection model is trained in a manner similar as the training of a grating parameter determination model as described in FIG. 3.
[0117] In some embodiments of the present disclosure, selecting a target depth information estimation algorithm by utilizing a model rather than by humans improves the accuracy of the selected target depth information estimation algorithm and improves the determination efficiency by reducing human intervention.
[0118] In some embodiments, the first training input and the first training label may be determined as follows: obtaining an image generation record; determining whether a historical 3D lenticular image in the image generation record satisfies a preset condition; and in response to determining that the historical 3D lenticular image satisfies the preset condition, designating a historical 2D image in the image generation record as the first training input and historical depth information estimation algorithm in the image generation record as the first training label.
[0119] The image generation record is a historical record relating to a 2D image transformed into a 3D lenticular image. Each image generation record includes a historical 2D image, a historical depth information estimation algorithm, and a historical 3D lenticular image, and the historical 3D lenticular image is generated based on the historical 2D image and the historical depth information estimation algorithm. In some embodiments, the processor may obtain the image generation record by accessing a storage device, designating a 2D image in the image generation record as the historical 2D image, designating a corresponding 3D lenticular image as the historical 3D lenticular image, and designating a depth information estimation algorithm used in a transformation process as the historical depth information estimation algorithm.
[0120] In some embodiments, the processor may determine, based on an off-screen distance of the historical 3D lenticular image and a preset off-screen distance, whether the historical 3D lenticular image satisfies the preset condition. The off-screen distance may be configured to measure the visual effect of the target object in the 3D lenticular image. More information about the off-screen distance can be found in step 330 and its related description. For example, the processor may determine a difference between the off-screen distance of the historical 3D lenticular image and the preset off-screen distance. In response to determining that the difference is less than a second difference threshold, the processor may determine that the historical 3D lenticular image satisfies the preset condition. The second difference threshold may be a default value, a preset value, or the like.
[0121] In some embodiments of the present disclosure, a historical 2D image and a historical depth information estimation algorithm corresponding to a historical 3D lenticular image that satisfies a preset condition are used as a first training sample, which realizes the automated determination of the first training sample and reduces the time of manual labeling, henceforth improving training efficiency. In addition, the historical 3D lenticular image is evaluated based on its off-screen distance, and the historical depth information estimation algorithm is designated as the first training label only if the historical 3D lenticular image satisfies the present condition. Therefore, the accuracy of the determined first training label can be improved.
[0122] In some embodiments, the candidate depth information estimation algorithms include a depth information estimation model. The depth information estimation model is a machine learning model for determining depth information of the 2D target image. An input into the depth information estimation model includes the 2D target image, and an output includes the depth information of the 2D target image. For example, the depth information estimation model may include a U-network incorporating an attention mechanism. Using the U-net network incorporating an attention mechanism enables the depth information estimation model to automatically recognize important regions in the 2D target image and utilize morecomputational resources to determine depth information in these regions.
[0123] In some embodiments, the input to the depth information estimation model includes the 2D target image, and the output includes the depth information of the 2D target image. In some embodiments, the depth information estimation model may be trained with a plurality of second training samples. Each second training sample includes a second training input and a second training label. The second training input includes a sample 2D target image of a sample object, and the second training label includes sample depth information corresponding to the sample 2D target image. The second training input may be determined based on historical data, and the second training label may be determined using a depth information estimation algorithm selected by humans. The depth information estimation model is trained in a similar manner to the grating parameter determination model described in FIG. 3.
[0124] In some embodiments, the second training input and the second training label may be determined as follows: obtaining an image generation record; determining whether a historical 3D lenticular image in the image generation record satisfies a preset condition; in response to determining that the historical 3D lenticular image satisfies the preset condition, a historical 2D image in the image generation record is designated as the second training input, and historical depth information in the image generation record is designated as the second training label.
[0125] The image generation record is a historical record relating to a 2D image transformed into a 3D lenticular image. Each image generation record includes a historical 2D image, historical depth information of the historical 2D image, and a historical 3D lenticular image corresponding to the historical 2D image. The historical 3D lenticular image is generated based on the historical depth information.
[0126] In some embodiments, the processor may determine whether the historical 3D lenticular image satisfies the preset condition based on an off-screen distance of the historical 3D lenticular image and a preset off-screen distance. For example, the processor may determine a difference between the off-screen distance of the historical 3D lenticular image and the preset off-screen distance. In response to determining that the difference is less than a second difference threshold, the processor may determine that the historical 3D lenticular image satisfies the preset condition. The second difference threshold may be a default value, a preset value, or the like.
[0127] In some embodiments of the present disclosure, a historical 2D image and historical depth information corresponding to a historical 3D lenticular image that satisfies a preset condition are used as a second training sample, which realizes the automated determination of the second training sample and reduces the time of manual labeling, henceforth improving training efficiency. In addition, the historical 3D lenticular image is evaluated based on its offscreen distance, and the historical depth information is designated as the second training labelonly if the historical 3D lenticular image satisfies the present condition. Therefore, the accuracy of the determined second training label can be improved.
[0128] One or more embodiments of the present disclosure provide a device for generating and producing a 3D lenticular image. The device comprises at least one processor and at least one storage device. The at least one storage device is configured to store computer instructions and the at least one processor is configured to execute at least a portion of the computer instructions to implement a method for generating and producing a 3D lenticular image as described in any one of the above embodiments.
[0129] One or more embodiments of the present disclosure provide a computer-readable storage medium, the storage medium is configured to store computer instructions. When a computer executes the computer instructions in the storage medium, the computer performs a method for generating and producing a 3D lenticular image as described in any one of the above embodiments.
[0130] The basic concepts have been described above, and it will be apparent to those skilled in the art that the foregoing detailed disclosure is intended as an example only and does not constitute a limitation of the present disclosure. While not expressly stated herein, a person skilled in the art may make various modifications, improvements, and amendments to the present disclosure. Those types of modifications, improvements, and amendments are suggested in the present disclosure, so those types of modifications, improvements, and amendments remain within the spirit and scope of the exemplary embodiments of the present disclosure.
[0131] Also, the present disclosure uses specific words to describe embodiments of the present disclosure, such as "one embodiment", "an embodiment", and / or "some embodiment" means a feature, structure, or characteristic associated with at least one embodiment of the present disclosure. Accordingly, it should be emphasized and noted that "an embodiment" or "one embodiment", or "an alternative embodiment" in different locations in the present disclosure do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics of one or more embodiments of the present disclosure may be suitably combined.
[0132] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numerical letters, or the use of other names described herein are not intended to limit the order of the processes and methods of the present disclosure. While some embodiments of the invention that are currently considered useful are discussed in the foregoing disclosure by way of various examples, it should be appreciated that such details serve only illustrative purposes, and that additional claims are not limited to the disclosed embodiments, rather, the claims are intended to cover all amendments and equivalent combinations that are consistent with the substance and scope of the embodiments of the present disclosure. Forexample, although the implementation of various components described above may be embodied in a hardware device, it may also be implemented as a software only solution, e.g., an installation on an existing server or mobile device.
[0133] Similarly, it should be noted that in order to simplify the presentation of the present disclosure, and thereby aid in the understanding of one or more embodiments of the invention, the foregoing descriptions of embodiments of the present disclosure sometimes group multiple features together in a single embodiment, accompanying drawings, or in a description thereof. However, this method of disclosure does not imply that the objects of the present disclosure require more features than those mentioned in the claims. Rather, claimed subject matter may lie in less than all features of a single foregoing disclosed embodiment.
[0134] Some embodiments use numbers describing the number of components, attributes, and it should be understood that such numbers used in the description of embodiments are modified in some examples by the modifiers "approximately", "nearly", or "substantially". Unless otherwise noted, the terms "approximately," "nearly," or "substantially" indicate that a ±20% variation in the stated number is allowed. Correspondingly, in some embodiments, the numerical parameters used in the present disclosure and claims are approximations, which are subject to change depending on the desired characteristics of individual embodiments. In some embodiments, the numerical parameters should take into account the specified number of valid digits and employ general place-keeping. While the numerical domains and parameters used to confirm the breadth of their ranges in some embodiments of the present disclosure are approximations, in specific embodiments, such values are set to be as precise as possible within a feasible range.
[0135] For each patent, patent application, patent application disclosure, and other material cited in the present disclosure, such as articles, books, manuals, publications, documents, etc., the entire contents of which are hereby incorporated by reference into the present disclosure. Application history documents that are inconsistent with or conflict with the contents of the present disclosure are excluded, as are documents (currently or hereafter appended to this specification) that limit the broadest scope of the claims of the present disclosure. It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or use of terms in the materials appended to the present disclosure and those set forth herein, the descriptions, definitions and / or use of terms in the present disclosure shall prevail.
[0136] Finally, it should be understood that the embodiments described in the present disclosure are only used to illustrate the principles of the embodiments of the present disclosure. Other deformations may also fall within the scope of the present disclosure. As such, alternative configurations of embodiments of the present disclosure may be viewed as consistentwith the teachings of the present disclosure as an example, not as a limitation.Correspondingly, the embodiments of the present disclosure are not limited to the embodiments expressly presented and described herein.
Claims
WHAT IS CLAIMED IS:
1. A method for generating and producing a 3D lenticular image, comprising: determining depth information of a 2D target image of a target object; generating a multi-viewpoint image corresponding to the 2D target image using a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique based on the depth information, and the multi -viewpoint image including a plurality of image frames corresponding to different shooting angles; and generating a 3D lenticular image of the target object based on the multi -viewpoint image.
2. The method of claim 1, wherein the determining depth information of a 2D target image of a target object includes: obtaining an algorithm selection model, the algorithm selection model being a trained machine learning model; utilizing the algorithm selection model to process the 2D target image to select a target depth information estimation algorithm from candidate depth information estimation algorithms; and determining the depth information of the 2D target image by processing the 2D target image using the target depth information estimation algorithm.
3. The method of claim 2, wherein the algorithm selection model is an adaptive depth perception network.
4. The method of claim 2, wherein a first training input and a first training label of the algorithm selection model are determined by: obtaining an image generation record, the image generation record including a historical 2D image, a historical depth information estimation algorithm, and a historical 3D lenticular image, the historical 3D lenticular image being generated based on the historical 2D image and the historical depth information estimation algorithm; and determining whether the historical 3D lenticular image in the image generation record satisfies a preset condition; and in response to determining that the historical 3D lenticular image satisfies the preset condition, designating the historical 2D image in the image generation record as the first training input and the historical depth information estimation algorithm in the image generation record as the first training label.
5. The method of claim 2, wherein the candidate depth information estimation algorithms include a depth information estimation model, and a second training input and a second training label of the depth information estimation model are determined by: obtaining an image generation record, the image generation record including a historical 2D image, historical depth information of the historical 2D image, and a historical 3D lenticular image corresponding to the historical 2D image, the historical 3D lenticular image being generated based on the historical depth information; and determining whether the historical 3D lenticular image in the image generation record satisfies a preset condition; and in response to determining that the historical 3D lenticular image satisfies the preset condition, designating the historical 2D image in the image generation record as the second training input and the historical depth information in the image generation record as the second training label.
6. The method of claim 1, wherein the generating a multi -viewpoint image corresponding to the 2D target image using a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique based on the depth information includes: generating an initial multi-viewpoint image using the 2D-3D image transformation algorithm based on the depth information and preset shooting parameters, the initial multiviewpoint image including a plurality of initial image frames corresponding to different initial shooting angles, and the preset shooting parameters including at least a preset shooting angle range and a preset shooting angle step; and processing at least a portion of the plurality of initial image frames using the dynamic viewpoint interpolation technique to generate the plurality of image frames of the multiviewpoint image, wherein shooting angles of the plurality of image frames are consistent with the preset shooting angle range and the preset shooting angle step.
7. The method of claim 1, wherein the generating a 3D lenticular image of the target object based on the multi -viewpoint image includes: determining arrangement information of each pixel of the plurality of image frames of the 3D lenticular image using a content-adaptive lenticular prints algorithm based on the multiviewpoint image; and generating the 3D lenticular image based on the arrangement information.
8. The method of claim 7, wherein the determining arrangement information of each pixelof the plurality of image frames of the 3D lenticular image using a content-adaptive lenticular prints algorithm based on the multi-viewpoint image includes: obtaining a grating device parameter; and determining the arrangement information of each pixel of the plurality of image frames of the 3D lenticular image using the content-adaptive lenticular prints algorithm based on the grating device parameter and the multi-viewpoint image.
9. The method of claim 8, wherein the obtaining a grating device parameter includes: obtaining the grating device parameter by processing category information of the target object and a target off-screen distance using a grating parameter determination model, the grating parameter determination model being a trained machine learning model.
10. The method of claim 1, wherein the generating a 3D lenticular image of the target object based on the multi -viewpoint image includes: generating the 3D lenticular image by processing the multi-viewpoint image using an image simulation model, wherein the image simulation model is a trained machine learning model, the image simulation model includes a feature extraction module and an image simulation module, the feature extraction module is configured to extract feature information of the multiviewpoint image; and the image simulation module is configured to generate the 3D lenticular image based on the feature information.
11. The method of claim 10, wherein an output of the image simulation model is designated as a candidate 3D lenticular image, and the method further comprises: determining whether the candidate 3D lenticular image satisfies a preset condition based on a target off-screen distance and an off-screen distance of the candidate 3D lenticular image; in response to determining that the candidate 3D lenticular image satisfies the preset condition, using the candidate 3D lenticular picture as the 3D lenticular image; and in response to determining that the candidate 3D lenticular image does not satisfy the preset condition, processing the candidate 3D lenticular image to generate the 3D lenticular image.
12. The method of claim 11, wherein the processing the candidate 3D lenticular image to generate the 3D lenticular image includes:generating the 3D lenticular image by processing the candidate 3D lenticular image, the target off-screen distance, and the off-screen distance of the candidate 3D lenticular image using an image transformation model, the image transformation model being a trained machine learning model.
13. The method of claim 1, further comprising: controlling a printing device to print the 3D lenticular image; obtaining a first real-time image of a printing result of the printing device, the first real-time image being captured by a first image-capturing device during a printing process; and processing the first real-time image using a quality problem monitoring model to monitor a quality problem during the printing process, the quality problem monitoring model being a trained machine learning model.
14. The method of claim 13, wherein the printing device includes a grating, and the processing the first real-time image using a quality problem monitoring model to monitor a quality problem during the printing process includes: obtaining a second real-time image of the grating, the second real-time image being captured by a second image-capturing device during the printing process; and processing the first real-time image and the second real-time image using the quality problem monitoring model to monitor the quality problem during the printing process.
15. A system for generating and producing a 3D lenticular image, comprising: a determination module, configured to determine depth information of a 2D target image of a target object; a first generation module, configured to generate a multi-viewpoint image corresponding to the 2D target image using a 2D-3D image transformation algorithm and a dynamic viewpoint interpolation technique based on the depth information, and the multi-viewpoint image including a plurality of image frames corresponding to different shooting angles; and a second generation module, configured to generate a 3D lenticular image of the target object based on the multi -viewpoint image.
16. A device for generating and producing a 3D lenticular image, comprising at least one processor and at least one storage device, wherein the at least one storage device is configured to store computer instructions; and the at least one processor is configured to execute at least a portion of the computerinstructions to implement a method for generating and producing a 3D lenticular image of any one of claims 1 to 14.
17. A computer-readable storage medium, wherein the storage medium stores computer instructions, and when a computer executes the computer instructions in the storage medium, the computer performs a method for generating and producing a 3D lenticular image.
Citation Information
Patent Citations
Holographic image multi-angle processing conversion, display methods and devices
CN108495117A
Producing 3D images from captured 2d video
US20120242794A1
Method of converting 2d video to 3D video using machine learning
US20170085863A1
System and method for generating light field images
US20210321081A1
Cited By
Real estate 3D modeling precision evaluation method and system based on unmanned aerial vehicle oblique photography
CN122312925A