Video generation method and device, storage medium and computer program product

By acquiring the user's visual attribute information, the pixel parallax information of each frame in the 3D video is determined, and personalized target images and 3D videos are generated. This solves the problem of mismatch between 3D videos and user visual characteristics and improves the user's viewing experience.

CN121509768APending Publication Date: 2026-02-10MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511556483.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies do not take into account the user's visual attributes when generating 3D videos, resulting in a mismatch between the generated 3D videos and the user's visual characteristics, causing problems such as dizziness and nausea, and affecting the user's viewing experience.

Method used

By acquiring the user's visual attribute information, including interpupillary distance and focal length, the parallax information of each pixel in the video to be processed is determined based on this information, and personalized target images and 3D videos are generated.

Benefits of technology

It improves the user experience of watching 3D videos, ensures that the generated 3D videos are suitable for each user's visual characteristics, and reduces dizziness and ghosting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509768A_ABST
    Figure CN121509768A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video generation method. The method comprises the steps of obtaining a to-be-processed video and visual attribute information of a user; the visual attribute information represents the focusing capability of eyes of the user on light; on the basis of the visual attribute information, determining parallax information of pixels of each frame of image in the to-be-processed video; determining a plurality of target images based on the parallax information and each frame of image; the target three-dimensional video is determined based on the plurality of target images, so that the problem that the generated 3D video is not matched with the visual characteristics of the user when the 3D video is generated in the related technology is solved, and the watching experience of the user is improved. The embodiment of the invention further discloses video generation equipment, a storage medium and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, and in particular to a video generation method, device, storage medium and computer program product. BACKGROUND

[0002] In the field of three dimensions (3D) video generation, usually, the image collected by an image collection device is processed to obtain a two dimensional (2D) video, then the disparity information of each pixel of each frame image of the 2D video is calculated, and then each frame image of the 2D video is processed according to the disparity information to convert the 2D video into a 3D video. However, in the related art, the visual attribute information such as the interpupillary distance of a user is not considered when the disparity information of a pixel is calculated, which may result in that the generated 3D video does not match the visual characteristics of the user, and further result in that the user may feel dizzy, nausea and other phenomena when watching the 3D video, thereby affecting the user's viewing experience. SUMMARY

[0003] To solve the above technical problems, the embodiments of the present application provide a video generation method, device, storage medium and computer program product, which solve the problem that the generated 3D video does not match the visual characteristics of the user in the related art when generating a 3D video, thereby improving the user's viewing experience.

[0004] To achieve the above object, the technical scheme of the embodiments of the present application is as follows: A video generation method, the method comprising: obtaining a to-be-processed video and visual attribute information of a user; wherein the visual attribute information represents the focusing ability of the eyes of the user to light; determining the disparity information of a pixel of each frame image in the to-be-processed video based on the visual attribute information; determining a plurality of target images based on the disparity information and the each frame image; determining a target three-dimensional video based on the plurality of target images and the each frame image.

[0005] In the above scheme, the visual attribute information of the user is obtained, comprising: obtaining a face image of the user, and processing the face image by using a target detection model to obtain the interpupillary distance of the user; obtaining a reference three-dimensional video corresponding to each reference focal length; wherein the reference three-dimensional video is a three-dimensional video whose clear degree corresponding to each reference focal length is higher than a target degree; The user's viewing experience for each reference 3D video is obtained, and the user's focal length is determined from multiple reference focal lengths based on the viewing experience; wherein, the visual attribute information includes the interpupillary distance and the focal length.

[0006] The method in the above scheme further includes: Store the interpupillary distance and the user's focal length in the target database.

[0007] In the above scheme, determining the disparity information of pixels in each frame of the video to be processed based on the visual attribute information includes: The target value is determined based on the interpupillary distance and the focal length; Determine the target depth value corresponding to the pixel in each frame of the image; wherein, the target depth value is the depth value of the target region of the object in each frame of the image; the target region is the region corresponding to the pixel in the object; Based on the target numerical value and the target depth value, the disparity value of the pixel is determined; wherein, the disparity information includes the disparity value.

[0008] In the above scheme, determining the target depth value corresponding to the pixels of each frame image includes: Determine a target image with a subtitle region from multiple frames of images; wherein, the subtitle region is the region containing subtitles; The target depth value corresponding to the pixel of the target image is determined to be a target threshold. Based on the target depth estimation algorithm and other images, the target depth values ​​corresponding to the pixels in the other images are determined; wherein, the other images are images other than the target image in the multi-frame images.

[0009] In the above scheme, determining the target depth value corresponding to the pixels in the other images based on the target depth estimation algorithm and other images includes: The target depth estimation algorithm is used to process the other images to obtain the first depth value corresponding to the pixels of the other images; The target time series stabilization algorithm is used to correct multiple first depth values ​​to obtain multiple corrected depth values; Determine the second depth value and the third depth value from the plurality of corrected depth values; Based on the second depth value, the third depth value, and the plurality of corrected depth values, the target depth value corresponding to the pixels of the other images is determined.

[0010] In the above scheme, determining multiple target images based on the disparity information and each frame image includes: Determine the first position of the pixel in each frame of the image; A target reverse mapping algorithm is used to process multiple first positions and the disparity information to obtain multiple second positions; The plurality of target images are determined based on the plurality of second positions.

[0011] In the above scheme, determining the target 3D video based on the plurality of target images and each frame image includes: A target optimization algorithm is used to optimize each target image, resulting in multiple optimized images; The multiple optimized images and each frame image are processed to obtain the target 3D video.

[0012] A video generation device includes: a processor and a memory for storing a computer program capable of running on the processor; The processor is used to execute the steps of the above method when running a computer program.

[0013] A storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above method.

[0014] A computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described method.

[0015] The video generation method, device, storage medium, and computer program product provided in this application can acquire the video to be processed and visual attribute information characterizing the user's eye's ability to focus light. Based on the visual attribute information, it determines the disparity information of pixels in each frame of the video to be processed. Then, based on the disparity information and each frame, it determines multiple target images, and then determines a target 3D video based on the multiple target images. In this way, the disparity information of pixels in each frame of the video to be processed (i.e., 2D video) can be determined according to the user's visual attribute information. Then, based on the disparity information and each frame of the video to be processed, a target 3D video can be generated. That is, the user's visual attribute information is taken into account in the process of determining the disparity information of pixels, thereby solving the problem of mismatch between the generated 3D video and the user's visual characteristics in the related technology, thus improving the user's viewing experience. Attached Figure Description

[0016] Figure 1 A flowchart illustrating a video generation method provided in an embodiment of this application; Figure 2 A flowchart illustrating another video generation method provided in an embodiment of this application; Figure 3 A schematic diagram of the structure of a video generation device provided for an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a video generation device provided for an embodiment of this application. Detailed Implementation

[0017] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0018] It should be understood that the phrases "embodiments of this application" or "foreign embodiments" throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "embodiments of this application" or "in the foreign embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0019] Unless otherwise specified, any step in the embodiments of this application performed by the electronic device may be executed by the processor of the electronic device. It is also worth noting that the embodiments of this application do not limit the order in which the electronic device performs the following steps. Furthermore, the methods used to process data in different embodiments may be the same or different methods. It should also be noted that any step in the embodiments of this application can be executed independently by the electronic device; that is, when the electronic device performs any step in the following embodiments, it may not depend on the execution of other steps.

[0020] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0021] This application provides a method, referring to... Figure 1 As shown, the method may include the following steps: Step 101: Obtain the visual attribute information of the video to be processed and the user.

[0022] Among them, visual attribute information represents the user's eye's ability to focus light.

[0023] In this embodiment, the video to be processed is typically a two-dimensional video, i.e., a 2D video; the user's visual attribute information may include the user's interpupillary distance and focal length, etc. The video to be processed can be a movie, a TV series, or a documentary; the specific type is not specifically limited here.

[0024] In this embodiment of the application, a user's face image can be acquired and processed to obtain the user's interpupillary distance. Furthermore, the user's focal length can be determined from multiple reference focal lengths based on the user's viewing experience with multiple reference 3D videos.

[0025] It should be noted that interpupillary distances are usually different for different users, while focal lengths may be the same for different users.

[0026] Step 102: Based on visual attribute information, determine the disparity information of pixels in each frame of the video to be processed.

[0027] In this embodiment of the application, each frame of the video to be processed can be referred to as the left image, i.e., the image seen by the left eye; the disparity information of a pixel refers to the horizontal offset of each pixel. Specifically, the depth value corresponding to the pixel of each frame image can be determined, and the disparity information of the pixel can be determined based on the depth value, the user's interpupillary distance, and the user's focal length.

[0028] It should be noted that the parallax information of each frame of an image is different for different users.

[0029] Currently, traditional 2D video to 3D video conversion methods typically use fixed parallax information. This means that the parallax information of pixels is the same for all users, resulting in a consistent 3D effect for all users. However, this 3D effect may not be suitable for every user. If a user is not comfortable with this 3D effect, problems such as ghosting and dizziness may occur. In this embodiment, by determining the parallax information corresponding to each user based on their visual attribute information, the 3D effect seen by the user can be made perfectly suitable for them, thereby improving the user's video viewing experience.

[0030] Step 103: Based on the parallax information and each frame of the image, determine multiple target images.

[0031] In this embodiment, the target image can be referred to as the right image, i.e., the image seen by the right eye. Specifically, the first position of a pixel in each frame of the image can be determined, and the second position of the pixel can be determined based on the disparity information and the first position. Then, multiple target images can be determined based on multiple second positions. The position of a pixel in the image can be represented by coordinates.

[0032] It should be noted that one left image corresponds to one right image, meaning that the number of left images and the number of right images are the same.

[0033] Step 104: Determine the target 3D video based on multiple target images and each frame image.

[0034] In this embodiment of the application, the target 3D video can refer to a 3D video. Specifically, after obtaining multiple target images, i.e., multiple right images, multiple left images and multiple right images can be processed by a video editing component to generate the target 3D video.

[0035] In this embodiment, the various steps are closely logically related. First, by acquiring the user's visual attribute information (such as interpupillary distance and focal length), a personalized basis is provided for subsequent parallax information calculation. Next, based on this information and each frame of the video to be processed, a parallax value suitable for the user is calculated. Then, a target image (i.e., the right image) is generated using the parallax information. Finally, the target image and multiple frames of the video to be processed are combined to generate a target 3D video. The entire process achieves efficient conversion from the original 2D video to a personalized 3D video, enhancing the immersive experience for users watching 3D videos.

[0036] The video generation method provided in this application can determine the disparity information of pixels in each frame of the video to be processed (i.e., 2D video) based on the user's visual attribute information, and then generate a target 3D video based on the disparity information and each frame of the video to be processed. That is, the user's visual attribute information is taken into account in the process of determining the disparity information of pixels, thereby solving the problem in related technologies where the generated 3D video does not match the user's visual characteristics, thus improving the user's viewing experience.

[0037] Based on the foregoing embodiments, embodiments of this application provide a video generation method, which can be applied to a video generation device, with reference to... Figure 2 As shown, the method may include the following steps: Step 201: The video generation device acquires the video to be processed.

[0038] In this embodiment, the video to be processed can be various types of two-dimensional video. Specifically, the video to be processed can be obtained directly from the target database of the video generation device.

[0039] Step 202: The video generation device acquires the user's face image and processes the face image using an object detection model to obtain the user's interpupillary distance.

[0040] In this embodiment, the user's interpupillary distance (IPD) refers to the horizontal distance between the centers of the pupils of both eyes; the target detection model can be a convolutional neural network model, or other models capable of determining the user's IPD, without specific limitations. Specifically, the user's face image can be directly input into the target detection model as an input parameter to obtain the user's IPD.

[0041] It should be noted that the user's eyes must be fully visible in the user's face image and cannot be obscured.

[0042] In one feasible approach, the facial image can be an image taken by a user holding a standard-sized object. This standard-sized object can be an ID card or bank card, etc. When taking the facial image, the object can be placed parallel to the user's eyes (covering the forehead) or parallel to the user's eyes (covering the nose). Specifically, the facial image can be processed using visual technology measurement algorithms to obtain the distance *d* between the centers of the user's two pupils, and the length *scale* of each pixel can be determined based on the size information of the bank card or ID card. Then, the distance and length can be calculated to obtain the user's interpupillary distance *B* = *d*scale.

[0043] Step 203: The video generation device acquires the reference 3D video corresponding to each reference focal length.

[0044] Among them, the reference 3D video is a 3D video with a higher level of clarity than the target at each reference focal length.

[0045] In this embodiment, the reference focal length can be multiple focal lengths ranging from 15mm to 30mm, increasing in 1mm increments, i.e., the reference focal lengths include 15mm, 16mm, 17mm...29mm, 30mm; the reference 3D video with a higher clarity than the target level can refer to a 3D video with higher clarity than expected for a reference user at each reference focal length. Preferably, the reference 3D video can refer to the 3D video with the highest viewing clarity selected by the user at each focal length from multiple 3D videos.

[0046] For example: If user A's focal length is 20mm and user A reports that 3D video A1 has the highest clarity, then 3D video A1 is the reference 3D video corresponding to the 20mm focal length; if user B's focal length is 18mm and user B reports that 3D video B3 has the highest clarity, then 3D video B3 is the reference 3D video corresponding to the 18mm focal length.

[0047] In one feasible approach, an initial two-dimensional video can be acquired, and the disparity information of the pixels in each frame of the initial two-dimensional video can be determined sequentially based on each reference focal length and the interpupillary distance of the reference user corresponding to each reference focal length. Based on this disparity information and each frame of the initial two-dimensional video, a three-dimensional video, i.e., a reference three-dimensional video, can be generated.

[0048] It should be noted that one reference focal length can correspond to multiple reference 3D videos.

[0049] Step 204: The video generation device acquires the user's viewing experience for each reference 3D video and determines the user's focal length from multiple reference focal lengths based on the viewing experience.

[0050] The visual attribute information includes interpupillary distance and focal length.

[0051] In this embodiment, focal length refers to the focal distance at which light converges on the retina when the eye is focused on an object; the viewing experience may include the clarity of the video image, whether the user feels dizzy, and whether ghosting occurs. Specifically, the 3D video with the optimal viewing experience can be determined from multiple reference 3D videos, and the reference focal length corresponding to that 3D video can be determined as the user's focal length.

[0052] For example, if users report that the viewing experience of reference 3D video B2 is the best, then the reference focal length corresponding to reference 3D video B3 will be determined as the user's focal length.

[0053] It should be noted that if a user does not select a single reference video with the best viewing experience, for example, if the user reports that the viewing experience of multiple reference 3D videos is similar, the user can perform a secondary screening of multiple reference 3D videos to determine the reference 3D video with the best viewing experience.

[0054] In this embodiment, each reference 3D video can be shown to the user in sequence, and the user's viewing experience of each reference 3D video can be recorded, such as whether they feel dizzy, whether the picture is clear, and whether there is obvious parallax discomfort. The user's focal length can be determined based on the viewing experience. This method of determining the user's focal length through an interactive feedback mechanism can make the determined pixel parallax information closer to the user's actual perception, thereby ensuring that the generated 3D video is both in line with scientific principles and meets the user's needs, thus significantly improving the user's viewing experience.

[0055] In other embodiments of this application, after determining the user's interpupillary distance and focal length, the following steps can be performed: A1. Store the user's interpupillary distance and focal length in the target database.

[0056] In this embodiment, the target database can be a local database of the video generation device, a cloud database, or a distributed database. The target database may contain fields such as the user's identifier (e.g., ID card number), the user's interpupillary distance, the user's focal length, and the creation time, for subsequent quick retrieval.

[0057] In this embodiment of the application, by persistently storing the user's visual attribute information, namely the user's interpupillary distance and focal length, reliable data support can be provided for subsequent personalized 3D video generation, while reducing the operational burden caused by repeated acquisition and improving the generation efficiency of 3D video.

[0058] Step 205: The video generation device determines the target value based on the interpupillary distance and focal length.

[0059] In this embodiment of the application, the user's interpupillary distance B and focal length F can be multiplied to obtain the target value S1=B*F.

[0060] Step 206: The video generation device determines the target depth value corresponding to the pixels of each frame of the image.

[0061] Here, the target depth value is the depth value of the target region of the object in each frame of the image; the target region is the region corresponding to the pixels in the object.

[0062] In the embodiments of this application, the object in each frame of the image can refer to a real object such as a flower, a dog, or a building in the image, the target area can refer to a part of the object, such as the dog's eye, and the depth value of the target area can refer to the distance between the target area and the camera when the image acquisition device (e.g., a camera) takes the image.

[0063] It should be noted that there is a one-to-one correspondence between pixels and target areas. That is, one pixel corresponds to one target area, and different pixels correspond to different target areas.

[0064] In this embodiment of the application, a target image with a subtitle region can be determined from multiple frames of the video to be processed, and the target depth value corresponding to the pixels of the target image and the target depth value corresponding to the pixels of other images in the multiple frames besides the target image can be determined respectively.

[0065] In the embodiments of this application, step 206 can be implemented by steps 206a to 206c.

[0066] Step 206a: The video generating device determines the target image with the subtitle region from the multi-frame images.

[0067] The subtitle area is the area containing subtitles.

[0068] In the embodiments of this application, subtitles are non-image content displayed in the video in the form of text. In one possible implementation, subtitles are usually text content, and the subtitle area refers to the area in the image that contains text content.

[0069] It should be noted that the subtitle area is usually embedded in the image.

[0070] In this embodiment, a Differentiable Binarization Network (DBNet) algorithm can be used to detect each frame of the video to be processed to determine whether there is a subtitle region in each frame. If there is no subtitle region in a certain frame, then the frame is not the target image. If there is a subtitle region in a certain frame, then the frame can be segmented by an image segmentation model to extract the target image with the subtitle region from the frame.

[0071] It should be noted that if a frame contains a target region, the target image can be a part of that frame or the entire image (e.g., the subtitle region covers the entire image).

[0072] In the embodiments of this application, there are many algorithms that can be used to detect whether there is a subtitle region in each frame of the image, and no specific limitation is made here.

[0073] Step 206b: The video generating device determines the target depth value corresponding to the pixel of the target image as the target threshold.

[0074] In this embodiment of the application, since the subtitle area is usually embedded in the image during the video generation process, the area itself does not have the feature of depth value. In this case, the target depth value corresponding to the pixel of the target image can be set to a fixed value, namely the target threshold.

[0075] In one feasible approach, to ensure that the subtitles float at the forefront of the image, the target threshold can be set to 0 to ensure that the subtitle area is always clearly visible and not obscured by non-subtitle areas.

[0076] It should be noted that the target threshold is preset according to the user's needs, and the target threshold can also be different for different videos.

[0077] In this embodiment, the depth values ​​corresponding to the pixels in the subtitle area are specially processed to ensure the naturalness and consistency of the 3D effect. The setting of the target threshold can ensure that the subtitle area has the same depth representation in all images, thereby improving the viewing experience of the generated 3D video.

[0078] Step 206c: The video generating device determines the target depth value corresponding to the pixels of other images based on the target depth estimation algorithm and other images.

[0079] Other images are those other than the target image in a multi-frame image set.

[0080] In the embodiments of this application, the target depth estimation algorithm may refer to the DepthAnything (DA) algorithm, or other monocular depth estimation algorithms such as the Monocular DepthEstimation via Mutual Information Disentangling and Aggregation Strategy (MIDAS) algorithm based on deep neural networks.

[0081] In the embodiments of this application, other images may refer to images in a multi-frame image that do not have a caption area, that is, images other than the target image.

[0082] In this embodiment of the application, step 206c can be implemented through steps 206c1-206c4.

[0083] Step 206c1: The video generation device uses a target depth estimation algorithm to process other images to obtain the first depth value corresponding to the pixels of the other images.

[0084] In this embodiment of the application, there may be multiple other images. Specifically, a target depth estimation algorithm can be used to process each other image to obtain a depth map corresponding to each other image. Then, the first depth value corresponding to each pixel can be directly obtained from the obtained depth map.

[0085] It should be noted that other images may include images from multiple frames that do not have a subtitle area at all, as well as parts of a frame other than the target image that has a subtitle area.

[0086] Step 206c2: The video generation device uses a target temporal stabilization algorithm to correct multiple first depth values, thereby obtaining multiple corrected depth values.

[0087] In this embodiment of the application, the target temporal stabilization algorithm may refer to a depth temporal stabilization algorithm. Specifically, in order to avoid the problem of inter-frame depth flicker, after obtaining multiple first depth values, a depth temporal stabilization algorithm can be used to correct each first depth value to obtain multiple corrected depth values.

[0088] In the embodiments of this application, the depth estimation algorithm usually processes a single frame image sequentially, which can cause the depth value of a certain pixel to change abruptly between two adjacent frames, resulting in "flickering" or "jittering" in the user's vision. However, by correcting each first depth value through the depth temporal stabilization algorithm, the abnormal jump in the depth value corresponding to the pixel can be eliminated, so that the corrected depth value is more in line with the changing trend of the actual scene, which helps to improve the adaptability of the generated 3D video to the user's visual characteristics.

[0089] Step 206c3: The video generating device determines a second depth value and a third depth value from a plurality of corrected depth values.

[0090] In this embodiment, the largest second depth value and the smallest third depth value can be determined from a plurality of corrected depth values. That is, the second depth value is the largest depth value among the plurality of corrected depth values; and the third depth value is the smallest depth value among the plurality of corrected depth values.

[0091] It should be noted that the first, second, and third depth values ​​are all absolute depth values.

[0092] Step 206c4: The video generating device determines the target depth value corresponding to the pixels of other images based on the second depth value, the third depth value, and multiple corrected depth values.

[0093] In the embodiments of this application, the difference between the second depth value and the third depth value can be determined according to the following formula (1).

[0094] Formula (1) Where MaxDisp represents the difference between the second and third depth values; Zmax represents the second depth value; and Zmin represents the third depth value.

[0095] Then, the relative depth value corresponding to each pixel can be obtained by performing the following formula (2) on each corrected depth value, the second depth value and the third depth value.

[0096] Formula (2) Where dz represents the difference between the second and third depth values; Zmax represents the second depth value; Zmin represents the third depth value; and z represents the corrected depth value.

[0097] Furthermore, the difference and the relative depth value corresponding to each pixel can be calculated according to the following formula (3) to obtain the target depth value corresponding to the pixels of other images.

[0098] Formula (3) Where z' represents the difference between the second and third depth values; dz represents the relative depth values ​​corresponding to pixels in other images; and MaxDisp represents the difference.

[0099] In this embodiment, the target image with subtitle regions can be identified from multiple frames of images first, and the depth value corresponding to the pixels of the target image can be set to a fixed value. Then, a depth estimation algorithm can be used to estimate the depth of the images in non-subtitle regions to obtain the depth values ​​corresponding to the pixels of other images. After that, the disparity information of the pixels can be determined based on the obtained depth values, and a 3D video can be generated based on the disparity information and each frame of the video to be processed. In this way, not only can the accuracy of the generated 3D video and its adaptability to the user's visual characteristics be improved, but the efficiency of the generated 3D video can also be improved, thereby reducing manual intervention, so that the method can be applied to large-scale 2D video to 3D video conversion scenarios.

[0100] Step 207: The video generating device determines the disparity value of the pixel based on the target value and the target depth value.

[0101] The disparity information includes the disparity value.

[0102] In this embodiment, the disparity value of each pixel in the video to be processed can be obtained by directly dividing the target value and the target depth value using the following formula: D = S1 / S2. Here, D represents the pixel disparity value; S1 represents the target value; and S2 represents the target depth value.

[0103] Step 208: The video generating device determines the first position of the pixel in each frame of the image.

[0104] In the embodiments of this application, the first position may refer to the first coordinate of each pixel in each frame of the image. This coordinate may be a coordinate in the image coordinate system.

[0105] In this embodiment, the first coordinates of a pixel in each frame of an image can be determined using existing image editing tools. It should be noted that the process of determining the first coordinates is existing technology and will not be described in detail here.

[0106] Step 209: The video generation device uses a target reverse mapping algorithm to process multiple first positions and disparity information to obtain multiple second positions.

[0107] In this embodiment, the second position may refer to the second coordinates of a pixel in the target image. Specifically, a reverse mapping algorithm can be used to process the first position (i.e., the first coordinates) of each pixel in the image and the disparity value of each pixel to obtain the second position (i.e., the second coordinates) of each pixel.

[0108] It should be noted that determining the second position of a pixel is equivalent to horizontally shifting the pixel from the first position to the second position.

[0109] Step 210: The video generation device determines multiple target images based on multiple second locations.

[0110] In this embodiment of the application, an initial blank image can be obtained. Then, multiple pixels are sequentially filled into the initial blank image according to the calculated multiple second positions. After that, the color value of each pixel of each frame image (i.e., the left image) in the video to be processed can be assigned to the corresponding pixel in the initial blank image, thereby obtaining multiple target images, i.e., multiple right images.

[0111] In this embodiment of the application, the above steps can automatically generate high-quality target images without human intervention, while taking into account individual differences of users, which not only improves processing efficiency but also significantly reduces costs.

[0112] Step 211: The video generation device uses a target optimization algorithm to optimize each target image to obtain multiple optimized images.

[0113] In this embodiment, the target optimization algorithm may include an occlusion repair algorithm and a super-resolution algorithm. Specifically, an occlusion repair algorithm can be used first to process each target image to fill in the hole areas where information is lost due to changes in viewing angle, thereby ensuring that the hole area maintains the same color and other information as the surrounding area, thus obtaining multiple repaired images. Then, a super-resolution algorithm can be used to process each repaired image to optimize the resolution of each repaired image, i.e., improve the resolution of each repaired image, thus obtaining multiple optimized images.

[0114] It should be noted that the occlusion repair algorithm can be the ProPainter algorithm.

[0115] In this embodiment of the application, by introducing target optimization algorithms (including occlusion repair algorithms and super-resolution algorithms), the original target image can be optimized before generating the target 3D video, so as to avoid the problems of difficulty in occlusion repair and user visual discomfort caused by poor target image quality, thereby improving the viewing experience and accuracy of the final generated target 3D video.

[0116] Step 212: The video generation device processes multiple optimized images and each frame to obtain the target 3D video.

[0117] In this embodiment, existing video editing tools can be used to process multiple optimized images and multiple frames of the video to be processed to generate a target 3D video. Different users will have different target 3D videos.

[0118] In one feasible approach, the video editing tool could be Adobe Premiere Pro, StereoPhotoMaker, or similar software.

[0119] In this embodiment of the application, the precise parallax of the pixels of each frame in the video to be processed can be calculated using a determined user interpupillary distance B and user focal length F. Then, a 3D video can be generated based on the precise parallax. In this way, personalized 3D videos can be generated for different users so that users can achieve the best viewing experience when watching 3D videos.

[0120] It should be noted that the descriptions of the same steps and contents as in other embodiments in this embodiment can be found in the descriptions in other embodiments, and will not be repeated here.

[0121] The video generation method provided in this application can determine the disparity information of pixels in each frame of the video to be processed (i.e., 2D video) based on the user's visual attribute information, and then generate a target 3D video based on the disparity information and each frame of the video to be processed. That is, the user's visual attribute information is taken into account in the process of determining the disparity information of pixels, thereby solving the problem in related technologies where the generated 3D video does not match the user's visual characteristics, thus improving the user's viewing experience.

[0122] Based on the foregoing embodiments, embodiments of this application provide a video generation apparatus that can be applied to... Figure 1 and Figure 2 In the corresponding embodiment of the video generation method, refer to Figure 3 As shown, the video generation device 3 may include: an acquisition unit 31, a first determination unit 32, a second determination unit 33, and a third determination unit 334, wherein: The acquisition unit 31 is used to acquire the video to be processed and the user's visual attribute information; wherein, the visual attribute information represents the user's eye's ability to focus light; The first determining unit 32 is used to determine the disparity information of pixels in each frame of the video to be processed based on visual attribute information. The second determining unit 33 is used to determine multiple target images based on disparity information and each frame image; The third determining unit 34 is used to determine the target 3D video based on multiple target images and each frame image.

[0123] In other embodiments of this application, the acquisition unit 31 is further configured to perform the following steps: The system acquires the user's facial image and processes it using an object detection model to obtain the user's interpupillary distance. Obtain the reference 3D video corresponding to each reference focal length; wherein, the reference 3D video is a 3D video with a higher degree of clarity than the target for each reference focal length; The system acquires the user's viewing experience for each reference 3D video and determines the user's focal length from multiple reference focal lengths based on the viewing experience; the visual attribute information includes interpupillary distance and focal length.

[0124] In other embodiments of this application, the acquisition unit 31 is further configured to perform the following steps: Store the interpupillary distance and the user's focal length in the target database.

[0125] In other embodiments of this application, the first determining unit 32 is further configured to perform the following steps: Determine the target value based on interpupillary distance and focal length; Determine the target depth value corresponding to each pixel in each frame of the image; where the target depth value is the depth value of the target region of the object in each frame of the image; the target region is the region corresponding to the pixel in the object; Based on the target numerical value and the target depth value, the disparity value of the pixel is determined; whereby the disparity information includes the disparity value.

[0126] In other embodiments of this application, the first determining unit 32 is further configured to perform the following steps: Identify the target image with a subtitle region from multiple frames of images; where the subtitle region is the area containing subtitles. The target depth value corresponding to the pixel in the target image is determined as the target threshold. Based on the target depth estimation algorithm and other images, determine the target depth values ​​corresponding to the pixels in the other images; where other images are images other than the target image in a multi-frame image set.

[0127] In other embodiments of this application, the first determining unit 32 is further configured to perform the following steps: The target depth estimation algorithm is used to process other images to obtain the first depth value corresponding to the pixels in the other images; The target time series stabilization algorithm is used to correct multiple first depth values ​​to obtain multiple corrected depth values; The second and third depth values ​​are determined from multiple corrected depth values; Based on the second depth value, the third depth value, and multiple corrected depth values, the target depth value corresponding to the pixels in other images is determined.

[0128] In other embodiments of this application, the second determining unit 33 is further configured to perform the following steps: Determine the first position of each pixel in each frame of the image; A target reverse mapping algorithm is used to process multiple first positions and disparity information to obtain multiple second positions; Multiple target images are determined based on multiple second locations.

[0129] In other embodiments of this application, the third determining unit 34 is further configured to perform the following steps: A target optimization algorithm is used to optimize each target image, resulting in multiple optimized images; Multiple optimized images and each frame of the image are processed to obtain the target 3D video.

[0130] It should be noted that the specific implementation process of the steps performed by each unit in the embodiments of this application can be referred to Figure 1 and Figure 2 The implementation process of the video generation method provided in the corresponding embodiment will not be described in detail here.

[0131] The video generation apparatus provided in the embodiments of this application can determine the disparity information of pixels in each frame of the video to be processed (i.e., 2D video) based on the user's visual attribute information, and then generate a target 3D video based on the disparity information and each frame of the video to be processed. That is, the user's visual attribute information is taken into account in the process of determining the disparity information of pixels, thereby solving the problem in the related technology that the generated 3D video does not match the user's visual characteristics, thereby improving the user's viewing experience.

[0132] Based on the foregoing embodiments, embodiments of this application provide a video generation device that can be applied to... Figure 1 and Figure 2 In the corresponding embodiment of the video generation method, refer to Figure 4 As shown, the video generation device 4 may include: a processor 41, a memory 42, and a communication bus 43, wherein: Communication bus 43 is used to realize the communication connection between processor 41 and memory 42; The processor 41 is used to execute the video generation program in the memory 42 to perform the following steps: Obtain the visual attribute information of the video to be processed and the user; whereby the visual attribute information represents the user's eye's ability to focus light; Based on visual attribute information, determine the disparity information of pixels in each frame of the video to be processed; Based on parallax information and each frame of image, multiple target images are determined; Based on multiple target images and each frame of the image, the target 3D video is determined.

[0133] In other embodiments of this application, processor 41 is used to execute a video generation program in memory 42 to perform the following steps: The system acquires the user's facial image and processes it using an object detection model to obtain the user's interpupillary distance. Obtain the reference 3D video corresponding to each reference focal length; wherein, the reference 3D video is a 3D video with a higher degree of clarity than the target for each reference focal length; The system acquires the user's viewing experience for each reference 3D video and determines the user's focal length from multiple reference focal lengths based on the viewing experience; the visual attribute information includes interpupillary distance and focal length.

[0134] In other embodiments of this application, processor 41 is used to execute a video generation program in memory 42 to perform the following steps: Store the interpupillary distance and the user's focal length in the target database.

[0135] In other embodiments of this application, processor 41 is used to execute a video generation program in memory 42 to perform the following steps: Determine the target value based on interpupillary distance and focal length; Determine the target depth value corresponding to each pixel in each frame of the image; where the target depth value is the depth value of the target region of the object in each frame of the image; the target region is the region corresponding to the pixel in the object; Based on the target numerical value and the target depth value, the disparity value of the pixel is determined; whereby the disparity information includes the disparity value.

[0136] In other embodiments of this application, processor 41 is used to execute a video generation program in memory 42 to perform the following steps: Identify the target image with a subtitle region from multiple frames of images; where the subtitle region is the area containing subtitles. The target depth value corresponding to the pixel in the target image is determined as the target threshold. Based on the target depth estimation algorithm and other images, determine the target depth values ​​corresponding to the pixels in the other images; where other images are images other than the target image in a multi-frame image set.

[0137] In other embodiments of this application, processor 41 is used to execute a video generation program in memory 42 to perform the following steps: The target depth estimation algorithm is used to process other images to obtain the first depth value corresponding to the pixels in the other images; The target time series stabilization algorithm is used to correct multiple first depth values ​​to obtain multiple corrected depth values; The second and third depth values ​​are determined from multiple corrected depth values; Based on the second depth value, the third depth value, and multiple corrected depth values, the target depth value corresponding to the pixels in other images is determined.

[0138] In other embodiments of this application, processor 41 is used to execute a video generation program in memory 42 to perform the following steps: Determine the first position of each pixel in each frame of the image; A target reverse mapping algorithm is used to process multiple first positions and disparity information to obtain multiple second positions; Multiple target images are determined based on multiple second locations.

[0139] In other embodiments of this application, processor 41 is used to execute a video generation program in memory 42 to perform the following steps: A target optimization algorithm is used to optimize each target image, resulting in multiple optimized images; Multiple optimized images and each frame of the image are processed to obtain the target 3D video.

[0140] It should be noted that a detailed description of the steps performed by the processor can be found in [reference needed]. Figure 1 and Figure 2 The video generation method provided in the corresponding embodiments will not be described in detail here.

[0141] The video generation device provided in the embodiments of this application can determine the disparity information of pixels in each frame of the video to be processed (i.e., 2D video) based on the user's visual attribute information, and then generate a target 3D video based on the disparity information and each frame of the video to be processed. That is, the user's visual attribute information is taken into account in the process of determining the disparity information of pixels, thereby solving the problem in the related technology that the generated 3D video does not match the user's visual characteristics, thereby improving the user's viewing experience.

[0142] Based on the foregoing embodiments, embodiments of this application provide a storage medium storing a computer program, which is implemented when executed by a processor. Figure 1 and 2 The corresponding embodiment provides the steps of the video generation method.

[0143] Based on the foregoing embodiments, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements... Figure 1and 2 The corresponding embodiment provides the steps of the video generation method.

[0144] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0145] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0146] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0147] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes ​ The steps of the function specified in one or more boxes.

[0148] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A video generation method, characterized in that, The method includes: The visual attribute information of the user and the video to be processed is obtained; wherein the visual attribute information characterizes the user's eye's ability to focus light. Based on the visual attribute information, the disparity information of pixels in each frame of the video to be processed is determined; Based on the disparity information and each frame of the image, multiple target images are determined; Based on the multiple target images and each frame image, a target 3D video is determined.

2. The method according to claim 1, characterized in that, The acquisition of the user's visual attribute information includes: The user's facial image is acquired, and the facial image is processed using an object detection model to obtain the user's interpupillary distance; Obtain a reference 3D video corresponding to each reference focal length; wherein, the reference 3D video is a 3D video with a higher degree of clarity than the target for each reference focal length; The user's viewing experience for each reference 3D video is obtained, and the user's focal length is determined from multiple reference focal lengths based on the viewing experience; wherein, the visual attribute information includes the interpupillary distance and the focal length; Accordingly, the method further includes: Store the interpupillary distance and the user's focal length in the target database.

3. The method according to claim 2, characterized in that, The step of determining the disparity information of pixels in each frame of the video to be processed based on the visual attribute information includes: The target value is determined based on the interpupillary distance and the focal length; Determine the target depth value corresponding to the pixel in each frame of the image; wherein, the target depth value is the depth value of the target region of the object in each frame of the image; the target region is the region corresponding to the pixel in the object; Based on the target numerical value and the target depth value, the disparity value of the pixel is determined; wherein, the disparity information includes the disparity value.

4. The method according to claim 3, characterized in that, Determining the target depth value corresponding to the pixels of each frame image includes: Determine a target image with a subtitle region from multiple frames of images; wherein, the subtitle region is the region containing subtitles; The target depth value corresponding to the pixel of the target image is determined to be a target threshold. Based on the target depth estimation algorithm and other images, the target depth values ​​corresponding to the pixels in the other images are determined; wherein, the other images are images other than the target image in the multi-frame images.

5. The method according to claim 4, characterized in that, The step of determining the target depth value corresponding to the pixels in the other images based on the target depth estimation algorithm and other images includes: The target depth estimation algorithm is used to process the other images to obtain the first depth value corresponding to the pixels of the other images; The target time series stabilization algorithm is used to correct multiple first depth values ​​to obtain multiple corrected depth values; Determine the second depth value and the third depth value from the plurality of corrected depth values; Based on the second depth value, the third depth value, and the plurality of corrected depth values, the target depth value corresponding to the pixels of the other images is determined.

6. The method according to claim 1, characterized in that, The determination of multiple target images based on the disparity information and each frame image includes: Determine the first position of the pixel in each frame of the image; A target reverse mapping algorithm is used to process multiple first positions and the disparity information to obtain multiple second positions; The plurality of target images are determined based on the plurality of second positions.

7. The method according to claim 1, characterized in that, The process of determining the target 3D video based on the plurality of target images and each frame image includes: A target optimization algorithm is used to optimize each target image, resulting in multiple optimized images; The multiple optimized images and each frame image are processed to obtain the target 3D video.

8. A video generation device, characterized in that, include: Processor and memory used to store computer programs that can run on the processor; When the processor is used to run a computer program, it executes the steps of the method according to any one of claims 1 to 7.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.