Method, apparatus, device and storage medium for converting pictures into videos

By segmenting the pictures and iteratively transforming the visual depth of the background area, the problem that mobile terminals cannot realize Hitchcock-style mobile zoom is solved, and the conversion from static pictures to dynamic videos is realized, which improves the convenience of album production.

CN114331828BActive Publication Date: 2025-07-29FACE CUTE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202011063249.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-30
Publication Date
2025-07-29
Estimated Expiration
2041-01-16

AI Technical Summary

Technical Problem

Existing mobile terminals cannot realize Hitchcock-style mobile zoom technology shooting, resulting in poor fun in static images.

Method used

By segmenting the original image, the foreground and background areas are obtained, the visual depth is iteratively transformed for the background areas, and the transformed image is stored as picture frames, and finally the multi-frame images are spliced into dynamic videos.

Benefits of technology

It realizes that static images can be converted into dynamic videos with foreground focus and background transformation effects without manual operations, improving the convenience of album production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114331828B_ABST
    Figure CN114331828B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, apparatus, device, and storage medium for converting pictures into videos. The method includes: segmenting an original picture to obtain a foreground area and a background area; performing iterative transformation of visual depth on the background area, and storing each transformed image as a frame to obtain multiple frames of images; wherein the iterative transformation includes at least two transformations of visual depth; and splicing the multiple frames of images to obtain a target video. The method for converting pictures into videos provided by the embodiments of the present disclosure splices multiple images produced by iterative transformation of visual depth on the images in the background area to obtain a video album with the effect of foreground image focusing and background image Hitchcock transformation, without manual operation, improving the convenience of album production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image processing technology, and in particular, to a method, apparatus, device, and storage medium for converting pictures into videos. Background Art

[0002] With the continuous popularization of smart devices, the camera function has become an indispensable function in mobile phones. Currently, the photos taken by mobile phones are just static pictures with poor interestingness.

[0003] The Hitchcock-style moving zoom technique is applied to film and television works. Through tracking shots and zoom lenses, the distance change between the subject and the background is captured to create a visual effect of picture and space distortion, leading the audience into the psychological state of the protagonist. The principle of the Hitchcock-style moving zoom technique is the change of the focal length during video shooting. On the premise of ensuring that the proportion of the subject in each frame of the video remains unchanged, it switches between the telephoto focal length and the wide-angle focal length. That is to say, relative to the subject, while zooming in or out the lens, the lens is zoomed for shooting. This technique generally requires the assistance of professional shooting equipment to perform stepless switching of the lens focal length during the process of zooming in or out. Currently, most of the lenses on mobile terminals are non-zoomable lenses or consist of several lenses with different focal lengths, and thus it is impossible to perform shooting using the Hitchcock-style moving zoom technique, and there are limitations in performing shooting using the Hitchcock-style moving zoom technique. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, device, and storage medium for converting pictures into videos, which can convert static pictures into dynamic videos, achieve foreground image focusing, and perform album production with Hitchcock transformation of the background image, without manual operation, improving the convenience of album production.

[0005] In a first aspect, embodiments of the present disclosure provide a method for converting pictures into videos, including:

[0006] Segmenting the original picture to obtain a foreground area and a background area;

[0007] Performing iterative transformation of the visual depth on the background area, and storing the image obtained by each transformation as a frame of the picture to obtain multiple frames of images; wherein the iterative transformation includes at least two transformations of the visual depth;

[0008] Stitching the multiple frames of images to obtain a target video.

[0009] In a second aspect, embodiments of the present disclosure further provide a device for converting pictures into videos, including:

[0010] An area acquisition module, configured to segment the original picture to obtain a foreground area and a background area;

[0011] A visual depth transformation module for iteratively transforming the visual depth of the background region and storing the image obtained from each transformation as a frame, resulting in multiple frames of images; wherein the iterative transformation includes at least two visual depth transformations.

[0012] A target video acquisition module for stitching the multiple frames of images to obtain a target video.

[0013] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes:

[0014] One or more processing devices;

[0015] A storage device for storing one or more instructions;

[0016] When the one or more instructions are executed by the one or more processing devices, the one or more processing devices implement the method for converting a picture to a video as described in the embodiments of the present disclosure.

[0017] In a fourth aspect, an embodiment of the present disclosure further discloses a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processing device, it implements the method for converting a picture to a video as described in the embodiments of the present disclosure.

[0018] The embodiments of the present disclosure disclose a method, apparatus, device, and storage medium for converting a picture to a video. First, the original picture is segmented to obtain a foreground region and a background region, then the visual depth of the background region is iteratively transformed, and the image obtained from each transformation is stored as a frame, resulting in multiple frames of images; wherein the iterative transformation includes at least two visual depth transformations, and finally, the multiple frames of images are stitched to obtain a target video. The method for converting a picture to a video provided by the embodiments of the present disclosure stitches multiple images produced by the iterative transformation of the visual depth of the images in the background region to obtain a video album with the foreground image in focus and the background image having a Hitchcock transformation effect, without manual operation, improving the convenience of album production. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flowchart of a method for converting a picture to a video in an embodiment of the present disclosure;

[0020] Figure 2 is a schematic structural diagram of a device for converting a picture to a video in an embodiment of the present disclosure;

[0021] Figure 3 is a schematic structural diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0023] It should be understood that the various steps described in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0024] As used herein, the term "comprising" and its variants are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0025] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0026] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0027] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0028] Figure 1 It is a flowchart of a method for converting pictures to videos provided for the embodiments of the present disclosure. This embodiment is applicable to the situation of converting static pictures into dynamic videos. This method can be executed by a picture-to-video device, which can be composed of hardware and / or software and is generally integrated in a device with picture-to-video function. The device can be an electronic device such as a server or a server cluster. As Figure 1 shown, the method specifically includes the following steps:

[0029] Step 110, segment the original picture to obtain a foreground region and a background region.

[0030] The original image can be an image input or selected by the user. The foreground area can be a target area to be identified, such as a portrait area, an animal area, or a building area. The background area is the remaining area excluding the foreground area. Segmenting the original image can be understood as separating the foreground area from the background area.

[0031] In this embodiment, the original image is segmented to obtain the foreground area and the background area by: identifying the target object in the original image; and segmenting the area where the target object is located as the foreground area to obtain the foreground area and the background area.

[0032] Specifically, an object recognition model can be used to identify the target object in the original image. For example, if the foreground image is a portrait, a portrait recognition model can be used for recognition; if the foreground image is an animal, an animal recognition model can be used for recognition. This embodiment does not limit the type of target object to be recognized. The area where the target object is located is cut out from the original image, and the foreground area and background area are separated to obtain the foreground area and background area.

[0033] In this embodiment, the area where the target object is located is segmented as the foreground area, and the method for obtaining the foreground area and the background area can also be: obtaining the depth of the center point of the foreground area; performing depth smoothing processing on the pixel points of the foreground area based on the depth of the center point, and performing edge depth sharpening processing on the background area and the foreground area respectively, to obtain the processed foreground area and background area.

[0034] Among them, the method for obtaining pixel depth can adopt focusing method, brightness and illumination method, relative depth or depth sorting method, etc., and the depth acquisition method is not limited here. The process of depth smoothing of pixels in the foreground area based on the depth of the center point can be understood as adjusting the depth of each pixel in the foreground area to the depth of the center point so that the image of the foreground area is at the same visual depth.

[0035] Specifically, after the foreground area is segmented, the pixels in the foreground area are depth-smoothed according to the depth of the center point of the foreground area, the edges are depth-sharpened, and non-continuous closed edges are obtained through depth information to remove the foreground area and retain the background area.

[0036] Step 120 , performing iterative visual depth transformation on the background area, and storing the image obtained by each transformation as a picture frame to obtain multiple frames of images.

[0037] Among them, the iterative transformation includes at least two visual depth transformations. Each transformation continues to perform the visual depth transformation based on the result of the previous transformation. The visual depth transformation includes a transformation from far to near or from near to far. The visual depth transformation can be understood as zooming in on the image. In this embodiment, first, a visual depth transformation range is determined, and then within this transformation range, the background area is iteratively transformed in terms of visual depth at a certain step size.

[0038] In this embodiment, the way to perform iterative visual depth transformation on the background area can be: based on a set machine learning model, the area of the background area after foreground removal is filled in with an image; and then iterative visual depth transformation is performed on the filled background area.

[0039] Among them, the set machine learning model can be a model obtained through training with a large number of samples. The samples can be images with a part removed, and the training is supervised with the complete image. For example: the background area is a building. In the original picture, a part of the building is blocked by the foreground image. In the foreground-removed area, the set machine learning model is used to fill in the building image in the background area.

[0040] In this embodiment, the way to perform iterative visual depth transformation on the background area can be: based on the center point depth, the depth of the pixel points in the background area is transformed from near to far at a first set step size.

[0041] Specifically, taking the center point depth as a reference, gradually make the depth of the pixel points in the background area farther. Exemplarily, the set step size is d. Then, at the first transformation, the visual depth becomes farther by d to obtain the first frame of the picture. At the second transformation, based on the visual depth of the first frame of the picture, it continues to become farther by d to obtain the second frame of the picture. In this way, the depth of the second picture is 2d relative to the depth of the original picture, and so on until multiple frames of pictures are obtained.

[0042] In this embodiment, the way to perform iterative visual depth transformation on the background area can be: based on the center point depth, the depth of the pixel points in the background area is transformed from far to near at a second set step size.

[0043] Among them, the second set step size can be the same as or different from the first set step size. Specifically, taking the center point depth as a reference, gradually make the depth of the pixel points in the background area closer. Exemplarily, the set step size is d. Then, at the first transformation, the visual depth becomes closer by d to obtain the first frame of the picture. At the second transformation, based on the visual depth of the first frame of the picture, it continues to become closer by d to obtain the second frame of the picture. In this way, the depth of the second picture is closer by 2d relative to the depth of the original picture, and so on until multiple frames of pictures are obtained.

[0044] Step 130, splice multiple frames of images to obtain the target video.

[0045] In this embodiment, the multi-frame images can be stitched according to the timestamps of the multi-frame images. The target video after stitching is a photo album with a Hitchcock effect.

[0046] Optionally, if the original pictures include at least two pictures, for each original picture, perform operations of segmenting the original picture to obtain a foreground region and a background region; performing iterative transformation of visual depth on the background region, and storing the image obtained by each transformation as a frame to obtain multi-frame images; stitching the multi-frame images to obtain a target video; and obtaining at least two target videos.

[0047] Optionally, after obtaining at least two target videos, the following steps are further included: sorting the at least two target videos in a set order; stitching the sorted at least two target videos by adding set transition effects between adjacent target videos; and rendering the stitched at least two target videos to obtain a final video.

[0048] Among them, the set order can be the order in which the user inputs the pictures, or the adjusted order by the user, which is not limited here. The set transition effects between adjacent target videos can be the same or different. The set transition effects can be pre-set and can be arbitrarily selected.

[0049] The technical solution of this embodiment first segments the original picture to obtain a foreground region and a background region, then performs iterative transformation of visual depth on the background region, and stores the image obtained by each transformation as a frame to obtain multi-frame images; among them, the iterative transformation includes at least two visual depth transformations, and finally stitches the multi-frame images to obtain a target video. The method for converting pictures to videos provided by the embodiments of the present disclosure stitches multiple images produced by iterative transformation of visual depth of the images in the background region to obtain a video photo album with the foreground image in focus and the background image with a Hitchcock transformation effect, without manual operation, improving the convenience of photo album production.

[0050] The method for converting pictures to videos provided by the embodiments of the present disclosure can be launched as a function of a video APP. This function can realize automatic editing, creation, and sharing of videos. In this scenario, the user selects the function of converting pictures to videos, the user selects pictures, the client uploads the pictures to the server, the server obtains the pictures uploaded by the client, generates Hitchcock video segments, and returns them to the client; the client decodes and clips the video, and performs automatic playback preview after rendering the picture and adding transition effects, and the user can share or publish the video. The solution of this application does not require the user to manually operate the video, only needs to upload the pictures, greatly reducing the cost of generating videos from pictures.

[0051] Figure 2 is a schematic structural diagram of a device for converting pictures into videos disclosed in an embodiment of the present disclosure. As Figure 2 shown, the device includes: a region acquisition module 210, a visual depth transformation module 220, and a target video acquisition module 230.

[0052] The region acquisition module 210 is configured to segment an original picture to obtain a foreground region and a background region;

[0053] The visual depth transformation module 220 is configured to perform iterative transformation of the visual depth on the background region, and store the image obtained by each transformation as a frame of the picture to obtain multiple frames of images; wherein, the iterative transformation includes at least two transformations of the visual depth;

[0054] The target video acquisition module 230 is configured to splice the multiple frames of images to obtain a target video.

[0055] Optionally, the region acquisition module 210 is further configured to:

[0056] Identify a target object in the original picture;

[0057] Segment the region where the target object is located as the foreground region to obtain a foreground region and a background region; wherein, the background region is the region in the original picture except for the region where the target object is located.

[0058] Optionally, the region acquisition module 210 is further configured to:

[0059] Obtain the center point depth of the foreground region;

[0060] Perform depth smoothing processing on the pixel points of the foreground region based on the center point depth, and perform edge depth sharpening processing on the background region and the foreground region respectively to obtain the processed foreground region and background region.

[0061] Optionally, the visual depth transformation module 220 is further configured to:

[0062] Based on a set machine learning model, fill in the image of the region in the background region where the foreground is removed;

[0063] Perform iterative transformation of the visual depth on the filled background region.

[0064] Optionally, the visual depth transformation module 220 is further configured to:

[0065] Based on the center point depth, transform the depth of the pixel points in the background region from near to far according to a first set step size.

[0066] Optionally, the visual depth transformation module 220 is further configured to:

[0067] Based on the depth of the center point, the depth of the pixel points in the background area is transformed from far to near according to the second set step size.

[0068] Optionally, the original picture is the picture input or selected by the user. If the original picture includes at least two pictures, then for each original picture, perform the operations of segmenting the original picture to obtain the foreground area and the background area; performing iterative transformation of the visual depth on the background area, and storing the image obtained by each transformation as a frame of the picture to obtain multiple frames of images; splicing the multiple frames of images to obtain the target video; obtaining at least two target videos.

[0069] Optionally, it further includes: a video splicing module, which is used for:

[0070] Sorting at least two target videos in a set order;

[0071] Splicing the sorted at least two target videos by adding a set transition effect between adjacent target videos;

[0072] Rendering the spliced at least two target videos to obtain the final video.

[0073] The above device can execute the methods provided in all the foregoing embodiments of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the above methods. For the technical details not described in detail in this embodiment, reference can be made to the methods provided in all the foregoing embodiments of the present disclosure.

[0074] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc., or various forms of servers, such as independent servers or server clusters. Figure 3 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0075] As Figure 3As shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only storage device (ROM) 302 or a program loaded from a storage device 305 into a random access storage device (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0076] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0077] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the word recommendation method. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 309, or installed from the storage device 305, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0078] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0079] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0080] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device.

[0081] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: segment an original picture to obtain a foreground area and a background area; perform iterative transformation of visual depth on the background area, and store the image obtained by each transformation as a frame of the picture to obtain multiple frames of images; wherein the iterative transformation includes at least two transformations of visual depth; splice the multiple frames of images to obtain a target video.

[0082] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).

[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0084] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Wherein, the name of the unit does not constitute a limitation to the unit itself in some cases.

[0085] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0086] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0087] According to one or more embodiments of the embodiments of the present disclosure, the embodiments of the present disclosure disclose a method for converting a picture into a video, including:

[0088] Segment the original picture to obtain a foreground region and a background region;

[0089] Perform iterative transformation of the visual depth on the background region, and store the image obtained by each transformation as a frame of the picture to obtain multiple frames of images; wherein, the iterative transformation includes at least two transformations of the visual depth;

[0090] Stitch the multiple frames of images together to obtain the target video.

[0091] Further, segmenting the original picture to obtain a foreground region and a background region includes:

[0092] Identify the target object in the original picture;

[0093] Segment the region where the target object is located as the foreground region to obtain a foreground region and a background region; wherein, the background region is the region of the original picture other than the region where the target object is located.

[0094] Further, segmenting the region where the target object is located as the foreground region to obtain a foreground region and a background region includes:

[0095] Obtain the depth of the center point of the foreground area;

[0096] Based on the depth of the center point, perform depth smoothing on the pixel points of the foreground area, and perform edge depth sharpening on the background area and the foreground area respectively to obtain the processed foreground area and background area.

[0097] Further, perform iterative transformation of visual depth on the background area, including:

[0098] Based on a set machine learning model, fill in the image of the area in the background area where the foreground is removed;

[0099] Perform iterative transformation of visual depth on the filled background area.

[0100] Further, perform iterative transformation of visual depth on the background area, including:

[0101] Based on the depth of the center point, transform the depth of the pixel points of the background area from near to far at a first set step size.

[0102] Further, perform iterative transformation of visual depth on the background area, including:

[0103] Based on the depth of the center point, transform the depth of the pixel points of the background area from far to near at a second set step size.

[0104] Further, the original picture is a picture input or selected by the user. If the original picture includes at least two pictures, for each original picture, perform the operations of segmenting the original picture to obtain a foreground area and a background area; performing iterative transformation of visual depth on the background area, and storing the image obtained by each transformation as a frame of the picture to obtain multiple frames of images; splicing the multiple frames of images to obtain a target video; obtaining at least two target videos;

[0105] After obtaining at least two target videos, it further includes:

[0106] Sort the at least two target videos in a set order;

[0107] Splice the sorted at least two target videos by adding a set transition effect between adjacent target videos;

[0108] Render the spliced at least two target videos to obtain the final video.

[0109] Further, the foreground area includes a portrait area.

[0110] Note that the above is only a preferred embodiment of the present disclosure and the technical principles applied. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present disclosure. Therefore, although the present disclosure has been described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments. Without departing from the concept of the present disclosure, more other equivalent embodiments can be included, and the scope of the present disclosure is determined by the scope of the appended claims.

Claims

1. A method for converting pictures into videos, characterized in that, including: Segmenting the original image to obtain a foreground region and a background region; The original image is an image input or selected by the user; Performing iterative transformation of visual depth on the background region, and storing the image obtained by each transformation as a frame of the picture to obtain multiple frames of images; wherein, the iterative transformation includes at least two transformations of visual depth, the iterative transformation of visual depth includes a transformation from far to near or from near to far, and the iterative transformation of visual depth is to perform zoom processing on the image; Stitching the multiple frames of images to obtain a target video; Performing iterative transformation of visual depth on the background region includes: determining the transformation range of visual depth, and within this transformation range, performing iterative transformation of visual depth on the background region according to a set step size; The method further includes: If the original image includes at least two images, for each original image, perform the operations of segmenting the original image to obtain a foreground region and a background region; performing iterative transformation of visual depth on the background region, and storing the image obtained by each transformation as a frame of the picture to obtain multiple frames of images; stitching the multiple frames of images to obtain a target video; obtaining at least two target videos; Sorting the at least two target videos in a set order; wherein, the set order includes: the order of the images input by the user, or the adjusted order by the user; Stitching the sorted at least two target videos by adding set transition effects between adjacent target videos; Rendering the stitched at least two target videos to obtain a final video.

2. The method according to claim 1, wherein Segmenting the original image to obtain a foreground region and a background region includes: Identifying the target object in the original image; Taking the region where the target object is located as the foreground region for segmentation to obtain a foreground region and a background region; wherein, the background region is the region in the original image other than the region where the target object is located.

3. The method according to claim 2, characterized in that, Taking the region where the target object is located as the foreground region for segmentation to obtain a foreground region and a background region includes: Obtaining the center point depth of the foreground region; Performing depth smoothing processing on the pixel points of the foreground region based on the center point depth, and performing edge depth sharpening processing on the background region and the foreground region respectively to obtain the processed foreground region and background region.

4. The method according to claim 3, wherein Performing iterative transformation of visual depth on the background region includes: Based on a set machine learning model, filling in the image of the region in the background region where the foreground is removed; Performing iterative transformation of visual depth on the filled background region.

5. The method according to claim 3, characterized in that Performing iterative transformation of visual depth on the background region includes: Based on the center point depth, transforming the depth of the pixel points of the background region from near to far according to a first set step size.

6. The method according to claim 3, wherein Performing iterative transformation of visual depth on the background region includes: Based on the center point depth, transforming the depth of the pixel points of the background region from far to near according to a second set step size.

7. According to the method according to any one of claims 1-6, characterized in that, The foreground region includes a portrait region.

8. An apparatus for converting pictures into videos, characterized in that, including: A region acquisition module for segmenting the original image to obtain a foreground region and a background region; The original image is an image input or selected by the user; A visual depth transformation module, which is used to perform iterative transformation of visual depth on the background area, store the image obtained by each transformation as a frame, and obtain multiple frames of images; wherein, the iterative transformation includes at least two transformations of visual depth, the iterative transformation of visual depth includes a transformation from far to near or from near to far, and the iterative transformation of visual depth is to perform zoom processing on the image; A target video acquisition module, which is used to splice the multiple frames of images to obtain a target video; The visual depth transformation module is specifically used to determine the transformation range of visual depth, and within this transformation range, perform iterative transformation of visual depth on the background area according to a set step size; The target video acquisition module is specifically used to: if the original picture includes at least two pictures, then for each original picture, perform the operations of segmenting the original picture to obtain a foreground area and a background area; performing iterative transformation of visual depth on the background area, storing the image obtained by each transformation as a frame, and obtaining multiple frames of images; splicing the multiple frames of images to obtain a target video; obtaining at least two target videos; The device further includes: a video splicing module, which is used to: Sort the at least two target videos in a set order; wherein, the set order includes: the order of the pictures input by the user, or the adjusted order of the user; Splice the sorted at least two target videos by adding a set transition effect between adjacent target videos; Render the spliced at least two target videos to obtain a final video.

9. An electronic device, characterized in that, The electronic device includes: One or more processing devices; A storage device for storing one or more instructions; When the one or more instructions are executed by the one or more processing devices, the one or more processing devices implement the method for converting pictures to videos as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the method for converting pictures to videos as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method of generating blur

    US20050220358A1

  • Converting 2d video into stereo video

    US20100111417A1

  • Multi-source video input

    US20160064035A1

  • Method of converting 2d video to 3D video using machine learning

    US20170085863A1

  • Synthesis of transformed image views

    US20170308990A1