Media data generation method and apparatus, device and storage medium
By generating media data containing foreground, medium and background elements of different depths of field and making their animation effects interrelated, users' needs for personalized editing of social media data are solved and the presentation effect of media data is improved.
Patent Information
- Application Number
- PCT/CN2024/138066
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2024-12-10
- Publication Date
- 2025-07-24
AI Technical Summary
Users' demand for personalized editing of social media data is increasing, and the prior art is difficult to generate media data with specific styles and themes, and lacks vivid presentation effects.
By obtaining the target image, media data containing foreground, medium scene and background elements are generated. Each element corresponds to different depths of field, and the dynamic effects of the foreground and medium scene elements are correlated to each other to present a dynamic effect.
Media data with specific topics is implemented in a rich and vivid way, improving user editing satisfaction.
Smart Images

Figure CN2024138066_24072025_PF_FP_ABST
Abstract
Description
Media data generation method, device, equipment and storage medium
[0001] This application claims priority to the Chinese invention patent application entitled “Media Data Generation Method, Apparatus, Device and Storage Medium” filed on January 19, 2024, with application number 202410081722.8, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to a method, apparatus, device, and computer-readable storage medium for generating media data. Background Art
[0003] With the development of social applications, more and more users use social applications to publish pictures or videos that show their daily lives or browse content published by other publishers in the applications. As a result, users' demand for personalized editing of the media data to be published is also increasing. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for generating media data is provided. The method includes acquiring a target image based on a preset operation of a user; and providing media data to the user, wherein the visual content of the media data includes a foreground element, a midground element, and a background element, wherein the foreground element, the midground element, and the background element correspond to different depths of field, wherein the foreground element is generated based on the target image, and wherein the foreground element exhibits a first preset motion effect, and the midground element and / or the background element exhibit a second preset motion effect, and wherein the first preset motion effect and the second preset motion effect are generated in association with each other.
[0005] In a second aspect of the present disclosure, a media data generation device is provided. The device includes an image acquisition module configured to acquire a target image based on a preset user operation; and a media data providing module configured to provide media data to the user, wherein the visual content of the media data includes a foreground element, a midground element, and a background element, wherein the foreground element, the midground element, and the background element correspond to different depths of field, wherein the foreground element is generated based on the target image, and wherein the foreground element exhibits a first preset motion effect, and the midground element and / or the background element exhibit a second preset motion effect, and wherein the first preset motion effect and the second preset motion effect are generated in association with each other.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the program is executed by a processor, the method of the first aspect is implemented.
[0008] It should be understood that the contents described in the summary of the present invention are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0011] FIG2 shows a schematic diagram of a media data generation process according to some embodiments of the present disclosure;
[0012] 3A to 3D are schematic diagrams illustrating a motion effect presentation process according to some embodiments of the present disclosure;
[0013] 4A and 4B are schematic diagrams showing a process of presenting a dynamic effect according to some embodiments of the present disclosure;
[0014] FIG5 shows a block diagram of an example implementation process according to some embodiments of the present disclosure;
[0015] FIG6 shows a flowchart of a process for generating media data according to some embodiments of the present disclosure;
[0016] FIG7 shows a block diagram of a media data generating apparatus according to some embodiments of the present disclosure; and
[0017] FIG8 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.
[0020] It is understandable that the image data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0021] It is understandable that before using the technical solutions disclosed in the various embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the data involved in the present disclosure should be informed to relevant users and authorization should be obtained from relevant users in an appropriate manner in accordance with relevant laws and regulations. The relevant users may include any type of right holders, such as individuals, enterprises, and groups.
[0022] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.
[0023] As mentioned above, users have increasingly high expectations for the presentation of media data. For example, users expect media data editing applications to generate media data with a specific style and theme based on the original media materials provided by the user. For example, they expect applications to add makeup effects corresponding to the specific style and theme to the target objects in the original media materials, and also expect the generated media data to present other materials corresponding to the specific style and theme in the generated media data. As a result, they expect the generated media data to provide viewers with a dazzling presentation effect.
[0024] The technical solution of the present disclosure provides a media data generation solution. According to the media data generation solution of the embodiment of the present disclosure, media data can be provided to the user after the target image is acquired. The visual content of the media data includes foreground elements, midground elements and background elements, which correspond to different depths of field respectively. The foreground elements are generated based on the target image. The foreground elements present a first preset motion effect and the midground elements and / or background elements present a second preset motion effect, and the first preset motion effect and the second preset motion effect are generated in association with each other. In this way, media data with a specific theme can be presented with a richer and more vivid presentation effect.
[0025] Sample Environment
[0026] First, reference is made to FIG1 , which schematically illustrates a diagram of an example environment 100 in which example implementations according to the present disclosure may be implemented.
[0027] 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example environment 100, a target application 120 is installed in an electronic device 110. A user 102 can interact with the target application 120 via the electronic device 110 and / or an attached device of the electronic device 110.
[0028] The target application 120 may be an application that can provide media data-related services to the user 102, including creation (e.g., shooting and / or editing), publishing, browsing, etc. of media data. In this document, "media data" may be content in various forms, including video, audio, image, image collection, text, etc.
[0029] For example, electronic device 110 can edit the received original media material through target application 120. In some embodiments, the original media material can be, for example, media material input by user 102. In some other embodiments, the original media material can also be media material collected by electronic device 110 using a collection device based on instructions from user 102. In some embodiments, the collection device can be configured to be connected to electronic device 110. In some other embodiments, the collection device can also be integrated within electronic device 110.
[0030] In the environment 100 of Figure 1, if the target application 120 is active, the electronic device 110 can present a page 140 of the target application 120 to the user 102. The page 140 can be any type of page that the target application 120 can provide, such as a media data presentation page, a content creation page, a content editing page, and the like.
[0031] In some embodiments, at least some of the functionality of the target application 120 may be implemented based on the model 131. The model 131 may be deployed, for example, in the server 130. For example, the electronic device 110 communicates with the server 130 to provide services for the target application 120. In other words, during the operation of the target application 120, the capabilities of one or more models (e.g., the model 131) may be invoked.
[0032] In some embodiments, the electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.). The server 130 is various types of computing systems / servers that can provide computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.
[0033] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0034] Animation presentation process
[0035] FIG2 shows a schematic diagram of a media data generation process 200 according to some embodiments of the present disclosure. The process 200 shown in FIG2 may be implemented in the electronic device 110. For example, the process 200 shown in FIG2 may be implemented by the target application 120 in the electronic device 110.
[0036] 2 , the electronic device 110 may acquire a target image, such as a target image 211 or a target image 212, based on a preset operation of the user 102. The target image may include, for example, a preset type of target object, such as a human face object.
[0037] In some embodiments, the target image may be a target image acquired based on an upload operation of the user 102. In some other embodiments, the target image may also be a target image acquired based on a shooting operation of the user 102.
[0038] Based on the acquired target image, the electronic device 110 may output media data 221 , such as media data 221 generated by editing the target image.
[0039] The visual content contained in the output media data may include foreground elements, midground elements, and background elements. These foreground elements, midground elements, and background elements may correspond to different depths of field. For example, from the perspective of human vision, foreground elements are the elements furthest forward from the user's perspective, while background elements are the elements furthest back from the user's perspective. Midground elements are located between foreground and background elements.
[0040] In the output media data, foreground elements are generated based on the target image captured by electronic device 110. Foreground elements may include a first portion of foreground elements generated based on a target object in the target image, for example, a first portion of foreground elements generated based on a human face. Furthermore, foreground elements may also include a second portion of foreground elements. For example, the second portion of foreground elements may include media material corresponding to a preset theme of the generated media data 221. For example, the media material may include clothing, styling, decorations, etc. that match the first portion of foreground elements.
[0041] It should be understood that the first portion of foreground elements generated based on the face object also corresponds to the preset theme of the media data 221. For example, the first portion of foreground elements may present an effect of adding makeup and hairstyle corresponding to the preset theme of the media data 221 behind the face object.
[0042] In some embodiments, the mid-ground element may include an image generated based on a preset theme. The background element may include text associated with the preset theme. The mid-ground element and background element may be selected from a set of candidate elements in a target application material library based on the acquired target image and the preset theme.
[0043] The generated media data 221 may be content having a certain playback duration and capable of presenting a predetermined motion effect. During playback of the media data 221, at least one of the foreground elements, midground elements, and background elements contained in the visual content thereof may present a predetermined motion effect.
[0044] The following further describes in detail the motion effects presented by the generated media data 221 in conjunction with FIG. 3A to FIG. 3D and FIG. 4A and FIG. 4B .
[0045] 3A to 3D illustrate schematic diagrams of a motion effect presentation process according to some embodiments of the present disclosure. Interfaces 300A to 300D shown in FIG3A to FIG3D may be presented by electronic device 110. For ease of discussion, the process shown in FIG3 will be described with reference to FIG1 and FIG2.
[0046] For example, in response to receiving the target image 211 in Figure 2, the electronic device 110 outputs the media data 221. The interfaces 300A to 300D shown in Figures 3A to 3D can be presented by the electronic device 110, for example, by the target application 120, during the playback of the media data 221. The interfaces 300A to 300D can be understood as images captured at different time points T1, T2, T3, and T4 on the timeline 301 of the playback of the media data 221.
[0047] It should be understood that when the media data 221 is played, the foreground element, the midground element and the background element are not necessarily presented at the same time. The foreground element, the midground element and the background element can be presented successively as the media data 221 is played.
[0048] 3A shows an image captured at time point T1 on the timeline 301. At time point T1, a background element 311 in the visual content of the media data 221 is presented. The background element 311 is text associated with a preset theme of the media data 221.
[0049] Interface 300B in FIG3B shows an image captured at time T2 on timeline 301. In addition to background element 311, foreground element 312 from the visual content of media data 221 is also presented at time T2. It can be seen that foreground element 312 and background element 311 have different depths of field. From the perspective of human vision, background element 311 is located behind foreground element 312 in interface 300B.
[0050] Although not shown, during the playback process from time point T1 to time point T2, the foreground element 312 can be presented from nothing to something using a predetermined animation effect. For example, the foreground element 312 can be presented with a flash-in effect. For another example, the foreground element 312 can be presented with a fade-in effect.
[0051] It can be seen that the foreground element 312 is generated based on the acquired target image 211. For example, the foreground element 312 may include a first portion generated based on the facial object of the target image 211 (i.e., the head portion superimposed with a specific hairstyle and accessories) and a second portion generated based on the preset theme of the media data 221 (i.e., the clothing and styling portion that matches the preset theme and the first portion).
[0052] Interface 300C in FIG3C shows an image captured at time T3 on timeline 301. In addition to background element 311 and foreground element 312, a mid-ground element 313 from the visual content of media data 221 is further presented at time T3. This mid-ground element is an image generated based on a preset theme. It can be seen that foreground element 312 and mid-ground element 313 have different depths of field. From the perspective of human vision, mid-ground element 313 is located behind foreground element 312 in interface 300C.
[0053] Interface 300D in FIG3D shows an image captured at time T4 on timeline 301. As shown in FIG3D , interface 300D presents foreground element 312, midground element 313, and background element 311 with different depths of field. From the perspective of human vision, midground element 313 is located behind foreground element 312 in interface 300D, while background element 311 is located behind foreground element 312 in interface 300D.
[0054] Comparing interface 300C shown in FIG3C with interface 300D shown in FIG3D , it can be seen that foreground element 312 in interface 300D is larger than foreground element 312 in interface 300C. This can be understood as the first preset animation effect of foreground element 312 being presented during the playback of media data 221 from time point T3 to time point T4. For example, foreground element 312 is enlarged.
[0055] It can also be seen that mid-ground element 313 in interface 300D is larger than mid-ground element 313 in interface 300C. Compared to interface 300C, where mid-ground element 313 is largely obscured by foreground element 312, mid-ground element 313 appears larger in interface 300D. From a human visual perspective, this gives the viewer of media data 211 the impression that mid-ground element 313 is poised to emerge from behind foreground element 312. This can also be understood as the presentation of the second preset motion effect for mid-ground element 313 during the playback of media data 221 from time point T3 to time point T4. For example, mid-ground element 313 is enlarged and shifted.
[0056] The first preset animation effect and the second preset animation effect may be generated in association with each other. For example, the first animation effect parameters that cause the foreground element 312 to present the first preset animation effect and the second animation effect parameters that cause the midground element 313 to present the second preset animation effect match each other. For example, there may be a corresponding relationship between the zoom factor that causes the foreground element 312 to present the first preset animation effect and the zoom factor that causes the midground element 313 to present the second preset animation effect.
[0057] For example, the motion effects presented by the foreground element 312, the midground element 313, and / or the background element 311 may be generated by applying a three-dimensional camera movement. The motion effects presented by the foreground element 312, the midground element 313, and / or the background element 311 may include, for example, a micro-motion effect of a portion of the element, an effect of moving the element, an effect of scaling the element, and / or an effect of rotating the element.
[0058] Therefore, the generation of the first preset motion effect and the second preset motion effect in association with each other may also include, for example, a corresponding relationship between the rotation direction and angle of the foreground element presenting the first preset motion effect and the rotation direction and angle of the midground element and / or background element presenting the second preset motion effect. Alternatively, a corresponding relationship between the movement direction and distance of the foreground element presenting the first preset motion effect and the movement direction and distance of the midground element and / or background element presenting the second preset motion effect.
[0059] In addition, the positional relationship between the mid-ground element and / or background element and the foreground element, as well as the motion effects implemented in the mid-ground element and / or background element, can be determined based on the posture information and / or motion effects of the foreground element. For example, in interfaces 300C and 300D, it can be seen that the background element 311 is located in the middle of the interface, and the mid-ground element 313 is also located in the middle, so as to create an effect that the image presented by the mid-ground element 313 slowly rises from behind the target object corresponding to the background element 311. The text presented by the background element 311 also happens to echo the image presented by the mid-ground element 313. Therefore, the generated media data 211 fits the preset theme in terms of overall effect, and the effect presented is harmonious and rich.
[0060] 4A and 4B illustrate schematic diagrams of a motion effect presentation process according to some embodiments of the present disclosure. Interfaces 400A and 400B shown in FIG4A and FIG4B may be presented by electronic device 110. For ease of discussion, the process shown in FIG3 will be described with reference to FIG1 and FIG2.
[0061] For example, in response to receiving the target image 212 in Figure 2, the electronic device 110 outputs the media data 221. The interfaces 400A and 400B shown in Figures 4A and 4B can be presented by the electronic device 110, for example, by the target application 120, during the playback of the media data 221. Interfaces 400A and 400B can be understood as images captured at different time points T1 and T2 on the timeline 401 of the playback of the media data 221. The dynamic effect presentation described in conjunction with Figures 3A to 3D will not be repeated below.
[0062] Comparing interface 400A shown in FIG4A with interface 400B shown in FIG4B , it can be seen that during the playback of media data 221 from time point T1 to time point T2, foreground element 411 exhibits a first preset motion effect of shifting and magnifying, while midground element 412 exhibits a second preset motion effect of rotating and magnifying. In this embodiment, there is a correspondence between the movement direction and zoom factor that causes foreground element 411 to exhibit the first preset motion effect and the rotation direction and zoom factor that causes midground element 412 to exhibit the second preset motion effect.
[0063] By generating the first preset motion effect and the second preset motion effect in association with each other, the foreground elements and the midground elements / or the background elements produce a coordinated motion effect from the perspective of human vision, thereby making the generated media data 211 more vivid and interesting, and improving the user's satisfaction with the editing effect.
[0064] Example implementation process
[0065] FIG5 shows a block diagram of an example implementation process 500 according to some embodiments of the present disclosure. The process of generating media data 211 is described in detail below in conjunction with FIG5 . In some embodiments, process 500 shown in FIG5 can be implemented at electronic device 110. In some other embodiments, process 500 shown in FIG5 can be implemented by server 130, specifically, by model 131 deployed at server 130. In process 500, the input obtained by electronic device 110 is, for example, target image 211, and the output of electronic device 110 is media data 221.
[0066] As shown in FIG5 , after acquiring the target image 211, the target image 211 is preprocessed in module 510. For example, module 510 may determine whether a human face is present in the target image 211. Module 510 may perform operations such as resizing and / or cropping the target image 211. Module 510 may also perform facial hair segmentation or, optionally, hairstyle removal (baldness) on the portrait subject in the target image 211. The preprocessed target image is provided to module 520 for stylized portrait processing. It should be understood that the preprocessed target image provided to module 520 should include at least a portion of the original input target image 211.
[0067] The operations at module 520 can be implemented, for example, by model 131. The model 131 can be deployed at server 130, for example, as shown in FIG1 . Guidance information (e.g., prompt words) can be provided to module 520. For example, the guidance information can include at least one constraint associated with the image to be generated. For example, the constraint can include a style constraint and / or a pose constraint associated with the image to be generated. In the present disclosure, the image processed by module 520 is also referred to as an intermediate image.
[0068] At modules 530 and 540, the intermediate image generated by module 520 can be fused again with the original input target image 211 and subjected to more fine-grained beautification to generate the visual elements contained in the media data 221. For example, facial objects in the intermediate image may become distorted during the stylization process. Through the refined processing at modules 530 and 540, the facial objects in the intermediate image can be replaced with the portrait objects in the original input target image 211, making the facial areas in the foreground elements 312 contained in the media data 211 appear more natural and realistic.
[0069] It should be understood that the processing process shown in Figure 5 is merely exemplary. The process 500 shown in Figure 5 may include more or fewer processing modules. The scope of the present disclosure is not limited in this respect.
[0070] Through the solution disclosed herein, a combined dynamic effect of multiple visual elements with different depths of field can be presented in media data, thereby improving the texture of the media data and giving users a refreshing visual experience.
[0071] Example Process
[0072] 6 shows a flow chart of a media data generation process 600 according to some embodiments of the present disclosure. The process 600 may be implemented at the electronic device 110. It should be understood that the process 600 may also be implemented at the server 130.
[0073] In block 610 , the electronic device 11 acquires a target image based on a preset operation of the user.
[0074] At block 620, the electronic device 110 provides media data to the user, wherein the visual content of the media data includes a foreground element, a midground element, and a background element. The foreground element, the midground element, and the background element correspond to different depths of field. The foreground element is generated based on the target image, wherein the foreground element exhibits a first preset motion effect, the midground element and / or the background element exhibit a second preset motion effect, and the first preset motion effect and the second preset motion effect are generated in association with each other.
[0075] In some embodiments, the target image includes a target object of a preset type, and the visual element in the foreground element includes a first portion corresponding to the target object.
[0076] In some embodiments, the visual element further includes a second portion generated based on the targeted object.
[0077] In some embodiments, the visual element is generated based on the following process: providing at least a portion of the target image to a target model to obtain an intermediate image generated by the target model; and generating the visual element by fusing the target image and the intermediate image.
[0078] In some embodiments, process 600 also includes: providing guidance information to the target model to guide the target model to generate the intermediate image, the guidance information including at least one constraint associated with the intermediate image to be generated, wherein the at least one constraint includes: a style constraint associated with the intermediate image; or a posture constraint associated with the intermediate image.
[0079] In some embodiments, generating the visual element by fusing the target image and the intermediate image includes: obtaining a first part corresponding to the target part from the target image; replacing a second part corresponding to the target part in the intermediate image with the first part; and generating the visual element based on the replaced intermediate image.
[0080] In some embodiments, the midground element includes an image generated based on a preset theme; and / or the background element includes text associated with the preset theme, wherein the midground element and / or the background element are determined from a set of candidate elements based on the target image.
[0081] In some embodiments, the positional relationship between the midground element and / or the background element and the foreground element is determined based on the posture information of the foreground element.
[0082] In some embodiments, the first preset motion effect and / or the second preset motion effect includes at least one of the following: a partial micro-motion effect of an element, a movement effect of an element, a scaling effect of an element, and a rotation effect of an element.
[0083] In some embodiments, the first preset motion effect and the second preset motion effect are generated in association with each other, including: making the first preset motion effect and the second preset motion effect have motion effect parameters that match each other.
[0084] Example devices and equipment
[0085] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 7 shows a schematic structural block diagram of a device 700 for data processing according to some embodiments of the present disclosure.
[0086] As shown in Figure 7, the device 700 may include an image acquisition module 710, which is configured to acquire a target image based on a preset operation of a user. The device 700 may also include a media data presentation module 720, which is configured to provide media data to the user. The visual content of the media data includes foreground elements, midground elements, and background elements. The foreground elements, the midground elements, and the background elements correspond to different depths of field. The foreground elements are generated based on the target image, and wherein the foreground elements present a first preset motion effect, the midground elements and / or the background elements present a second preset motion effect, and the first preset motion effect and the second preset motion effect are generated in association with each other.
[0087] In some embodiments, the target image includes a target object of a preset type, and the visual element in the foreground element includes a first portion corresponding to the target object.
[0088] In some embodiments, the visual element further includes a second portion generated based on the targeted object.
[0089] In some embodiments, the visual element is generated based on the following process: providing at least a portion of the target image to a target model to obtain an intermediate image generated by the target model; and generating the visual element by fusing the target image and the intermediate image.
[0090] In some embodiments, the device 700 also includes: a guidance information providing module, configured to provide guidance information to the target model for guiding the target model to generate the intermediate image, the guidance information including at least one constraint associated with the intermediate image to be generated, wherein the at least one constraint includes: a style constraint associated with the intermediate image; or a posture constraint associated with the intermediate image.
[0091] In some embodiments, generating the visual element by fusing the target image and the intermediate image includes: obtaining a first part corresponding to the target part from the target image; replacing a second part corresponding to the target part in the intermediate image with the first part; and generating the visual element based on the replaced intermediate image.
[0092] In some embodiments, the midground element includes an image generated based on a preset theme; and / or the background element includes text associated with the preset theme, wherein the midground element and / or the background element are determined from a set of candidate elements based on the target image.
[0093] In some embodiments, the positional relationship between the midground element and / or the background element and the foreground element is determined based on the posture information of the foreground element.
[0094] In some embodiments, the first preset motion effect and / or the second preset motion effect includes at least one of the following: a partial micro-motion effect of an element, a movement effect of an element, a scaling effect of an element, and a rotation effect of an element.
[0095] In some embodiments, the first preset motion effect and the second preset motion effect are generated in association with each other, including: making the first preset motion effect and the second preset motion effect have motion effect parameters that match each other.
[0096] The units included in the device 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 700 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0097] Figure 8 shows a block diagram of a computing device / server 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the computing device / server 800 shown in Figure 8 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0098] As shown in FIG8 , computing device / server 800 is in the form of a general-purpose computing device. Components of computing device / server 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 860, and one or more output devices 860. Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device / server 800.
[0099] The computing device / server 800 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device / server 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device / server 800.
[0100] The computing device / server 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 8 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0101] The communication unit 840 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device / server 800 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device / server 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0102] Input device 850 may be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 may be one or more output devices, such as a display, speaker, printer, etc. Computing device / server 800 may also communicate with one or more external devices (not shown) via communication unit 840 as needed, such as storage devices, display devices, etc., with one or more devices that allow a user to interact with computing device / server 800, or with any device that allows computing device / server 800 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0103] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, wherein the one or more computer instructions are executed by a processor to implement the method described above.
[0104] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0105] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0106] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0107] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0108] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
1. A method for generating media data, comprising: Obtaining a target image based on a preset operation of a user; And Providing media data to the user, wherein the visual content of the media data includes foreground elements, middle - ground elements, and background elements, the foreground elements, the middle - ground elements, and the background elements corresponding to different depths of field, wherein the foreground elements are generated based on the target image, and wherein The foreground elements present a first preset animation effect, the middle - ground elements and / or the background elements present a second preset animation effect, and the first preset animation effect and the second preset animation effect are generated in an associated manner with each other.
2. The method according to claim 1, wherein the target image includes a target object of a preset type, and the visual elements in the foreground elements include a first part corresponding to the target object.
3. The method according to claim 2, wherein the visual elements further include a second part generated based on the target object.
4. The method according to claim 2, wherein the visual elements are generated based on the following process: Providing at least a part of the target image to a target model to obtain an intermediate image generated by the target model; and Generating the visual elements by fusing the target image and the intermediate image.
5. The method according to claim 4, further comprising: Providing guidance information to the target model for guiding the target model to generate the intermediate image, the guidance information including at least one constraint associated with the intermediate image to be generated, wherein the at least one constraint includes: A style constraint associated with the intermediate image; or A pose constraint associated with the intermediate image.
6. The method according to claim 4, wherein generating the visual elements by fusing the target image and the intermediate image includes: Obtaining a first part corresponding to a target part from the target image; Using the first part to replace a second part corresponding to the target part in the intermediate image; And Generating the visual elements based on the replaced intermediate image.
7. The method according to claim 1, wherein: The middle - ground elements include an image generated based on a preset theme; and / or The background elements include text associated with the preset theme, wherein the middle - ground elements and / or the background elements are determined from a set of candidate elements based on the target image.
8. The method according to claim 1, wherein the positional relationship between the middle - ground elements and / or the background elements and the foreground elements is determined based on the pose information of the foreground elements.
9. The method according to claim 1, wherein the first preset animation effect and / or the second preset animation effect include at least one of the following: A partial micro - movement effect of an element, A movement effect of an element, A scaling effect of an element, A rotation effect of an element.
10. The method according to claim 1, wherein generating the first preset animation effect and the second preset animation effect in an associated manner with each other includes: Making the first preset animation effect and the second preset animation effect have mutually - matching animation parameters.
11. A media data generation device, comprising: An image acquisition module, configured to acquire a target image based on a preset operation of a user; and a media data providing module, configured to provide media data to the user, wherein visual content of the media data includes a foreground element, a middle ground element, and a background element, the foreground element, the middle ground element, and the background element corresponding to different depths of field, wherein the foreground element is generated based on the target image, and wherein the foreground element presents a first preset animation effect, the middle ground element and / or the background element present a second preset animation effect, and the first preset animation effect and the second preset animation effect are generated in association with each other.
12. An electronic device, comprising: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, having stored thereon a computer program, which when executed by a processor, implements the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Video generation method, device and equipment and medium
CN112887796A
Dynamic effect generation method and device based on image, equipment and storage medium
CN117078809A
Image editing method and device, equipment and storage medium
CN117201883A
Image dynamic effect generation method and device, storage medium and terminal
CN117218241A
Automatic generation of perceived real depth animation
US20210133928A1