Media data generation method and device, equipment and storage medium

By generating media data containing elements of different depths of field and using correlation effects, users' needs for personalized editing of social media data are solved, expressive media data presentation is achieved, and user experience is improved.

CN120358370APending Publication Date: 2025-07-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410081722.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Users' demand for personalized editing of social media data is increasing, and it is difficult for the prior art to generate expressive media data with specific topics.

Method used

By acquiring the target image, media data containing foreground, medium and background elements are generated. Each element corresponds to different depths of field and is presented through associated preset animations to enhance the vividness and theme consistency of the media data.

Benefits of technology

It realizes rich and vivid presentation effects of media data, meets users' needs for personalized editing, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358370A_ABST
    Figure CN120358370A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a media data generation method and device, equipment and a storage medium. The method comprises the steps of obtaining a target image based on a preset operation of a user; and providing media data to the user, visual content of the media data including foreground elements, medium scene elements and background elements, the foreground elements, the medium scene elements and the background elements corresponding to different depths of field, the foreground elements being generated based on the target image, and the background elements being generated based on the target image. And wherein the foreground element presents a first preset dynamic effect, the medium scene element and / or the background element presents a second preset dynamic effect, and the first preset dynamic effect and the second preset dynamic effect are generated in association with each other. In this way, media data with a specific subject can be presented with a richer and more vivid presentation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for generating media data. Background Art

[0002] With the development of social applications, more and more users use social applications to post pictures or videos for showing their daily lives or browse the content posted by other publishers in the applications. Accordingly, users' personalized editing requirements for the media data to be posted are also increasing day by day. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for generating media data is provided. The method includes obtaining a target image based on a preset operation of a user; and providing the user with media data, the visual content of the media data including foreground elements, middle-ground elements, and background elements, the foreground elements, the middle-ground elements, and the background elements corresponding to different depths of field, wherein the foreground elements are generated based on the target image, and wherein the foreground elements present a first preset animation effect, the middle-ground elements and / or the background elements present a second preset animation effect, and the first preset animation effect and the second preset animation effect are generated in an associated manner with each other.

[0004] In a second aspect of the present disclosure, a device for generating media data is provided. The device includes an image obtaining module configured to obtain a target image based on a preset operation of a user; and a media data providing module configured to provide the user with media data, the visual content of the media data including foreground elements, middle-ground elements, and background elements, the foreground elements, the middle-ground elements, and the background elements corresponding to different depths of field, wherein the foreground elements are generated based on the target image, and wherein the foreground elements present a first preset animation effect, the middle-ground elements and / or the background elements present a second preset animation effect, and the first preset animation effect and the second preset animation effect are generated in an associated manner with each other.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to execute the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, and when the program is executed by a processor, the method of the first aspect is implemented.

[0007] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 FIG. shows a schematic diagram of an exemplary environment in which embodiments of the present disclosure can be implemented;

[0010] Figure 2 FIG. shows a schematic diagram of a media data generation process according to some embodiments of the present disclosure;

[0011] Figures 3A to 3D FIG. shows a schematic diagram of an animation rendering process according to some embodiments of the present disclosure;

[0012] Figure 4A and Figure 4B FIG. shows a schematic diagram of an animation rendering process according to some embodiments of the present disclosure;

[0013] Figure 5 FIG. shows a block diagram of an exemplary implementation process according to some embodiments of the present disclosure;

[0014] Figure 6 FIG. shows a flowchart of a media data generation process according to some embodiments of the present disclosure;

[0015] Figure 7 FIG. shows a block diagram of a media data generation device according to some embodiments of the present disclosure; and

[0016] Figure 8 FIG. shows a block diagram of a device capable of implementing multiple embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not used to limit the protection scope of the present disclosure.

[0018] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0019] It is understandable that the image data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0020] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scopes of use, usage scenarios, etc. of the data involved in the present disclosure should be informed to relevant users and their authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. The relevant users may include any type of rights holders, such as individuals, enterprises, and groups.

[0021] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multi-layer processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0022] As mentioned above, users have increasingly higher requirements for the presentation effects of media data. For example, users expect that a media data editing application can generate media data with a specific style and theme based on the original media material provided by the user, such as adding a makeup effect corresponding to the specific style and theme to the target object in the original media material, and presenting other materials corresponding to the specific style and theme in the generated media data. Therefore, it is expected that the generated media data will bring a dazzling presentation effect to the viewer.

[0023] The technical solution of the present disclosure provides a media data generation solution. According to the media data generation solution of the embodiments of the present disclosure, after obtaining a target image, media data can be provided to a user. The visual content of the media data includes foreground elements, middle-ground elements, and background elements, which respectively correspond to different depths of field. The foreground element is generated based on the target image. The foreground element presents a first preset dynamic effect and the middle-ground element and / or the background element present a second preset dynamic effect, and the first preset dynamic effect and the second preset dynamic effect are generated in an associated manner with each other. In this way, media data with a specific theme can be presented with a richer and more vivid presentation effect.

[0024] Example environment

[0025] First, refer to Figure 1 , which schematically shows a schematic diagram of an example environment 100 in which exemplary implementations according to the present disclosure can be implemented.

[0026] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, a target application 120 is installed in an electronic device 110. A user 102 can interact with the target application 120 via the electronic device 110 and / or an attached device of the electronic device 110.

[0027] The target application 120 can be an application capable of providing services related to media data to the user 102, including creation (e.g., shooting and / or editing), publishing, browsing, etc. of media data. Herein, "media data" can be various forms of content, including video, audio, image, image set, text, etc.

[0028] For example, the electronic device 110 can edit the received original media material through the target application 120. In some embodiments, the original media material can be, for example, media material input by the user 102. In some other embodiments, the original media material can also be media material collected by the electronic device 110 based on an instruction of the user 102 through a collection device. In some embodiments, the collection device can be configured to be connected to the electronic device 110 with each other. In some other embodiments, the collection device can also be integrated inside the electronic device 110.

[0029] In Figure 1 's environment 100, if the target application 120 is in an active state, the electronic device 110 can present a page 140 of the target application 120 to the user 102. The page 140 can be various pages that the target application 120 can provide, such as a presentation page of media data, a content creation page, a content editing page, etc.

[0030] In some embodiments, the implementation of at least some functions of the target application 120 can be based on the model 131. The model 131 can be deployed in the server 130, for example. For example, the electronic device 110 communicates with the server 130 to implement the supply of services for the target application 120. That is to say, during the operation of the target application 120, the capabilities of one or more models (such as the model 131) can be invoked.

[0031] In some embodiments, the electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for users (such as "wearable" circuits, etc.). The server 130 is various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0032] It should be understood that the structure and functions of the environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.

[0033] Dynamic effect presentation process

[0034] Figure 2 A schematic diagram of a media data generation process 200 according to some embodiments of the present disclosure is shown. Figure 2 The illustrated process 200 can be implemented at the electronic device 110. For example, Figure 2 The illustrated process 200 can be implemented by the target application 120 at the electronic device 110.

[0035] As Figure 2 shown, the electronic device 110 can obtain a target image based on a preset operation of the user 102. For example, the target image 211 or the target image 212. The target image can include a target object of a preset type, such as a face object, for example.

[0036] In some embodiments, the target image can be a target image obtained based on an upload operation of the user 102. In some other embodiments, the target image can also be a target image obtained based on a shooting operation of the user 102.

[0037] Based on the obtained target image, the electronic device 110 can output media data 221, such as media data 221 generated by editing the target image.

[0038] The visual content included in the output media data may include foreground elements, middle-ground elements, and background elements. The foreground elements, middle-ground elements, and background elements may correspond to different depths of field. For example, from the perspective of human eye vision, the foreground element is the element closest to the user's viewing angle in the front, while the background element is the element farthest from the user's viewing angle in the back. The middle element is between the foreground element and the background element.

[0039] In the output media data, the foreground element is generated based on the target image obtained by the electronic device 110. Regarding the foreground element, it may include a first part of the foreground element generated based on the target object in the target image. For example, a first part of the foreground element generated based on a face object. In addition, the foreground element may further include a second part of the foreground element. For example, the second part of the foreground element may include media materials corresponding to the preset theme of the generated media data 221. For example, the media materials may include clothing, styling, decorations, etc. that match the first part of the foreground element.

[0040] It should be understood that the first part of the foreground element generated based on the face object also corresponds to the preset theme of the media data 221. For example, the first part of the foreground element may present the effect of adding makeup and hairstyle corresponding to the preset theme of the media data 221 to the face object.

[0041] In some embodiments, the middle-ground element may include an image generated based on the preset theme. The background element may include text associated with the preset theme. The middle-ground element and the background element may be selected, for example, from a set of candidate elements in the target application material library based on the obtained target image and the preset theme.

[0042] The generated media data 221 may be content with a certain playing duration and capable of presenting a predetermined dynamic effect. During the playback of the media data 221, at least one of the foreground element, middle-ground element, and background element included in its visual content may present a preset dynamic effect.

[0043] The following combines Figures 3A to 3D and Figure 4A and Figure 4B to further elaborate in detail on the dynamic effects presented by the generated media data 221.

[0044] Figures 3A to 3D FIG. shows a schematic diagram of the dynamic effect presentation process according to some embodiments of the present disclosure. Figures 3A to 3DThe interfaces 300A to 300D shown can be presented by the electronic device 110. For the sake of discussion, reference will be made to Figure 1 and Figure 2 to describe the process shown in FIG. 3.

[0045] For example, in response to receiving the Figure 2 target image 211 in, the electronic device 110 outputs media data 221. Figures 3A to 3D The interfaces 300A to 300D shown can be presented by the electronic device 110, such as by the target application 120, during the process of playing the media data 221. The interfaces 300A to 300D can be understood as images captured at different time points T1, T2, T3, and T4 on the timeline 301 of playing the media data 221.

[0046] It should be understood that when starting to play the media data 221, the foreground element, middle - ground element, and background element are not necessarily presented simultaneously. The foreground element, middle - ground element, and background element can be presented successively as the media data 221 is played.

[0047] Figure 3A The interface 300A of shows the image captured at the time point T1 on the timeline 301. At this time point T1, the background element 311 in the visual content of the media data 221 is presented. The background element 311 is text associated with the preset theme of the media data 221.

[0048] Figure 3B The interface 300B of shows the image captured at the time point T2 on the timeline 301. In addition to the background element 311, the foreground element 312 in the visual content of the media data 221 is further presented at the time point T2. It can be seen that the foreground element 312 and the background element 311 have different depths of field. From the perspective of human eye vision, the background element 311 is behind the foreground element 312 in the interface 300B.

[0049] Although not shown, during the playback process from time point T1 to time point T2, the foreground element 312 can be presented from non - existent to existent with a predetermined animation effect. For example, the foreground element 312 can be presented with a flash - in effect. Another example is that the foreground element 312 can also be presented with a fade - in effect.

[0050] It can be seen that the foreground element 312 is generated based on the acquired target image 211. For example, the foreground element 312 can include a first part generated based on the face object of the target image 211 (i.e., the head part with a specific hairstyle and ornaments superimposed) and a second part generated based on the preset theme of the media data 221 (i.e., the clothing and styling part that matches the preset theme and the first part).

[0051] Figure 3C The interface 300C shows an image captured at time point T3 on the time axis 301. In addition to the background element 311 and the foreground element 312, a middle-ground element 313 in the visual content of the media data 221 is further presented at time point T3. This middle-ground element is an image generated based on a preset theme. It can be seen that the foreground element 312 and the middle-ground element 313 have different depths of field. From the perspective of human eye vision, this middle-ground element 313 is behind the foreground element 312 in the interface 300C.

[0052] Figure 3D The interface 300D shows an image captured at time point T4 on the time axis 301. As Figure 3D shown, the interface 300D presents that the foreground element 312, the middle-ground element 313, and the background element 311 have different depths of field. From the perspective of human eye vision, this middle-ground element 313 is behind the foreground element 312 in the interface 300D, and this background element 311 is behind the foreground element 312 in the interface 300D.

[0053] Comparing Figure 3C the interface 300C shown and Figure 3D the interface 300D shown, it can be seen that the foreground element 312 in the interface 300D is larger than the foreground element 312 in the interface 300C. This can be understood as presenting a first preset dynamic effect of the foreground element 312 during the playback of the media data 221 from time point T3 to time point T4. For example, a zoom-in operation is performed on the foreground element 312.

[0054] It can also be seen that the middle-ground element 313 in the interface 300D has also become larger compared to the middle-ground element 313 in the interface 300C. And compared to the situation where the middle-ground element 313 is mostly covered by the foreground element 312 in the interface 300C, the middle-ground element 313 presents a larger range in the interface 300D. From the perspective of human eye vision, it enables the viewer of the media data 211 to feel the effect that the middle-ground element 313 emerges from behind the foreground element 312. This can also be understood as presenting a second preset dynamic effect of the middle-ground element 313 during the playback of the media data 221 from time point T3 to time point T4. For example, a zoom-in and shift operation is performed on the middle-ground element 313.

[0055] The first preset dynamic effect and the second preset dynamic effect can be generated in an associated manner, for example. For example, the first dynamic effect parameter that causes the foreground element 312 to present the first preset dynamic effect and the second dynamic effect parameter that causes the middle-ground element 313 to present the second preset dynamic effect are matched with each other. For example, there is a corresponding relationship between the zoom factor that causes the foreground element 312 to present the first preset dynamic effect and the zoom factor that causes the middle-ground element 313 to present the second preset dynamic effect.

[0056] For example, the dynamic effects presented by the foreground element 312, the middle-ground element 313, and / or the background element 311 can be generated by applying three-dimensional camera movements. The dynamic effects presented by the foreground element 312, the middle-ground element 313, and / or the background element 311 can include, for example, partial slight movement effects of the elements, movement effects of the elements, scaling effects of the elements, and / or rotation effects of the elements, and so on.

[0057] Therefore, the generation of the first preset dynamic effect and the second preset dynamic effect in association with each other can also include that there is a corresponding relationship between the rotation direction and angle for the foreground element to present the first preset dynamic effect and the rotation direction and angle for the middle-ground element and / or the background element to present the second preset dynamic effect. Alternatively, there is a corresponding relationship between the movement direction and distance for the foreground element to present the first preset dynamic effect and the movement direction and distance for the middle-ground element and / or the background element to present the second preset dynamic effect.

[0058] In addition, the positional relationship between the middle-ground element and / or the background element and the foreground element and the dynamic effects implemented on the middle-ground element and / or the background element can be determined based on the pose information and / or dynamic effects of the foreground element. For example, as can be seen in interfaces 300C and 300D, the background element 311 is located in the middle position of the interface, and the middle-ground element 313 is also located in the middle position to create an effect that the image presented by the middle-ground element 313 slowly rises from behind the target object corresponding to the background element 311. Moreover, the text presented by the background element 311 exactly corresponds to the image presented by the middle-ground element 313. Therefore, the generated media data 211 conforms to the preset theme in terms of the overall effect, and the presented effect is harmonious and rich.

[0059] Figure 4A and Figure 4B shows a schematic diagram of the dynamic effect presentation process according to some embodiments of the present disclosure. Figure 4A and Figure 4B The shown interfaces 400A and 400B can be presented by the electronic device 110. For the convenience of discussion, reference will be made to Figure 1 and Figure 2 to describe the process shown in FIG. 3.

[0060] For example, in response to receiving Figure 2 the target image 212 in, the electronic device 110 outputs the media data 221. Figure 4A and Figure 4B The shown interfaces 400A and 400B can be presented by the electronic device 110, such as by the target application 120, during the playback of the media data 221. The interfaces 400A and 400B can be understood as images captured at different time points T1 and T2 on the timeline 401 of the playback of the media data 221. It has been combined with Figures 3A to 3DThe described dynamic effects will not be elaborated below.

[0061] Compare Figure 4A the shown interface 400A and Figure 4B the shown interface 400B. It can be seen that during the playback of the media data 221 from time point T1 to time point T2, the foreground element 411 presents a first preset dynamic effect of displacement and magnification, while the middle - ground element 412 presents a second preset dynamic effect of rotation and magnification. In this embodiment, there is a corresponding relationship between the moving direction and magnification factor for the foreground element 411 to present the first preset dynamic effect and the rotation direction and magnification factor for the middle - ground element 412 to present the second preset dynamic effect.

[0062] By generating the first preset dynamic effect and the second preset dynamic effect in an associated manner, it makes the foreground element and the middle - ground element / or background element produce a coordinated movement effect from the perspective of human vision, thereby making the generated media data 211 more vivid and interesting and improving the user's satisfaction with the editing effect.

[0063] Example implementation process

[0064] Figure 5 The block diagram of an example implementation process 500 according to some embodiments of the present disclosure is shown. The following combines Figure 5 to describe in detail the process of generating the media data 211. In some embodiments, Figure 5 the shown process 500 can be implemented at the electronic device 110. In some other embodiments, Figure 5 the shown process 500 can be implemented by the server 130, particularly by the model 131 deployed at the server 130. In the process 500, the input obtained by the electronic device 110 is, for example, the target image 211, and the output of the electronic device 110 is the media data 221.

[0065] As Figure 5 shown, after obtaining the target image 211, the target image 211 is pre - processed in module 510. For example, module 510 can determine whether there is a face object in the target image 211. Module 510 can perform operations such as changing the size and / or cropping on the target image 211. Module 510 can also perform facial hair segmentation processing on the portrait object in the target image 211 or optionally perform hairstyle removal (shaved head) processing. The pre - processed target image is provided to module 520 for stylized portrait processing. It should be understood that the pre - processed target image provided to module 520 should at least include a part of the original input target image 211.

[0066] The operations at module 520 can be implemented, for example, by the model 131. This model 131 can be, for example, asFigure 1 As shown in, it is deployed at server 130. Boot information (e.g., prompt words) can be provided to module 520. For example, the boot information can include at least one constraint associated with the image to be generated. For example, the constraint can include, for example, a style constraint and / or a pose constraint associated with the image to be generated. In the present disclosure, the image processed by module 520 is also referred to as an intermediate image.

[0067] At modules 530 and 540, the intermediate image generated by module 520 can be fused again with the original input target image 211 and refined more finely to generate the visual elements included in media data 221. For example, the face object in the intermediate image may become distorted during the stylization process. Through the refinement process at modules 530 and 540, the face object in the intermediate image can be replaced with the portrait object in the original input target image 211, which can make the face region in the foreground element 312 included in media data 211 more natural and realistic.

[0068] It should be understood that the Figure 5 processing process shown is merely exemplary. Figure 5 The process 500 shown can include more or fewer processing modules. The scope of the present disclosure is not limited in this regard.

[0069] Through the solution of the present disclosure, it is possible to present a combined dynamic effect of multiple visual elements with different depths of field in media data, thereby enhancing the texture of the media data and giving users a refreshing visual experience.

[0070] Example process

[0071] Figure 6 FIG. shows a flowchart of a media data generation process 600 according to some embodiments of the present disclosure. The process 600 can be implemented at the electronic device 110. It should be understood that the process 600 can also be implemented at the server 130.

[0072] At block 610, the electronic device 11 obtains a target image based on a preset operation of the user.

[0073] At block 620, the electronic device 110 provides the user with media data, the visual content of which includes a foreground element, a middle - ground element, and a background element. The foreground element, the middle - ground element, and the background element correspond to different depths of field. The foreground element is generated based on the target image, and wherein the foreground element presents a first preset dynamic effect, the middle - ground element and / or the background element present a second preset dynamic effect, and the first preset dynamic effect and the second preset dynamic effect are generated in an associated manner with each other.

[0074] In some embodiments, the target image includes a target object of a preset type, and the visual elements in the foreground elements include a first part corresponding to the target object.

[0075] In some embodiments, the visual elements further include a second part generated based on the target object.

[0076] In some embodiments, the visual elements are generated based on the following process: providing at least a part of the target image to a target model to obtain an intermediate image generated by the target model; and generating the visual elements by fusing the target image and the intermediate image.

[0077] In some embodiments, process 600 further includes: providing guiding information to the target model for guiding the target model to generate the intermediate image, where the guiding information includes at least one constraint associated with the intermediate image to be generated, and the at least one constraint includes: a style constraint associated with the intermediate image; or a pose constraint associated with the intermediate image.

[0078] In some embodiments, generating the visual elements by fusing the target image and the intermediate image includes: obtaining a first part corresponding to a target part from the target image; using the first part to replace a second part corresponding to the target part in the intermediate image; and generating the visual elements based on the replaced intermediate image.

[0079] In some embodiments, the middle scene elements include an image generated based on a preset theme; and / or the background elements include text associated with the preset theme, where the middle scene elements and / or the background elements are determined from a set of candidate elements based on the target image.

[0080] In some embodiments, the positional relationship between the middle scene elements and / or the background elements and the foreground elements is determined based on the pose information of the foreground elements.

[0081] In some embodiments, the first preset dynamic effect and / or the second preset dynamic effect include at least one of the following: a partial micro-movement effect of an element, a movement effect of an element, a scaling effect of an element, a rotation effect of an element.

[0082] In some embodiments, the first preset dynamic effect and the second preset dynamic effect are generated in an associated manner, including: making the dynamic effect parameters of the first preset dynamic effect and the second preset dynamic effect match each other.

[0083] Example device and equipment

[0084] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 7 FIG. 700 is a schematic structural block diagram of an apparatus for data processing according to some embodiments of the present disclosure.

[0085] As Figure 7 shown, the apparatus 700 may include an image acquisition module 710 configured to acquire a target image based on a preset operation of a user. The apparatus 700 may further include a media data presentation module 720 configured to provide media data to the user. The visual content of the media data includes foreground elements, middle-ground elements, and background elements. The foreground elements, the middle-ground elements, and the background elements correspond to different depths of field. The foreground elements are generated based on the target image, and wherein the foreground elements present a first preset dynamic effect, the middle-ground elements and / or the background elements present a second preset dynamic effect, and the first preset dynamic effect and the second preset dynamic effect are generated in an associated manner with each other.

[0086] In some embodiments, the target image includes a target object of a preset type, and the visual elements in the foreground elements include a first part corresponding to the target object.

[0087] In some embodiments, the visual elements further include a second part generated based on the target object.

[0088] In some embodiments, the visual elements are generated based on the following process: providing at least a part of the target image to a target model to obtain an intermediate image generated by the target model; and generating the visual elements by fusing the target image and the intermediate image.

[0089] In some embodiments, the apparatus 700 further includes: a guidance information providing module configured to provide guidance information to the target model for guiding the target model to generate the intermediate image, the guidance information including at least one constraint associated with the intermediate image to be generated, wherein the at least one constraint includes: a style constraint associated with the intermediate image; or a pose constraint associated with the intermediate image.

[0090] In some embodiments, generating the visual elements by fusing the target image and the intermediate image includes: obtaining a first part corresponding to a target part from the target image; using the first part to replace a second part corresponding to the target part in the intermediate image; and generating the visual elements based on the replaced intermediate image.

[0091] In some embodiments, the middle-ground element includes an image generated based on a preset theme; and / or the background element includes text associated with the preset theme, wherein the middle-ground element and / or the background element are determined from a set of candidate elements based on the target image.

[0092] In some embodiments, the positional relationship between the middle-ground element and / or the background element and the foreground element is determined based on the pose information of the foreground element.

[0093] In some embodiments, the first preset dynamic effect and / or the second preset dynamic effect include at least one of the following: a partial micro-motion effect of an element, a moving effect of an element, a scaling effect of an element, a rotation effect of an element.

[0094] In some embodiments, the first preset dynamic effect and the second preset dynamic effect are generated in an associated manner, including: making the dynamic effect parameters of the first preset dynamic effect and the second preset dynamic effect match each other.

[0095] The units included in device 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units in device 700 can be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip (SOC), complex programmable logic devices (CPLD), and so on.

[0096] Figure 8 A block diagram of a computing device / server 800 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 8 The computing device / server 800 shown is merely exemplary and should not impose any limitation on the functions and scope of the embodiments described herein.

[0097] As Figure 8As shown, the computing device / server 800 is in the form of a general-purpose computing device. The components of the computing device / server 800 may include, but are not limited to, one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 860, and one or more output devices 860. The processing unit 810 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing ability of the computing device / server 800.

[0098] The computing device / server 800 typically includes multiple computer storage media. Such media can be any accessible media available to the computing device / server 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be removable or non-removable media and can include machine-readable media, such as a flash drive, a magnetic disk, or any other media that can be used to store information and / or data (such as training data for training) and can be accessed within the computing device / server 800.

[0099] The computing device / server 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 8 a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0100] The communication unit 840 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device / server 800 can be implemented in a single computing cluster or multiple computer machines that are capable of communicating via a communication connection. Thus, the computing device / server 800 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0101] The input device 850 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 860 can be one or more output devices, such as a display, a speaker, a printer, etc. The computing device / server 800 can also communicate with one or more external devices (not shown) as needed through the communication unit 840. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the computing device / server 800, or communicate with any device that enables the computing device / server 800 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0102] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, and wherein the one or more computer instructions are executed by a processor to implement the method described above.

[0103] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0104] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that when the instructions are executed by the processing unit of the computer or other programmable data processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing apparatus, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.

[0105] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operation steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0107] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field of the present technology without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the field of the present technology to understand the implementations disclosed herein.

Claims

1. A method for generating media data, comprising: Obtaining a target image based on a preset operation of a user; And Providing media data to the user, wherein the visual content of the media data includes foreground elements, middle - ground elements, and background elements, the foreground elements, the middle - ground elements, and the background elements correspond to different depths of field, wherein the foreground elements are generated based on the target image, and wherein The foreground elements present a first preset animation effect, the middle - ground elements and / or the background elements present a second preset animation effect, and the first preset animation effect and the second preset animation effect are generated in an associated manner with each other.

2. The method according to claim 1, wherein the target image includes a target object of a preset type, and the visual elements in the foreground elements include a first part corresponding to the target object.

3. The method according to claim 2, wherein the visual elements further include a second part generated based on the target object.

4. The method according to claim 2, wherein the visual elements are generated based on the following process: Providing at least a part of the target image to a target model to obtain an intermediate image generated by the target model; and Generating the visual elements by fusing the target image and the intermediate image.

5. The method according to claim 4, further comprising: Providing guidance information to the target model for guiding the target model to generate the intermediate image, the guidance information including at least one constraint associated with the intermediate image to be generated, wherein the at least one constraint includes: A style constraint associated with the intermediate image; or A pose constraint associated with the intermediate image.

6. The method according to claim 4, wherein generating the visual elements by fusing the target image and the intermediate image includes: Obtaining a first part corresponding to a target part from the target image; Using the first part to replace a second part corresponding to the target part in the intermediate image; And Generating the visual elements based on the replaced intermediate image.

7. The method according to claim 1, wherein: The middle - ground elements include images generated based on a preset theme; and / or The background elements include text associated with the preset theme, wherein the middle - ground elements and / or the background elements are determined from a set of candidate elements based on the target image.

8. The method according to claim 1, wherein the positional relationship between the middle - ground elements and / or the background elements and the foreground elements is determined based on the pose information of the foreground elements.

9. The method according to claim 1, wherein the first preset animation effect and / or the second preset animation effect include at least one of the following: A partial micro - motion effect of an element, A moving effect of an element, A scaling effect of an element, A rotating effect of an element.

10. The method according to claim 1, wherein generating the first preset animation effect and the second preset animation effect in an associated manner with each other includes: Making the first preset animation effect and the second preset animation effect have mutually - matching animation parameters.

11. A media data generation device, comprising: An image acquisition module, configured to acquire a target image based on a preset operation of a user; and a media data providing module, configured to provide media data to the user, wherein visual content of the media data includes a foreground element, a middle-ground element, and a background element, the foreground element, the middle-ground element, and the background element corresponding to different depths of field, wherein the foreground element is generated based on the target image, and wherein the foreground element presents a first preset dynamic effect, the middle-ground element and / or the background element present a second preset dynamic effect, and the first preset dynamic effect and the second preset dynamic effect are generated in an associated manner with each other.

12. An electronic device, comprising: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.