Multimedia content generation method, electronic equipment and storage medium

By obtaining the multimedia material and description information input by the user, combining image and video features, multimedia content that meets user needs is solved, and the problem of users' difficulty in shooting multimedia content that meets expectations is achieved, and accurate multimedia content generation is achieved.

CN119992398APending Publication Date: 2025-05-13HUAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311515387.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

It is difficult for users to shoot multimedia content that meets expectations under external conditions, such as videos of kittens hurdles, because kittens lack athletic ability or are unwilling to cooperate.

Method used

By obtaining the subject reference material, content reference material and description information input by the user, combining image and video features, the target subject and attribute are determined using the multimedia content generation method to generate multimedia content that meets user needs.

Benefits of technology

It realizes the generation of accurate multimedia content based on user input, ensures that the target subject and attributes are consistent with user expectations, and meets user customization needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992398A_ABST
    Figure CN119992398A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, in particular to a multimedia content generation method, electronic equipment and a storage medium. The multimedia content generation method comprises the following steps: acquiring a subject reference material including a target subject input by a user, such as a picture including a cat, and a content reference material including a target attribute input by the user, such as a picture including a hurdling action, and acquiring description information, such as a text describing the hurdling action in the content reference material; multimedia content is generated having a target subject and a target attribute, for example, an image or video in which the subject is a cat and the target attribute is a hurdle. According to the scheme of the invention, the main body can be accurately selected according to the instruction of the user, and the generated multimedia content can be constrained based on the content reference material input by the user, for example, the target main body is reloaded, so that an image or a video meeting the requirement of the user can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and specifically to a multimedia content generation method, electronic device and storage medium. Background Art

[0002] With the rapid development of social platforms, people are keen to share personalized multimedia content such as pictures and short videos on the Internet. Whether the device can quickly shoot multimedia content that meets user needs has become one of the important factors that users pay attention to. However, due to external conditions, such as limited space, uncooperative subjects, and difficulty in achieving special shapes, it is difficult for users to shoot the expected multimedia content. For example, a user wants to shoot a video of a kitten jumping over hurdles, but because the kitten is not athletic enough or unwilling to cooperate, the video is difficult to shoot. Summary of the invention

[0003] The purpose of this application is to provide a multimedia content generation method, an electronic device and a storage medium.

[0004] In a first aspect, an embodiment of the present application provides a multimedia content generation method, comprising: obtaining a subject reference material, a first content reference material and description information input by a user; generating multimedia content, wherein the subject of the multimedia content includes: a target subject in the subject reference material, and the multimedia content has a target attribute, the target attribute includes a first attribute and a second attribute, the first attribute is an attribute in the first content reference material constrained by the description information, and the second attribute is an attribute outside the first content reference material constrained by the description information.

[0005] That is, in an embodiment of the present application, the subject reference material may be an image and / or video containing a target subject, for example, including one image and one video, or two images, or three videos, and so on. The target subject may be an object such as a person, a pet, a building, or food. The first content reference material may be an image or video containing target attributes required by the user, such as a video of a specific action, or a picture of a specific outfit. The description information may be a description text, which may include description information for the target subject, and description information for the target attribute, such as description information constraining the first attribute in the first content reference material. The target attribute may be an attribute of the target subject, such as an action attribute, or may be an overall attribute of a picture or video.

[0006] Through a multimedia content generation method of an embodiment of the present application, the target subject in the generated content can be accurately determined by combining the subject reference material with the relevant content in the description information, and the first content reference material and the relevant content in the description information can be combined to constrain the attributes of the generated multimedia content, such as changing the target subject's clothes, changing the picture style, etc., so as to generate images or videos that meet user needs.

[0007] In a possible implementation of the first aspect above, the method further includes: acquiring a second content reference material input by a user, wherein the second attribute is an attribute of the second content reference material constrained by the description information.

[0008] That is, in the embodiment of the present application, the second content reference material may be an image or video containing the target attributes required by the user, such as a video of a specific action or a picture of a specific garment.

[0009] In a possible implementation of the first aspect above, the first attribute and the second attribute are respectively at least one of an action attribute, a background attribute, a color attribute, a style attribute, a clothing attribute, an expression attribute, and a weather attribute.

[0010] In a possible implementation of the first aspect above, the main body of the multimedia content includes: a target main body described in the description information.

[0011] That is, in the embodiment of the present application, the description information may include information describing the target subject, such as describing the characteristics of the target subject, such as clothing characteristics, appearance characteristics, gender characteristics, etc., or describing the position of the target subject in the subject reference material, such as the first one from the left, the first one from the right. In this way, the subject can be determined based on the description information, thereby increasing the accuracy of subject determination.

[0012] In a possible implementation of the first aspect above, the subject of the multimedia content includes: a target subject whose definition is greater than a preset definition in the subject reference material.

[0013] That is, in this embodiment of the application, clarity greater than the preset clarity indicates that the target subject is in the focused part of the subject reference material, rather than the blurred part or the unclear part outside the depth of field, that is, the target subject is the view focus object in the subject reference material.

[0014] In a possible implementation of the first aspect above, the method further includes: displaying multiple candidate subjects; and taking a first candidate subject selected by a user from the multiple candidate subjects as a target subject, wherein a similarity between the candidate subject and the target subject described by the description information is greater than a similarity threshold.

[0015] That is, in the embodiment of the present application, the similarity with the target subject described by the description information is greater than the similarity threshold, indicating that each of the multiple candidate subjects is similar to the target subject. In this case, determining the target subject based on the user's selection can prevent the wrong target subject from being selected.

[0016] In a possible implementation of the first aspect above, the number of subject reference materials is at least one, and each subject reference material contains an image of a target subject; generating multimedia content includes: determining image features of the target subject based on the image of the target subject of each subject reference material in at least one subject reference material; determining visual semantic features of the target attribute based on the description information and the first content reference material; generating multimedia content based on the image features of the target subject and the visual semantic features of the target attribute.

[0017] That is, in the embodiment of the present application, the image features of the target subject may be features obtained by the encoder performing feature extraction on the image, also known as visual latent features; the visual semantic features of the target attributes may represent the visual meaning of the target attributes, and are features determined based on multimodal information (at least one of text, image, and video) input by the user, that is, features determined based on the descriptive information of the constrained multimedia content and the first content reference material, also known as generated constrained multimodal information representation.

[0018] In a possible implementation of the first aspect above, image features of a target subject are determined based on an image of the target subject in each subject reference material in at least one subject reference material, including: determining visual features of the target subject corresponding to each subject reference material based on the image of the target subject in each subject reference material in at least one subject reference material; and determining image features of the target subject based on the visual features of the target subject corresponding to each subject reference material.

[0019] That is, in the embodiment of the present application, the image of the target subject in each subject reference material is the image portion in each subject reference material that only includes the target subject and does not include other view contents such as the background; the visual features of the target subject corresponding to each subject reference material are the visual features of the image of the target subject in each subject reference material.

[0020] In a possible implementation of the first aspect above, based on the image of the target subject in each subject reference material in at least one subject reference material, determining the target subject visual features corresponding to each subject reference material includes: inputting the image of each target subject in at least one subject reference material into an image encoder for feature extraction to obtain the target subject visual features corresponding to each subject reference material.

[0021] That is, in the embodiment of the present application, the image encoder may be a visual geometry group (VGG)-16 encoder.

[0022] In a possible implementation of the first aspect above, the image features of the target subject are determined based on the visual features of the target subject corresponding to each subject reference material, including: inputting the visual features of the target subject corresponding to each subject reference material into a variational encoder for feature extraction to obtain the image features of the target subject.

[0023] That is, in the embodiment of the present application, the variational encoder may be an encoder based on a Transformer network.

[0024] In a possible implementation of the first aspect above, the visual semantic features of the target attribute are determined based on the description information and the first content reference material, including: inputting the description information of the target attribute and the first content reference material into a visual semantic feature extraction model for feature extraction to obtain the visual semantic features of the target attribute.

[0025] That is, in the embodiment of the present application, the visual semantic feature extraction model can be a multimodal pre-trained model, that is, a trained multimodal model.

[0026] In a possible implementation of the first aspect above, multimedia content is generated based on image features of a target subject and visual semantic features of target attributes, including: performing noise processing on the image features of the target subject to obtain extended image features, where the extended image features include noise features; performing denoising processing on noise features in the extended image features based on visual semantic features of the target attributes to obtain target visual features; and generating multimedia content based on the target visual features.

[0027] That is, in the embodiment of the present application, the noise addition process is to add Gaussian noise until the data becomes random noise. It can be understood that the expanded image features include noise features that conform to the Gaussian distribution. The denoising process is to denoise the noise features of the Gaussian distribution. It can be understood that the target visual features correspond to the target subject and the target attribute, and are used to generate multimedia content.

[0028] In a possible implementation of the first aspect, the image features of the target subject are subjected to noise processing to obtain extended image features, including: inputting the image features of the target subject into a diffusion model for noise processing to obtain extended image features; and based on the visual semantic features of the target attributes, denoising the noise features in the extended image features to obtain the target visual features, including: inputting the visual semantic features of the target attributes and the extended image features into a denoising model, denoising the noise features in the extended image features to obtain the target visual features.

[0029] That is, in the embodiment of the present application, the denoising model can adopt a semantic segmentation (unity networking, UNet) network.

[0030] In a possible implementation of the first aspect above, the multimedia content includes one of a target image and a target video; generating the multimedia content based on the target visual features includes: inputting the target visual features into an image generation model to generate a target image; or; inputting the target visual features into a video generation model to generate a target video.

[0031] That is, in the embodiment of the present application, the image generation model may be a variational encoder for image generation, and the video generation model may be a variational encoder for video generation.

[0032] In a possible implementation of the first aspect above, the multimedia content includes a target video; generating the multimedia content based on the target visual features includes: inputting the target visual features into an image generation model to generate a first frame of the target video; inputting the first frame of the target video and the visual semantic features of the target attributes into a video generation model to generate the target video.

[0033] That is, in the embodiment of the present application, the first frame of the target video can be generated by the variational encoder for image generation, and then the target video can be generated by the variational encoder for video generation. In this way, the constraint degree in the process of generating the video can be increased by generating the video step by step, so that the generated target video can better meet the target subject and target attributes required by the user.

[0034] In a possible implementation of the first aspect above, the main reference material includes at least one of an image and a video.

[0035] In a possible implementation of the first aspect above, the first content reference material includes at least one of an image and a video.

[0036] In a second aspect, an embodiment of the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing instructions of the multimedia content generation method of the first aspect.

[0037] In a third aspect, an embodiment of the present application provides a readable medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device executes the multimedia content generation method of the first aspect.

[0038] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising: a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium contains a computer program code for executing the multimedia content generation method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 According to the present application, a schematic diagram of an interface of a client 10 is shown;

[0040] Figure 2A According to the present application, a schematic diagram of an image input by a user is shown;

[0041] Figure 2B According to the present application, a schematic diagram of a reference video input by a user is shown;

[0042] Figure 2C A schematic diagram of a generated video is shown according to the present application;

[0043] Figure 3 According to the present application, a schematic diagram of a process of generating a video is shown;

[0044] Figure 4A According to an embodiment of the present application, a schematic diagram of a first operation interface of an electronic device 100 is shown;

[0045] Figure 4B According to an embodiment of the present application, a schematic diagram of a second operation interface of an electronic device 100 is shown;

[0046] Figure 4C According to an embodiment of the present application, a third operation interface schematic diagram of an electronic device 100 is shown;

[0047] Figure 4D According to an embodiment of the present application, a schematic diagram of a fourth operation interface of an electronic device 100 is shown;

[0048] Figure 4E According to an embodiment of the present application, a fifth operation interface schematic diagram of an electronic device 100 is shown;

[0049] Figure 5 According to an embodiment of the present application, a first flow chart of a method for generating multimedia content is shown;

[0050] Fig. 6A According to an embodiment of the present application, a schematic diagram of a subject reference material is shown;

[0051] Figure 6B A schematic diagram of a content reference material is shown according to an embodiment of the present application;

[0052] Figure 6C A schematic diagram showing another content reference material according to an embodiment of the present application;

[0053] Fig.6D A schematic diagram of multimedia content is shown according to an embodiment of the present application;

[0054] Figure 7 According to an embodiment of the present application, a second flow diagram of a method for generating multimedia content is shown;

[0055] Figure 8 A third flow chart of a multimedia content generation method is shown according to an embodiment of the present application;

[0056] Fig. 9 A fourth flow chart of a multimedia content generation method is shown according to an embodiment of the present application;

[0057] Fig.10 According to an embodiment of the present application, a schematic diagram of the application effect of a method for generating multimedia content is shown;

[0058] Fig.11 According to an embodiment of the present application, a hardware structure diagram of an electronic device 100 is shown;

[0059] Fig.12 According to an embodiment of the present application, a software structure block diagram of an electronic device 100 is shown. DETAILED DESCRIPTION

[0060] The illustrative embodiments of the present application include, but are not limited to, a multimedia content generation method, an electronic device, and a storage medium.

[0061] As mentioned above, users have the need to take personalized pictures or personalized videos. However, due to the limitations of external conditions, such as limited space, non-cooperation of the subject, and difficulty in achieving special shapes, users have difficulty in shooting. At present, in some technical solutions, if users have difficulty shooting, user instructions input by the user can be received, and multimedia content expected by the user can be generated according to the user instructions. For example, according to the user instruction "generate a video of a kitten hurdling", a video is generated, but the hurdle in the generated video is not a kitten, or the kitten is not hurdling, that is, the visual subject is not necessarily close to the kitten that the user wants, and the hurdle action is not necessarily consistent with the user's needs. Therefore, this method has the problem that the quality of the generated multimedia content is poor and does not meet the user's expectations.

[0062] In order to solve the problem of poor quality of multimedia content expected by the user generated according to the user's instructions, in some optional embodiments, if the user wants to generate multimedia content of a target subject performing a target movement, the user can input to the client an image or video content containing the target subject, as well as an image or video of other subjects performing a target movement; then, the client generates multimedia content of the target subject performing a target movement based on the user's input. The user input can be a picture or a video, such as a picture of a kitten or a video of a puppy jumping over hurdles, and the client can generate a video of a kitten jumping over hurdles based on the picture of the kitten and the video of the puppy jumping over hurdles.

[0063] According to some embodiments, the client may obtain an image containing a target subject and a reference video containing target motion input by a user, determine a motion pattern of a reference object in the reference video (i.e., a pattern of target motion), and generate an output video with the input image as a starting frame, so that the target object in the output video has the motion pattern of the reference object. Figure 1 As shown, a picture input control 11, a video input control 12 and a customized multimedia content generation button 13 are displayed on the screen of the client 10. The user can click the picture input control 11, select a picture containing a subject from the album, and complete the image input operation; click the video input control 11, select a reference video from the album, and complete the reference video input operation; finally, click the customized multimedia content generation button 13, so that the client generates the user's expected multimedia content based on the user-input image and the reference video, that is, the multimedia content of the target subject performing the target movement.

[0064] Optionally, the user-entered image can be referenced Figure 2A , i.e., picture 21 of the kitten 01 sitting next to the hurdle; the reference video entered by the user can refer to Figure 2B , i.e., video 22 of puppy 02 jumping over hurdles. The client 10 can Figure 2A The picture 21 and Figure 2B Video 22 shows the generation of Figure 2C The user is expected to see the video 23 of the kitten 01 jumping over the hurdle. It can be understood that Figure 2C The first frame of the video 23 shown is Figure 2A Picture 21 shown.

[0065] The specific implementation method of the above solution is introduced below.

[0066] Specific reference Figure 3The client 10 performs semantic segmentation on the reference video (301) to obtain semantic objects (303) in the frame-by-frame images, and performs semantic segmentation on the input image (302) to obtain semantic objects (304) in the input image, and matches the semantic objects (303) in the frame-by-frame images and the semantic objects (304) in the input image based on predetermined rules such as shape and color (305) to obtain video matching semantic objects. The client 10 uses optical flow estimation or convolutional neural network (CNN) to obtain motion patterns (306) between semantic object frames according to the semantic objects (303) in the frame-by-frame images, wherein the motion patterns (306) between semantic object frames are used for subsequent output video generation, that is, the client 10 adjusts the image corresponding target (307) based on the motion patterns (306) between video semantic object frames to obtain the next few frames (308) of the input image, and then obtains the output video. It can be understood that the output video consists of the input image (302) and the next few frames (308) of the input image.

[0067] In the above embodiment, the visual content of the subject in video 23 cannot be changed, and it is not applicable to the scene where the input image 21 contains multiple subjects. Specifically, the first frame image in the generated video 23 must be picture 21, that is, the dress of kitten 01 in video 23 must be consistent with picture 21. However, if the user has the need to change the clothes of kitten 01, such as letting kitten 01 wear sportswear, the above method cannot be implemented. Moreover, the above embodiment does not have a process for identifying the subject of the input image. When there are multiple subjects in the input image, it may not be possible to accurately locate the target subject required by the user. For example, if there is another kitten in addition to kitten 01 in image 2A, because the client 10 does not have a process for screening the target subject required by the user from multiple subjects, it may mistakenly select another kitten as the subject in video 23 for generation.

[0068] In order to solve the problem that when there are multiple subjects in the input image, the target subject required by the user may not be accurately located, the present application proposes a multimedia content generation method. In this method, when the user expects to generate multimedia content of the target subject, in addition to inputting the subject reference material (such as an image or video containing the target subject) and the content reference material (such as an image or video containing the target movement), the user is also required to input description information, wherein the description information is used to specify the target subject in the subject reference material, and the definition of the target attribute of the target subject in the generated content. For example, the description information "the first kitten on the right side of the subject reference material wears a scarf, and the hurdle action is the same as the puppy in the content reference material" specifies that the target subject is the kitten on the right side of the subject reference material, the clothing is a scarf, and the hurdle action is the action of the puppy in the content reference material. According to the description information, the visual features of the target subject can be determined, such as the visual features of the kitten in the subject reference material, and then the multimedia content is generated according to the definition of the target subject in the generated content in the description information, such as the definition of the kitten's clothing and action. It can be understood that the generated multimedia content is consistent with the description information, that is, the kitten's clothing is a scarf, and the hurdle action is consistent with the puppy in the content reference material. Through the above method, not only can the subject be accurately selected according to the user's specification to ensure that the subject in the generated content is consistent with the subject expected by the user, but the generated multimedia content can also be constrained based on the content reference material input by the user, such as changing clothes, changing colors, etc., so as to meet the user's customized needs in all aspects and generate images or videos that meet the user's needs.

[0069] In some embodiments, the content reference material provided by the user may include one or more of images and videos, such as pictures containing actions that the user wants to generate, videos with a specific motion pattern, etc., which are used to limit the target attributes of the generated multimedia content, and the target attributes may include action attributes, background attributes, clothing attributes, etc. In this way, the client can directly generate content based on the material provided by the user, avoiding the problem of non-existence and incompatibility of the material caused by querying the material library; and because the content reference material is real material, such as a video of a puppy jumping over hurdles in reality, the video generated based on it will not be out of touch with reality, for example, in the generated video, the kitten jumps too high, the jumping direction does not match the direction of the runway, etc., so that multimedia content that meets the user's expectations can be generated.

[0070] In some embodiments, the subject reference material provided by the user may include one or more of images and videos, wherein each subject reference material may include multiple items. For example, the user may provide two images of view materials containing the subject, and provide descriptive information corresponding to each of the two images, such as "the first kitten on the right in image 1" and "the first kitten on the left in image 2", to assist in determining the subject and ensure that the subject in the generated content is the subject expected by the user.

[0071] In some embodiments, the client can perform target subject detection on the subject reference material provided by the user based on the target detection model. If the number of detected candidate subject objects is 1, the candidate subject object is directly determined as the subject; if the number of detected candidate subject objects is greater than 1, a graphic semantic matching model can be used to calculate semantic similarity based on the description information corresponding to the subject reference material and the visual features of each candidate subject object, and then determine the subject based on the calculated similarity result. For example, the candidate subject object with the highest similarity is determined as the subject, which can improve the accuracy of subject recognition.

[0072] In some embodiments, the client can input the description information corresponding to the content reference material and the content reference material into the multimodal pre-training model to obtain a constrained multimodal information representation, that is, input text, images or text, and video into the multimodal pre-training model to obtain a constrained multimodal information representation. It can be understood that in the case where the content reference material includes multiple, the constrained multimodal information representation can also be multiple. For example, corresponding to the reference multimedia content including pictures and videos, the corresponding description information includes "the clothing is consistent with the clothing of the characters in the picture" and "the action is consistent with the action of the characters in the video". The multimodal pre-training model can output the constrained multimodal information representation corresponding to the picture-clothing and video-action respectively. Then, the client inputs the constrained multimodal information representation into the denoising model, obtains the denoised visual potential representation, and uses the visual potential representation to generate multimedia content through a variational decoder to obtain an image or video that meets the user's needs. Through the above steps, when there is a large amount of depth displacement in the content reference material, for example, when there are a large amount of body movement details in the dog's hurdle jumping action, the characteristics of these depth displacements can be retained in the visual potential representation (i.e., the large amount of body movement details in the dog's hurdle jumping action are retained), so that the generated multimedia content can fully restore the user's expected action, that is, reflect the above-mentioned body movement details.

[0073] It is understandable that the client applicable to the technical solution of the present application can be a multimedia application, or an electronic device 100 capable of running the multimedia application, such as a mobile phone, a tablet computer, a wearable device, a notebook computer, a personal computer (PC), a netbook, etc. The present application does not impose any restrictions on the specific type of the electronic device 100.

[0074] Combine the following Figure 4A-4E The operating interface of the electronic device 100 in the embodiment of the present application is introduced.

[0075] refer to Figure 4A, the electronic device 100 screen displays prompts of “Please enter the view material containing the subject”, “Please enter the customized guidance information”, “Please enter the guidance audio information”, and “Please enter the guidance view information”, wherein a first upload control 401A is provided below “Please enter the view material containing the subject”, a view generation guidance text box 402A is provided below “Please enter the customized guidance information”, an audio upload button 403A is provided next to “Please enter the guidance audio information”, a second upload control 404A is provided below “Please enter the guidance view information”, and a customized multimedia content generation button 405A is provided at the very bottom of the page.

[0076] It can be understood that the user can click one or more of the first upload control 401A, the view generation guide text box 402A, the audio upload button 403A, and the second upload control 404A to complete the input of multimodal information, that is, the input of at least one of text, image, video, and audio, and finally click the customized multimedia content generation button 405A to complete the user input operation and wait for the multimedia content generation to be completed.

[0077] Among them, after clicking the first upload control 401A and the second upload control 404A, the screen of the electronic device 100 can display its own gallery for the user to select pictures and videos, that is, select the main reference material and content reference material for uploading respectively.

[0078] After the click view generates the guide text box 402A, a text input box may be displayed on the screen of the electronic device 100 for the user to input text, ie, to input description information.

[0079] Optionally, after single-clicking or long-pressing the audio upload button 403A, the electronic device 100 can record the user's voice, that is, the user can input the description information to the electronic device 100 by voice.

[0080] It can be understood that the above-mentioned upload operations do not limit the number of uploaded contents. For example, you can select an image or video to upload each time you click the first upload control 401A, thereby uploading multiple images and videos through multiple clicks and upload operations; you can also click the first upload control 401A once, and then select multiple images and videos in the gallery to upload. The embodiment of the present application does not limit this.

[0081] Figure 4BThe figure shows another optional operation interface of the electronic device 100. The screen of the electronic device 100 displays prompts of "Please enter the view material containing the subject", "Please enter the customized guidance information", "Please enter the guidance audio information", and "Please enter the guidance view information", wherein a first upload control 401B is provided below "Please enter the view material containing the subject", a view generation guidance text box 402B is provided below "Please enter the customized guidance information", an audio upload button 403B is provided next to "Please enter the guidance audio information", a second upload control 404B is provided below "Please enter the guidance view information", and a customized multimedia content generation button 405B is provided at the bottom of the page.

[0082] Among them, for Figure 4B For the introduction of the view generation guide text box 402B, the audio upload button 403B, the second upload control 404B, and the customized multimedia content generation button 405A, please refer to the above description for Figure 4A The introduction of the middle view generation guide text box 402A, the audio upload button 403A, the second upload control 404A, and the customized multimedia content generation button 405A will not be repeated here.

[0083] It can be understood that after the user clicks the first upload control 401B, the camera application of the electronic device 100 can be awakened, that is, the operation interface of the camera application and the real-time shooting image of the camera are displayed on the screen, so that the user can shoot images or videos in real time as input of the first upload control 401B. Figure 4C , the camera application interface of the electronic device 100 displays a real-time image, a shooting button, a video button, and a completion selection button; Figure 4D , the user can select the first object 41 as the visual subject, and a selection box of the first object 41 is displayed on the screen of the electronic device 100, indicating that the selection is successful; Figure 4E The user can further select the second object 42 as the second visual subject, and then click the Finish Selection button to complete the selection.

[0084] The following takes the client as an electronic device 100 as an example to introduce the multimedia content generation method mentioned in the embodiment of the present application. Figure 5 , an exemplary process of a multimedia content generation method according to an embodiment of the present application includes:

[0085] S501: Acquire subject reference material, content reference material and description information input by the user.

[0086] It can be understood that when the user inputs the subject reference material and the content reference material, the description information may include description information corresponding to the subject reference material and the content reference material. That is, the description information is used to define the subject in the generated multimedia content based on the subject reference material; and is used to define the corresponding attributes in the generated multimedia content based on the characteristics of the content reference material.

[0087] In some embodiments, the subject reference material provided by the user may include one or more of images and videos, wherein each subject reference material may include multiple items. For example, the user may provide two images of view materials containing the subject, and provide descriptive information corresponding to each of the two images, such as "the first kitten on the right in subject reference material 1" and "the first kitten on the left in subject reference material 2", to assist in determining the subject and ensure that the subject in the generated content is the subject expected by the user.

[0088] In some embodiments, the content reference material provided by the user may include one or more images and videos, each of which may include multiple items. For example, the user may provide multiple videos and pictures, and provide different attributes respectively, such as at least one of the background attributes, color attributes, action attributes, expression attributes, weather attributes, background attributes, clothing attributes, etc., and provide descriptive information corresponding to different content reference materials, such as "the background of content reference material 1", "the clothing of content reference material 2", and "the action of content reference material 3".

[0089] Combine the following Figure 6A-6C An optional implementation of the subject reference material and content reference material input by the user is introduced.

[0090] Fig. 6A According to an embodiment of the present application, a schematic diagram of a subject reference material is shown. Figure 6B According to an embodiment of the present application, a schematic diagram of a content reference material is shown. Figure 6C According to an embodiment of the present application, another schematic diagram of content reference material is shown. Optionally, the user can Figure 4A or Figure 4B The operation interface of the electronic device 100 shown in FIG. Fig. 6A , Figure 6B , Figure 6C Enter as main reference material 1, content reference material 1, content reference material 2, and enter Fig. 6A , Figure 6B , Figure 6C corresponding description information, and then click Figure 4A , 4BThe customized multimedia content generation button shown. Optionally, the description information may be "generate a video of the kitten in subject reference material 1, change the action to the action of content reference material 1, and change the background to the background of content reference material 2".

[0091] S502: Generate multimedia content.

[0092] Among them, the main body of the multimedia content includes: the target main body in the main reference material described in the description information; and the multimedia content can have a target attribute, wherein the target attribute can include a first attribute and a second attribute, the first attribute is the first attribute in the content reference material constrained by the description information, and the second attribute is an attribute outside the content reference material constrained by the description information.

[0093] In an optional embodiment, the first attribute may be an action feature, and the second attribute may be a background feature. Figure 6A-6C In the illustrated embodiment, the description information can be used to define the background attribute and action attribute of the multimedia content to be consistent with the content reference material 1 and the content reference material 2, respectively. It should be noted that the present application does not limit the types of the first attribute and the second attribute, and in addition to the first attribute and the second attribute, the description information can also be used to define other attributes, and the content reference material can also provide references for other elements in the generation of multimedia content, such as the third attribute and the fourth attribute in addition to the first attribute and the second attribute, and so on.

[0094] In some embodiments, the electronic device 100 can perform target subject detection on the subject reference material provided by the user based on the target detection model. When the number of detected candidate subject objects is 1, the candidate subject object is directly determined as the subject; when the number of detected candidate subject objects is greater than 1, a graphic semantic matching model can be used to calculate semantic similarity based on the description information corresponding to the subject reference material and the visual features of each candidate subject object, and then determine the subject based on the calculated similarity result. For example, the candidate subject object with the highest similarity is determined as the target subject, which can improve the accuracy of subject recognition.

[0095] Optionally, after determining the target subject, the electronic device 100 can generate a visual latent representation (i.e., image features of the target subject) based on the image of the target subject in the subject reference material. For example, the image of the target subject is input into a pre-trained image encoder and a pre-trained variational encoder to obtain the image features of the target subject, i.e., the visual latent representation. Then, the electronic device 100 can perform noise processing on the visual latent representation to generate noise features (i.e., expanded image features) that conform to the Gaussian distribution. It can be understood that noise addition is to add Gaussian noise to the visual latent representation until the data becomes random noise, and a noise feature that conforms to the Gaussian distribution is obtained. After obtaining the noise features of the Gaussian distribution, the denoising model can be used to perform a reverse generation process based on the noise features of the Gaussian distribution, that is, by denoising the random noise, the target visual features corresponding to the multimedia content to be generated are obtained.

[0096] In some embodiments, the electronic device 100 may input the description information corresponding to the content reference material and the content reference material into the multimodal pre-training model to obtain a generation-constrained multimodal information representation (i.e., visual semantic features), that is, input text, images or text, and video into the multimodal pre-training model (i.e., visual semantic feature extraction model) to obtain a generation-constrained multimodal information representation. Among them, the generation-constrained multimodal information represents the attributes that the user wants to limit in the generated multimedia content. In the case where the user wants to limit the action features, the generation-constrained multimodal representation may include the features of depth displacement. It can be understood that in the case where the content reference material includes multiple, the generation-constrained multimodal information representation may also be multiple. For example, corresponding to the reference multimedia content including pictures and videos, the corresponding description information includes "clothing is consistent with the clothing of the characters in the picture" and "action is consistent with the action of the characters in the video". The multimodal pre-training model can output the constraint multimodal information representations corresponding to the picture-clothing and video-action respectively.

[0097] Then, the electronic device 100 can input the generation constrained multimodal information representation and the noise features that conform to the Gaussian distribution into the denoising model, and denoise the noise features that conform to the Gaussian distribution based on the generation constrained multimodal information representation. After obtaining the target visual features (i.e., the denoised visual potential representation mentioned above), the target visual features are used to generate multimedia content through an image generator (i.e., an image generation model) or a video generator (i.e., a video generation model). The generated multimedia content has a target subject and target attributes, that is, an image or video that meets user needs is obtained.

[0098] Based on the user input described above Fig. 6A , Figure 6B , Figure 6C In the embodiment, the electronic device 100 may generate in step S502 Fig.6D , that is, generate Fig. 6A Medium Kitten Based Figure 6B Actions, Figure 6C background hurdle video to obtain multimedia content that meets user needs.

[0099] The embodiments of the present application can directly generate content based on the materials provided by the user, avoiding the problem of non-existence or incompatibility of materials caused by querying the material library; and, since the content reference material is real material, such as a video of a puppy hurdling in reality, the video generated based on it will not be out of touch with reality, for example, in the generated video, the kitten jumps too high, the jumping direction does not match the direction of the runway, the background does not match the reality, etc., thereby being able to generate multimedia content that meets the user's expectations, and, when there are a large amount of depth displacement in the content reference material, for example, when there are a large number of limb movement details in the hurdle jumping action of a puppy, by adding noise to the visual latent representation and denoising it based on the generated constrained multimodal information representation that includes the depth displacement features, these depth displacement features can be retained in the denoised visual latent representation, so that the generated multimedia content can fully restore the user's expected actions, that is, reflect the above-mentioned limb movement details.

[0100] The following is based on Figure 7 A multimedia content generation method according to an embodiment of the present application is further introduced.

[0101] S701: Acquire subject reference material.

[0102] It can be understood that the subject reference material may be an image or video containing a target subject, which is used to assist in determining the target subject and provide visual features of the target subject.

[0103] In an optional embodiment, the user can operate the electronic device 100, that is, Figure 4A or Figure 4B Click the upload icon below the prompt "Please enter the view material containing the subject" shown, and select an image or video containing the target subject from the gallery of the electronic device 100. It can be understood that the image or video can include one or more, for example, one image and one video, or two images, or three videos, etc. The target subject can be an object such as a person, a pet, a building, or food.

[0104] According to one example, if the target subject that the user wants is a plainclothes man, then the user can first click Figure 4A Upload the first image containing the plainclothes man, then click the first icon below the "Visual material containing the subject" prompt shown in the figure. Figure 4A Click the second icon below the "Visual material containing the subject" prompt shown, and continue to upload the second picture containing the plainclothes man.

[0105] According to one example, a user can enter Figure 4A In the interface shown in FIG. 1 , select a view content from the gallery to upload. According to another example, the user can enter the Figure 4B In the interface shown, click the camera icon to enter the shooting preview interface, for example Figure 4C The user can then click Figure 4C The photo or video button at the bottom of the interface shown starts real-time shooting. It can be understood that whether in the shooting process, or in the preview before shooting, or after shooting, the user can select the target subject by clicking on the subject object. Optionally, for the same object, when the user clicks for the first time, the selection box of the object can be displayed, and when the user clicks for the second time, the selected object can be determined as the subject. At this time, if it is still in the preview display interface before shooting, the solid line detection frame can be used to intelligently track the subject. If the user needs to calibrate more than one target subject object, multiple objects can be clicked and confirmed separately to add target subjects. After the user completes the selection of the target subject, he can click the Complete Target Subject Selection button to complete the operation.

[0106] S702: Obtain content reference material.

[0107] It can be understood that the content reference material includes images or videos required by the user, which are used to assist in generating multimedia content that meets the user's needs. The content reference material may include a first content reference material, a second content reference material, etc., and this application does not limit the number of content reference materials.

[0108] In an optional embodiment, the user can operate the electronic device 100, that is, Figure 4A or Figure 4B Click the upload icon below the "Please enter the guidance view information" shown, and select the image or video containing the user's needs from the gallery of the electronic device 100. It can be understood that the image and video can include one or more, and the present application does not limit the number and type of images or videos. The content of the image or video may include the action that the user wants to generate, such as splits, and the image or video may be a video with a specific motion mode subject, such as dancing, running, etc.; it may also be a view content containing specific elements, such as a specific place background, specific weather, specific picture, specific clothing, specific picture style, etc., that is, a view content with the same or similar attributes as the multimedia content that the user wants to generate. Optionally, the specific picture style may include oil painting style, black and white, etc.

[0109] According to one example, in the generated multimedia content, if the user wants the target subject to push a car against a natural scenery background, the user can first click Figure 4A or Figure 4BThe first icon below the "Please enter the guide view information" prompt is shown, upload the first picture, in which the subject is pushing the cart, and then click Figure 4A or Figure 4B Click the second icon below the "Please enter the guide view information" prompt and continue to upload the second picture, which has a natural scenery background.

[0110] S703: Obtain customized requirement indication information.

[0111] It can be understood that the customized requirement indication information is the description information mentioned above.

[0112] In an optional embodiment, the user can operate the electronic device 100, that is, Figure 4A or Figure 4B In the View Generation Guide text box shown, click the button and enter a description, such as guide text, or Figure 4A or Figure 4B Click the audio icon on the right side of the generated guide audio to input the audio. It can be understood that after the electronic device 100 receives the audio input by the user, it can perform voice recognition on the audio to obtain the generated guide text, that is, the description information. Optionally, voice recognition can be continued based on the voice recognition model.

[0113] The customized requirement indication information may correspond to the image or video input by the user.

[0114] In an optional embodiment, the user inputs subject reference material 1 including a plainclothes man, subject reference material 2 including a plainclothes man, content reference material 1 in which the subject is pushing a cart, and content reference material 2 with a natural scenery background. Accordingly, the user can input the following customized requirement instruction information: "Generate a picture, the picture includes subject reference material 1 and the man in subject reference material 2, the action is changed to the action in content reference material 1, and the picture background is consistent with content reference material 2."

[0115] In an optional embodiment, the user inputs a subject reference material 1 including a man and a content reference material 1 including a woman skating. Accordingly, the user can input the following customized requirement instruction information: "Change the man on the right side of the subject reference material 1 to wear a white shirt, a black vest and trousers to skate on the ice rink, and the skating action is the same as the action of the woman in the content reference material 1" or "Change the action to the same as the guide Figure 1 The women do the same skating moves.”

[0116] It should be noted that the material names such as main reference material 1, main reference material 2, content reference material 1, content reference material 2, etc. mentioned in this application are only examples. In other optional embodiments, the material input by the user may also have other names, such as guide material 1, guide material 2, etc. Figure 1 , Guide Video 1, Reference Figure 1 ,etc.

[0117] S704: Divide the customized requirement indication information into target subject selection auxiliary information and generation constraint information.

[0118] Optionally, in response to the user clicking the customized view generation button, the customized requirement indication information can be divided into visual subject selection auxiliary information and generation constraint information. The target subject selection auxiliary information can be information in the description information that is used to specify the target subject in the subject reference material, such as information describing the position and features of the target subject; the generation constraint information can be information in the description information that limits the attributes of the target subject in the generated content, such as information describing the actions and clothing of the target subject.

[0119] As an example, the customized requirement indication information is "generate a picture, the picture includes the man in subject reference material 1 and subject reference material 2, the action is changed to the action in content reference material 1, and the background of the picture is consistent with content reference material 2", then the target subject selection auxiliary information obtained by division can be "the man in subject reference material 1 and subject reference material 2", and the generation constraint information can be "the action is changed to the action in content reference material 1, and the background is consistent with content reference material 2".

[0120] In an optional implementation, the name of the subject reference material can be used as a keyword for identification, such as subject reference material 1 and subject reference material 2, and identified in the customized demand information to obtain the target subject auxiliary information. Alternatively, the division can be performed according to the word combination of [adjective][character noun][verb][adjective][character noun] in the description information. Most of the [verbs] are verbs with modification semantics such as "change to", "draw into", "become", and "adjust to". Therefore, verbs with modification semantics (i.e., part-of-speech patterns), such as the [adjective][noun] content before "change to", can be divided into auxiliary information for selecting the main view image, that is, "the man in the subject reference material 1 and the subject reference material 2" is divided into auxiliary information for selecting the main view image, and the content after "change to" is divided into generation constraint information, that is, "the action is changed to the action in the content reference material 1, and the background is consistent with the content reference material 2" is divided into generation constraint information.

[0121] In the embodiment where the customized requirement information is "Change the man on the right in the main reference material 1 to wear a white shirt, a black vest and trousers and skate on the ice rink, and the skating movements are the same as the woman's movements in the content reference material 1", based on the same reason, the content before "Change to" can be divided into view main image selection auxiliary information, and the content after "Change to" can be divided into generation constraint information.

[0122] In some embodiments, if the target subject selected auxiliary information is not detected in the customized requirement indication information, that is, the keyword of the subject reference material is not detected, the customized requirement indication information can be directly determined as the generation constraint information.

[0123] S705: Perform target subject detection on the subject reference material to determine the target subject.

[0124] In an optional embodiment, the candidate target subject can be detected based on the video and image in all subject reference materials provided by the user. Among them, there may be one or more subjects in the subject reference material provided by the user, including the target subject. The following describes how to determine the target subject specified by the user from one or more subjects. For example, the video containing the target subject input by the user is converted into a video frame, and each video frame and each image containing the target subject are subjected to significant subject detection, thereby obtaining the significant subject objects that exist simultaneously in each subject reference material (i.e., each image or video containing the target subject) as the candidate target subject. If there are no significant subject objects that exist simultaneously, such as subject reference material 1 includes man 1 and woman 1, and subject reference material 2 includes man 2 and woman 2, and there are no overlapping significant subject objects between the two, then the user is prompted that a suitable target subject object has not been found, such as displaying a corresponding prompt message on the display interface of the electronic device 100.

[0125] The following describes an exemplary process for determining a target subject in multimedia content based on candidate target subjects:

[0126] 1) If the number of candidate target subjects is 1, the candidate target subject is directly determined as the target subject, and the process ends;

[0127] If the number of candidate target subjects is greater than 1, the view focus object at the shooting moment is calculated by analyzing the depth information and focal length information, and the candidate target subjects that do not belong to the view focus object, that is, the candidate target subjects that are blurred and not focused during shooting, are deleted to obtain the focused target subject, and go to step 2).

[0128] 2) If the number of the focused target subject is 1, the focused target subject is directly determined as the target subject, and the process ends;

[0129] Corresponding to the number of focused target subjects being greater than 1, if there is auxiliary information for selecting the target subject, a text-image semantic matching model is used to calculate semantic similarity based on the text features of the auxiliary information for selecting the target subject and the visual features of each focused target subject, and the focused target subjects with similarities lower than the similarity threshold are deleted to obtain similar target subjects, and then go to step 3). For example, the similarity threshold may be 0.5. It should be noted that the embodiment of the present application does not limit the similarity threshold, and the similarity threshold may also be other optional values ​​besides 0.5.

[0130] 3) If the number of similar target subjects is 1, the similar target subject is directly determined as the target subject, and the process ends;

[0131] If the number of similar target subjects is greater than 1, then the ranking of similar target subjects is obtained based on the calculated similarity, the number of times the similar target subject appears in the subject reference material, and the imaging quality score of the similar target subject in the picture. The higher the ranking, the higher the comprehensive score of the above evaluation factors. Then, for each similar target subject, the best quality image or video frame of the similar target subject is found through the imaging evaluation algorithm, and then the image of the similar target subject is segmented from the image or video frame through the image segmentation algorithm, and displayed in the pop-up page according to the ranking, that is, the images of each similar target subject are displayed in sequence according to the ranking of the similar target subjects, and the user is prompted to manually select and determine the target subject object, and then the target subject is determined according to the user's selection, and the process ends.

[0132] It can be understood that after the target subject is determined, the image of the target subject in each subject reference material can be determined at the same time, and the image of the target subject can be processed in subsequent steps.

[0133] S706: Input the image of the target subject into a pre-trained image encoder to obtain visual features of the target subject.

[0134] The target subject visual features may be the target subject visual features corresponding to each subject reference material, and specifically may be image feature vectors of the target subject images in each subject reference material, which are used to describe or represent information of the target subject image, and may be in the form of a multi-dimensional feature vector, such as 4096 dimensions. It can be understood that the target subject image is the part of the target subject in the subject reference material, such as a picture containing the target subject.

[0135] It can be understood that the image of the target subject is an image of the target subject in the subject reference material whose imaging quality meets the preset requirements. Optionally, an imaging evaluation algorithm can be used to find a preset number of images or video frames with the best quality of the target subject from all subject reference materials containing the target subject, such as a preset number of images or video frames with the highest clarity of the target subject, and then segment the images or video frames using an image segmentation algorithm to obtain an image for the target subject.

[0136] In an optional embodiment, a preset number of target subject images can be respectively input into a trained visual geometry group (VGG)-16 encoder to extract image features of each target subject image, and obtain a preset number of fixed-length target subject visual features. Optionally, the preset number can be 3, and the fixed length can be 4096 dimensions, which is not limited in the embodiment of the present application.

[0137] S707: Input the target subject's visual features into the pre-trained variational encoder to obtain a visual latent representation.

[0138] The variational encoder may be a variational autoencoder composed of an encoder and a decoder. In the embodiment of the present application, the variational encoder supports multi-feature input, which is used to further extract features from multiple features to obtain image features of the target subject, that is, visual potential representation.

[0139] In an optional embodiment, a preset number of target subject visual features obtained may be input into a pre-trained variational encoder that supports multi-feature input to generate a visual latent representation.

[0140] Optionally, the variational encoder may use an encoder based on a Transformer network, and the input supports a maximum length of N. In this embodiment, if the number of images of the target subject exceeds N, for example, the subject reference material is a video containing N+1 frames, in which there are N+1 images of the target subject, the images of the target subject are scored based on the imaging quality assessment algorithm, and the top N target subject images with the highest scores are selected, and their corresponding target subject visual features are input into the variational encoder to obtain a visual latent representation.

[0141] It is understood that the scoring standard may be imaging quality, that is, the imaging quality of the image of the target subject is evaluated to obtain an imaging quality score. Specifically, the evaluation method may include evaluating color saturation, evaluating contrast, determining whether there is a clear visual focus in the image, evaluating clarity, and the like.

[0142] It should be noted that the embodiment of the present application does not limit the type of variational encoder. In other optional embodiments, the variational encoder may also adopt other networks or other models.

[0143] S708: Perform noise processing based on the visual potential representation to obtain noise features that conform to the Gaussian distribution.

[0144] Optionally, the visual latent representation can be subjected to multiple rounds of step-size noise addition through a diffusion model to generate noise features that conform to the Gaussian distribution, namely, diffusion image features.

[0145] It can be understood that noise addition can be achieved through the noise addition module in the diffusion model. As an image generation model, the diffusion model includes a noise addition module, which is used to perform the forward diffusion process, that is, noise addition, adding Gaussian noise to the visual potential representation until the data becomes random noise, and obtaining noise features that conform to the Gaussian distribution. After obtaining the noise features of the Gaussian distribution, the denoising model can be used to perform the reverse generation process, that is, by denoising the random noise, the target visual features corresponding to the multimedia content to be generated can be obtained.

[0146] It can be understood that Gaussian noise is white noise, which is a random noise that follows a normal distribution. In deep learning, Gaussian noise is added to the data during training to improve the robustness and generalization ability of the model and achieve data expansion. By adding noise to the input data, the model is forced to learn features that are robust to small changes in the input, which can help improve model performance.

[0147] S709: Input the generation constraint information and content reference material into the multimodal pre-training model to obtain a generation constraint multimodal information representation.

[0148] It can be understood that the generated constraint information and content reference material are input into the multimodal pre-training model, that is, text + image or text + video are input into the multimodal training model, wherein the text corresponds to the image or video in a one-to-one relationship, and the text is used to describe the target attribute in the image or video. Specifically, through secondary analysis technology, based on the name of the content reference material, such as "content reference material 1" and "content reference material 2", each content reference material can be matched with the subsequent part-of-speech pattern content such as [adjective] [noun] that follows.

[0149] For example, in an embodiment where the generation constraint information is "the action is changed to the action in content reference material 1, and the background is consistent with content reference material 2", content reference material 1 can be paired with "action", and content reference material 2 can be paired with "background", and the multimodal pre-trained model can be input to output the generation constraint multimodal information representation.

[0150] It can be understood that in the above embodiment, the output of the multimodal pre-training model is a generation-constrained multimodal information representation of the target attribute. The target attribute includes a first attribute, namely, an action attribute, and a second attribute, namely, a background attribute.

[0151] For example, the generated constraint information is "Change to guide Figure 1 The clothing of the man in the video is used to guide the dance movements of the man in video 2 to generate a video", then the part-of-speech analysis technology is needed based on the labeling word "guide Figure 1 " and "Guide video 2" to match the corresponding view content with the following part-of-speech pattern content such as [adjective] [noun], for example, guide Figure 1 The guide video 1 will be paired with "men's clothing" and the guide video 2 will be paired with "men's dance movements" and input into the multimodal pre-training model. It can be understood that in the above embodiment, the output of the multimodal pre-training model is a generation-constrained multimodal information representation of the target attribute. The target attribute includes a first attribute, namely the clothing attribute, and a second attribute, namely the action attribute. It can be understood that the multimodal representation learning model can be trained based on multiple tasks such as visual question answering and semantic segmentation, and the corresponding visual semantic features can be generated based on the text content through methods such as cross-attention mechanisms, so that the visual semantic features of elements such as background, action, weather, clothing, etc. in the view content can be extracted to generate constrained multimodal information representation.

[0152] S710: Perform denoising processing based on noise features that conform to Gaussian distribution and generate constrained multimodal information representation to obtain target visual features.

[0153] Among them, denoising is reverse generation, that is, based on the generative constraint multimodal information representation, the noise features that conform to the Gaussian distribution are denoised to obtain image features close to the generative constraint multimodal information representation, that is, the target visual features (that is, the denoised visual potential representation mentioned above). It can be understood that by executing steps S708 and S710, the target visual features are generated based on the visual potential representation and the generative constraint multimodal information, and the specific implementation method is noise addition and denoising.

[0154] In an optional embodiment, the noise features that conform to the Gaussian distribution can be subjected to multiple rounds of denoising through a denoising model based on the generative constrained multimodal information representation to obtain the target visual features. Optionally, the denoising model can be a semantic segmentation (UNet) network.

[0155] It should be noted that the embodiments of the present application do not limit the type of denoising model. In some other optional embodiments, the denoising model may also adopt other models.

[0156] S711: Generate customized multimedia content based on target visual features.

[0157] The method of generating multimedia content may include: inputting the target visual features into a generation model, and outputting the multimedia content. Optionally, the generation model may be a generation model of an image or video, such as a variational encoder for image generation or video generation.

[0158] Among them, multimedia content can be generated according to the user's generation format intention. For example, if the customized demand indication information input by the user includes the generation format intention of generating a picture, such as "generate picture", then a customized picture can be generated; if the customized demand indication information input by the user includes the intention of generating a video, such as "generate video", then a customized video can be generated.

[0159] As an example, the generation format intention of the user in the customized demand indication information can be judged based on technologies such as keyword matching and grammatical analysis. For example, "generate a picture with the same action as in content reference material 1" can be judged based on the noun subject "picture" in the object "picture with the same action as in content reference material 1" that the view format that the user expects to generate is a picture. It can be understood that after determining the generation format intention, a variational decoder for picture generation or video generation can be selected based on the generation format intention to generate view content to obtain user customized multimedia content.

[0160] It can be understood that the subject of the multimedia content generated in the embodiment of the present application is the target subject and has target attributes, and the target attributes include the first attribute and the second attribute mentioned above. It should be noted that the present application does not limit the number of attributes and attribute categories included in the target attributes.

[0161] The following is based on Figure 8 The following further introduces an exemplary process of a multimedia content generation method according to an embodiment of the present application. Figure 8 As shown in the figure, the multimodal guidance information analysis module (802) receives multimodal guidance information (801) such as text, voice, image, etc., that is, receives content reference materials, subject reference materials and description information, and obtains subject selection auxiliary information (803) and multimodal generation constraint information.

[0162] It can be understood that the subject selection auxiliary information (803) can be the description information related to the subject reference material in the description information. For the obtained subject selection auxiliary information (803), based on the subject selection auxiliary information (803) and the input image and video (804), the visual subject visual features (805) can be obtained through algorithms such as visual semantic segmentation and multimodal semantic matching, and then the visual potential representation (806) is generated through the variational encoder, and the visual potential representation is denoised with a step length of t rounds based on the diffusion model (807) to obtain the noise feature.

[0163] It can be understood that the multimodal generation constraint information can be obtained by inputting the content reference material and the description information related to the content reference material into the multimodal guidance information analysis module (802). The obtained multimodal generation constraint information can be input into the multimodal pre-training model (809) to output the generation constraint multimodal information representation. Then, based on the UNet model and the like, the noise features and the generation constraint multimodal information representation are denoised with a step length of t rounds (808), and finally the view content is generated through the variational encoder (810) to obtain the customized view content (811), that is, the multimedia content that the user wants to generate.

[0164] In some optional embodiments, the customized view content (811) is a video, and the first frame of the video can be generated first, and then all frames of the video can be generated to obtain the customized video. Fig. 9 First, the multimodal guidance information (901) such as text, speech, and image is input into the multimodal pre-training model (902) to obtain the generative constraint multimodal information representation, and then the generative constraint multimodal information representation and the noise features obtained after t rounds of noise addition are denoised with a step length of t rounds based on models such as UNet (903), and then the view content is generated through a variational encoder (904) to obtain a customized video start frame (905). Finally, the generative constraint multimodal information representation and the customized video start frame (905) are input into the video generation model based on the start frame (906) to obtain a customized video (907). Through the above steps, the first frame of the video can be generated first, and then the video can be generated, so that the constraint degree of video generation is increased, so that the generated customized video is more in line with user needs, that is, it is closer to the content reference material and the main reference material provided by the user, and the quality of the generated video picture can be guaranteed.

[0165] According to one example, combining Figure 8 ,refer to Fig.10 , text, voice, image and other multimodal guidance information (801) may include guidance Figure 1 (i.e. the content reference material mentioned above), guide image 2 (i.e. the content reference material mentioned above), guide text (i.e. the description information mentioned above), and the input image and video (804) may include main view material 1 and main view material 2. Among them, the guide text is "changed to guide Figure 1 It can be understood that if the guide text does not include a description of the visual subject, the object that exists in both the subject view material 1 and the subject view material 2, namely, the man 03, can be identified as the visual subject. Then, by executing Figure 8 The illustrated process generates customized images, i.e. Fig.10 The figure in the figure is based on man 03, action and guidance Figure 1 Consistent, background and guide Fig. 2 Consistent generation effect diagram.

[0166] Fig.11 A schematic diagram of the hardware structure of the electronic device 100 is shown.

[0167] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0168] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0169] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0170] A memory may also be provided in the processor 110 for storing instructions and data corresponding to a multimedia content generation method provided in an embodiment of the present application. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or circulated. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. Repeated access is avoided, the waiting time of the processor 110 is reduced, and the efficiency of the system is improved.

[0171] In some embodiments, the processor 110 of the electronic device 100 calls program instructions stored in the memory to execute the multimedia content generation method mentioned in the present application according to the obtained program instructions, for example: obtaining the main reference material, content reference material and description information input by the user; generating multimedia content, wherein the main body of the multimedia content includes: the target subject described in the description information and in the main reference material, and the multimedia content has a target attribute, the target attribute includes a first attribute and a second attribute, the first attribute is the first attribute in the content reference material constrained by the description information, and the second attribute is a feature outside the content reference material constrained by the description information.

[0172] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0173] The electronic device 100 implements the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.

[0174] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0175] The electronic device 100 can realize the shooting function through ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0176] ISP is used to process the data fed back by camera 193. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 193.

[0177] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0178] Video codecs are used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. Thus, the electronic device 100 may play or record videos in a variety of coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0179] The key 190 includes a power key, a volume key, etc. The key 190 may be a mechanical key or a touch key, such as a camera application shooting key, etc. The electronic device 100 may receive key input and generate key signal input related to user settings and function control of the electronic device 100.

[0180] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present invention, the Android system of the layered architecture is taken as an example to exemplify the software structure of the electronic device 100.

[0181] Fig.12 1 is a software structure block diagram of the electronic device 100 according to an embodiment of the present invention.

[0182] The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system library, and the kernel layer.

[0183] The application layer can include a series of application packages. Figure 8 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, dual SIM and mobile network. Among them, users can wake up the camera application to take images and record videos in real time.

[0184] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0185] like Fig.12 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0186] Android Runtime includes core libraries and virtual machines. Android runtime is responsible for scheduling and management of the Android system.

[0187] The system library may include multiple functional modules, such as surface manager, media libraries, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0188] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, and sensor driver.

[0189] Accordingly, an embodiment of the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor for executing instructions of the above-mentioned multimedia content generation method.

[0190] Accordingly, an embodiment of the present application provides a readable medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device executes the above-mentioned multimedia content generation method.

[0191] This specification provides method or process operation steps as shown in the embodiments or flow charts, but more or fewer operation steps may be included based on conventional or non-creative work. The order of steps listed in the embodiments is only one of many execution orders and does not represent the only execution order. In actual execution, the method or process shown in the embodiments or drawings may be executed in sequence or in parallel (for example, in a parallel controller or multi-threaded processing environment).

[0192] The various embodiments disclosed in the present application may be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application may be implemented as a computer program or program code executed on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0193] Program code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0194] Program code can be implemented with high-level programming language or object-oriented programming language to communicate with the processing system. When necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.

[0195] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including, but not limited to, floppy disks, optical disks, optical disks, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Therefore, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a machine (e.g., computer) readable form.

[0196] As used herein, the term "module" may refer to, be part of, or include: a memory (shared, dedicated, or group) for running one or more software or firmware programs, an application-specific integrated circuit (ASIC), an electronic circuit and / or processor (shared, dedicated, or group), a combinational logic circuit, and / or other suitable components that provide the described functionality.

[0197] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order is not required. Instead, in some embodiments, these features may be described in a manner and / or order different from that shown in the illustrative drawings. In addition, the structural or method features included in a specific figure do not mean that all embodiments need to include such features. In some embodiments, these features may not be included, or these features may be combined with other features.

[0198] The embodiments of the present application are described in detail above in conjunction with the accompanying drawings, but the use of the technical solution of the present application is not limited to the various applications mentioned in the embodiments of the present patent. Various structures and variations can be easily implemented with reference to the technical solution of the present application to achieve the various beneficial effects mentioned herein. Various changes made within the knowledge of ordinary technicians in the field without departing from the purpose of the present application should all fall within the scope of the patent application.

Claims

1. A method for generating multimedia content, characterized in that: include: Acquire the subject reference material, the first content reference material and the description information input by the user; Generate multimedia content, wherein the subject of the multimedia content includes: a target subject in the subject reference material, and The multimedia content has a target attribute, which includes a first attribute and a second attribute. The first attribute is an attribute in the first content reference material constrained by the description information, and the second attribute is an attribute outside the first content reference material constrained by the description information.

2. The method according to claim 1, characterized in that Also includes: Obtain the second content reference material input by the user, The second attribute is an attribute of the second content reference material constrained by the description information.

3. The method according to claim 1, characterized in that The first attribute and the second attribute are respectively at least one of an action attribute, a background attribute, a color attribute, a style attribute, a clothing attribute, an expression attribute, and a weather attribute.

4. The method according to claim 1, characterized in that: The main body of the multimedia content includes: the target main body described in the description information.

5. The method according to claim 1, characterized in that The main body of the multimedia content includes: a target main body whose definition is greater than a preset definition in the main body reference material.

6. The method according to claim 1, characterized in that Also includes: Display multiple alternative subjects; The first candidate subject selected by the user from the plurality of candidate subjects is used as the target subject. The similarity between the candidate subject and the target subject described by the description information is greater than a similarity threshold.

7. The method according to claim 1, characterized in that The number of the subject reference material is at least one, and each of the subject reference materials contains an image of the target subject; The generating of multimedia content comprises: determining an image feature of the target subject according to an image of a target subject of each subject reference material in at least one of the subject reference materials; Determining the visual semantic features of the target attribute based on the description information and the first content reference material; The multimedia content is generated based on the image features of the target subject and the visual semantic features of the target attributes.

8. The method according to claim 7, characterized in that The determining the image feature of the target subject according to the image of the target subject of each subject reference material in at least one of the subject reference materials comprises: Based on an image of a target subject in each of the subject reference materials in at least one of the subject reference materials, determining a visual feature of the target subject corresponding to each of the subject reference materials; Based on the visual features of the target subject corresponding to the respective subject reference materials, the image features of the target subject are determined.

9. The method according to claim 8, characterized in that The determining, based on the image of the target subject in each of the subject reference materials in at least one of the subject reference materials, the target subject visual features corresponding to each of the subject reference materials comprises: The image of each target subject in at least one of the subject reference materials is input into an image encoder for feature extraction to obtain visual features of the target subject corresponding to each subject reference material.

10. The method according to claim 8, characterized in that The determining of the image features of the target subject based on the visual features of the target subject corresponding to each subject reference material includes: The visual features of the target subject corresponding to each subject reference material are input into the variational encoder for feature extraction to obtain the image features of the target subject.

11. The method according to claim 7, characterized in that The determining of the visual semantic features of the target attribute based on the description information and the first content reference material includes: The description information of the target attribute and the first content reference material are input into a visual semantic feature extraction model to perform feature extraction to obtain visual semantic features of the target attribute.

12. The method according to claim 7, characterized in that The generating the multimedia content based on the image features of the target subject and the visual semantic features of the target attribute includes: Performing noise processing on the image features of the target subject to obtain extended image features, wherein the extended image features include noise features; Based on the visual semantic features of the target attributes, denoising the noise features in the extended image features to obtain target visual features; The multimedia content is generated based on the target visual feature.

13. The method according to claim 12, characterized in that The step of performing noise processing on the image features of the target subject to obtain expanded image features includes: Inputting the image features of the target subject into a diffusion model for noise addition processing to obtain the expanded image features; The step of performing denoising on the noise feature in the extended image feature based on the visual semantic feature of the target attribute to obtain the target visual feature includes: The visual semantic features of the target attributes and the extended image features are input into a denoising model, and the noise features in the extended image features are denoised to obtain the target visual features.

14. The method according to claim 12, characterized in that The multimedia content includes one of a target image and a target video. The generating the multimedia content based on the target visual feature comprises: Inputting the target visual features into an image generation model to generate the target image; or; The target visual features are input into a video generation model to generate the target video.

15. The method according to claim 12, characterized in that The multimedia content includes a target video, The generating the multimedia content based on the target visual feature comprises: Inputting the target visual features into an image generation model to generate a first frame of the target video; The first frame of the target video and the visual semantic features of the target attributes are input into a video generation model to generate the target video.

16. The method according to any one of claims 1 to 15, characterized in that: The main reference material includes at least one of an image and a video.

17. The method according to any one of claims 1 to 15, characterized in that: The first content reference material includes at least one of an image and a video.

18. An electronic device, characterized in that: include: a memory for storing instructions to be executed by one or more processors of the electronic device, and A processor, configured to execute instructions of the multimedia content generating method according to any one of claims 1 to 17.

19. A readable medium, characterized in that The readable medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the multimedia content generation method according to any one of claims 1 to 17.

Citation Information

Cited By

  • Content generation method and device, medium, electronic equipment and program product

    CN121304842A

  • Multimedia content generation method and device, electronic equipment and storage medium

    CN121527240A