Multimedia resource generation method and device, electronic equipment and storage medium

By recognizing the reference image, obtaining image description information and subject form information, and generating multimedia resources with text prompt words, solving the problem of special effects blocking key content and achieving high-quality multimedia resource generation.

CN120448564AActive Publication Date: 2025-08-08BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510941836.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-08-08
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

In the prior art, when generating images, special effects easily obscur the key content in the reference image, resulting in a lower quality of the generated image.

Method used

By recognizing the reference image, obtaining image description information and subject form information, and generating multimedia resources with text prompt words, ensuring that the location of the special effects is different from that of the subject object and avoiding occlusion.

Benefits of technology

It improves the quality of multimedia resources, avoids special effects to block key content, is simple to operate, and improves generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448564A_ABST
    Figure CN120448564A_ABST
Patent Text Reader

Abstract

The invention provides a multimedia resource generation method and device, electronic equipment and a storage medium, and belongs to the technical field of multimedia. The method comprises the following steps: acquiring an input reference image and a text cue word; image recognition is conducted on the reference image through the image processing model, image description information and main body form information of the reference image are obtained, the image description information comprises the category of at least one main body object in the reference image, and the main body form information is used for indicating the position of the at least one main body object in the reference image; and based on the image description information, the subject form information, the reference image and the text cue word, generating a multimedia resource, the multimedia resource comprising at least one subject object and a special effect, and the position of the special effect in the multimedia resource being different from the position of the at least one subject object. According to the method, the generated special effect can be prevented from shielding key contents such as a main object in the reference image, and the quality of multimedia resources is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of multimedia technology, and in particular to a method, device, electronic device, and storage medium for generating multimedia resources. Background Art

[0002] With the continuous development of multimedia technology, text-to-image generation technology and image-to-image generation technology have become commonly used image generation methods. How to generate high-quality images is the focus of research in this field.

[0003] In related technologies, a commonly used approach is to input a reference image and text prompt words provided by a user into an image generation model, and then add special effects indicated by the text prompt words to the reference image through the image generation model to generate a new image.

[0004] However, the aforementioned technical solution simply adds special effects directly to the reference image without considering the specific content of the reference image. This results in the special effects in the generated image obscuring key content of the reference image, resulting in lower quality. For example, if the reference image is a person sitting on a chair and the prompt is "flowers growing," the flowers in the generated image are located on the person's face and obscure the portrait, creating the uncomfortable effect of the portrait being transformed into a flower. Summary of the Invention The present disclosure provides a method, device, electronic device, and storage medium for generating multimedia resources, which can prevent generated special effects from obscuring key content such as main objects in a reference image, thereby improving the quality of multimedia resources. The technical solution of the present disclosure is as follows.

[0005] According to one aspect of an embodiment of the present disclosure, a method for generating multimedia resources is provided, comprising: Acquire an input reference image and a text prompt word, wherein the reference image is used to provide a main object for generating a multimedia resource, and the text prompt word is used to indicate a special effect to be generated in the multimedia resource; performing image recognition on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image, wherein the image description information includes a category of at least one subject object in the reference image, and the subject morphology information is used to indicate a position of the at least one subject object in the reference image; The multimedia resource is generated based on the image description information, the subject morphology information, the reference image and the text prompt word, wherein the multimedia resource includes the at least one subject object and the special effect, and the position of the special effect in the multimedia resource is different from the position of the at least one subject object. The multimedia resource is an image or a video.

[0006] According to another aspect of an embodiment of the present disclosure, there is provided a device for generating multimedia resources, including: A first acquisition unit is configured to acquire an input reference image and a text prompt word, wherein the reference image is used to provide a main object for generating a multimedia resource, and the text prompt word is used to indicate a special effect to be generated in the multimedia resource; a recognition unit configured to perform image recognition on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image, wherein the image description information includes a category of at least one subject object in the reference image, and the subject morphology information is used to indicate a position of the at least one subject object in the reference image; A generation unit is configured to generate the multimedia resource based on the image description information, the subject morphology information, the reference image, and the text prompt word, wherein the multimedia resource includes the at least one subject object and the special effect, and the position of the special effect in the multimedia resource is different from the position of the at least one subject object, and the multimedia resource is an image or a video.

[0007] In some embodiments, the generating unit includes: A first generating subunit is configured to generate resource description information based on the image description information, the subject morphology information, and the text prompt word, wherein the resource description information is used to indicate a positional relationship to be satisfied between the special effect and the at least one subject object; The second generating subunit is configured to generate the multimedia resource based on the resource description information, the reference image and the subject morphology information.

[0008] In some embodiments, the first generating subunit is configured to process the image description information, the subject morphology information, and the text prompt word through a text fusion model to obtain the resource description information; The second generating subunit is configured to execute processing on the resource description information, the reference image and the subject morphology information based on the resource generation model to obtain the multimedia resource.

[0009] In some embodiments, the apparatus further comprises: A second acquiring unit is configured to acquire resource prompt information, where the resource prompt information is used to indicate a condition that the multimedia resource must satisfy; The first generating subunit is configured to execute a text fusion model to process the image description information, the subject morphology information, the text prompt word and the resource prompt information to obtain the resource description information.

[0010] In some embodiments, the second acquiring unit is configured to perform any of the following: Determining the resource prompt information based on the reference image and the text prompt word; Determining the resource prompt information based on the style of the multimedia resource; In response to the prompt application instruction, the input resource prompt information corresponding to the prompt application instruction is obtained.

[0011] In some embodiments, the second acquisition unit is configured to determine the resource prompt information based on the category of the at least one main object in the reference image and the category of the special effect indicated by the text prompt word.

[0012] In some embodiments, the apparatus further comprises: An output unit, configured to output the resource description information; The second generating subunit is configured to execute an application instruction in response to the resource description information, and generate the multimedia resource based on the resource description information, the reference image, and the subject morphology information.

[0013] In some embodiments, the resource description information includes multiple items, and the positional relationship between the special effect and the at least one main object in different resource description information is different; The output unit is configured to output multiple resource description information; The second generating subunit is configured to execute an application instruction in response to any one of the multiple pieces of resource description information, and generate the multimedia resource based on the resource description information, the reference image and the main form information.

[0014] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, the electronic device including: one or more processors; a memory for storing program codes executable by the processor; The processor is configured to execute the program code to implement the above-mentioned method for generating multimedia resources.

[0015] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When a program code in the computer-readable storage medium is executed by a processor of an electronic device, the electronic device can execute the above-mentioned method for generating multimedia resources.

[0016] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program / instruction, which implements the above-mentioned method for generating multimedia resources when executed by a processor.

[0017] The solution provided by the embodiment of the present disclosure, in the process of generating multimedia resources based on a reference image and text prompt words, first performs image recognition on the reference image to determine the image description information and subject morphology information of the reference image, and then generates the multimedia resources through the category of the subject object indicated by the image description information, the position of the subject object indicated by the subject morphology information in the reference image, the reference image and the text prompt words. In the process of generating multimedia resources, the position of the subject object can be avoided, and the special effects indicated by the text prompt words can be generated at other positions outside the subject object, thereby avoiding the generated special effects from obscuring key content such as the subject object in the reference image, and improving the quality of the generated multimedia resources. In addition, the user does not need to input the descriptive text for indicating the position of the special effects, but only needs to provide a simple text prompt word to indicate the special effects, so as to generate high-quality multimedia resources. The operation is simple, which is conducive to improving the generation efficiency of multimedia resources.

[0018] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0020] Figure 1 The figure is a schematic diagram showing an implementation environment of a method for generating multimedia resources according to an exemplary embodiment.

[0021] Figure 2 The figure is a flowchart of a method for generating multimedia resources according to an exemplary embodiment.

[0022] Figure 3 The figure is a flowchart of another method for generating multimedia resources according to an exemplary embodiment.

[0023] Figure 4 It is a diagram showing a framework for generating multimedia resources according to an exemplary embodiment.

[0024] Figure 5 It is a block diagram of a device for generating multimedia resources according to an exemplary embodiment.

[0025] Figure 6It is a block diagram of a terminal according to an exemplary embodiment.

[0026] Figure 7 The figure is a block diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION

[0027] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0028] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0029] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the reference images and text prompts involved in this disclosure were obtained with full authorization.

[0030] Figure 1 FIG. 1 is a schematic diagram of an implementation environment of a method for generating multimedia resources according to an exemplary embodiment. Taking the electronic device as a server as an example, see Figure 1 , the implementation environment specifically includes: a terminal 101 and a server 102.

[0031] Terminal 101 is at least one of a smartphone, smartwatch, desktop computer, laptop, MP3 player, MP4 player, and portable computer. Terminal 101 runs an application that supports multimedia resource generation. This application can be an editing application, a multimedia application, or an AI (Artificial Intelligence) assistant application, etc., though this embodiment of the present disclosure is not limiting in this regard. Users can log in to this application through terminal 101 to access the services provided by this application. Users can upload reference images and text prompts to this application, and use the services provided by this application to generate multimedia resources that match these reference images and text prompts. These multimedia resources can be images or videos, though this embodiment of the present disclosure is not limiting in this regard. Terminal 101 can connect to server 102 via a wireless or wired network, and can then send the reference images and text prompts to server 102, which then generates the multimedia resources.

[0032] Terminal 101 generally refers to one of multiple terminals. This embodiment uses terminal 101 as an example. Those skilled in the art will appreciate that the number of terminals may be greater or lesser. For example, there may be a few terminals, or dozens, hundreds, or even more. This embodiment does not limit the number or device type of terminals.

[0033] Server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 can be connected to terminal 101 and other terminals via a wireless network or a wired network. Server 102 can receive reference images and text prompts sent by terminal 101, process the reference images and text prompts, generate multimedia resources that match the reference images and text prompts, and then send the multimedia resources to terminal 101, which then displays the generated multimedia resources to the user. In some embodiments, the number of servers described above can be greater or lesser, and this is not limited in the present disclosure. Of course, server 102 also includes other functional servers to provide more comprehensive and diverse services.

[0034] Figure 2 is a flowchart of a method for generating multimedia resources according to an exemplary embodiment. Figure 2 The method for generating multimedia resources is applied in a server and includes the following steps.

[0035] In step 201, the server obtains input reference images and text prompt words, the reference images are used to provide a main object for the generation of multimedia resources, and the text prompt words are used to indicate the special effects to be generated in the multimedia resources.

[0036] In the embodiment of the present disclosure, both the reference image and the text prompt words are input by the user. The reference image includes at least one main object. The main object can be a person, an animal, a plant, or any object (such as a building, a table), etc., and the embodiment of the present disclosure does not limit this. The text prompt words can reflect the special effects required by the subsequent user. For example, if the text prompt word is "blooming", the subsequently generated multimedia resources will show blooming flowers compared to the reference image. In short, the reference image and text prompt words are used to indicate the style of the multimedia resources required by the user. The embodiment of the present disclosure does not limit the style of the reference image and the specific content of the text prompt words. The server receives the reference image and text prompt words sent by the terminal.

[0037] In step 202, the server performs image recognition on the reference image through an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of at least one subject object in the reference image.

[0038] In the disclosed embodiments, an image processing model is deployed in the server. The disclosed embodiments do not limit the architecture of the image processing model. The image processing model is used to identify image content. After receiving a reference image, the server inputs the reference image into the image processing model, which then identifies the reference image and obtains image description information and subject morphology information.

[0039] Among them, the image description information is used to describe the picture content of the reference image. The image description information includes the category of each subject object in the reference image. The category of the subject object may include people, cats, dogs, trees, flowers, buildings, etc., which is not limited in the embodiment of the present disclosure. The subject morphology information is used to describe the position of each subject object in the reference image. The subject morphology information can be a mask image of the subject object, or it can be a morphology description text (such as "Object 1 is located on the left side of the reference image"), etc. The embodiment of the present disclosure does not limit the form of the subject morphology information.

[0040] In step 203, the server generates multimedia resources based on the image description information, the subject morphology information, the reference image and the text prompt word. The multimedia resources include at least one subject object and special effects, and the position of the special effects in the multimedia resources is different from the position of the at least one subject object. The multimedia resources are images or videos.

[0041] In the disclosed embodiment, a server processes image description information, subject morphology information, a reference image, and text prompts to obtain a multimedia resource. The multimedia resource can be an image or a video. The subject object in the multimedia resource is the subject object in the reference image. The special effects in the multimedia resource are the special effects indicated by the text prompts. The area occupied by the special effects in the multimedia resource does not overlap with the area occupied by the subject object. In other words, the special effects in the multimedia resource do not obscure the display of the subject object.

[0042] The disclosed embodiment provides a method for generating multimedia resources. In the process of generating multimedia resources based on a reference image and text prompt words, the reference image is firstly subjected to image recognition to determine the image description information and subject morphology information of the reference image. Then, the multimedia resources are generated by using the category of the subject object indicated by the image description information, the position of the subject object indicated by the subject morphology information in the reference image, the reference image and the text prompt words. In the process of generating multimedia resources, the position of the subject object can be avoided, and the special effects indicated by the text prompt words can be generated at other positions outside the subject object, thereby avoiding the generated special effects from obscuring key content such as the subject object in the reference image, and improving the quality of the generated multimedia resources. Moreover, the user does not need to input the descriptive text for indicating the position of the special effects, but only needs to provide a simple text prompt word to indicate the special effects, thereby generating high-quality multimedia resources. The operation is simple, which is conducive to improving the generation efficiency of multimedia resources.

[0043] In some embodiments, generating multimedia resources based on image description information, subject morphology information, reference images, and text prompts includes: Generate resource description information based on the image description information, the subject shape information, and the text prompt word, where the resource description information is used to indicate a positional relationship to be satisfied between the special effect and at least one subject object; Generate multimedia resources based on resource description information, reference images and subject morphology information.

[0044] The solution provided by the embodiment of the present disclosure first performs a comprehensive analysis on the image description information, the subject morphology information and the text prompt words to generate resource description information, so as to determine the positional relationship between the special effects and the subject object in the subsequently generated multimedia resources. Then, the multimedia resources are generated based on the reference image and the subject morphology information of the resource description information. In the process of generating the multimedia resources, the position of the subject object can be avoided, and the special effects indicated by the text prompt words can be generated at other positions outside the subject object, thereby preventing the generated special effects from obscuring key content such as the subject object in the reference image, and improving the quality of the generated multimedia resources.

[0045] In some embodiments, generating resource description information based on image description information, subject morphology information, and text prompt words includes: Through the text fusion model, the image description information, subject morphology information and text prompt words are processed to obtain resource description information; Generate multimedia resources based on resource description information, reference images, and subject morphology information, including: Based on the resource generation model, the resource description information, reference image and subject morphology information are processed to obtain multimedia resources.

[0046] The solution provided by the embodiment of the present disclosure analyzes and processes image description information, subject morphology information and text prompt words through a text fusion model to generate resource description information, thereby ensuring the accuracy of the resource description information, that is, the rationality of the display position between the subject object and the special effects indicated by the resource description information. Then, the resource description information, reference image and subject morphology information are analyzed and processed through a resource generation model, which can avoid the position of the subject object in the process of generating multimedia resources, and generate the special effects indicated by the text prompt words at other positions outside the subject object, thereby avoiding the generated special effects from obscuring key content such as the subject object in the reference image, and improving the quality of the generated multimedia resources. In addition, the content recognition of the reference image, the integration of the resource description information and the generation of multimedia resources in this solution are all processed using their own models, which is conducive to ensuring the accuracy of the output of each model, thereby ensuring the quality of the finally generated multimedia resources.

[0047] In some embodiments, the method further comprises: Obtain resource prompt information, which is used to indicate the conditions that multimedia resources must meet; The text fusion model processes image description information, subject morphology information, and text prompt words to obtain resource description information, including: Based on the text fusion model, the image description information, subject morphology information, text prompt words and resource prompt information are processed to obtain resource description information.

[0048] The solution provided by the embodiment of the present disclosure, before generating multimedia resources, in addition to image description information, subject form information and text prompt words, will also obtain additional resource prompt information to prompt the conditions that the subsequently generated multimedia resources must meet, which can ensure the accuracy of the resource description information, that is, the resource description information can accurately describe the style of the multimedia resources, thereby helping to improve the quality of the multimedia resources and meet the user's resource needs.

[0049] In some embodiments, obtaining resource prompt information includes any of the following: Determine resource prompt information based on reference images and text prompt words; Determine resource prompt information based on the style of multimedia resources; In response to the prompt application instruction, the input resource prompt information corresponding to the prompt application instruction is obtained.

[0050] The solution provided by the embodiments of the present disclosure can determine resource prompt information based on reference images and text prompt words, so that the resource prompt information matches the reference images and text prompt words. Since the reference images and text prompt words are information actively input by the user, they can reflect the user's demand for multimedia resources to a certain extent. This method can not only ensure that the resource prompt information meets the user's resource needs, but also ensure that the picture content displayed by the multimedia resource meets the characteristics of the main object and special effects. In other words, it can ensure that the pictures in the multimedia resource are more natural and reasonable, which is conducive to improving the quality of the multimedia resource.

[0051] Alternatively, the resource prompt information can be determined according to the style of the multimedia resource so that the resource prompt information matches the style of the multimedia resource, that is, it can ensure that the picture content of the subsequently generated multimedia resource conforms to the specified style, which is conducive to improving the quality and accuracy of the multimedia resource.

[0052] Alternatively, the resource prompt information may also be actively input by the user, which helps to ensure that the subsequently generated multimedia resources meet the user's resource needs and can improve the quality and accuracy of the multimedia resources.

[0053] In some embodiments, determining resource prompt information based on the reference image and the text prompt word includes: The resource prompt information is determined based on the category of at least one main object in the reference image and the category of the special effect indicated by the text prompt word.

[0054] The solution provided by the embodiments of the present disclosure can determine resource prompt information based on the category of the main object in the reference image and the category of the special effect indicated by the text prompt word, so that the resource prompt information matches the main object and the special effect. Since the reference image and the text prompt word are information actively input by the user, they can reflect the user's demand for multimedia resources to a certain extent. This method can not only ensure that the resource prompt information meets the user's resource needs, but also ensure that the picture content displayed by the multimedia resource meets the characteristics such as the category of the main object and the category of the special effect. That is, it can ensure that the picture in the multimedia resource is more natural and reasonable, which is conducive to improving the quality of the multimedia resource.

[0055] In some embodiments, the method further comprises: Output resource description information; Generate multimedia resources based on resource description information, reference images, and subject morphology information, including: In response to the application instruction of the resource description information, a multimedia resource is generated based on the resource description information, the reference image and the subject morphology information.

[0056] The solution provided by the embodiments of the present disclosure can output resource description information to the user before generating multimedia resources, so that the user can, to a certain extent, know in advance the screen content of the subsequently generated multimedia resources. After confirming that the user applies the resource description information, the multimedia resources are generated based on the resource description information, thereby ensuring that the multimedia resources meet the user's resource requirements and improving the accuracy and quality of the multimedia resources.

[0057] In some embodiments, there are multiple pieces of resource description information, and the positional relationship between the special effect and at least one main object in different resource description information is different; Output resource description information, including: Output multiple resource description information; In response to the application instruction of the resource description information, generating a multimedia resource based on the resource description information, the reference image, and the subject morphology information, including: In response to an application instruction of any one of the plurality of resource description information, a multimedia resource is generated based on the resource description information, the reference image and the main form information.

[0058] The solution provided by the embodiment of the present disclosure can output multiple resource description information to the user at one time before generating multimedia resources, so that the user can select a resource description information according to his or her own needs to generate multimedia resources. This not only ensures that the multimedia resources meet the user's resource needs and improves the accuracy and quality of the multimedia resources, but also improves the user's controllability of multimedia resource generation, thereby helping to increase the utilization rate of this solution.

[0059] above Figure 2 The following is only a basic process of the present disclosure. The solution provided by the present disclosure is further described based on a specific implementation method. Figure 3 FIG. 1 is a flow chart of another method for generating multimedia resources according to an exemplary embodiment. Taking the electronic device as a server as an example, see Figure 3 , the method includes the following steps.

[0060] In step 301, the server obtains input reference images and text prompt words, the reference images are used to provide a main object for the generation of multimedia resources, and the text prompt words are used to indicate the special effects to be generated in the multimedia resources.

[0061] In an embodiment of the present disclosure, before generating a multimedia resource, a user can input a reference image and text prompt words for generating the multimedia resource into a terminal. The terminal then sends a resource generation instruction to a server. The resource generation instruction includes the reference image, text prompt words, and the type of multimedia resource to be generated (video or image). The server can determine the reference image and text prompt words from the resource generation instruction. During the subsequent multimedia resource generation process, the server generates the multimedia resource based on the main object in the reference image and the special effects indicated by the text prompt words.

[0062] In step 302, the server performs image recognition on the reference image through an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of at least one subject object in the reference image.

[0063] In the disclosed embodiment, after acquiring a reference image, the server inputs the reference image into an image processing model, which processes the reference image to identify the image description and subject morphology information of the reference image. The disclosed embodiment does not limit this, and the disclosed embodiment does not limit this, such as a large language model of any architecture, or a medium or small architecture model that supports image recognition.

[0064] The image description information may include the category, behavior, and state (such as decoration) of each subject object in the reference image, as well as scene and other screen content, which is not limited in the present embodiment. The subject morphological information may be a mask image (mask) of the subject object in the reference image, or may be morphological description text, etc., which is not limited in the present embodiment. Each subject object in the reference image may correspond to a mask image, or the positions of all subject objects in the reference image may be located in the same mask image, which is not limited in the present embodiment.

[0065] For example, Figure 4 FIG. 1 is a diagram showing a framework for generating multimedia resources according to an exemplary embodiment. Figure 4 ,The server inputs the reference image input by the user into the image ,processing model, recognizes the reference image through the image ,processing model, and outputs the image description information and ,subject morphology information of the reference image.

[0066] In step 303, the server generates resource description information based on the image description information, the subject shape information, and the text prompt word. The resource description information is used to indicate the positional relationship to be satisfied between the special effect and at least one subject object.

[0067] In the disclosed embodiment, the server integrates the image description information, subject morphology information, and text prompts to generate resource description information, which is used to determine the positional relationship between special effects and subject objects in the subsequently generated multimedia resource. In other words, the server can integrate the image description information, subject morphology information, and text prompts into a more detailed resource description to indicate the subsequent generation of multimedia resources.

[0068] For example, the image description information includes a person (subject category), a child (subject object), and wearing a hairpin (subject status); the subject morphology information includes the position of the child and the position of the hairpin; the text prompt word is blooming; then the resource description information generated by the server can be "several real small flowers grow on both ends of the hairpin on the child's head, which is in line with the style of childhood children's accessories, and many tall flowers grow on the ground around the child (outside the range of the subject object), covering the entire ground."

[0069] The server can generate resource description information based on the characteristics of the main object in the image description information, the position of the main object in the main form information, and the characteristics of the special effects indicated by the text prompt word. The positional relationship indicated by the resource description information is consistent with the characteristics of the main object and the special effects, which facilitates the subsequent generation of natural and reasonable multimedia resources.

[0070] In some embodiments, resource description information can be generated by a model. Accordingly, the server processes the image description information, subject morphology information, and text prompt words through a text fusion model to obtain resource description information. The text fusion model can be a large language model of any architecture, or other medium or small architecture models that support image recognition, etc., and the embodiments of the present disclosure are not limited to this. The solution provided by the embodiments of the present disclosure generates resource description information by analyzing and processing the image description information, subject morphology information, and text prompt words through a text fusion model, thereby ensuring the accuracy of the resource description information, that is, the rationality of the display position between the subject object and the special effect indicated by the resource description information.

[0071] For example, see Figure 4 For the image description information and subject morphology information output by the image processing model, the server inputs the image description information and subject morphology information as well as the text prompt words input by the user into the text fusion model, processes them through the text fusion model, and outputs the resource description information.

[0072] In some embodiments, the server can also obtain resource prompt information. Resource prompt information is used to indicate the conditions that multimedia resources must meet. In the process of generating resource description information through the model, the server processes the image description information, subject morphology information, text prompt words and resource prompt information based on the text fusion model to obtain resource description information. Among them, the resource prompt information can be used to guide the focus of text writing in the resource description information. For example, the resource prompt information is "focus on the combination of the character and the input word, such as dressing accessories, and don't try to change the facial features of the person." It can be seen that the requirement of multimedia resources is to copy the main object in the reference image and cannot change the appearance of the main object. The embodiment of this disclosure does not limit the specific content of the resource prompt information.

[0073] The solution provided by the embodiment of the present disclosure, before generating multimedia resources, in addition to image description information, subject form information and text prompt words, will also obtain additional resource prompt information to prompt the conditions that the subsequently generated multimedia resources must meet, which can ensure the accuracy of the resource description information, that is, the resource description information can accurately describe the style of the multimedia resources, thereby helping to improve the quality of the multimedia resources and meet the user's resource needs.

[0074] For example, see Figure 4 The input of the text fusion model includes image description information, subject morphology information, text prompt words and resource prompt information; the output of the text fusion model is resource description information.

[0075] The embodiment of the present disclosure does not limit the method for obtaining the resource prompt information. The following exemplifies three methods for obtaining the resource prompt information, but is by no means limited thereto.

[0076] In the first approach, the server determines resource prompt information based on a reference image and text prompts. The solution provided by the disclosed embodiments can determine resource prompt information based on the reference image and text prompts, ensuring that the resource prompt information matches the reference image and text prompts. Because the reference image and text prompts are user-initiated information and can, to a certain extent, reflect the user's demand for multimedia resources, this method not only ensures that the resource prompt information meets the user's resource needs but also ensures that the displayed content of the multimedia resource conforms to the characteristics of the main object and special effects. This ensures that the images in the multimedia resource are more natural and reasonable, thereby improving the quality of the multimedia resource.

[0077] Among them, the server can determine the resource prompt information based on at least one of the picture content, style, category of the main object, status of the main object, scene (or background) of the reference image and the special effects indicated by the text prompt word, which is not limited in this embodiment of the present disclosure.

[0078] Optionally, the server determines resource prompt information based on the category of at least one main object in the reference image and the category of special effects indicated by the text prompt word. The solution provided by the embodiment of the present disclosure can determine resource prompt information based on the category of the main object in the reference image and the category of special effects indicated by the text prompt word, so that the resource prompt information matches the main object and the special effects. Since the reference image and the text prompt word are information actively input by the user, they can reflect the user's demand for multimedia resources to a certain extent. This method can not only ensure that the resource prompt information meets the user's resource needs, but also ensure that the picture content displayed by the multimedia resource meets the characteristics such as the category of the main object and the category of the special effects. That is, it can ensure that the picture in the multimedia resource is more natural and reasonable, which is conducive to improving the quality of the multimedia resource.

[0079] In the second approach, the server determines resource prompt information based on the style of the multimedia resource. The style of the multimedia resource can be specified by the user or determined based on at least one of a user-provided reference image and a text prompt, which is not limited in this embodiment. The solution provided in this embodiment can also determine resource prompt information based on the style of the multimedia resource, ensuring that the resource prompt information matches the style of the multimedia resource. This ensures that the image content of the subsequently generated multimedia resource conforms to the specified style, thereby improving the quality and accuracy of the multimedia resource.

[0080] In the third method, the server responds to the prompt application instruction and obtains the input resource prompt information corresponding to the prompt application instruction. In addition to providing users with an upload entry for reference images and an upload entry for text prompts, this solution can also provide users with an upload entry for resource prompt information, so that users can actively input resource prompt information while inputting reference images and text prompt words. Then, when the user confirms to use the resource prompt information, the server obtains the resource prompt information to indicate the generation of subsequent multimedia resources. In the solution provided by the embodiment of the present disclosure, the resource prompt information can also be actively input by the user, which is conducive to ensuring that the subsequently generated multimedia resources meet the user's resource requirements and can improve the quality and accuracy of multimedia resources.

[0081] The resource prompt information may be generated in real time or selected from a plurality of pre-set candidate prompt information. For example, the server determines the resource prompt information from a plurality of candidate prompt information based on a reference image and text prompt words. This is not limited in the embodiments of the present disclosure.

[0082] In step 304, the server generates multimedia resources based on the resource description information, the reference image, and the subject morphology information.

[0083] In the embodiment of the present disclosure, the server first performs a comprehensive analysis of the image description information, the subject morphology information and the text prompt words to generate resource description information to determine the positional relationship between the special effects and the subject object in the subsequently generated multimedia resources, and then generates the multimedia resources based on the reference image and the subject morphology information of the resource description information. In the process of generating the multimedia resources, the position of the subject object can be avoided, and the special effects referred to by the text prompt words can be generated at other positions outside the subject object, thereby avoiding the generated special effects from obscuring key content such as the subject object in the reference image, and improving the quality of the generated multimedia resources.

[0084] For example, see Figure 4 ,The server inputs the subject morphology information output by the image ,processing model, the resource description information output by the text fusion ,model, and the reference image provided by the user into the resource ,generation model, which processes them and outputs multimedia resources.

[0085] In some embodiments, multimedia resources can be generated by a model. Accordingly, the server processes the resource description information, reference image and subject morphology information based on the resource generation model to obtain multimedia resources. The resource generation model can be a large language model of any architecture, or other medium and small architecture models that support image recognition, etc., which is not limited by the embodiments of the present disclosure. The solution provided by the embodiments of the present disclosure uses a resource generation model to analyze and process the resource description information, reference image and subject morphology information. It can avoid the location of the subject object in the process of generating multimedia resources, and generate the special effects referred to by the text prompt words at other locations outside the subject object, thereby avoiding the generated special effects from blocking key content such as the subject object in the reference image, and can improve the quality of the generated multimedia resources. Moreover, the content recognition of the reference image, the integration of the resource description information and the generation of multimedia resources in this solution are all processed using their own models, which is conducive to ensuring the accuracy of the output of each model, thereby ensuring the quality of the multimedia resources finally generated.

[0086] In some embodiments, the server outputs resource description information. That is, the server sends the resource description information to the terminal. Then, in response to the application instruction of the resource description information, the server generates multimedia resources based on the resource description information, the reference image, and the subject morphology information. The solution provided by the embodiment of the present disclosure can output multiple resource description information to the user at one time before generating multimedia resources, so that the user can select a resource description information to generate multimedia resources according to their own needs. This not only ensures that the multimedia resources meet the user's resource needs and improves the accuracy and quality of the multimedia resources, but also improves the user's controllability of multimedia resource generation, thereby helping to increase the utilization rate of this solution.

[0087] Among them, there can be multiple resource description information, and the positional relationship between the special effects and at least one main object in different resource description information is different. Accordingly, the server outputs multiple resource description information. Then, in response to the application instruction of any resource description information among the multiple resource description information, the server generates multimedia resources based on the resource description information, the reference image and the main morphological information. The solution provided by the embodiment of the present disclosure can output multiple resource description information to the user at one time before generating multimedia resources, so that the user can select a resource description information to generate multimedia resources according to his own needs. It can not only ensure that the multimedia resources meet the user's resource needs and improve the accuracy and quality of multimedia resources, but also improve the user's controllability of multimedia resource generation, thereby helping to improve the utilization rate of this solution.

[0088] Alternatively, the server may first generate corresponding multimedia resources based on multiple pieces of resource description information, and send the generated multiple multimedia resources to the terminal for the user to select.

[0089] In the case where the user inputs multiple reference images, the server can generate corresponding multimedia resources for each reference image and the corresponding text prompt word using the above method, thereby realizing batch generation of multimedia resources.

[0090] The disclosed embodiment provides a method for generating multimedia resources. In the process of generating multimedia resources based on a reference image and text prompt words, the reference image is firstly subjected to image recognition to determine the image description information and subject morphology information of the reference image. Then, the multimedia resources are generated by using the category of the subject object indicated by the image description information, the position of the subject object indicated by the subject morphology information in the reference image, the reference image and the text prompt words. In the process of generating multimedia resources, the position of the subject object can be avoided, and the special effects indicated by the text prompt words can be generated at other positions outside the subject object, thereby avoiding the generated special effects from obscuring key content such as the subject object in the reference image, and improving the quality of the generated multimedia resources. Moreover, the user does not need to input the descriptive text for indicating the position of the special effects, but only needs to provide a simple text prompt word to indicate the special effects, thereby generating high-quality multimedia resources. The operation is simple, which is conducive to improving the generation efficiency of multimedia resources.

[0091] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0092] Figure 5 FIG1 is a block diagram of a device for generating multimedia resources according to an exemplary embodiment. Figure 5 The device includes: a first acquisition unit 501, an identification unit 502 and a generation unit 503.

[0093] The first acquisition unit 501 is configured to acquire an input reference image and a text prompt word, wherein the reference image is used to provide a main object for generating a multimedia resource, and the text prompt word is used to indicate a special effect to be generated in the multimedia resource; The recognition unit 502 is configured to perform image recognition on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image, where the image description information includes a category of at least one subject object in the reference image, and the subject morphology information indicates a position of the at least one subject object in the reference image. The generation unit 503 is configured to generate multimedia resources based on image description information, subject morphology information, reference images, and text prompt words. The multimedia resources include at least one subject object and special effects, and the position of the special effects in the multimedia resources is different from the position of the at least one subject object. The multimedia resources are images or videos.

[0094] In some embodiments, the generating unit 503 includes: The first generating subunit is configured to generate resource description information based on the image description information, the subject shape information and the text prompt word, where the resource description information is used to indicate a positional relationship to be satisfied between the special effect and at least one subject object; The second generating subunit is configured to generate multimedia resources based on the resource description information, the reference image and the subject morphology information.

[0095] In some embodiments, the first generating subunit is configured to process the image description information, the subject morphology information, and the text prompt words through a text fusion model to obtain resource description information; The second generation subunit is configured to execute a resource generation model based on the resource description information, the reference image and the main form information to obtain the multimedia resource.

[0096] In some embodiments, the apparatus further comprises: A second acquiring unit is configured to acquire resource prompt information, where the resource prompt information is used to indicate conditions that the multimedia resource must meet; The first generating subunit is configured to execute a text fusion model to process the image description information, the subject morphology information, the text prompt words and the resource prompt information to obtain the resource description information.

[0097] In some embodiments, the second acquiring unit is configured to perform any of the following: Determine resource prompt information based on reference images and text prompt words; Determine resource prompt information based on the style of multimedia resources; In response to the prompt application instruction, the input resource prompt information corresponding to the prompt application instruction is obtained.

[0098] In some embodiments, the second acquisition unit is configured to determine the resource prompt information based on the category of at least one main object in the reference image and the category of the special effect indicated by the text prompt word.

[0099] In some embodiments, the apparatus further comprises: An output unit, configured to output resource description information; The second generating subunit is configured to execute an application instruction in response to the resource description information, and generate multimedia resources based on the resource description information, the reference image, and the subject morphology information.

[0100] In some embodiments, there are multiple pieces of resource description information, and the positional relationship between the special effect and at least one main object in different resource description information is different; An output unit, configured to output a plurality of resource description information; The second generating subunit is configured to execute an application instruction in response to any one of the multiple resource description information, and generate multimedia resources based on the resource description information, the reference image and the main form information.

[0101] The disclosed embodiment provides a device for generating multimedia resources. In the process of generating multimedia resources based on a reference image and text prompt words, the reference image is firstly subjected to image recognition to determine the image description information and subject morphology information of the reference image. Then, the multimedia resources are generated by using the category of the subject object indicated by the image description information, the position of the subject object indicated by the subject morphology information in the reference image, the reference image and the text prompt words. In the process of generating multimedia resources, the position of the subject object can be avoided, and the special effects indicated by the text prompt words can be generated at other positions outside the subject object, thereby avoiding the generated special effects from obscuring key content such as the subject object in the reference image, and improving the quality of the generated multimedia resources. Moreover, the user does not need to input the descriptive text for indicating the position of the special effects, but only needs to provide a simple text prompt word to indicate the special effects, thereby generating high-quality multimedia resources. The operation is simple, which is conducive to improving the generation efficiency of multimedia resources.

[0102] It should be noted that the multimedia resource generation device provided in the above embodiment only uses the division of the above functional units as an example when generating multimedia resources. In actual applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the multimedia resource generation device provided in the above embodiment and the multimedia resource generation method embodiment are of the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0103] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0104] When an electronic device is provided as a terminal, Figure 6 FIG. 6 is a block diagram of a terminal 600 according to an exemplary embodiment. Figure 6 The following is a block diagram of a terminal 600 according to an exemplary embodiment of the present disclosure. Terminal 600 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 600 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.

[0105] Typically, the terminal 600 includes a processor 601 and a memory 602 .

[0106] Processor 601 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 601 may be implemented in hardware using at least one of the following: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). Processor 601 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content displayed on the display screen. In some embodiments, processor 601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0107] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 602 is used to store at least one computer program, which is executed by the processor 601 to implement the method for generating multimedia resources provided in the method embodiment of the present application.

[0108] In some embodiments, terminal 600 may optionally include a peripheral device interface 603 and at least one peripheral device. Processor 601, memory 602, and peripheral device interface 603 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 603 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, and a power supply 608.

[0109] The peripheral device interface 603 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0110] The RF circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 604 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. In some embodiments, the RF circuit 604 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The RF circuit 604 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 604 may also include circuitry related to Near Field Communication (NFC), although this application does not limit this.

[0111] Display screen 605 is used to display a user interface (UI). This UI can include graphics, text, icons, videos, or any combination thereof. If display screen 605 is a touchscreen display, it can also detect touch signals on or above the surface of display screen 605. These touch signals can be input as control signals to processor 601 for processing. Display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 605, located on the front panel of terminal 600. In other embodiments, there can be at least two display screens 605, located on different surfaces of terminal 600 or in a foldable design. In still other embodiments, display screen 605 can be a flexible display, located on a curved or foldable surface of terminal 600. Display screen 605 can also be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 605 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0112] The camera assembly 606 is used to capture images or videos. In some embodiments, the camera assembly 606 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 606 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0113] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 601 for processing, or input into the radio frequency circuit 604 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 607 may also include a headphone jack.

[0114] Power supply 608 is used to power various components in terminal 600. Power supply 608 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 608 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is charged via a wired line, while a wireless rechargeable battery is charged via a wireless coil. The rechargeable battery can also support fast charging technology.

[0115] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation on the terminal 600, and the terminal 600 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0116] When an electronic device is provided as a server, Figure 7 This is a block diagram of a server 700 according to an exemplary embodiment. This server 700 may vary significantly depending on its configuration or performance. It may include one or more processors (CPUs) 701 and one or more memories 702. The memories 702 store at least one program code, which is loaded and executed by the processor 701 to implement the multimedia resource generation methods provided in the various method embodiments described above. Of course, the server may also include components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The server 700 may also include other components for implementing device functions, which are not described in detail here.

[0117] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as memory 602 or memory 702 including instructions. The instructions may be executed by processor 601 of terminal 600 or processor 701 of server 700 to implement the above-described method for generating multimedia resources. Alternatively, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, or the like.

[0118] A computer program product includes a computer program / instruction, which implements the above-mentioned method for generating multimedia resources when executed by a processor.

[0119] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0120] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for generating multimedia resources, characterized in that: The method comprises: Acquire an input reference image and a text prompt word, wherein the reference image is used to provide a main object for generating a multimedia resource, and the text prompt word is used to indicate a special effect to be generated in the multimedia resource; performing image recognition on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image, wherein the image description information includes a category of at least one subject object in the reference image, and the subject morphology information is used to indicate a position of the at least one subject object in the reference image; The multimedia resource is generated based on the image description information, the subject morphology information, the reference image and the text prompt word, wherein the multimedia resource includes the at least one subject object and the special effect, and the position of the special effect in the multimedia resource is different from the position of the at least one subject object. The multimedia resource is an image or a video.

2. The method for generating multimedia resources according to claim 1, wherein: The generating of the multimedia resource based on the image description information, the subject morphology information, the reference image, and the text prompt word includes: Generate resource description information based on the image description information, the subject shape information, and the text prompt word, wherein the resource description information is used to indicate a positional relationship to be satisfied between the special effect and the at least one subject object; The multimedia resource is generated based on the resource description information, the reference image, and the subject morphology information.

3. The method for generating multimedia resources according to claim 2, wherein: The generating of resource description information based on the image description information, the subject shape information and the text prompt word includes: Processing the image description information, the subject morphology information, and the text prompt words through a text fusion model to obtain the resource description information; The generating the multimedia resource based on the resource description information, the reference image, and the subject morphology information includes: Based on a resource generation model, the resource description information, the reference image, and the subject morphology information are processed to obtain the multimedia resource.

4. The method for generating multimedia resources according to claim 3, wherein: The method further comprises: Acquiring resource prompt information, where the resource prompt information is used to indicate conditions that the multimedia resource must meet; The image description information, the subject morphology information, and the text prompt word are processed by a text fusion model to obtain the resource description information, including: Based on a text fusion model, the image description information, the subject morphology information, the text prompt words and the resource prompt information are processed to obtain the resource description information.

5. The method for generating multimedia resources according to claim 4, characterized in that: The acquisition of resource prompt information includes any of the following: Determining the resource prompt information based on the reference image and the text prompt word; Determining the resource prompt information based on the style of the multimedia resource; In response to the prompt application instruction, the input resource prompt information corresponding to the prompt application instruction is obtained.

6. The method for generating multimedia resources according to claim 5, characterized in that: The determining the resource prompt information based on the reference image and the text prompt word includes: The resource prompt information is determined based on the category of the at least one main object in the reference image and the category of the special effect indicated by the text prompt word.

7. The method for generating multimedia resources according to claim 2, characterized in that: The method further comprises: Outputting the resource description information; The generating the multimedia resource based on the resource description information, the reference image, and the subject morphology information includes: In response to the application instruction of the resource description information, the multimedia resource is generated based on the resource description information, the reference image and the main form information.

8. The method for generating multimedia resources according to claim 7, characterized in that: There are multiple pieces of resource description information, and the positional relationship between the special effect and the at least one main object in different resource description information is different; The outputting of the resource description information includes: Output multiple resource description information; The generating of the multimedia resource based on the resource description information, the reference image, and the subject form information in response to the application instruction of the resource description information includes: In response to an application instruction of any one of the plurality of resource description information, the multimedia resource is generated based on the resource description information, the reference image, and the main form information.

9. A device for generating multimedia resources, characterized in that: The device comprises: A first acquisition unit is configured to acquire an input reference image and a text prompt word, wherein the reference image is used to provide a main object for generating a multimedia resource, and the text prompt word is used to indicate a special effect to be generated in the multimedia resource; a recognition unit configured to perform image recognition on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image, wherein the image description information includes a category of at least one subject object in the reference image, and the subject morphology information is used to indicate a position of the at least one subject object in the reference image; A generation unit is configured to generate the multimedia resource based on the image description information, the subject morphology information, the reference image, and the text prompt word, wherein the multimedia resource includes the at least one subject object and the special effect, and the position of the special effect in the multimedia resource is different from the position of the at least one subject object, and the multimedia resource is an image or a video.

10. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the method for generating multimedia resources according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for generating multimedia resources according to any one of claims 1 to 8.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating multimedia resources according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Method for generating scene graph and electronic equipment

    CN118115627A

  • Image generation method and device

    CN119295297A

  • Media generation method and device, equipment and medium

    CN119579453A

  • Masking method for augmented reality effects

    US20210103730A1

  • Information generation method and apparatus, device, computer readable medium and program product

    WO2025130676A1