Methods, apparatuses, electronic devices and storage media for generating multimedia resources

By performing image recognition on reference images to obtain descriptive and morphological information, the location of the main object is avoided when generating multimedia resources, thus solving the problem of special effects occlusion and improving image quality and generation efficiency.

CN120448564BActive Publication Date: 2025-10-31BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510941836.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-31
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

In existing technologies, special effects can easily obscure key content in the reference image during image generation, resulting in lower image quality.

Method used

By performing image recognition on the reference image, image description information and subject shape information are obtained. Based on this information and text prompts, multimedia resources are generated to ensure that the position of the special effects is different from the main object and to avoid obscuring key content.

Benefits of technology

It improves the quality of multimedia resources, avoids special effects obscuring the main object, is easy to operate, and has high generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448564B_ABST
    Figure CN120448564B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, electronic device, and storage medium for generating multimedia resources, belonging to the field of multimedia technology. The method includes: acquiring an input reference image and text prompts; performing image recognition on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information indicates the position of at least one subject object in the reference image; generating multimedia resources based on the image description information, subject morphology information, the reference image, and the text prompts. The multimedia resources include at least one subject object and special effects, and the position of the special effects in the multimedia resources differs from the position of at least one subject object. This method can avoid the generated special effects obscuring key content such as the subject object in the reference image, thus improving the quality of the multimedia resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of multimedia technology, and in particular to a method, apparatus, electronic device, and storage medium for generating multimedia resources. Background Technology

[0002] With the continuous development of multimedia technology, text-to-image (TTO) generation technology and image-to-image (IPE) generation technology have become commonly used image generation methods. How to generate high-quality images is the focus of research in this field.

[0003] In related technologies, the common approach is to input a user-provided reference image and text prompts into an image generation model, and then use the image generation model to add the effects indicated by the text prompts to the reference image to generate a new image.

[0004] However, the above technical solutions simply add effects directly to the reference image without considering the specific content of the reference image. This results in the effects obscuring the key content of the reference image in the generated image, leading to a still low quality of the generated image. For example, if the reference image is "a person sitting in a chair" and the prompt is "flowers grow," the generated image might have flowers on the person's face, obscuring the person's image and creating the effect of the person turning into a flower, which is unpleasant. Summary of the Invention

[0005] This disclosure provides a method, apparatus, electronic device, and storage medium for generating multimedia resources, which can avoid the generated special effects from obscuring key content such as the main object in the reference image, thereby improving the quality of multimedia resources. The technical solution of this disclosure is as follows.

[0006] According to one aspect of the embodiments of this disclosure, a method for generating multimedia resources is provided, comprising:

[0007] The system obtains an input reference image and a text prompt, wherein the reference image is used to provide the main object for the generation of multimedia resources, and the text prompt is used to indicate the special effects to be generated in the multimedia resources;

[0008] Image recognition is performed on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of the at least one subject object in the reference image.

[0009] Based on the image description information, the subject shape information, the reference image, and the text prompt, the multimedia resource is generated. The multimedia resource includes at least one subject object and the special effect, and the position of the special effect in the multimedia resource is different from the position of the at least one subject object. The multimedia resource is an image or a video.

[0010] According to another aspect of the embodiments of this disclosure, a multimedia resource generation apparatus is provided, comprising:

[0011] The first acquisition unit is configured to acquire a reference image and a text prompt as input, wherein the reference image is used to provide a subject object for the generation of multimedia resources, and the text prompt is used to indicate the special effects to be generated in the multimedia resources;

[0012] The recognition unit is configured to perform image recognition on the reference image through an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of the at least one subject object in the reference image.

[0013] The generation unit is configured to generate the multimedia resource based on the image description information, the subject shape information, the reference image, and the text prompt. The multimedia resource includes the at least one subject object and the special effect, and the position of the special effect in the multimedia resource is different from the position of the at least one subject object. The multimedia resource is an image or a video.

[0014] In some embodiments, the generating unit includes:

[0015] The first generation subunit is configured to generate resource description information based on the image description information, the subject shape information, and the text prompt, wherein the resource description information is used to indicate the positional relationship to be satisfied between the special effect and the at least one subject object;

[0016] The second generation subunit is configured to generate the multimedia resource based on the resource description information, the reference image, and the subject morphology information.

[0017] In some embodiments, the first generation subunit is configured to perform processing of the image description information, the subject morphology information, and the text prompt words through a text fusion model to obtain the resource description information;

[0018] The second generation subunit is configured to perform a resource generation model to process the resource description information, the reference image, and the subject morphology information to obtain the multimedia resource.

[0019] In some embodiments, the apparatus further includes:

[0020] The second acquisition unit is configured to acquire resource prompt information, the resource prompt information being used to indicate the conditions that the multimedia resource must meet;

[0021] The first generation subunit is configured to execute a text fusion model to process the image description information, the subject morphology information, the text prompts, and the resource prompts to obtain the resource description information.

[0022] In some embodiments, the second acquisition unit is configured to perform any of the following:

[0023] The resource prompt information is determined based on the reference image and the text prompt words;

[0024] Based on the style of the multimedia resources, determine the resource prompt information;

[0025] In response to a prompt application instruction, the system obtains the input resource prompt information corresponding to the prompt application instruction.

[0026] In some embodiments, the second acquisition unit is configured to determine the resource prompt information based on the category of the at least one subject object in the reference image and the category of the effect indicated by the text prompt word.

[0027] In some embodiments, the apparatus further includes:

[0028] The output unit is configured to output the resource description information.

[0029] The second generation subunit is configured to execute application instructions in response to the resource description information, and generate the multimedia resource based on the resource description information, the reference image, and the subject morphology information.

[0030] In some embodiments, there are multiple resource description information, and the positional relationship between the special effects and the at least one main object is different in different resource description information;

[0031] The output unit is configured to output multiple resource description information.

[0032] The second generation subunit is configured to execute an application instruction in response to any one of the multiple resource description information, and to generate the multimedia resource based on the resource description information, the reference image, and the subject morphology information.

[0033] According to another aspect of the embodiments of this disclosure, an electronic device is provided, the electronic device comprising:

[0034] One or more processors;

[0035] Memory used to store the executable program code of the processor;

[0036] The processor is configured to execute the program code to implement the aforementioned method for generating multimedia resources.

[0037] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which, when the program code in the computer-readable storage medium is executed by a processor of an electronic device, enables the electronic device to perform the above-described method for generating multimedia resources.

[0038] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the above-described method for generating multimedia resources.

[0039] The solution provided in this disclosure, in the process of generating multimedia resources based on a reference image and text prompts, first performs image recognition on the reference image to determine its image description information and subject shape information. Then, it generates multimedia resources by using the category of the subject object indicated by the image description information, the position of the subject object indicated by the subject shape information in the reference image, the reference image, and the text prompts. This method avoids the location of the subject object during the generation of multimedia resources, generating the special effects indicated by the text prompts in locations other than the subject object. This prevents the generated special effects from obscuring key content such as the subject object in the reference image, thereby improving the quality of the generated multimedia resources. Furthermore, users do not need to input descriptive text to indicate the location of the special effects; they only need to provide simple text prompts to indicate the special effects to generate high-quality multimedia resources. The operation is simple and helps to improve the efficiency of multimedia resource generation.

[0040] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0042] Figure 1 This is a schematic diagram illustrating the implementation environment of a method for generating multimedia resources according to an exemplary embodiment.

[0043] Figure 2 This is a flowchart illustrating a method for generating multimedia resources according to an exemplary embodiment.

[0044] Figure 3 This is a flowchart illustrating another method for generating multimedia resources according to an exemplary embodiment.

[0045] Figure 4 This is a framework diagram illustrating the generation of multimedia resources according to an exemplary embodiment.

[0046] Figure 5 This is a block diagram illustrating a multimedia resource generation apparatus according to an exemplary embodiment.

[0047] Figure 6 This is a block diagram illustrating a terminal according to an exemplary embodiment.

[0048] Figure 7 This is a block diagram illustrating a server according to an exemplary embodiment. Detailed Implementation

[0049] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0050] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0051] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the reference images and text prompts involved in this disclosure were obtained with full authorization.

[0052] Figure 1 This is a schematic diagram illustrating an implementation environment for a method of generating multimedia resources according to an exemplary embodiment. Taking an electronic device provided as a server as an example, see [link to example]. Figure 1 The implementation environment specifically includes: terminal 101 and server 102.

[0053] Terminal 101 is at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, MP3 player, MP4 player, and laptop computer. Terminal 101 runs an application that supports multimedia resource generation. This application can be a video editing application, a multimedia application, or an AI (Artificial Intelligence) assistant application, etc., and this embodiment does not limit the specific application. Users can log in to the application through terminal 101 to access the services provided by the application. Users can upload reference images and text prompts in the application and use the services provided by the application to generate multimedia resources that match the reference images and text prompts. The multimedia resources can be images or videos, and this embodiment does not limit the specific multimedia resources. Terminal 101 can connect to server 102 via a wireless or wired network, and can then send reference images and text prompts to server 102 for server 102 to generate multimedia resources.

[0054] Terminal 101 generally refers to one of a plurality of terminals; this embodiment uses terminal 101 as an example. Those skilled in the art will understand that the number of terminals can be more or less. For example, there may be several terminals, or dozens or hundreds of terminals, or even more. This disclosure does not limit the number of terminals or the type of device.

[0055] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 can connect to terminal 101 and other terminals via a wireless or wired network. Server 102 can receive reference images and text prompts sent by terminal 101, process the reference images and text prompts, generate multimedia resources matching the reference images and text prompts, and then send the multimedia resources to terminal 101, which displays the generated multimedia resources to the user. In some embodiments, the number of servers can be more or less, and this disclosure does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services.

[0056] Figure 2 This is a flowchart illustrating a method for generating multimedia resources according to an exemplary embodiment. See also... Figure 2 The method for generating multimedia resources is applied to a server and includes the following steps.

[0057] In step 201, the server obtains the input reference image and text prompts. The reference image is used to provide the main object for the generation of multimedia resources, and the text prompts are used to indicate the special effects to be generated in the multimedia resources.

[0058] In this embodiment, both the reference image and the text prompt are input by the user. The reference image includes at least one main object. The main object can be a person, animal, plant, or any object (such as a building or table), etc., and this embodiment does not limit this. The text prompt reflects the special effects desired by the user later. For example, if the text prompt is "flowering," the subsequently generated multimedia resource will show blooming flowers compared to the reference image. In short, the reference image and text prompt are used to indicate the style of the multimedia resource desired by the user, and this embodiment does not limit the style of the reference image or the specific content of the text prompt. The server receives the reference image and text prompt sent by the terminal.

[0059] In step 202, the server performs image recognition on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of at least one subject object in the reference image.

[0060] In this embodiment of the disclosure, an image processing model is deployed in the server, and the architecture of the image processing model is not limited in this embodiment. The image processing model is used to identify image content. After receiving a reference image, the server inputs the reference image into the image processing model, and the image processing model identifies the reference image to obtain image description information and subject shape information of the reference image.

[0061] The image description information describes the content of the reference image. It includes the category of each subject object in the reference image. The category of the subject object can include various types such as people, cats, dogs, trees, flowers, and buildings; this embodiment does not limit this. The subject shape information describes the position of each subject object in the reference image. The subject shape information can be a mask image of the subject object or shape description text (such as "object 1 is located on the left side of the reference image"). This embodiment does not limit the form of the subject shape information.

[0062] In step 203, the server generates multimedia resources based on image description information, subject shape information, reference images, and text prompts. The multimedia resources include at least one subject object and special effects, and the position of the special effects in the multimedia resources is different from the position of at least one subject object. The multimedia resources are images or videos.

[0063] In this embodiment, the server processes image description information, subject shape information, reference image, and text prompts to obtain multimedia resources. These multimedia resources can be images or videos. The main object in the multimedia resource is the main object in the reference image. The special effects in the multimedia resource are those indicated by the text prompts. The area occupied by the special effects in the multimedia resource does not overlap with the area occupied by the main object. That is, the special effects in the multimedia resource will not obscure the display of the main object.

[0064] This disclosure provides a method for generating multimedia resources. In the process of generating multimedia resources based on a reference image and text prompts, image recognition is first performed on the reference image to determine its image description information and subject shape information. Then, multimedia resources are generated using the category of the subject object indicated by the image description information, the position of the subject object in the reference image indicated by the subject shape information, the reference image, and the text prompts. This method avoids obscuring the subject object's location during multimedia resource generation, generating the effects indicated by the text prompts in locations other than the subject object. This prevents the generated effects from obscuring key content such as the subject object in the reference image, thus improving the quality of the generated multimedia resources. Furthermore, users do not need to input descriptive text to indicate the effect's location; they only need to provide simple text prompts to indicate the effect, which generates high-quality multimedia resources. The method is simple to operate and improves the efficiency of multimedia resource generation.

[0065] In some embodiments, multimedia resources are generated based on image description information, subject morphology information, reference images, and text prompts, including:

[0066] Based on image description information, subject shape information, and text prompts, resource description information is generated. The resource description information is used to indicate the positional relationship that the special effect and at least one subject object must satisfy.

[0067] Multimedia resources are generated based on resource description information, reference images, and subject morphology information.

[0068] The solution provided in this embodiment first comprehensively analyzes image description information, subject shape information, and text prompts to generate resource description information, thereby determining the positional relationship between special effects and the subject object in the subsequently generated multimedia resource. Then, the multimedia resource is generated based on the resource description information, the reference image, and the subject shape information. This method avoids the location of the subject object during the generation of the multimedia resource, generating the special effects indicated by the text prompts in locations other than the subject object. This prevents the generated special effects from obscuring key content such as the subject object in the reference image, thereby improving the quality of the generated multimedia resource.

[0069] In some embodiments, resource description information is generated based on image description information, subject shape information, and text prompts, including:

[0070] By using a text fusion model, image description information, subject morphology information, and text prompts are processed to obtain resource description information;

[0071] Based on resource description information, reference images, and subject morphology information, multimedia resources are generated, including:

[0072] Based on the resource generation model, the resource description information, reference images, and subject morphology information are processed to obtain multimedia resources.

[0073] The solution provided in this disclosure analyzes and processes image description information, subject shape information, and text prompts using a text fusion model to generate resource description information. This ensures the accuracy of the resource description information, specifically the rationality of the display position between the subject object indicated by the resource description information and the special effect. Then, the resource generation model analyzes and processes the resource description information, reference image, and subject shape information. This avoids the location of the subject object during the generation of multimedia resources, generating the special effect indicated by the text prompt in a location other than the subject object. This prevents the generated special effect from obscuring key content such as the subject object in the reference image, thus improving the quality of the generated multimedia resources. Furthermore, the content recognition of the reference image, the integration of resource description information, and the generation of multimedia resources in this solution all use their respective models for processing, which helps to ensure the accuracy of the output of each model, thereby ensuring the quality of the final generated multimedia resources.

[0074] In some embodiments, the method further includes:

[0075] Obtain resource hint information, which indicates the conditions that multimedia resources must meet;

[0076] By using a text fusion model, image description information, subject morphology information, and text prompts are processed to obtain resource description information, including:

[0077] Based on the text fusion model, image description information, subject morphology information, text prompts, and resource prompts are processed to obtain resource description information.

[0078] The solution provided in this embodiment, before generating multimedia resources, in addition to image description information, subject shape information, and text prompts, also acquires additional resource prompt information to indicate the conditions that the subsequently generated multimedia resources must meet. This ensures the accuracy of the resource description information, that is, the resource description information can accurately describe the style of the multimedia resources, thereby improving the quality of the multimedia resources and meeting the user's resource needs.

[0079] In some embodiments, obtaining resource prompt information includes any of the following:

[0080] Based on reference images and text prompts, determine resource prompt information;

[0081] Based on the style of multimedia resources, determine the resource prompt information;

[0082] In response to a prompt application command, retrieve the resource prompt information that has been entered corresponding to the prompt application command.

[0083] The solution provided in this embodiment can determine resource prompt information based on reference images and text prompts, so that the resource prompt information matches the reference images and text prompts. Since the reference images and text prompts are information actively input by the user, they can reflect the user's needs for multimedia resources to a certain extent. This method can not only ensure that the resource prompt information meets the user's resource needs, but also ensure that the content of the multimedia resources displayed conforms to the characteristics of the main object and special effects. That is, it can ensure that the images in the multimedia resources are more natural and reasonable, which is conducive to improving the quality of multimedia resources.

[0084] Alternatively, resource prompts can be determined based on the style of the multimedia resources, ensuring that the prompts match the style of the multimedia resources. This guarantees that the content of the subsequently generated multimedia resources conforms to the specified style, thereby improving the quality and accuracy of the multimedia resources.

[0085] Alternatively, resource prompts can be entered by the user, which helps ensure that the generated multimedia resources meet the user's needs and improves the quality and accuracy of the multimedia resources.

[0086] In some embodiments, resource prompt information is determined based on a reference image and text prompt words, including:

[0087] Resource cues are determined based on the category of at least one subject object in the reference image and the category of the effect indicated by the text cues.

[0088] The solution provided in this disclosure can determine resource prompt information based on the category of the main object in the reference image and the category of the special effect indicated by the text prompt, so that the resource prompt information matches the main object and the special effect. Since the reference image and the text prompt are information actively input by the user, they can reflect the user's needs for multimedia resources to a certain extent. This method can not only ensure that the resource prompt information meets the user's resource needs, but also ensure that the content of the multimedia resource displays conforms to the characteristics such as the category of the main object and the category of the special effect. That is, it can ensure that the picture in the multimedia resource is more natural and reasonable, which is conducive to improving the quality of the multimedia resource.

[0089] In some embodiments, the method further includes:

[0090] Output resource description information;

[0091] Based on resource description information, reference images, and subject morphology information, multimedia resources are generated, including:

[0092] In response to application commands that provide resource description information, multimedia resources are generated based on the resource description information, reference images, and subject shape information.

[0093] The solution provided in this disclosure can output resource description information to the user before generating multimedia resources, so that the user can know the screen content of the subsequently generated multimedia resources in advance to a certain extent. After confirming that the user applies the resource description information, the multimedia resources are generated based on the resource description information, which ensures that the multimedia resources meet the user's resource needs and can improve the accuracy and quality of the multimedia resources.

[0094] In some embodiments, there are multiple resource description information, and the positional relationship between the special effects and at least one main object is different in different resource description information.

[0095] Output resource description information, including:

[0096] Output multiple resource descriptions;

[0097] In response to application commands based on resource description information, reference images, and subject shape information, multimedia resources are generated, including:

[0098] In response to an application instruction that includes any one of the multiple resource description information, a multimedia resource is generated based on the resource description information, a reference image, and the subject's morphological information.

[0099] The solution provided in this disclosure can output multiple resource description information to the user at one time before generating multimedia resources, so that the user can select one resource description information to generate multimedia resources according to their own needs. This not only ensures that the multimedia resources meet the user's resource needs and improves the accuracy and quality of the multimedia resources, but also improves the user's control over the generation of multimedia resources, thereby improving the utilization rate of this solution.

[0100] The above Figure 2 The diagram shown is merely the basic process of this disclosure. The following section will further elaborate on the solution provided in this disclosure based on a specific implementation method. Figure 3 This is a flowchart illustrating another method for generating multimedia resources according to an exemplary embodiment. Taking an electronic device provided as a server as an example, see [link to example]. Figure 3 The method includes the following steps.

[0101] In step 301, the server obtains the input reference image and text prompts. The reference image is used to provide the main object for the generation of multimedia resources, and the text prompts are used to indicate the special effects to be generated in the multimedia resources.

[0102] In this embodiment of the disclosure, before generating multimedia resources, the user can input a reference image and text prompts on the terminal. The terminal then sends a resource generation instruction to the server. The resource generation instruction includes the reference image, text prompts, and the type of multimedia resource to be generated (video or image). The server can determine the reference image and text prompts from the resource generation instruction. During the subsequent generation of the multimedia resource, the server generates the multimedia resource based on the main object in the reference image and the effects indicated by the text prompts.

[0103] In step 302, the server performs image recognition on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of at least one subject object in the reference image.

[0104] In this embodiment of the disclosure, after acquiring the reference image, the server inputs the reference image into the image processing model, and the image processing model processes the reference image to identify the image description information and subject morphology information of the reference image. The image processing model can be any large language model or other medium or small architecture models that support image recognition; this embodiment of the disclosure does not limit it in this regard.

[0105] The image description information may include the category, behavior, and state (such as decoration) of each subject object in the reference image, as well as scene content, etc., which are not limited in this embodiment. The subject shape information may be a mask image of the subject object in the reference image, or it may be shape description text, etc., which are not limited in this embodiment. Each subject object in the reference image may correspond to a mask image, or the positions of all subject objects in the reference image may be located in the same mask image, which are not limited in this embodiment.

[0106] For example, Figure 4 This is a framework diagram illustrating the generation of multimedia resources according to an exemplary embodiment. See also... Figure 4 The server inputs the reference image from the user into the image processing model, which then identifies the reference image and outputs the image description information and subject morphology information of the reference image.

[0107] In step 303, the server generates resource description information based on image description information, subject shape information, and text prompts. The resource description information is used to indicate the positional relationship that the special effect and at least one subject object must satisfy.

[0108] In this embodiment, the server fuses image description information, subject shape information, and text prompts to obtain resource description information, thereby determining the positional relationship between special effects and the main object in the subsequently generated multimedia resources. That is, the server can integrate image description information, subject shape information, and text prompts into a more detailed resource description to instruct the generation of subsequent multimedia resources.

[0109] For example, the image description information includes a person (subject category), a child (subject object), and wearing a hair clip (subject state); the subject shape information includes the child's position and the position of the hair clip; the text prompt is "flowering"; then the resource description information generated by the server can be "Several real small flowers grow from both ends of the hair clip on the child's head, which is in line with the style of childhood children's accessories, and many tall flowers grow on the ground around the child (outside the range of the subject object), covering the entire ground."

[0110] The server can generate resource description information based on the characteristics of the main object in the image description information, the position of the main object in the main object shape information, and the characteristics of the special effects indicated by the text prompts. The positional relationship indicated by this resource description information conforms to the characteristics of the main object and the special effects, which is conducive to the subsequent generation of natural and reasonable multimedia resources.

[0111] In some embodiments, resource description information can be generated by a model. Accordingly, the server processes image description information, subject morphology information, and text prompts using a text fusion model to obtain resource description information. The text fusion model can be any large language model or other medium or small architecture models that support image recognition; this disclosure does not limit this. The solution provided in this disclosure generates resource description information by analyzing and processing image description information, subject morphology information, and text prompts using a text fusion model, ensuring the accuracy of the resource description information, that is, the rationality of the display position between the subject object indicated by the resource description information and the special effect.

[0112] For example, see continue. Figure 4 For the image description information and subject shape information output by the image processing model, the server inputs the image description information, subject shape information and text prompts input by the user into the text fusion model. The text fusion model processes the text and outputs resource description information.

[0113] In some embodiments, the server can also obtain resource hint information. Resource hint information is used to indicate the conditions that multimedia resources must meet. During the process of generating resource description information through a model, the server processes image description information, subject morphology information, text hints, and resource hint information based on a text fusion model to obtain resource description information. The resource hint information can be used to guide the focus of the text writing in the resource description information. For example, the resource hint information might be "Focus on the connection between the person and the input words, such as in clothing and accessories; do not attempt to change the person's facial features." Therefore, the requirement for multimedia resources is to directly copy the main object in the reference image and not alter its appearance. This disclosure does not limit the specific content of the resource hint information.

[0114] The solution provided in this embodiment, before generating multimedia resources, in addition to image description information, subject shape information, and text prompts, also acquires additional resource prompt information to indicate the conditions that the subsequently generated multimedia resources must meet. This ensures the accuracy of the resource description information, that is, the resource description information can accurately describe the style of the multimedia resources, thereby improving the quality of the multimedia resources and meeting the user's resource needs.

[0115] For example, see continue. Figure 4 The input to the text fusion model includes image description information, subject morphology information, text prompts, and resource prompts; the output of the text fusion model is resource description information.

[0116] This disclosure does not limit the method of obtaining the above-mentioned resource prompt information. Three methods are described below as examples, but are by no means limited thereto.

[0117] In the first method, the server determines resource prompt information based on reference images and text prompts. The solution provided in this disclosure can determine resource prompt information based on reference images and text prompts, ensuring that the resource prompt information matches the reference images and text prompts. Since the reference images and text prompts are information actively input by the user, they can reflect the user's needs for multimedia resources to a certain extent. This method not only ensures that the resource prompt information meets the user's resource needs but also ensures that the content displayed in the multimedia resources conforms to the characteristics of the main subject and special effects, thus ensuring that the visuals in the multimedia resources are more natural and reasonable, thereby improving the quality of the multimedia resources.

[0118] The server can determine the resource prompt information based on at least one of the following: the content and style of the reference image, the category of the main object, the state of the main object, the scene (or background), and the special effects indicated by the text prompt words. This embodiment of the disclosure does not limit this.

[0119] Optionally, the server determines resource prompt information based on the category of at least one main object in the reference image and the category of the special effect indicated by the text prompt. The solution provided in this disclosure can determine resource prompt information according to the category of the main object in the reference image and the category of the special effect indicated by the text prompt, ensuring that the resource prompt information matches the main object and the special effect. Since the reference image and text prompt are information actively input by the user, they can reflect the user's needs for multimedia resources to a certain extent. This method not only ensures that the resource prompt information meets the user's resource needs but also ensures that the content displayed in the multimedia resources conforms to the characteristics such as the category of the main object and the category of the special effect, thus ensuring that the images in the multimedia resources are more natural and reasonable, which is conducive to improving the quality of the multimedia resources.

[0120] In the second approach, the server determines resource prompt information based on the style of the multimedia resource. The style of the multimedia resource can be specified by the user, or it can be determined based on at least one of the reference images and text prompts provided by the user; this embodiment does not limit this. The solution provided in this embodiment can also determine resource prompt information based on the style of the multimedia resource, ensuring that the resource prompt information matches the style of the multimedia resource. This guarantees that the content of the subsequently generated multimedia resource conforms to the specified style, thus improving the quality and accuracy of the multimedia resource.

[0121] The third method involves the server responding to an application prompt instruction and retrieving the corresponding input resource prompt information. In this solution, besides providing upload entry points for reference images and text prompts, it also provides an upload entry point for resource prompt information, allowing users to actively input resource prompt information simultaneously with the reference image and text prompt. Then, upon user confirmation of using the resource prompt information, the server retrieves it to instruct on the subsequent generation of multimedia resources. The solution provided in this disclosure allows for user-inputted resource prompt information, ensuring that the generated multimedia resources meet the user's resource needs and improving the quality and accuracy of the multimedia resources.

[0122] The aforementioned resource prompt information can be generated in real time or selected from multiple pre-set candidate prompt information. For example, the server determines the resource prompt information from multiple candidate prompt information based on a reference image and text prompt words. This embodiment of the disclosure does not limit this.

[0123] In step 304, the server generates multimedia resources based on resource description information, reference images, and subject morphology information.

[0124] In this embodiment, the server first performs a comprehensive analysis of image description information, subject shape information, and text prompts to generate resource description information to determine the bitwise relationship between special effects and the subject object in the subsequently generated multimedia resources. Then, the server generates multimedia resources based on the resource description information, the reference image, and the subject shape information. This method avoids the location of the subject object during the generation of multimedia resources and generates the special effects indicated by the text prompts in other locations outside the subject object. This prevents the generated special effects from obscuring key content such as the subject object in the reference image and improves the quality of the generated multimedia resources.

[0125] For example, see continue. Figure 4 The server inputs the subject shape information output by the image processing model, the resource description information output by the text fusion model, and the reference image provided by the user into the resource generation model. The resource generation model processes the data and outputs multimedia resources.

[0126] In some embodiments, multimedia resources can be generated by a model. Accordingly, the server processes resource description information, reference images, and subject morphology information based on the resource generation model to obtain the multimedia resources. This resource generation model can be a large language model of any architecture, or other medium or small architecture models that support image recognition, etc., and this disclosure does not limit this. The solution provided by this disclosure analyzes and processes resource description information, reference images, and subject morphology information through a resource generation model. This allows the system to avoid the location of the main object during the generation of multimedia resources, generating the special effects indicated by the text prompts in locations other than the main object. This prevents the generated special effects from obscuring key content such as the main object in the reference image, thus improving the quality of the generated multimedia resources. Furthermore, the content recognition of the reference image, the integration of resource description information, and the generation of multimedia resources in this solution all use their respective models for processing, which helps ensure the accuracy of the outputs of each model, thereby guaranteeing the quality of the final generated multimedia resources.

[0127] In some embodiments, the server outputs resource description information. That is, the server sends resource description information to the terminal. Then, in response to application instructions based on the resource description information, the server generates multimedia resources based on the resource description information, reference images, and subject shape information. The solution provided by this disclosure can output multiple resource description information to the user at once before generating multimedia resources, so that the user can select one resource description information to generate multimedia resources according to their own needs. This not only ensures that the multimedia resources meet the user's resource needs and improves the accuracy and quality of the multimedia resources, but also improves the user's control over the generation of multimedia resources, thereby increasing the utilization rate of this solution.

[0128] The resource description information can be multiple, and the positional relationship between the special effects and at least one main object differs in different resource description information. Accordingly, the server outputs multiple resource description information. Then, in response to an application instruction from any one of the multiple resource description information, the server generates a multimedia resource based on the resource description information, a reference image, and the main object's shape information. The solution provided in this embodiment can output multiple resource description information to the user at once before generating the multimedia resource, allowing the user to select one according to their needs. This not only ensures that the multimedia resource meets the user's resource requirements, improving the accuracy and quality of the multimedia resource, but also enhances the user's control over the generation of the multimedia resource, thereby increasing the utilization rate of this solution.

[0129] Alternatively, the server can first generate corresponding multimedia resources based on multiple resource description information, and then send the generated multimedia resources to the terminal for the user to choose from.

[0130] When a user inputs multiple reference images, the server can generate corresponding multimedia resources for each reference image and its corresponding text prompt using the method described above, thereby achieving batch generation of multimedia resources.

[0131] This disclosure provides a method for generating multimedia resources. In the process of generating multimedia resources based on a reference image and text prompts, image recognition is first performed on the reference image to determine its image description information and subject shape information. Then, multimedia resources are generated using the category of the subject object indicated by the image description information, the position of the subject object in the reference image indicated by the subject shape information, the reference image, and the text prompts. This method avoids obscuring the subject object's location during multimedia resource generation, generating the effects indicated by the text prompts in locations other than the subject object. This prevents the generated effects from obscuring key content such as the subject object in the reference image, thus improving the quality of the generated multimedia resources. Furthermore, users do not need to input descriptive text to indicate the effect's location; they only need to provide simple text prompts to indicate the effect, which generates high-quality multimedia resources. The method is simple to operate and improves the efficiency of multimedia resource generation.

[0132] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0133] Figure 5 This is a block diagram illustrating a multimedia resource generation apparatus according to an exemplary embodiment. See also... Figure 5 The device includes: a first acquisition unit 501, an identification unit 502, and a generation unit 503.

[0134] The first acquisition unit 501 is configured to acquire reference images and text prompts as input. The reference images are used to provide the main object for the generation of multimedia resources, and the text prompts are used to indicate the special effects to be generated in the multimedia resources.

[0135] The recognition unit 502 is configured to perform image recognition on a reference image through an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of at least one subject object in the reference image.

[0136] The generation unit 503 is configured to generate multimedia resources based on image description information, subject shape information, reference images and text prompts. The multimedia resources include at least one subject object and special effects, and the position of the special effects in the multimedia resources is different from the position of at least one subject object. The multimedia resources are images or videos.

[0137] In some embodiments, the generation unit 503 includes:

[0138] The first generation subunit is configured to generate resource description information based on image description information, subject shape information and text prompts. The resource description information is used to indicate the positional relationship to be satisfied between the special effect and at least one subject object.

[0139] The second generation subunit is configured to generate multimedia resources based on resource description information, reference images, and subject morphology information.

[0140] In some embodiments, the first generation subunit is configured to process image description information, subject morphology information and text prompts through a text fusion model to obtain resource description information;

[0141] The second generation subunit is configured to execute a resource generation model to process resource description information, reference images, and subject morphology information to obtain multimedia resources.

[0142] In some embodiments, the apparatus further includes:

[0143] The second acquisition unit is configured to acquire resource prompt information, which is used to indicate the conditions that the multimedia resources must meet.

[0144] The first generation subunit is configured to execute a text fusion model to process image description information, subject morphology information, text prompts, and resource prompts to obtain resource description information.

[0145] In some embodiments, the second acquisition unit is configured to perform any of the following:

[0146] Based on reference images and text prompts, determine resource prompt information;

[0147] Based on the style of multimedia resources, determine the resource prompt information;

[0148] In response to a prompt application command, retrieve the resource prompt information that has been entered corresponding to the prompt application command.

[0149] In some embodiments, the second acquisition unit is configured to determine resource hint information based on the category of at least one subject object in the reference image and the category of the effect indicated by the text prompt.

[0150] In some embodiments, the apparatus further includes:

[0151] The output unit is configured to output resource description information.

[0152] The second generation subunit is configured to execute application instructions in response to resource description information, and generate multimedia resources based on resource description information, reference images, and subject shape information.

[0153] In some embodiments, there are multiple resource description information, and the positional relationship between the special effects and at least one main object is different in different resource description information.

[0154] The output unit is configured to output multiple resource description information.

[0155] The second generation subunit is configured to execute application instructions that respond to any one of the multiple resource description information, and generate multimedia resources based on the resource description information, reference image, and subject shape information.

[0156] This disclosure provides a multimedia resource generation apparatus. In the process of generating multimedia resources based on a reference image and text prompts, image recognition is first performed on the reference image to determine its image description information and subject shape information. Then, multimedia resources are generated using the category of the subject object indicated by the image description information, the position of the subject object in the reference image indicated by the subject shape information, the reference image, and the text prompts. This method avoids obscuring the subject object's location during multimedia resource generation, generating the effects indicated by the text prompts in locations other than the subject object. This prevents the generated effects from obscuring key content such as the subject object in the reference image, thus improving the quality of the generated multimedia resources. Furthermore, users do not need to input descriptive text to indicate the effect's location; they only need to provide simple text prompts to indicate the effect, which generates high-quality multimedia resources. The operation is simple and improves the efficiency of multimedia resource generation.

[0157] It should be noted that the multimedia resource generation apparatus provided in the above embodiments is only illustrated by the division of the above functional units when generating multimedia resources. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the multimedia resource generation apparatus and the multimedia resource generation method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0158] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0159] When an electronic device is provided as a terminal, Figure 6 This is a block diagram illustrating a terminal 600 according to an exemplary embodiment. The terminal... Figure 6 A structural block diagram of a terminal 600 provided in an exemplary embodiment of this disclosure is shown. The terminal 600 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 600 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0160] Typically, terminal 600 includes a processor 601 and a memory 602.

[0161] Processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 601 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 601 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0162] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 are used to store at least one computer program, which is executed by the processor 601 to implement the multimedia resource generation method provided in the method embodiments of this application.

[0163] In some embodiments, the terminal 600 may optionally include a peripheral device interface 603 and at least one peripheral device. The processor 601, memory 602, and peripheral device interface 603 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 603 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, and a power supply 608.

[0164] Peripheral interface 603 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 601 and memory 602. In some embodiments, processor 601, memory 602 and peripheral interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 601, memory 602 and peripheral interface 603 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0165] The radio frequency (RF) circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 604 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. In some embodiments, the RF circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 604 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 604 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0166] Display screen 605 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 605 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 601 for processing. In this case, display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, disposed on the front panel of terminal 600; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 600 or in a folded design; in other embodiments, display screen 605 may be a flexible display screen, disposed on a curved or folded surface of terminal 600. Furthermore, display screen 605 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 605 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0167] The camera assembly 606 is used to acquire images or videos. In some embodiments, the camera assembly 606 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 606 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0168] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 601 for processing, or input to the radio frequency circuit 604 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal 600. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 607 may also include a headphone jack.

[0169] Power supply 608 is used to power the various components in terminal 600. Power supply 608 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 608 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0170] Those skilled in the art will understand that Figure 6 The structure shown does not constitute a limitation on terminal 600, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0171] When electronic devices are provided as servers, Figure 7 This is a block diagram illustrating a server 700 according to an exemplary embodiment. The server 700 can vary significantly due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 701 and one or more memories 702. The memories 702 store at least one line of program code, which is loaded and executed by the processor 701 to implement the multimedia resource generation method provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 700 may also include other components for implementing device functions, which will not be elaborated here.

[0172] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as memory 602 or memory 702 including instructions. These instructions can be executed by processor 601 of terminal 600 or processor 701 of server 700 to complete the multimedia resource generation method described above. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.

[0173] A computer program product includes a computer program / instructions that, when executed by a processor, implement the aforementioned method for generating multimedia resources.

[0174] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0175] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for generating multimedia resources, characterized in that, The method includes: The system obtains an input reference image and a text prompt, wherein the reference image is used to provide the main object for the generation of multimedia resources, and the text prompt is used to indicate the special effects to be generated in the multimedia resources; Image recognition is performed on the reference image using an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of the at least one subject object in the reference image. Based on the image description information, the subject shape information, the reference image, and the text prompt, the multimedia resource is generated. The multimedia resource includes at least one subject object and the special effect, and the position of the special effect in the multimedia resource is different from the position of the at least one subject object. The multimedia resource is an image or a video.

2. The method for generating multimedia resources according to claim 1, characterized in that, The process of generating the multimedia resource based on the image description information, the subject morphology information, the reference image, and the text prompts includes: Based on the image description information, the subject shape information, and the text prompts, resource description information is generated, which is used to indicate the positional relationship between the special effect and the at least one subject object. The multimedia resource is generated based on the resource description information, the reference image, and the subject morphology information.

3. The method for generating multimedia resources according to claim 2, characterized in that, The step of generating resource description information based on the image description information, the subject morphology information, and the text prompts includes: The resource description information is obtained by processing the image description information, the subject morphology information, and the text prompts using a text fusion model. The process of generating the multimedia resource based on the resource description information, the reference image, and the subject morphology information includes: Based on the resource generation model, the resource description information, the reference image, and the subject morphology information are processed to obtain the multimedia resource.

4. The method for generating multimedia resources according to claim 3, characterized in that, The method further includes: Obtain resource prompt information, which indicates the conditions that the multimedia resource must meet; The process of using a text fusion model to process the image description information, the subject morphology information, and the text prompts to obtain the resource description information includes: Based on a text fusion model, the image description information, the subject morphology information, the text prompts, and the resource prompts are processed to obtain the resource description information.

5. The method for generating multimedia resources according to claim 4, characterized in that, The resource acquisition prompt information includes any one of the following: The resource prompt information is determined based on the reference image and the text prompt words; Based on the style of the multimedia resources, determine the resource prompt information; In response to a prompt application instruction, the system obtains the input resource prompt information corresponding to the prompt application instruction.

6. The method for generating multimedia resources according to claim 5, characterized in that, The step of determining the resource prompt information based on the reference image and the text prompt words includes: The resource prompt information is determined based on the category of at least one main object in the reference image and the category of the effect indicated by the text prompt.

7. The method for generating multimedia resources according to claim 2, characterized in that, The method further includes: Output the resource description information; The process of generating the multimedia resource based on the resource description information, the reference image, and the subject morphology information includes: In response to the application instruction of the resource description information, the multimedia resource is generated based on the resource description information, the reference image, and the subject morphology information.

8. The method for generating multimedia resources according to claim 7, characterized in that, The resource description information includes multiple items, and the positional relationship between the special effects and the at least one main object is different in different resource description information; The output of the resource description information includes: Output multiple resource descriptions; The application instruction responding to the resource description information generates the multimedia resource based on the resource description information, the reference image, and the subject morphology information, including: In response to an application instruction for any one of the multiple resource description information, the multimedia resource is generated based on the resource description information, the reference image, and the subject morphology information.

9. A multimedia resource generation device, characterized in that, The device includes: The first acquisition unit is configured to acquire a reference image and a text prompt as input, wherein the reference image is used to provide a subject object for the generation of multimedia resources, and the text prompt is used to indicate the special effects to be generated in the multimedia resources; The recognition unit is configured to perform image recognition on the reference image through an image processing model to obtain image description information and subject morphology information of the reference image. The image description information includes the category of at least one subject object in the reference image, and the subject morphology information is used to indicate the position of the at least one subject object in the reference image. The generation unit is configured to generate the multimedia resource based on the image description information, the subject shape information, the reference image, and the text prompt. The multimedia resource includes the at least one subject object and the special effect, and the position of the special effect in the multimedia resource is different from the position of the at least one subject object. The multimedia resource is an image or a video.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the method for generating multimedia resources as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method for generating multimedia resources as described in any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for generating multimedia resources according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for generating scene graph and electronic equipment

    CN118115627A

  • Image generation method and device

    CN119295297A