Content generation method and apparatus, and electronic device and storage medium

By obtaining the content to be processed and generating the optimized target text and pictures, combining the text and image generation model, the problems of low efficiency and poor quality of content generation in the existing technology are solved, and efficient and diversified content creation is achieved.

WO2025103008A1PCT designated stage expired Publication Date: 2025-05-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/123101
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-13
Filing Date
2024-09-30
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

The prior art is less efficient and has poor content quality when generating different types of content, making it difficult to meet users' diversified needs for content creation.

Method used

By obtaining the input pending content, the optimized target text description information is determined, and the key text description information is determined based on the information, combining the text generation model and the image generation model, alternately synergistically generates graphic and text results, and the optimized target picture and target text description information are generated.

Benefits of technology

It improves the efficiency and quality of content generation, meets users' needs for diversified, rich and high-quality content, and improves the efficiency and user experience of content creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024123101_22052025_PF_FP_ABST
    Figure CN2024123101_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A content generation method and apparatus, and an electronic device and a storage medium. The method comprises: acquiring input content to be processed; determining optimized target text description information corresponding to said content, and determining corresponding key text description information on the basis of the target text description information, wherein the target text description information is generated in an alternately collaborative manner on the basis of a text generation model and an image generation model, and represents optimized text description information corresponding to said content; generating a first image-text result of a first type or a second image-text result of a second type on the basis of the key text description information and the target text description information, wherein different types of image-text results represent different degrees of importance of image forms and text forms, and image-form target images in different types of image-text results are generated on the basis of the key text description information or the target text description information.
Need to check novelty before this filing date? Find Prior Art

Description

Content generation method, device, electronic device and storage medium

[0001] This application claims priority to the Chinese invention patent application entitled “A content generation method, device, electronic device and storage medium” and application number 2023115102045, filed on November 13, 2023. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of computer technology, and in particular to a content generation method, device, electronic device, and storage medium. Background Art

[0003] With the development of computer technology, users can create content more conveniently. Usually, users have different content creation needs, for example, creating pictures and picture copywriting, and then uploading the created content to the platform to share with other users. However, the copywriting or pictures created by users may be relatively simple and single. In related technologies, a certain type of content can also be converted into other types to improve the richness of the content. However, when generating other types of content in related technologies, the efficiency is low and the quality of the generated content is also poor.

[0004] Summary of the Invention

[0005] The embodiments of the present disclosure at least provide a content generation method, device, electronic device, and storage medium.

[0006] In a first aspect, an embodiment of the present disclosure provides a content generation method, comprising: obtaining input content to be processed; determining optimized target text description information corresponding to the content to be processed, and determining corresponding key text description information based on the target text description information, wherein the target text description information is generated based on a text generation model and an image generation model in an alternating and collaborative manner to represent the optimized text description information corresponding to the content to be processed; generating a first image-text result of a first type, or a second image-text result of a second type, based on the key text description information and the target text description information, wherein different types of image-text results represent different degrees of importance of image genres and text genres, and the target images of the image genres included in the different types of image-text results are generated based on the key text description information or the target text description information.

[0007] In the second aspect, the embodiment of the present disclosure also provides a content generation device, including: an acquisition module for acquiring input content to be processed; a determination module for determining the optimized target text description information corresponding to the content to be processed, and determining the corresponding key text description information based on the target text description information, wherein the target text description information is generated in an alternating collaborative manner based on a text generation model and an image generation model; a generation module for generating a first image-text result of a first type, or a second image-text result of a second type based on the key text description information and the target text description information, wherein different types of image-text results represent different degrees of importance of image genres and text genres, and the target images of image genres in different types of image-text results are generated based on the key text description information or the target text description information.

[0008] In a third aspect, an optional implementation of the present disclosure further provides an electronic device, comprising a processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor executes the steps in the possible implementation of the above-mentioned first aspect.

[0009] In a fourth aspect, an optional implementation of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the possible implementation of the above-mentioned first aspect when the computer program is executed by a processor.

[0010] For a description of the effects of the above-mentioned content generation device, electronic device, and computer-readable storage medium, please refer to the description of the above-mentioned content generation method, which will not be repeated here.

[0011] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure.

[0012] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.

[0014] FIG1 shows a flow chart of a content generation method provided by an embodiment of the present disclosure;

[0015] FIG2 is a schematic diagram showing an implementation effect of a content generation method provided by an embodiment of the present disclosure;

[0016] FIG3 shows a schematic diagram showing a text optimization process principle in the content generation method provided by an embodiment of the present disclosure;

[0017] FIG4 shows a schematic diagram of an implementation effect of another content generation method provided by an embodiment of the present disclosure;

[0018] FIG5 shows another schematic diagram of the text optimization process principle in the content generation method provided by an embodiment of the present disclosure;

[0019] FIG6 shows a schematic diagram of an implementation effect of another content generation method provided by an embodiment of the present disclosure;

[0020] FIG7 shows a flowchart of another content generation method provided by an embodiment of the present disclosure;

[0021] FIG8 shows a schematic diagram of a content generation device provided by an embodiment of the present disclosure;

[0022] FIG9 shows a schematic diagram of another content generation device provided by an embodiment of the present disclosure;

[0023] FIG10 shows a schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.

[0026] Research has found that users usually have different content creation needs, for example, creating pictures and picture copy, and then uploading the created content to the platform to share with other users. However, the copy or pictures created by users may be relatively simple and single. In order to make the content more attractive, how to more effectively generate richer and higher-quality content is an urgent problem to be solved.

[0027] The defects in the above solutions are the results obtained by the inventors after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by this disclosure for the above problems below should be the contributions made by the inventors to this disclosure during the disclosure process.

[0028] Based on the above research, the present disclosure provides a content generation method. Specifically, in one method, the input content to be processed is obtained; the optimized target text description information corresponding to the content to be processed is determined, and the corresponding key text description information is determined based on the target text description information, wherein the target text description information is generated in an alternating collaborative manner based on the text generation model and the image generation model; based on the key text description information and the target text description information, a first type of first graphic result, or a second type of second graphic result is generated, wherein different types of graphic results represent different degrees of importance of the image genre and the text genre, and the target image of the image genre in different types of graphic results is generated based on the key text description information or the target text description information. In this way, for the input text to be processed or the image to be processed, through graphic generation interaction, richer and higher-quality target images and target text description information for the target image can be generated, which improves efficiency and makes content creation more diversified.

[0029] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0030] To facilitate understanding of this embodiment, a content generation method disclosed in an embodiment of the present disclosure is first introduced in detail. The execution subject of the content generation method provided in the embodiment of the present disclosure is generally an electronic device with certain computing capabilities, such as a terminal device, a server, or other processing device. The terminal device can be a user equipment (UE), a mobile device, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. Among them, a personal digital assistant is a handheld electronic device that has some functions of an electronic computer and can be used to manage personal information, browse the Internet, send and receive emails, etc. It is generally not equipped with a keyboard and can also be called a handheld computer. In some possible implementations, the content generation method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0031] The content generation method can be executed in any execution entity with computing capabilities, for example, the above method can be executed at a terminal device or a server device. Referring to FIG1 , a flow chart of a content generation method provided in an embodiment of the present disclosure is shown, and the method includes:

[0032] S101: Obtain input content to be processed, where the content to be processed includes text to be optimized or images to be optimized.

[0033] In the embodiments of the present disclosure, mainly for content creation scenarios based on machine learning, users may need to obtain different, richer and high-quality content through expansion, imagination, beautification, etc. on the basis of the initial content, and taking into account the diversity of content types and diversified needs of content creation, for example, for the creative content uploaded to the platform, users usually hope that it is not just text or pictures alone, but a combination of text and pictures to increase traffic. Therefore, in the embodiments of the present disclosure, the user only needs to input the initial text to be optimized or the picture to be optimized to obtain the optimized picture-text pair.

[0034] S102: Determine the optimized target text description information corresponding to the content to be processed, and determine the corresponding key text description information based on the target text description information, wherein the target text description information is generated based on the text generation model and the image generation model in an alternating collaborative manner.

[0035] When executing step S102, the present disclosure provides possible implementations:

[0036] s1. Based on the text generation model, generate first text description information of the first image corresponding to the optimized target text description information, wherein the first image is generated based on the image generation model and the text to be processed, or the first image is a target frame image included in the video to be processed.

[0037] In the disclosed embodiments, the content to be processed includes text to be processed or video to be processed. If the user input is text to be optimized, a first image can be generated, and first text description information of the first image can be obtained. Then, based on the text generation model, optimized target text description information can be generated. If the user input is a picture to be optimized, the first text description information of the picture to be optimized can be obtained, and based on the text generation model, optimized target text description information can be generated.

[0038] In an embodiment of the present disclosure, based on a text generation model, first text description information of a first image is generated corresponding to optimized target text description information, including: when the content to be processed is a video to be processed, determining a target frame image according to a key frame of the video to be processed, or determining a corresponding target frame image according to a preset frame interval; based on a picture description model, taking the target frame image as input, obtaining first text description information of the target frame image; determining target optimization guide information; inputting the target optimization guide information and the first text description information into a text generation model, performing semantic analysis on the first text description information according to the target optimization guide information, and generating optimized target text description information corresponding to the first text description information.

[0039] In an embodiment of the present disclosure, based on a text generation model, first text description information of a first image is generated corresponding to optimized target text description information, including: when the content to be processed is a text to be processed, based on the image generation model, taking the text to be processed as input, generating a first image corresponding to the text to be processed; based on the image description model, taking the first image as input, obtaining first text description information of the first image; determining target optimization guidance information; inputting the target optimization guidance information and the first text description information into the text generation model, performing semantic analysis on the first text description information according to the target optimization guidance information, and generating optimized target text description information corresponding to the first text description information.

[0040] In an embodiment of the present disclosure, determining the target optimization guidance information includes: determining that the target optimization guidance information is preset optimization guidance sentence information; or determining that the content to be processed corresponds to the target style that needs to be optimized, and determining the target optimization guidance information corresponding to the target style based on the target style.

[0041] In an embodiment of the present disclosure, determining the target style that the content to be processed corresponds to and needs to be optimized includes: in response to a selection operation for each preset style, determining the preset style corresponding to the selection operation as the target style that the content to be processed corresponds to and needs to be optimized; or, based on the content to be processed, extracting the target style that the content to be processed corresponds to and needs to be optimized.

[0042] S2. Determine corresponding key text description information according to the target text description information.

[0043] S103. Generate a first graphic result of a first type or a second graphic result of a second type based on the key text description information and the target text description information, wherein different types of graphic results represent different levels of importance of the image genre and the text genre, and the target images of the image genre in the different types of graphic results are generated based on the key text description information or the target text description information.

[0044] In an embodiment of the present disclosure, a first type of first graphic and text result is generated based on the key text description information and the target text description information, including: based on the image generation model, performing semantic analysis on the key text description information according to the key text description information to generate an optimized target image; and obtaining the first type of first graphic and text result based on the target text description information and the target image.

[0045] In an embodiment of the present disclosure, a first graphic result of the second type is generated based on the key text description information and the target text description information, including: obtaining multiple sub-target text description information after segmentation based on the target text description information; based on the image generation model, performing semantic analysis on the multiple sub-target text description information based on the multiple sub-target text description information, and generating multiple corresponding optimized target images; obtaining a second graphic result of the second type based on the key text description information and the multiple target images.

[0046] In an embodiment of the present disclosure, based on an image generation model, semantic analysis is performed on multiple sub-target text description information according to multiple sub-target text description information, and multiple corresponding optimized target images are generated, including: based on an image generation model, multiple sub-target text description information and target style are used as input, semantic analysis is performed on multiple sub-target text description information and target style, and multiple corresponding optimized target images are generated.

[0047] In an embodiment of the present disclosure, based on an image generation model, a semantic analysis is performed on the key text description information according to the key text description information to generate an optimized target image, including: based on an image generation model, with the key text description information and the target style as input, a semantic analysis is performed on the key text description information and the target style to generate an optimized target image.

[0048] In one possible implementation, in the embodiment of the present disclosure, during free optimization, that is, when the target optimization guidance information does not include the target style, then based on the text generation model, more vivid and high-quality target text description information has been obtained. Therefore, the target text description information can be input into the image generation model, and the target text description information can be subjected to sentence analysis to generate an optimized target image.

[0049] In one possible implementation, when guiding the optimization towards a specific target style, optimized target text description information is obtained based on the target style. In order to further enhance the target style generation effect of the image, in an embodiment of the present disclosure, the target style can be substituted into the image generation process again. Specifically, the present disclosure provides a possible implementation method, which is based on an image generation model, takes the target text description information and the target style as input, performs semantic analysis on the target text description information and the target style, obtains text features, and obtains image features associated with the text features based on the text features, and generates an optimized target image based on the image features.

[0050] Furthermore, in an embodiment of the present disclosure, an optimized target image is generated. In order to improve the relevance between the text and the target image, text description information of the target image can also be generated based on the target image. Specifically, the present disclosure provides a possible implementation method, which inputs the target image into an image description model to obtain second text description information of the target image; inputs the target optimization guidance information and the second text description information into a text generation model, performs semantic analysis on the second text description information, and generates optimized second text description information to return the target image and the optimized second target text description information.

[0051] Furthermore, in the embodiment of the present disclosure, there is no limit on the number of optimized target images and target text description information. It can be a preset target number or a target number requirement input by the user. Specifically, based on the text generation model, the first text description information of the first image is generated corresponding to the optimized target number of target text description information; for each target text description information, the optimized target image is generated based on the image generation model and the target text description information.

[0052] In the embodiment of the present disclosure, the target text description information is the description information for the target image, and the target image and the target text description information are image-text pairs, and the number is the same.

[0053] In the embodiment of the present disclosure, the optimized target text description information and target image are returned to the terminal device, and the optimized target text description information and target image can be displayed to the user on the terminal device.

[0054] In addition, in order to further improve the accuracy, the process in the above steps S102-S103 can be executed in a loop multiple times. Specifically, the process of generating the target text description information and generating the target image, that is, generating the optimized target text description information based on the text generation model, and generating the optimized target image based on the image generation model and the target text description information, and then optimizing the text description information of the target image again for the target image, and generating the optimized target image again based on the optimized text description information, until the preset number of cycles is reached or other preset conditions are met. In this way, the accuracy can be improved, the content of the final target image and the target text description information can be richer and higher quality, and the relevance of the target text description information to the target image can also be improved.

[0055] In the embodiment of the present disclosure, the input content to be processed is obtained; the optimized target text description information corresponding to the content to be processed is determined, and the corresponding key text description information is determined based on the target text description information, wherein the target text description information is generated in an alternating collaborative manner based on the text generation model and the image generation model; based on the key text description information and the target text description information, a first graphic result of the first type or a second graphic result of the second type is generated, wherein different types of graphic results represent different degrees of importance of the image genre and the text genre, and the target images of the image genre in different types of graphic results are generated based on the key text description information or the target text description information. In this way, by connecting the text generation model and the image generation model, the text generation and image generation are alternately cycled for the content to be optimized, and diversified image-text pair generation is achieved, and the input content to be optimized can be optimized to generate optimized target images and target text description information, so that the content is richer and higher quality, the content quality is improved, and the content creation efficiency is also improved.

[0056] In the embodiment of the present disclosure, in step S102, the optimized target text description information corresponding to the content to be processed is determined, and the corresponding key text description information is determined based on the target text description information, wherein the target text description information is generated in an alternating collaborative manner based on the text generation model and the image generation model. Accordingly, in step S102, the target text description information can be generated in an alternating collaborative manner based on the text generation model and the image generation model. In this process, possible implementation methods are also provided:

[0057] In one possible embodiment, generating first text description information of a first image corresponding to optimized target text description information based on a text generation model, wherein the first image is generated based on the image generation model and the text to be processed, or the first image is a target frame image included in the video to be processed, includes:

[0058] 1) When the content to be processed is a video to be processed, a target frame image is determined according to a key frame of the video to be processed, or a corresponding target frame image is determined according to a preset frame interval; based on an image description model, the target frame image is used as input to obtain first text description information of the target frame image.

[0059] Alternatively and / or additionally, when the content to be processed is text to be processed, based on the image generation model, the text to be processed is used as input to generate a first image corresponding to the text to be processed; based on the image description model, the first image is used as input to obtain first text description information of the first image.

[0060] For example, the user only has a certain picture, which may be relatively simple or only include the user's simple picture concept. It is necessary to create some pictures with richer content based on the picture, and more vivid text description information for the picture. At this time, the user can input the simple picture, and then the simple picture can be further optimized as the picture to be optimized.

[0061] In the embodiment of the present disclosure, after obtaining the image to be optimized, the image feature vector of the image to be optimized can be extracted based on the image description model, and then decoded based on the image feature vector to obtain the first text description information.

[0062] Among them, the first text description information generated by the image description model is usually relatively simple and basic description information. For example, if the image to be optimized is a picture containing a cat, the first text description information obtained by the image description model may be "a cat", which is relatively simple. Therefore, in the embodiment of the present disclosure, it is necessary to further optimize and expand the first text description information to obtain richer and more vivid image description information.

[0063] 2) Determine target optimization guidance information.

[0064] In the disclosed embodiment, determining target optimization guidance information includes: determining that the target optimization guidance information is preset optimization guidance sentence information; or determining that the content to be processed corresponds to a target style to be optimized, and determining target optimization guidance information corresponding to the target style based on the target style. Specifically, when optimizing images or text, different optimization directions are provided, and different optimization directions correspond to different target optimization guidance information. Specifically, possible implementation methods are provided:

[0065] One way: determine the target optimization guidance information as preset optimization guidance statement information.

[0066] In the embodiment of the present disclosure, optimization guidance sentence information can be preset. For example, the preset optimization guidance sentence information is: "Optimize the language of the following text, require it to be more vivid, and add imagination and emoticons appropriately: "In the embodiment of the present disclosure, the preset optimization guidance sentence information tends to be free optimization for pictures or texts, and does not specify the optimization direction or style. It only needs to optimize richer, more vivid and high-quality content.

[0067] Another way: determine the target style that the content to be optimized corresponds to, and based on the target style, determine the target optimization guidance information corresponding to the target style.

[0068] In the disclosed embodiment, a target style can also be determined to guide images or texts to be optimized toward the target style, which can meet the personalized needs of users and provide more selectivity. Determining the target style required for the content to be optimized includes: responding to operations for each preset style, determining the preset style corresponding to the operation as the target style required to be optimized for the content to be processed; or, based on the content to be processed, extracting the target style required to be optimized for the content to be processed.

[0069] For example, in an embodiment of the present disclosure, some selectable preset styles can be provided to the user, such as a funny style, a warm style, a horror style, etc. When the user inputs the content to be optimized, he or she can select the target style to be optimized. For example, if the user selects a horror style, the picture or text can be optimized to the horror style.

[0070] For another example, in order to meet more diverse needs of users, users can also be supported to input other target styles that need to be optimized. For example, when the user enters the text to be optimized, the user can also enter the words of the target style that they want to optimize at the same time. For example, when the user enters the picture to be optimized, the user can also enter the words of the target style at the same time. Or the user can also identify the target style expressed in the picture to be optimized by identifying the picture to be optimized. This is not limited in the embodiments of the present disclosure.

[0071] After determining the target style, the target optimization guidance information corresponding to the target style can be determined. Each target style can have corresponding target optimization guidance information. For example, if the target style is horror style, the target optimization guidance information of the horror style is: "Rewrite the following description, you can add appropriate imagination, and the content must be scary:".

[0072] 3) Inputting the target optimization guidance information and the first text description information into the text generation model, performing semantic analysis on the first text description information according to the target optimization guidance information, and generating optimized target text description information corresponding to the first text description information.

[0073] Among them, the target optimization guidance information and the first description information can be combined and converted in a certain way, and then input into the text generation model. For example, the combination conversion method is target optimization guidance information + first text description information. For example, the input into the text generation model is: "Optimize the language of the following text, make it more vivid, and add imagination and emoticons appropriately: a cat."

[0074] In the disclosed embodiment, through the text generation model, the relatively simple and basic first text description information can be expanded and optimized according to the target optimization guidance information to generate target text description information with richer content and higher quality. Among them, the text generation model is, for example, a large language model (LLM), and the LLM model can be fine-tuned (Fine Tune) based on internal data samples.

[0075] In another possible embodiment, generating optimized target text description information corresponding to the first text description information of the first image based on the text generation model includes:

[0076] 1) When the content to be optimized is a text to be optimized, based on the image generation model, the text to be optimized is used as input to generate a first image corresponding to the text to be optimized.

[0077] In the embodiment of the present disclosure, it is also possible to support users to input only the text to be optimized, and generate a first image of the text to be optimized based on the image generation model. The image generation model is used for text-generated images, for example, it can be a stable diffusion model, and it can also be fine-tuned based on LoRA and internal data samples to make it more in line with the requirements.

[0078] 2) Based on the image description model, taking the first image as input, obtaining first text description information of the first image.

[0079] 3) Determine target optimization guidance information.

[0080] 4) The target optimization guidance information and the first text description information are input into the text generation model, and the first text description information is semantically analyzed according to the target optimization guidance information to generate the optimized target text description information corresponding to the first text description information.

[0081] In the embodiment of the present disclosure, when the input is text to be optimized, a first image is generated first, and then the first text description information of the first image is generated, and the text is optimized to generate the target text description information. In this way, by alternately generating images and texts for optimization, the accuracy can be improved and the possibility of large deviations between the final optimized text and image can be reduced.

[0082] Of course, the text to be optimized can also be directly optimized based on the text generation model, and the target image can be generated based on the optimized text description information, which is not limited in the embodiments of the present disclosure.

[0083] The following uses specific application scenarios to illustrate. For ease of explanation, the following application scenarios are introduced separately:

[0084] Application scenario 1: Taking the input content to be optimized as the picture to be optimized, and free optimization without a specific target style as an example, refer to Figure 2, which is a schematic diagram of the implementation effect of a content generation method in an embodiment of the present disclosure. As shown in Figure 2, the user inputs the picture to be optimized, for example, the picture to be optimized is a picture of a dish. For the case where the input is the picture to be optimized, in the embodiment of the present disclosure, it is necessary to first generate a picture description and optimize the generated picture description, and then based on the optimized picture description, realize the optimized generation of the picture.

[0085] Among them, referring to Figure 3, there is a schematic diagram of the principle of a text optimization process in the content generation method of an embodiment of the present invention. For the picture to be optimized, based on the picture description model, the first text description information of the picture to be optimized is obtained, for example, "a dish of stir-fried vegetables". The first text description information is relatively simple, so based on the text generation model, according to the first text description information and the target optimization guidance information, extended optimization is performed to generate target text description information with more vivid and rich content.

[0086] For example, the target optimization guidance information is: "Optimize the language of the following text, require it to be more vivid, and appropriately add imagination and emoticons:", then input "Optimize the language of the following text, require it to be more vivid, and appropriately add imagination and emoticons: a dish of stir-fried vegetables" into the text generation model, and then based on the text generation model, the optimized target text description information can be generated. For example, the target number of image-text pair candidates can be pre-set to provide to the user. For example, as shown in Figure 2, the target number is 4, then four different target text description information can be generated, namely: 1. "A dish of meat and vegetables on the table, cabbage, garlic slices and green peppers, served in a small bowl, 1. "A plate of meat and vegetables on the table. The plate of grilled meat and vegetables fills the table." 2. "A plate of meat and vegetables on the table seems like a container, containing our happiness and harvest. They are the foundation of our lives, allowing us to feel the abundance of food and the beauty of life. Happiness on the table, a dish that allows you to feel the beauty of life, full of harvest and enjoyment!" 4. "Let the deliciousness of food be presented in a more relaxed and interesting way! The delicious meat and vegetables spread on the table appear even more tempting. Food should be enjoyed in a relaxed and pleasant atmosphere. Enjoying food together can enhance relationships and add happiness!"

[0087] Furthermore, the optimized target text description information can be input into the image generation model, and the corresponding optimized target image can be generated for each target text description information. For example, as shown in Figure 2, the generated target image has more diverse and richer content than the previous image to be optimized.

[0088] It should be noted that the target text description information and target image obtained in the embodiment of the present disclosure at this time have been optimized to be more vivid and rich, and can be directly returned to the user. In addition, in order to obtain text description information that is more highly correlated with the target image, the embodiment of the present disclosure can also use the generated target image as input to repeat the above-mentioned text optimization process as shown in Figure 3, that is, based on the image description model, a simple second text description information is first generated, and then based on the text generation model and the target optimization guidance information, the second text description information of the target image is optimized to generate the final optimized text description information for return to the user.

[0089] In this way, in the embodiment of the present disclosure, by inputting the image to be optimized, the optimized target image and target text description information can be generated, and a variety of different styles of content can be freely optimized to improve content creation efficiency and enhance user experience.

[0090] Application scenario 2: Taking the input content to be optimized as a picture to be optimized, and guided optimization according to a specific target style as an example, refer to Figure 4, which is a schematic diagram of the implementation effect of another content generation method in an embodiment of the present disclosure. As shown in Figure 4, the user inputs a picture to be optimized, for example, the picture to be optimized is a picture of a cat, and the embodiment of the present disclosure provides several different preset styles for the user to choose from, such as funny style, warm style, horror style and magic world style, and then according to the selected target style, the picture to be optimized is optimized to generate a target picture and target text description information that conform to the target style.

[0091] Among them, referring to Figure 5, there is a schematic diagram of another text optimization process principle in the content generation method of an embodiment of the present disclosure. The image to be optimized is input into the image description model to obtain the first text description information of the image to be optimized, such as "a cat". Then, based on the text generation model, according to the first text description information, target optimization guidance information and target style, as shown in Figure 5, the gray square represents the target style selected by the user, and extended optimization is performed to generate target text description information with more vivid and rich content.

[0092] In the disclosed embodiment, corresponding target optimization guidance information can be set for each style. For example, the horror style is: "Rewrite the following description, but add appropriate imagination to make the content scary:", the horror warmth style is: "Rewrite the following description, but add appropriate imagination to make the content warm:", etc. For example, if the target style selected by the user is the horror style, "Rewrite the following description, but add appropriate imagination to make the content scary: a cat" can be input into the text generation model, and then based on the text generation model, an optimized target text description information that meets the target style can be generated.

[0093] As shown in Figure 4, if the funny style is selected, for example, the target text description information generated is: "A cat looking at the camera suddenly screams and runs back to the bed to hide because it finds a mouse in the photo playing TV"; if the warm style is selected, for example, the target text description information generated is: "A cat looking at the camera seems to see a beautiful world in its mind, its eyes reveal excitement and a hint of tenderness, it seems to see an infinitely beautiful adventure, and it begins to enjoy this moment"; if the horror style is selected, for example, the target text description information generated is: "A cat looking at the camera, a mysterious light flashes in its eyes, which is creepy"; if the magical world style is selected, for example, the target text description information generated is: "A cat stares at the camera, stunned, and then it suddenly turns into a lively snake, its tail swaying to the rhythm, very happy. It has a bold idea, so it turns on the camera and turns back time. The cat and snake's bodies can be perfectly entwined. They hear a sound, which brings all their memories back to the past."

[0094] Furthermore, the optimized target text description information can be input into the image generation model, and corresponding optimized target images can be generated for the target text description information. For example, as shown in Figure 4, based on the target text description information of different target styles, target images that conform to the target style are generated accordingly, and the image content is also more vivid and diversified.

[0095] In this way, in the embodiments of the present disclosure, text and image content can be generated according to a specific target style, which can guide the realization of personalized optimization needs and enhance the user experience.

[0096] Application scenario three: Taking the input content to be optimized as the text to be optimized as an example, refer to Figure 6, which is a schematic diagram of the implementation effect of another content generation method in the embodiment of the present disclosure. As shown in Figure 6, the user inputs the text to be optimized, for example, the text to be optimized is: "My cat". In the embodiment of the present disclosure, for the text to be optimized, the first picture of the text to be optimized is first generated based on the picture generation model, and then the optimized target text description information and target picture are generated for the first picture. The specific processing operation for the first picture is the same as the processing operation for the picture to be optimized in the above embodiment, so it will not be repeated here.

[0097] As shown in FIG6 , the text to be optimized is: “My cat”, which can be optimized based on different modes of free optimization or specific target style optimization. For example, as shown in FIG6 , a first picture is generated for the input “My cat”, and then the first text description information of the first picture is obtained based on the picture description model. According to the text generation model, the target optimization guide information and the first text description information, or according to the text generation model, the target optimization guide information, the target style and the first text description information, four target text description information as shown in FIG6 are generated, for example: 1. “A cute cat is looking at the camera, Its focused eyes instantly endear it. Its clear and lively gaze seems to be gazing into the photographer's heart." 2. "The cat is looking at the camera. Every day, it sneaks in a corner and stares, trying to see if we notice it's thinking." 3. "A cat is fiddling with the camera, but it looks like it's in a good mood! Move the camera away and let the cat have fun!" 4. "The cat is sitting under a tree, squinting quietly. It pokes its head out and looks at the blue sky, as if trying to capture something." Then, based on the image generation model, a corresponding target image is generated for each target text description.

[0098] In this way, in the embodiment of the present disclosure, by connecting the image generation model and the text generation model, image and text optimization can also be performed on the text to be optimized, and the optimized target text description information and target image can be obtained to meet the user's needs for creating various types of content and improve efficiency.

[0099] Another content generation method provided by an embodiment of the present disclosure is described below by taking a terminal device as an example of an execution subject.

[0100] 7 is a flowchart of another content generation method provided by an embodiment of the present disclosure, the method comprising:

[0101] S701: receiving input content to be optimized, wherein the content to be optimized includes text to be optimized or a picture to be optimized.

[0102] S702: Obtain the optimized target image and target text description information corresponding to the content to be optimized, wherein the target text description information is generated based on a text generation model and first text description information of a first image, the first image is the image to be optimized, or the first image is generated based on the text to be optimized, and the target image is generated based on the image generation model and the target text description information.

[0103] In the embodiment of the present disclosure, the method of generating the optimized target image and target text description information of the content to be optimized is the same as the implementation method in the above embodiment, and will not be repeated here.

[0104] For example, the input text to be optimized is: "I have a sweeping robot". In the absence of a specific target style, target text description information with various styles can be freely optimized. For example, the generated target text description information is: "We can imagine using this robot to clean the house. It is not only clean but also makes the house more comfortable and warm." In addition, in the embodiment of the present disclosure, target style instructions can be added to guide the optimization in the direction of the target style. For example, the target style is a storytelling style, and the generated target text description information is, for example: "When the sweeping robot was sweeping the floor, suddenly a bird bumped into it, scaring it to give up the cleaning work. Then, the bird lay on the ground and sighed, 'You sweeping robot, you always don't know that I also have the need to sweep the floor. It's too much!'".

[0105] Furthermore, in the embodiment of the present disclosure, a target image corresponding to the optimized target text description information may be generated, and the terminal device may obtain the optimized target image and the target text description information for the target image.

[0106] S703: Display the optimized target text description information and target image.

[0107] In the disclosed embodiment, input content to be optimized is received, and the optimized target image and target text description information corresponding to the content to be optimized can be obtained, and then the target text description information and target image can be displayed. In this way, the optimized target image and target text description information can be quickly obtained for the content to be optimized, thereby improving the efficiency of content creation, and through optimization, more diversified, richer and high-quality content can be obtained, thereby improving the quality of content creation.

[0108] In an embodiment of the present disclosure, a content generation method is further provided. The method includes: obtaining input content to be optimized, wherein the content to be optimized includes text to be optimized or an image to be optimized; generating optimized target text description information corresponding to first text description information of a first image based on a text generation model; and generating an optimized target image based on the image generation model and the target text description information, wherein the first image is the image to be optimized or the first image is generated based on the text to be optimized; and returning the optimized target text description information and the target image.

[0109] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0110] In the disclosed embodiment, through image and text generation interaction, image and text optimization is performed, thereby generating richer and higher-quality target image and target text description information, improving content creation efficiency, and enhancing content quality and diversity.

[0111] Based on the same inventive concept, the embodiment of the present disclosure also provides a content generation device corresponding to the content generation method. Since the principle of solving the problem by the device in the embodiment of the present disclosure is similar to the above-mentioned content generation method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0112] Referring to Figure 8, which is a schematic diagram of a content generation device provided in an embodiment of the present disclosure, the device includes: an acquisition module 81, used to acquire input content to be processed; a determination module 82, used to determine the optimized target text description information corresponding to the content to be processed, and determine the corresponding key text description information based on the target text description information, wherein the target text description information is generated in an alternating collaborative manner based on the text generation model and the image generation model; a generation module 83, used to generate a first image-text result of a first type, or a second image-text result of a second type, based on the key text description information and the target text description information, wherein different types of image-text results represent different degrees of importance of image genres and text genres, and the target images of the image genres in different types of image-text results are generated based on the key text description information or the target text description information.

[0113] In an embodiment of the present disclosure, the content to be processed includes text to be processed or video to be processed, and the determination module includes: a text description information generation module, which is used to generate first text description information of a first image corresponding to optimized target text description information based on a text generation model, wherein the first image is generated based on the image generation model and the text to be processed, or the first image is a target frame image included in the video to be processed; an information determination module, which is used to determine the corresponding key text description information based on the target text description information.

[0114] In an embodiment of the present disclosure, a text description information generation module includes: a picture determination module for determining a target frame picture according to a key frame of a video to be processed, or determining a corresponding target frame picture according to a preset frame interval, when the content to be processed is a video to be processed; an acquisition module for obtaining first text description information of a target frame picture based on a picture description model and taking the target frame picture as input; a guide information determination module for determining target optimization guide information; an analysis module for inputting the target optimization guide information and the first text description information into a text generation model, performing semantic analysis on the first text description information according to the target optimization guide information, and generating optimized target text description information corresponding to the first text description information.

[0115] In an embodiment of the present disclosure, the text description information generation module includes: an image generation module, which is used to generate a first image corresponding to the text to be processed based on an image generation model and taking the text to be processed as input when the content to be processed is the text to be processed; an acquisition module, which is used to obtain first text description information of the first image based on an image description model and taking the first image as input; a guidance information determination module, which is used to determine target optimization guidance information; and an analysis module, which is used to input the target optimization guidance information and the first text description information into the text generation model, perform semantic analysis on the first text description information according to the target optimization guidance information, and generate optimized target text description information corresponding to the first text description information.

[0116] In an embodiment of the present disclosure, the guidance information determination module includes: a sentence information determination module, which is used to determine that the target optimization guidance information is preset optimization guidance sentence information; or, a style-based determination module, which is used to determine the target style that needs to be optimized corresponding to the content to be processed, and determine the target optimization guidance information corresponding to the target style based on the target style.

[0117] In the embodiment of the present disclosure, a style determination module is further included, including: a selection module for responding to a selection operation for each preset style, and determining the preset style corresponding to the selection operation as the target style that needs to be optimized for the content to be processed; or, an extraction module for extracting the target style that needs to be optimized for the content to be processed based on the content to be processed.

[0118] In an embodiment of the present disclosure, the generation module includes: a first optimization module, which is used to generate a model based on the image, perform semantic analysis on the key text description information according to the key text description information, and generate an optimized target image; a first image and text result acquisition module, which is used to obtain a first image and text result of a first type according to the target text description information and the target image.

[0119] In the embodiment of the present disclosure, the generation module includes: a segmentation module for obtaining multiple sub-target text description information after segmentation based on the target text description information; a second optimization module for performing semantic analysis on the multiple sub-target text description information based on the image generation model, respectively, according to the multiple sub-target text description information, to generate corresponding multiple optimized target images; a second image and text result acquisition module for obtaining a second type of second image and text result based on the key text description information and the multiple target images.

[0120] In the embodiment of the present disclosure, the first optimization module includes: a first analysis module, which is used to generate a model based on an image, take key text description information and target style as input, perform semantic analysis on the key text description information and target style, and generate an optimized target image.

[0121] In an embodiment of the present disclosure, the second optimization module includes: a second analysis module, which is used to generate a model based on an image, take multiple sub-target text description information and target style as input, perform semantic analysis on the multiple sub-target text description information and target style, and generate corresponding multiple optimized target images.

[0122] Referring to Figure 9, which is a schematic diagram of another content generation device provided in an embodiment of the present disclosure, the device includes: an acquisition module 91, used to acquire input content to be optimized, wherein the content to be optimized includes text to be optimized or a picture to be optimized; a generation module 92, used to generate first text description information of a first picture corresponding to optimized target text description information based on a text generation model, and generate an optimized target picture based on the picture generation model and the target text description information, wherein the first picture is the picture to be optimized, or the first picture is generated based on the text to be optimized; a return module 93, used to return the optimized target text description information and the target picture.

[0123] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference can be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.

[0124] The present disclosure also provides an electronic device, as shown in FIG10 , which is a schematic diagram of the structure of the electronic device provided by the present disclosure, including:

[0125] Processor 101 and memory 102; the memory 102 stores machine-readable instructions executable by the processor 101, and the processor 101 is configured to execute the machine-readable instructions stored in the memory 102. When the machine-readable instructions are executed by the processor 101, the processor 101 is configured to perform the following steps:

[0126] Obtaining input content to be optimized, wherein the content to be optimized includes text to be optimized or an image to be optimized;

[0127] Based on the text generation model, generating optimized target text description information corresponding to first text description information of the first image, and generating an optimized target image based on the image generation model and the target text description information, wherein the first image is the image to be optimized, or the first image is generated based on the text to be optimized;

[0128] Return the optimized target text description information and the target image.

[0129] Alternatively, the processor 101 is configured to perform the following steps:

[0130] Receive input content to be optimized, wherein the content to be optimized includes text to be optimized or a picture to be optimized; obtain optimized target pictures and target text description information corresponding to the content to be optimized, wherein the target text description information is generated based on a text generation model and first text description information of a first picture, the first picture is the picture to be optimized, or the first picture is generated based on the text to be optimized, and the target picture is generated based on a picture generation model and the target text description information; display the optimized target text description information and the target picture.

[0131] The above-mentioned memory 102 includes internal memory 1021 and external memory 1022; the memory 1021 here is also called internal memory, which is used to temporarily store the calculation data in the processor 101, as well as the data exchanged with the external memory 1022 such as the hard disk. The processor 101 exchanges data with the external memory 1022 through the memory 1021.

[0132] The specific execution process of the above instructions can refer to the steps of the content generation method described in the embodiment of the present disclosure, and will not be repeated here.

[0133] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes the steps of the content generation method described in the above method embodiment. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0134] The embodiments of the present disclosure further provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the content generation method described in the above method embodiment. For details, please refer to the above method embodiment and will not be repeated here.

[0135] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0136] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0137] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0138] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0139] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling an electronic device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0140] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure shall be subject to the scope of protection of the claims.

Claims

1. A content generation method, comprising: Get the input content to be processed; Determine the optimized target text description information corresponding to the to-be-processed content, and determine the corresponding key text description information according to the target text description information, wherein the target text description information is generated based on the text generation model and the image generation model in an alternating collaborative manner; Based on the key text description information and the target text description information, a first graphic result of a first type, or a second graphic result of a second type, is generated, wherein different types of graphic results represent different levels of importance of image genre and text genre, and target images of image genres in different types of graphic results are generated based on the key text description information or the target text description information.

2. The method according to claim 1, wherein the content to be processed includes text to be processed or video to be processed, and determining that the content to be processed corresponds to the optimized target text description information comprises: Based on the text generation model, generate optimized target text description information corresponding to first text description information of a first picture, wherein the first picture is generated based on the picture generation model and the text to be processed, or the first picture is a target frame picture included in the video to be processed; According to the target text description information, corresponding key text description information is determined.

3. The method according to claim 2, wherein generating the first text description information of the first image corresponding to the optimized target text description information based on the text generation model comprises: In the case where the content to be processed is the video to be processed, determining a target frame image according to a key frame of the video to be processed, or determining a corresponding target frame image according to a preset frame interval; Based on the picture description model, taking the target frame picture as input, obtaining first text description information of the target frame picture; Determine target optimization guidance information; The target optimization guidance information and the first text description information are input into the text generation model, and semantic analysis is performed on the first text description information according to the target optimization guidance information to generate optimized target text description information corresponding to the first text description information.

4. The method according to claim 2, wherein generating the first text description information of the first image corresponding to the optimized target text description information based on the text generation model comprises: In a case where the content to be processed is the text to be processed, based on the image generation model, taking the text to be processed as input, generating a first image corresponding to the text to be processed; Based on the picture description model, taking the first picture as input, obtaining first text description information of the first picture; Determine target optimization guidance information; The target optimization guidance information and the first text description information are input into the text generation model, and semantic analysis is performed on the first text description information according to the target optimization guidance information to generate optimized target text description information corresponding to the first text description information.

5. The method according to claim 3 or 4, wherein the determining target optimization guidance information comprises: Determining the target optimization guidance information as preset optimization guidance sentence information; or, Determine the target style that needs to be optimized corresponding to the to-be-processed content, and determine target optimization guidance information corresponding to the target style based on the target style.

6. The method according to claim 5, wherein determining the target style that the content to be processed corresponds to and needs to be optimized comprises: In response to a selection operation for each preset style, determining the preset style corresponding to the selection operation as a target style to be optimized corresponding to the to-be-processed content; or, According to the content to be processed, a target style that needs to be optimized corresponding to the content to be processed is extracted.

7. The method according to any one of claims 1 to 6, wherein generating a first type of first graphic result according to the key text description information and the target text description information comprises: Based on the image generation model, and according to the key text description information, performing semantic analysis on the key text description information to generate the optimized target image; A first image-text result of a first type is obtained according to the target text description information and the target image.

8. The method according to any one of claims 1 to 6, wherein generating a first graphic result of a second type according to the key text description information and the target text description information comprises: According to the target text description information, obtaining a plurality of segmented sub-target text description information; Based on the image generation model, semantic analysis is performed on the plurality of sub-target text description information according to the plurality of sub-target text description information respectively, to generate the corresponding plurality of optimized target images; A second image-text result of a second type is obtained according to the key text description information and the multiple target images.

9. The method according to claim 8, wherein the step of performing semantic analysis on the plurality of sub-goal text description information based on the image generation model and generating the plurality of corresponding optimized target images according to the plurality of sub-goal text description information respectively comprises: Based on the image generation model, the plurality of sub-target text description information and the target style are respectively used as input. Input, semantic analysis is performed on the multiple sub-target text description information and the target style to generate corresponding multiple optimized target images.

10. The method according to claim 7, wherein the step of performing semantic analysis on the key text description information based on the image generation model and generating the optimized target image according to the key text description information comprises: Based on the image generation model, the key text description information and the target style are taken as input, semantic analysis is performed on the key text description information and the target style, and an optimized target image is generated.

11. A content generating device, comprising: An acquisition module is used to obtain input content to be processed; A determination module, used to determine the optimized target text description information corresponding to the to-be-processed content, and determine the corresponding key text description information according to the target text description information, wherein the target text description information is generated based on the text generation model and the image generation model in an alternating collaborative manner; A generation module is used to generate a first image-text result of a first type, or a second image-text result of a second type, based on the key text description information and the target text description information, wherein different types of image-text results represent different levels of importance of image genre and text genre, and target images of image genres in different types of image-text results are generated based on the key text description information or the target text description information.

12. An electronic device comprising: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor performs the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image generation method and device and storage medium

    CN111415396A

  • Bidirectional text image generation method and system based on semantic consistency

    CN113361250A

  • Image-text content generation method and device, equipment and storage medium

    CN117032869A

  • Content generation method and device, electronic equipment and storage medium

    CN117523020A

  • Techniques for generating optimized video segments utilizing a visual search

    US11763564B1