Content generation method and device, server, terminal and storage medium

By determining and using the supplementary description of the target text fragment and its first material in the literary graphics model, the problem of low accuracy in generating unseen concepts or images with high timeliness in the prior art is solved, and a more efficient content generation process is achieved.

CN119938952APending Publication Date: 2025-05-06HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311464468.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When existing literary and biographical graphics models process unseen concepts or inputs with high timeliness, the accuracy of generating images is low, and users need to debug the input prompt words multiple times to generate images that meet expectations, resulting in wasting time and computing resources.

Method used

By determining the target text fragment and its corresponding first material on the basis of text, using these materials as supplementary descriptions, the expected content is then generated. The method includes constructing a preset text fragment set to calibrate knowledge blind spots in the content generation model and enhancing difficult-to-understand text fragments by retrieval.

Benefits of technology

Improve the accuracy of generated content, reduce the number of debugging of text, and save user time and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938952A_ABST
    Figure CN119938952A_ABST
Patent Text Reader

Abstract

The invention provides a content generation method and device, a server, a terminal and a storage medium, and belongs to the field of artificial intelligence content generation. In the method, when expected content described by a text is generated, at least one first material corresponding to a target text segment in the text is determined, a target material in the at least one first material is used as supplementary description of the expected content, and then the expected content is generated based on the text and the target material. In the method, the target material is added on the basis of the text to describe the expected content, so that more abundant description information can be provided for the text which is difficult to understand, the generated content better conforms to expectation, the accuracy of the generated content can be improved, the debugging frequency of the text can be reduced, and the debugging efficiency of the text is improved. Therefore, computing resources and time cost caused by debugging are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence content generation, and in particular to a content generation method, device, server, terminal and storage medium. Background Art

[0002] With the development of artificial intelligence generated content (AIGC) technology, many content generation applications have emerged, such as Wenshengtu applications, which are applications that generate images based on the input natural language description of the expected image (that is, prompt words). Compared with traditional art creation methods such as manual drawing or taking photos and then modifying them, Wenshengtu applications can generate images within a few minutes based on the input natural language, and the quality of the generated images is comparable to that of images generated by traditional methods, which greatly improves the efficiency of art creation and greatly saves costs such as manpower, equipment and site construction.

[0003] In the related art, the Wensheng graph application mainly relies on the Wensheng graph model for image generation, wherein the Wensheng graph model is pre-trained, the prompt word input by the user is obtained, the prompt word is input into the Wensheng graph model, the Wensheng graph model processes the prompt word, and outputs the image corresponding to the prompt word. In some related technologies, the Wensheng graph application also supports transformation of the generated image or further detailed characterization.

[0004] However, due to the limitations of knowledge capacity and timeliness of the text-based graph model, when the input prompt word is a concept that has not been seen during training or a concept that appears after the model training is completed, the probability that the generated image meets expectations is low; and generating an image that meets expectations usually requires inputting a long and carefully designed prompt word, so the user needs to debug the prompt word multiple times before an image that meets expectations may be generated, resulting in a waste of user time and computing resources. Summary of the invention

[0005] The embodiments of the present application provide a content generation method, device, server, terminal and storage medium, which can improve the accuracy of generated content and reduce the number of text debugging times, thereby saving user time and computing resources. The technical solution is as follows.

[0006] In a first aspect, a content generation method is provided, which is applied to a server, and the method comprises:

[0007] When generating the expected content described by the text, at least one first material corresponding to the target text segment in the text is determined, and the target material in the at least one first material is used as a supplementary description of the expected content, and then the expected content is generated based on the text and the target material.

[0008] The expected content may be at least one of an image, text, video, and audio.

[0009] In the above method, since target materials are added to the text to describe the expected content, for difficult-to-understand texts, richer descriptive information can be provided, so that the generated content is more in line with expectations, the accuracy of the generated content can be improved, and the number of debugging of the text can be reduced, thereby saving computing resources and the time cost caused by debugging.

[0010] Optionally, the determining of the target text segment in the text includes at least one of the following:

[0011] In response to a marking operation performed by the terminal on any text fragment in the text, the text fragment corresponding to the marking operation is determined as the target text fragment; based on a preset text fragment set, the text is detected, and the text fragment in the text belonging to the preset text fragment set is determined as the target text fragment, and the preset text fragment set is a knowledge blind spot whitelist of a content generation model used to generate the expected content.

[0012] Among them, the preset text fragment set is used to calibrate the knowledge blind spots of the content generation model. That is, if a text fragment belongs to the preset text fragment set, it means that the content generation model cannot understand the text fragment well, and thus cannot accurately generate the content corresponding to the text fragment.

[0013] In the above method, by constructing a preset text fragment set, the knowledge blind spots of the model are fully calibrated, and then when the text includes a text fragment in the knowledge blind spot, the text fragment is retrieved and enhanced, that is, the text fragment is described and explained through the retrieved information, so as to more fully describe the expected content, so that the content generation model can understand the text, so that the generated content is more in line with expectations, and the influence of the knowledge blind spots and knowledge timeliness of the content generation model on the generation effect can be avoided, and the availability of the content generation service is improved; in addition, the text fragment marked by the user is determined as the target text fragment, and the text fragment that the user considers to be more important in the text to be processed can be determined, and then the server retrieves the text fragment and obtains the corresponding first material to supplement the description of the text fragment, which is conducive to generating content that is more in line with user expectations and improving the accuracy of the generated content.

[0014] Optionally, after determining the text segment in the text that belongs to the preset text segment set as the target text segment, the method further includes:

[0015] Prompt information about the target text segment is sent to the terminal, where the prompt information is used to prompt that the target text segment belongs to the preset text segment set and prompt to modify the target text segment.

[0016] The prompt information may be displayed in a manner of: marking the target text segment in red and displaying a prompt box next to the target text segment; or in a manner of: displaying a dialog box and displaying the prompt information in the dialog box, etc. The prompt information display method may be determined according to actual needs and is not limited thereto.

[0017] In the above method, by displaying prompt information for the target text segment, the user can be prompted to modify the target text segment to modify the target text segment into a text segment that is easier to understand, thereby facilitating the generation of expected content and improving the accuracy of content generation.

[0018] Optionally, the at least one first material includes a text material and an image material, and determining the at least one first material corresponding to the target text segment includes:

[0019] The target text segment is searched to obtain at least one search result of the target text segment, wherein the search result includes a text search result, an image search result and a video search result; based on the text search result, the text material in the at least one first material is determined, and the text material in the first material is used to supplement the text description of the expected content; the video key frame in the video search result is determined, and based on the image search result and the video key frame, the image material in the at least one first material is determined, and the image material in the first material is used to supplement the image description of the expected content.

[0020] The server extracts a keyword from the text search result and determines the keyword as a text material in the first material.

[0021] Optionally, the method further comprises:

[0022] Performing image retrieval on the text material in the first material to obtain image retrieval results corresponding to the text material;

[0023] The method of determining the image material in the at least one first material based on the image retrieval result and the video key frame includes: determining the image material in the at least one first material based on the image retrieval result corresponding to the text material in the first material, the image retrieval result corresponding to the target text segment, and the video key frame.

[0024] In the above method, performing image retrieval again based on the determined text material can enrich the image material, and since the text material is the keyword extracted from the text retrieval result, the image material retrieved based on the text material is more accurate and easier to understand.

[0025] Optionally, the process of determining the target material includes at least one of the following:

[0026] In response to a selection operation of at least one of the first materials by the terminal, the first material corresponding to the selection operation is determined as the target material; and the first material in the at least one first material that is sorted in a front target position is determined as the target material.

[0027] Optionally, the target material includes a text material and an image material, and the generating the expected content corresponding to the text based on the text and the target material in the at least one first material includes:

[0028] The text and the text material in the target material are encoded to obtain text features; the image material in the target material is encoded to obtain image features; the text features and the image features are processed through a content generation model to obtain the expected content corresponding to the text.

[0029] Among them, the content generation model is a text-based image model, for example, a deep learning model for image generation, a stable diffusion model, etc., and the content generation model is not limited. The content generation model can generate images based on pure text, generate images based on text and images, or generate images based on text, images and control information, that is, the content generation model is a multi-mode multi-stream generation model. Among them, the content generation model is a multi-mode multi-stream generation model, which means that the input of the content generation model includes multi-modal information such as text, images and control information, and each mode includes at least one data stream. In some embodiments, the server matches and aligns multiple data streams in the multi-modal information, and inputs each matched and aligned data stream into the content generation model, so that the content generation model can better understand the input multi-modal information, and then generate more accurate and high-quality content. For example, the text to be processed is "Draw a fruit plate with an apple and a banana in it, the apple is on the left and the banana is on the right", and "apple" and "banana" in the text are target text segments, and the target materials include text materials "red, oblate spherical", "yellow, long strip", multiple images of apples, and multiple images of bananas. The server matches and aligns the text segment "apple", the text material "red, oblate spherical", multiple images of apples, and the control information "left" in the text, and matches and aligns the text segment "banana", the text material "yellow, long strip", multiple images of bananas, and the control information "right"; the server inputs the matched and aligned data streams into the content generation model.

[0030] Optionally, the method further comprises:

[0031] Acquire control information, where the control information is used to control the generation of the expected content; encode the control information to obtain control features; and process the text features and the image features through a content generation model to obtain the expected content corresponding to the text, including: based on the control features, process the text features and the image features through the content generation model to obtain the expected content corresponding to the text.

[0032] Optionally, the acquiring of the control information includes at least one of the following:

[0033] The text is identified to obtain control information in the text, where the control information in the text is used to indicate the relationship between the elements in the expected content; in response to a selection operation of a target element in the target material based on the terminal, control information corresponding to the selection operation is obtained, where the control information corresponding to the selection operation indicates to use the target element when generating the expected content.

[0034] The relationship between the elements in the expected content includes combination and fusion, etc., wherein the combination is, for example, element a is on the left of element b; and the fusion is, for example, element a is superimposed on element b.

[0035] In response to the selection operation of the target element in the target material, the terminal generates control information (i.e., an interactive signal) corresponding to the selection operation, and sends the control information corresponding to the selection operation to the server. The selection operation includes at least one of point, box, and mask, which is not limited.

[0036] In the above method, the input text and the multimodal information obtained through retrieval are integrated, and the integrated information is input into the content generation model. That is, the input of the content generation model combines multimodal information such as text, images and interactive signals at the same time, which enables the content generation model to complete content generation more accurately and with high quality, thereby ensuring the knowledge extensibility of the content generation model and the accuracy of the generation effect.

[0037] Optionally, the text is based on terminal input, and the method further includes:

[0038] Acquire an input text fragment, and predict the input intention based on the input text fragment; acquire at least one second material corresponding to the predicted input intention from the Internet and / or a database, wherein the at least one second material is used to complete the input text fragment, and the database stores text and content generated based on the text; and send the at least one second material to the terminal.

[0039] In the above method, in the stage of acquiring the text to be processed, the next input intention is predicted based on the input text fragments, and materials are recommended to the user in the form of images and texts. Compared with only text recommendations, users can more intuitively feel the generated content corresponding to different text fragments, thereby helping users to accurately align future generation effects.

[0040] Optionally, the method further comprises:

[0041] In response to a selection operation of any of the second materials by the terminal, a text segment corresponding to the second material is determined, and the text segment corresponding to the second material is used to complete the input text segment; and the text segment corresponding to the second material is sent to the terminal.

[0042] In the above method, filling in text based on the material selected by the user can help the user quickly construct the desired text and improve the efficiency of text input.

[0043] Optionally, the at least one second material includes a text material and an image material, and determining the text segment corresponding to the second material includes:

[0044] If the second material is a text material, determining a text segment corresponding to the predicted input intention from the second material;

[0045] If the second material is an image material and the image material comes from the Internet, the second material is converted into text by using an image-to-text model, and a text segment corresponding to the predicted input intention is determined from the text converted from the second material;

[0046] If the second material is an image material and the image material comes from the database, the text corresponding to the second material is obtained from the database, and the text segment corresponding to the predicted input intention is determined from the text corresponding to the second material.

[0047] In a second aspect, a content generation method is provided, which is applied to a terminal, and the method includes:

[0048] Obtain a text to be processed and send the text to a server, where the text is used to describe expected content; receive the expected content corresponding to the text returned by the server, and display the expected content, where the expected content is generated based on the text and a target material, where the target material is a material in at least one first material corresponding to a target text segment in the text, and the first material is used to supplement the description of the expected content.

[0049] In the above method, since target materials are added to the text to describe the expected content, for difficult-to-understand texts, richer descriptive information can be provided, so that the generated content is more in line with expectations, the accuracy of the generated content can be improved, and the number of debugging of the text can be reduced, thereby saving computing resources and the time cost caused by debugging.

[0050] Optionally, the method further comprises:

[0051] In response to a marking operation on any text segment in the text, the text segment is marked as the target text segment.

[0052] Optionally, the method further comprises:

[0053] Prompt information for the target text segment is displayed, where the prompt information is used to prompt that the target text segment belongs to a preset text segment set and prompt to modify the target text segment, where the preset text segment set is a knowledge blind spot whitelist of a content generation model used to generate the expected content.

[0054] Optionally, the method further comprises:

[0055] Displaying the at least one first material corresponding to the target text segment;

[0056] In response to a selection operation on at least one of the first materials, the first material corresponding to the selection operation is determined as the target material, and / or the first material in the at least one first material that is sorted in a front target position is determined as the target material.

[0057] In the above method, the terminal displays the first material used to supplement the description of the target text segment, and the user can select the desired first material by himself. Then the server generates the expected content based on the text and the first material selected by the user, which can make the generated content more in line with the user's expectations and improve the controllability and availability of the content generation service.

[0058] Optionally, the method further comprises:

[0059] In response to a selection operation on a target element in the target material, control information corresponding to the selection operation is generated, wherein the control information corresponding to the selection operation indicates that the target element is used when generating the expected content; and the control information corresponding to the selection operation is sent to the server.

[0060] Optionally, the text to be processed is obtained, including:

[0061] In response to the input operation, based on the input text segment, display at least one second material, the second material is used to complete the input text segment, the second material corresponds to the predicted input intention, and the predicted input intention is determined based on the input text segment;

[0062] In response to a selection operation on any of the second materials, a text segment corresponding to the second material is displayed after the input text segment.

[0063] In a third aspect, a content generation device is provided, which is applied to a server. The device includes at least one functional module, and the at least one functional module is used to execute the content generation method provided by the first aspect or any possible implementation method of the first aspect.

[0064] In a fourth aspect, a content generation device is provided, which is applied to a terminal. The device includes at least one functional module, and the at least one functional module is used to execute the content generation method provided by the second aspect or any possible implementation method of the second aspect.

[0065] In a fifth aspect, a server is provided, the server comprising a processor and a memory;

[0066] The processor is used to execute instructions stored in the memory so that the server executes the content generation method provided in the first aspect or any optional manner of the first aspect.

[0067] In a sixth aspect, a server cluster is provided, comprising at least one server, each server comprising a processor and a memory;

[0068] The processor of the at least one server is used to execute instructions stored in the memory of the at least one server, so that the server cluster executes the content generation method provided in the first aspect or any optional manner of the first aspect.

[0069] In a seventh aspect, a terminal is provided, the terminal including a processor and a memory;

[0070] The processor is used to execute instructions stored in the memory, so that the terminal executes the content generation method provided in the above-mentioned second aspect or any optional manner of the second aspect.

[0071] In an eighth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a server, the server executes the content generation method provided in the first aspect or any optional manner in the first aspect.

[0072] In a ninth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a server cluster, the server cluster executes the content generation method provided in the first aspect or any optional method in the first aspect.

[0073] In a tenth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a terminal, the terminal executes a content generation method as provided in the second aspect or any optional method in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 It is a schematic diagram of an implementation environment of a content generation method provided in an embodiment of the present application;

[0075] Figure 2 is a flow chart of a content generation method provided by an embodiment of the present application;

[0076] Figure 3 This is a schematic diagram of a multimodal recommendation completion process provided by an embodiment of the present application;

[0077] Figure 4 It is a schematic diagram of a process of multimodal knowledge recall and multimodal knowledge optimization provided by an embodiment of the present application;

[0078] Figure 5 It is a schematic diagram of a process of information integration and content generation provided by an embodiment of the present application;

[0079] Figure 6 is a system architecture diagram of a content generation system provided by an embodiment of the present application;

[0080] Figure 7 It is a flowchart of a content generation method provided by an embodiment of the present application;

[0081] Figure 8 is a structural diagram of a content generation device provided in an embodiment of the present application;

[0082] Fig. 9 is a structural diagram of a content generation device provided in an embodiment of the present application;

[0083] Fig.10 It is a structural diagram of a server provided in an embodiment of the present application;

[0084] Fig.11 is a schematic diagram of a server cluster provided in an embodiment of the present application;

[0085] Fig.12 It is a schematic diagram of a possible implementation method of a server cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0086] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0087] The embodiments of the present application relate to artificial intelligence services and content generation technologies in the cloud field. For ease of understanding, the relevant concepts in artificial intelligence services and content generation technologies in the cloud field are introduced below.

[0088] First, we introduce the relevant concepts of artificial intelligence services in the cloud field.

[0089] 1. Artificial intelligence (AI): The basic principle of AI is to combine massive data with super-powerful computing and processing capabilities and intelligent algorithms to build an AI model that solves specific problems, so that the AI ​​model can automatically summarize and learn potential patterns or features from the data, thereby achieving a way of thinking close to that of humans. AI models are also AI algorithms (or AI operators), which are a general term for mathematical algorithms built on the principles of AI. They are also the basis for using AI to solve specific problems, such as deep learning (DL) models.

[0090] 2. Training and reasoning of AI models: Any AI model needs to be trained before solving specific technical problems. The training method of an AI model refers to the process of using a specified initial model to calculate the training data, and adjusting the parameters in the initial model using a certain method based on the calculation results, so that the model gradually learns certain rules and has specific functions. After training, an AI model with stable functions can be used for reasoning. The training of an AI model is the process of using a trained AI model to calculate the input data and obtain predicted reasoning results.

[0091] 3. AI Basic Development Platform: The AI ​​Basic Development Platform is a one-stop AI development platform for users. The AI ​​Basic Development Platform can provide various capabilities in the entire AI development process, including data preprocessing, model building and training, model management, model deployment, data optimization, and model optimization updates. Users can complete the development of AI models and the deployment and management of AI applications based on the AI ​​Basic Development Platform. The various capabilities in the AI ​​Basic Development Platform can be integrated for users to use throughout the AI ​​process, or they can provide users with independent functions separately.

[0092] 4. AI services in the cloud field: There are two mainstream types of AI services in the cloud field, one is the AI ​​basic development platform service of the platform-as-a-service (PaaS) type, and the other is the AI ​​application cloud service of the software-as-a-service (SaaS) type. For the first type of AI basic development platform service, public cloud service providers provide customers with AI basic development platforms with the support of sufficient underlying resources and upper-level AI algorithm capabilities. The AI ​​development framework and various AI algorithms built into the AI ​​basic development platform can be used by customers to quickly build and develop AI models or AI applications that meet personalized needs on the AI ​​basic development platform. For the second type of AI application cloud service, public cloud service providers provide ready-to-use, general AI application cloud services through cloud platforms, allowing customers to use AI capabilities in various application scenarios with zero barriers.

[0093] The following introduces relevant concepts in content generation technology.

[0094] Query: The text entered by the user in the search box.

[0095] Prompt: Prompt word, input by the user when generating content, usually refers to a text prompt word. It should be noted that Query and Prompt are both texts, and they are essentially the same, except that the text is called Query in the retrieval stage and Prompt in the content generation stage.

[0096] Modality: Every source or form of information can be called a modality.

[0097] Cross-modal retrieval: The demand for information retrieval is often not just data of a single modality for the same event. Data from other modalities may also be needed to enrich our understanding of the same thing or event. In this case, cross-modal retrieval is needed to achieve retrieval between data in different modalities.

[0098] Multimodal retrieval: Multimodal retrieval often refers to the retrieval input or retrieval output containing data in multiple modalities, and is often accompanied by cross-modal retrieval formulas.

[0099] Multimodal content generation: Similar to multimodal retrieval, in the embodiments of the present application, it refers to content generation using data from multiple modalities as input.

[0100] Multi-source fusion: Integrate data from various sources, absorb the characteristics of different data sources, and then extract unified information from them that is better and richer than single data.

[0101] Multi-mode and multi-stream: Multi-mode refers to multi-modality, that is, it includes multiple modes; multi-stream refers to multiple data streams, and in the embodiments of the present application, it specifically refers to multiple groups of data streams under the same mode.

[0102] The implementation environment of the embodiments of the present application is introduced below.

[0103] Figure 1 Schematic diagram of an implementation environment of a content generation method provided in an embodiment of the present application. Figure 1 As shown, the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are directly or indirectly connected via wired or wireless communication, which is not limited.

[0104] Among them, the terminal 101 can be at least one of a personal computer (PC), a desktop computer, a laptop, a mobile terminal, a smart phone, a tablet computer, a smart watch, a virtual reality terminal, an augmented reality terminal, a wireless terminal and a laptop computer. The terminal 101 has a communication function and can access a wired network or a wireless network. The terminal 101 can generally refer to one of a plurality of terminals, and the embodiment of the present application is only illustrated by the terminal 101. Those skilled in the art can know that the number of the above terminals can be more or less. Exemplarily, a target application is running on the terminal 101, and the target application provides a content generation function and can call the content generation service provided by the server 102. For example, the terminal 101 obtains a text to be processed based on the target application, and the text is used to describe the expected content; the terminal 101 sends the text to the server 102, and after the server 102 receives the text, it generates the expected content corresponding to the text and returns the expected content to the terminal 101; the terminal 101 displays the expected content. Wherein, in some embodiments, the target application provides a text input function, and the terminal 101 obtains the text to be processed in response to the input operation based on the target application; in other embodiments, the target application provides a text import function, and the terminal 101 obtains the text to be processed in response to the text import operation based on the target application; in some other embodiments, the target application provides a text import function and a text input function, and the terminal 101 obtains the imported text in response to the text import operation based on the target application, and obtains the text to be processed in response to the input operation implemented on the basis of the imported text, and the embodiment of the present application does not limit the way in which the terminal 101 obtains the text to be processed. Wherein, the expected content can be at least one of an image, text, video, and audio, which is not limited. The target application can also provide content editing (for example, image editing, video editing, text editing, or audio editing) function and a save function, which is not limited in the embodiment of the present application. The target application can be an application that specifically provides content generation functions, such as a text generation application, an image generation application, a video generation application, a text generation application, a composite content generation application (for example, a generation application for books to be illustrated); the target application can also be a functional module or plug-in on an image editing application, a social application, or a video playback application, etc., without limitation. The target application can take the form of a web application, a microprogram, or a client, etc., and the present application is not limited thereto. Among them, a microprogram refers to a program that relies on other applications to run, such as a mini-program.

[0105] Among them, server 102 can be an independent physical server, or a server cluster or distributed file system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The number of the above-mentioned servers 102 can be more or less, and the embodiments of the present application are not limited to this. Among them, server 102 can include a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a memory and an object storage service (OBS), etc., and the embodiments of the present application are not limited to this. Exemplarily, a trained content generation model is deployed on server 102, which can provide content generation services to generate the expected content corresponding to the text. For example, server 102 receives a text to be processed sent by terminal 101, and the text is used to describe the expected content; server 102 processes the text through a content generation model to generate the expected content corresponding to the text; server 102 returns the expected content to terminal 101, and terminal 101 displays the expected content; for another example, server 102 receives a text to be processed and an image to be processed sent by terminal 101, and the text is used to describe the content generated based on the image; server 102 processes the text and the image, and generates the content corresponding to the text based on the image; server 102 returns the image to terminal 101, and the terminal displays the image. In some embodiments, the server 102 also provides an object storage service, which can store the generated content and the text corresponding to the content, that is, store historical generation data, which is not limited in the embodiments of the present application; in other embodiments, the server 102 has a third-party search function, and can access a third-party search engine to meet the search requirements for multimodal information such as text, images, and videos.

[0106] In some embodiments, the wired network or wireless network uses standard communication technology and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a dedicated network, or any combination of a virtual private network. In some embodiments, technologies and / or formats including hypertext markup language (HTML), extensible markup language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as secure socket sayer (SSL), transport layer security (TLS), virtual private network (VPN), and Internet protocol security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0107] The embodiment of the present application provides a content generation method, in which, when generating the expected content described by the text, at least one first material corresponding to the target text segment in the text is determined, and the target material in the at least one first material is used as a supplementary description of the expected content, and then the expected content is generated based on the text and the target material. In the above method, since the target material is added on the basis of the text to describe the expected content, for the text that is difficult to understand, there can be richer description information, so that the generated content is more in line with expectations, the accuracy of the generated content can be improved, and the number of debugging of the text can be reduced, thereby saving computing resources and the time cost caused by debugging.

[0108] Among them, the expected content can be at least one of images, text, video, and audio, that is, the content generation method provided in the embodiment of the present application can be applied to various content generation applications that are mainly text-driven, such as text generation, image generation, video generation or audio generation, etc.; it can also be applied to composite content generation, for example, the generation of books to be illustrated, which is not limited in the embodiment of the present application.

[0109] The following takes image generation as an example to introduce the process of the content generation method. Figure 2is a flow chart of a content generation method provided by an embodiment of the present application, such as Figure 2 As shown, taking the interaction between a terminal and a server as an example, the method includes the following steps 201 to 2010.

[0110] 201. In response to an input operation, the terminal displays at least one second material based on an input text segment, where the second material is used to complete the input text segment, and the second material corresponds to a predicted input intention, where the predicted input intention is determined based on the input text segment.

[0111] Among them, the text fragment can be in the form of any natural language, for example, Chinese or English, etc., and the embodiment of the present application does not limit this. The at least one second material includes text material and image material. The predicted input intention refers to the information type of the text fragment that may be input after the input text fragment, and the information type includes structural description, style description, quality description, and physical description, etc. The embodiment of the present application does not limit the information type. The correspondence between the second material and the predicted input intention means that the information described by the second material belongs to the information type corresponding to the predicted input intention. For example, if the information type corresponding to the predicted input intention is style description, the second material is a text describing the style ("realistic style" or "abstract school", etc.) and an image with a distinct style (an image with a fairy tale color or a pixel style image, etc.); if the information type corresponding to the predicted input intention is a structural description, the second material is a text describing the structure ("symmetrical structure" or "diagonal structure", etc.) and an image with a clear structure.

[0112] In some embodiments, the predicted input intent is predicted by the server based on the input text segment, and the process includes: the server obtains the input text segment; the server encodes the input text segment, and processes the encoded text segment through the intent prediction model to obtain the predicted input intent. Among them, the intent prediction model includes a decoding layer and an output layer, and the decoding layer can predict the elements that may appear in the next position based on the input text segment, and iterate until the prediction of the entire text segment is completed; the output layer is used to determine the probability that the predicted text segment belongs to different information types, and the information type with the highest probability is used as the predicted input intent. It should be noted that the above description of the prediction process and intent prediction model of the input intent is only exemplary, and the embodiments of the present application do not limit the prediction process and intent prediction model of the input intent.

[0113] In some embodiments, the at least one second material is determined by the server based on the predicted input intent, and the process includes: the server obtains at least one second material corresponding to the predicted input intent from the Internet and / or the database; the server sends the at least one second material to the terminal. Wherein, the server obtains the second material from the Internet means: the server retrieves based on the predicted input intent and obtains at least one text and image corresponding to the predicted input intent. Wherein, the database stores text and images generated based on the text, that is, historical generation data, and the server obtains text and images corresponding to the predicted input intent from the database. In some embodiments, the database stores the text historically input by each user in the user community and the images generated based on the text, and the server obtains at least one text and image from the historical generation data of the user community. It should be noted that the above description of the source of the second material is only exemplary, and second materials from other sources can be used according to actual needs, and the embodiments of the present application do not limit this.

[0114] 202. In response to a selection operation on any second material, the terminal displays a text segment corresponding to the second material after the input text segment to obtain a text to be processed, where the text is used to describe the expected content.

[0115] The text segment corresponding to the second material is generated by the server based on the second material and sent to the terminal by the server. The process of the server determining the text segment corresponding to the second material includes the following three situations.

[0116] The first case: the second material corresponding to the selected operation is a text material, and the server determines the text segment corresponding to the predicted input intent from the second material. The server extracts the keyword corresponding to the predicted input intent from the second material, and uses the keyword as the text segment corresponding to the predicted input intent. For example, the predicted input intent is a style description, and the selected second material is "realistic style images have rich descriptions of real details", and the keyword extracted by the server from the text is "realistic style", then the text segment corresponding to the predicted input intent is "realistic style".

[0117] The second case: the second material corresponding to the selected operation is an image material and the image material comes from the Internet. The server converts the second material into text through the image-to-text model, and determines the text segment corresponding to the predicted input intent from the text converted from the second material. Among them, the image-to-text model can be a convolutional neural network (CNN) or other model that can be used for image-to-text conversion, and the embodiment of the present application does not limit this. The process of the server determining the text segment corresponding to the predicted input intent from the text converted from the second material is the same as the first case above, and will not be repeated.

[0118] The third case: the second material corresponding to the selected operation is an image material and the image material comes from a database. The server obtains the text corresponding to the second material from the database, and determines the text segment corresponding to the predicted input intent from the text corresponding to the second material. The database stores text and images generated based on the text. If the second material is an image material and comes from a database, the server obtains the text corresponding to the image material from the database based on the image material. The obtained text includes multiple text segments, and the server identifies the text segment corresponding to the predicted input intent from the multiple text segments. For example, the predicted input intent is a style description, and the text corresponding to the selected image material is "draw a pixel-style village image, there is a small stream next to the village, and there is a big tree next to the stream." The server recognizes that the text segment belonging to the style description in the text is "pixel style", and the text segment corresponding to the predicted input intent is "pixel style".

[0119] It should be noted that the above step 202 is explained by taking the selection of one second material as an example, and the selection of multiple second materials is the same as the selection of one second material, which will not be described in detail. In addition, after the text segment corresponding to the second material is displayed after the input text segment, the above steps 201 and 202 can be repeated until the text input is completed.

[0120] The above steps 201 and 202 are to obtain the text to be processed by multi-modal recommendation completion. Figure 3 The process shown in the above step 201 and step 202 is described with an example. Figure 3 is a schematic diagram of a multimodal recommendation completion process provided by an embodiment of the present application, such as Figure 3As shown, step ①: in response to the input operation, the terminal displays the input text segment, which is “draw a picture”; step ②: the server predicts the input intention based on “draw a picture”, and determines that the probability of the input intention being quality description, physical description and style description is 0.1, 0.3 and 0.6 respectively, and the server determines the input intention with the highest probability, namely style description, as the predicted input intention; step ③: the server obtains the second material related to the style description, and the terminal displays the second material, the second material includes image material and text material, wherein the image material includes images from the Internet and images from the database; step ④: in response to the selection operation of the image material from the Internet, the server converts the image material into text through the image-to-text model, and determines that the text segment corresponding to the style description in the text is “fairy tale color”, and in response to the selection operation of the image material from the database (that is, the user's work), the server obtains the text corresponding to the image material, and extracts the text segment “pixel style” related to the style description from the text (that is, performs prompt word analysis); step ⑤: the terminal displays “fairy tale color, pixel style” after “draw a picture”.

[0121] In the above method, in the stage of acquiring the text to be processed, the next input intention is predicted based on the input text fragments, and materials are recommended to the user in the form of images and texts. Compared with only text recommendations, users can more intuitively feel the generated content corresponding to different text fragments, thereby helping users to accurately align future generation effects. In addition, text filling based on the materials selected by the user can help users quickly construct the desired text and improve text input efficiency.

[0122] It should be noted that the above-mentioned steps 201 and 202 are optional steps. In some embodiments, this step may not be performed, and the text to be processed is obtained based on other methods. For example, the terminal directly obtains the input text in response to the input operation; for another example, the terminal obtains the imported text in response to the text import operation, and the imported text is the text to be processed; for another example, the terminal converts the input voice into text in response to the voice input operation, and the converted text is the text to be processed. The embodiments of the present application do not limit the method of obtaining the text to be processed.

[0123] 203. The terminal sends the text to the server.

[0124] 204. The server detects the text based on a preset text segment set, and determines the text segments in the text that belong to the preset text segment set as target text segments. The preset text segment set is a knowledge blind spot whitelist of a content generation model used to generate the expected content.

[0125] The preset text segment set is used to calibrate the knowledge blind spots of the content generation model, that is, if a text segment belongs to the preset text segment set, it means that the content generation model cannot understand the text segment well, and thus cannot accurately generate the content corresponding to the text segment. The preset text segment set is constructed by the server, and the construction method includes at least one of the following.

[0126] Method 1: The server performs word segmentation on each training corpus of the content generation model, counts the frequency of occurrence of each word segmentation, and adds the word segmentation with a frequency less than the preset frequency to the preset text segmentation set. For example, a training corpus is "draw a fruit plate with an apple and a banana in it", and the word segmentation results are "draw", "one", "fruit plate", "banana", "apple", "... in", and the frequency of occurrence of "apple" in all training corpus is less than the preset frequency, then the server adds "apple" to the preset text segmentation set. In the above method, the word segmentation with a lower frequency in the training corpus of the content generation model is classified into the knowledge blind spot, which can fully calibrate the knowledge blind spot of the content generation model and reduce the impact of the uneven distribution of the training corpus on the generated content.

[0127] Method 2: The server inputs the training corpus into the trained content generation model, obtains the shallow feature response corresponding to each training corpus, determines the training corpus whose shallow feature response is less than the preset value, extracts the keywords in the training corpus, and adds the keywords to the preset text segment set. The smaller the shallow feature response, the more incorrect the understanding of the training corpus by the first few layers of the content generation model is, and the less accurate the content generated by the content generation model will be.

[0128] Method 3: The server inputs the training corpus into the trained content generation model to obtain the content generated by the content generation model. The server determines the relevance between the content generated by the content generation model (which can be understood as the predicted value) and the content corresponding to the training corpus (which can be understood as the true value), determines the training corpus with a relevance less than the preset relevance, and extracts the keywords from the training corpus and adds the keywords to the preset text segment set. Among them, the lower the relevance between the content generated by the content generation model and the content corresponding to the training corpus, the less expected the generation effect of the content generation model on the training corpus is.

[0129] Method 4: When the server generates content based only on the input text, the server analyzes the user's modification operations on the text and the export operations on the generated content, and determines the text in which the number of modification operations is greater than the preset number and the generated content is not exported. The server extracts keywords from the text and adds the keywords to the preset text fragment set.

[0130] It should be noted that the server for constructing the preset text segment set and the server for providing content generation services may be the same server or may not be the same server, and this embodiment of the present application does not limit this.

[0131] In the above method, by constructing a preset text fragment set, the knowledge blind spots of the model are fully calibrated, and then when the text includes a text fragment in the knowledge blind spot, the text fragment is retrieved and enhanced, that is, the text fragment is described and explained through the retrieved information, so as to more fully describe the expected content, so that the content generation model can understand the text, thereby making the generated content more in line with expectations, avoiding the influence of the knowledge blind spots and knowledge timeliness of the content generation model on the generation effect, and improving the availability of the content generation service.

[0132] In some embodiments, after determining the target text segment in the text, the server sends a prompt message about the target text segment to the terminal, and the prompt message is used to prompt that the target text segment belongs to the preset text segment set and prompt to modify the target text segment; after receiving the prompt message, the terminal displays the prompt message. Among them, the display method of the prompt message can be: marking the target text segment in red and displaying a prompt box next to the target text segment; or: displaying a dialog box, displaying the prompt message in the dialog box, etc. The display method of the prompt message can be determined according to actual needs, and the embodiment of the present application does not limit this. In the above method, by displaying the prompt message for the target text segment, the user can be prompted to modify the target text segment, so as to modify the target text segment into a text segment that is easier for the content generation model to understand, thereby facilitating the generation of content that meets expectations and improving the accuracy of content generation.

[0133] 205. The server determines at least one first material corresponding to the target text segment, where the first material is used to supplement the description of the expected content.

[0134] The first material includes text material and image material. The process of the server determining at least one first material corresponding to the target text segment includes the following steps 2051 to 2053.

[0135] 2051. Search the target text segment to obtain at least one search result corresponding to the target text segment, where the search result includes a text search result, an image search result, and a video search result.

[0136] The server performs multimodal retrieval on the target text segment to obtain retrieval results in different modalities such as text, image and video, and each modality includes at least one retrieval result.

[0137] 2052. Based on the text search result, determine at least one text material in the first material, where the text material in the first material is used to supplement the text description of the expected content.

[0138] Among them, the server extracts keywords from the text search results and determines the keywords as the text material in the first material. In some embodiments, the server identifies the keywords in the text search results through a large language model (LLM) to obtain text materials. For example, the text to be processed is "Draw a fruit plate with an apple and a banana in it", the target text segment is "apple" (that is, the knowledge blind spot is "apple"), and a text search result is "Apple is a flat spherical fruit", then the server extracts the keyword "flat spherical" in the text search result, and the obtained text material is "flat spherical". In some embodiments, the server extracts keywords from multiple text search results to obtain multiple text materials, so that the user can select the text material in the first material by himself later, which is conducive to generating content that better meets the user's expectations; in other embodiments, the server extracts keywords from multiple text search results, merges multiple keywords, and obtains a text material, which can make the content of the text material richer and the supplementary description of the target text segment more detailed, which is more conducive to the content generation model to understand the target text segment.

[0139] 2053. Determine a video key frame in the video retrieval result, and based on the image retrieval result and the video key frame, determine at least one image material in the first material, where the image material in the first material is used to supplement the image description of the expected content.

[0140] It should be noted that the above-mentioned step 2052 and step 2053 are introduced by taking the order of first executing step 2052 and then executing step 2053 as an example. In some embodiments, step 2053 is executed first and then step 2052. In other embodiments, step 2052 and step 2053 are executed in parallel. The embodiments of the present application do not limit the execution order of step 2052 and step 2053.

[0141] In some embodiments, the server first determines the text material based on the text retrieval results, and then performs an image retrieval based on the text material to obtain an image retrieval result corresponding to the text material; the server determines the image material based on the image retrieval result corresponding to the text material, the image retrieval result corresponding to the target text segment, and the video key frame corresponding to the video retrieval result. In the above embodiments, performing an image retrieval again based on the determined text material can enrich the image material, and since the text material is a keyword extracted from the text retrieval results, the image material retrieved based on the text material is more accurate and easier to understand for the content generation model.

[0142] In some embodiments, after determining the text material and image material in at least one first material, the server sorts the text material and the image material. Wherein, in the case of multiple text materials, the server can sort the text material based on the relevance between the text material and the target text segment. In some embodiments, the text retrieval results recalled by the server based on the target text segment have been sorted according to the relevance, then the server only needs to determine the text material corresponding to the text retrieval result ranked in the front target position, and the sorting of the text retrieval results is the sorting of the text material. Wherein, for multiple image materials, the server quantifies the aesthetics, resolution, popularity with users and other indicators of each image material, and then fuses the quantified indicators (for example, weighted summation) to obtain the score of each image material, and then sorts the image material based on the score. It should be noted that the above description of the sorting method of text materials and image materials is only exemplary, and the sorting method can be determined according to actual needs. The embodiment of the present application does not limit this.

[0143] In the above method, by constructing a preset text fragment set, the knowledge blind spots of the model are fully calibrated, and then when the text includes a text fragment in the knowledge blind spot, the text fragment is retrieved and enhanced, that is, the text fragment is described and explained through the retrieved information, so as to more fully describe the expected content, so that the content generation model can understand the text, thereby making the generated content more in line with expectations, avoiding the influence of the knowledge blind spots and knowledge timeliness of the content generation model on the generation effect, and improving the availability of the content generation service.

[0144] It should be noted that the above steps 204 and 205 are a method of determining the target text segment in the text and determining at least one first material corresponding to the target text segment. This process is to perform retrieval enhancement on the text segment in the knowledge blind spot of the content generation model. In some embodiments, the process is also implemented based on other methods. For example, the server responds to the marking operation of any text segment in the text based on the terminal, determines the text segment corresponding to the marking operation as the target text segment, and the server determines at least one first material corresponding to the target text segment. Among them, the marking operation can be to select the text segment and then click the right button of the mouse, select the text segment, double-click the text segment, etc. The embodiment of the present application does not limit the marking operation. In the above embodiment, the text segment marked by the user is determined as the target text segment, which can determine the text segment that the user considers to be more important in the text to be processed, and then the server retrieves the text segment and obtains the corresponding first material to supplement the description of the text segment, which is conducive to generating content that is more in line with user expectations and improving the accuracy of the generated content.

[0145] 206. The server sends the at least one first material to the terminal.

[0146] In some embodiments, the server sorts the text material and the image material in the at least one first material, and then the server sends the sorted at least one first material to the terminal.

[0147] 207. The terminal displays the at least one first material, and in response to a selection operation on the at least one first material, determines the first material corresponding to the selection operation as a target material.

[0148] In some embodiments, if the number of first materials corresponding to the selected operation is less than a preset number, the terminal determines a target number of first materials in the front order among at least one first material except the first material corresponding to the selected operation as the target material.

[0149] In the above method, the terminal displays the first material used to supplement the description of the target text segment, and the user can select the desired first material by himself, and then the server generates the expected content based on the text and the first material selected by the user, which can make the generated content more in line with the user's expectations and improve the controllability and availability of the content generation service. In addition, the first material is displayed in a preferred order, which is convenient for users to select high-quality first materials, thereby improving the quality of the generated content.

[0150] It should be noted that the above steps 206 and 207 are optional steps. In some embodiments, steps 206 and 207 are not executed. For example, the server directly determines the first material ranked in the front target position among the at least one first material as the target material, that is, the server determines the first material ranked in the front as the first material by default, and does not send the first material to the terminal for display, that is, it does not interact with the user in the target material determination link, which can reduce the number of user operations and improve the content generation efficiency.

[0151] The above steps 205 to 207 are the process of multimodal knowledge recall and multimodal knowledge optimization based on the target text segment. Figure 4 The process shown in step 205 to step 207 is described with an example. Figure 4 is a schematic diagram of a process of multimodal knowledge recall and multimodal knowledge optimization provided by an embodiment of the present application, such as Figure 4 As shown, the text to be processed is "Draw a fruit plate with an apple and a banana in it", the server determines that the target text segment is "apple", and the server performs multimodal retrieval based on "apple" to obtain text retrieval results ("apple is a flat spherical fruit"), image retrieval results (Figure A, Figure B and Figure C) and video retrieval results (Video 1, Video 2 and Video 3); the server extracts the keyword "flat spherical" from the text retrieval results, that is, it performs information summary on the text retrieval results, and the server retrieves images based on "flat spherical" to obtain Figures D, E and F; the server extracts the key frames in the video retrieval results to obtain Figure 1 , Figure 2 and Figure 3 The server sorts the images based on aesthetics, resolution, and popularity (the short arrows above the images indicate that the images are sorted, "←" indicates that the corresponding images are moved forward, and "→" indicates that the corresponding images are moved backward), and sends the sorted text and image materials to the terminal, that is, performing material multi-channel aggregation; the terminal displays the sorted text and image materials. It should be noted that Figure 4 The shown is exemplary only. Figure 4 The examples in the description do not limit the embodiments of the present application.

[0152] 208. The server generates expected content corresponding to the text based on the text and the target material.

[0153] The server generates at least one expected content corresponding to the text. In some embodiments, the process of the server generating the expected content based on the text and the target material includes: encoding the text and the text material in the target material to obtain text features; encoding the image material in the target material to obtain image features; and processing the text features and the image features through a content generation model to obtain the expected content corresponding to the text.

[0154] In other embodiments, the server controls the generation of the expected content through control information, and the process includes: obtaining control information, the control information is used to control the generation of the expected content; encoding the control information to obtain control features; and processing the text features and image features based on the control information through a content generation model to obtain the expected content corresponding to the text. The server obtains the control information in at least one of the following ways.

[0155] Method 1: Identify the text and obtain control information in the text, where the control information in the text is used to indicate the relationship between the elements in the expected content. The relationship between the elements in the expected content includes combination and fusion, where combination is, for example, element a is on the left of element b; fusion is, for example, element a is superimposed on element b.

[0156] Method 2: In response to a selection operation of a target element in a target material based on a terminal, control information corresponding to the selection operation is obtained, and the control information corresponding to the selection operation indicates that the target element is used when generating the expected content. Among them, the terminal generates control information (that is, an interactive signal) corresponding to the selection operation in response to the selection operation of the target element in the target material, and the terminal sends the control information corresponding to the selection operation to the server. The selection operation includes at least one of click (point), box selection (box) and smear (mask), which is not limited in the embodiment of the present application.

[0157] Among them, the content generation model is a text-based graph model, for example, a deep learning model for image generation, a stable diffusion model (stable diffusion), etc., and the embodiments of the present application do not limit the content generation model. The content generation model can generate images based on plain text, generate images based on text and images, or generate images based on text, images and control information, that is, the content generation model is a multi-mode multi-stream generation model. Among them, the content generation model is a multi-mode multi-stream generation model, which means that the input of the content generation model includes multi-modal information such as text, images and control information, and each mode includes at least one data stream. In some embodiments, the server matches and aligns multiple data streams in the multi-modal information, and inputs each matched and aligned data stream into the content generation model, so that the content generation model can better understand the input multi-modal information, and then generate more accurate and high-quality content. For example, the text to be processed is "Draw a fruit plate with an apple and a banana in it, the apple is on the left and the banana is on the right", and "apple" and "banana" in the text are target text segments, and the target materials include text materials "red, oblate spherical", "yellow, long strip", multiple images of apples, and multiple images of bananas. The server will match and align the encoded text segment "apple", the text material "red, oblate spherical", multiple images of apples, and the control information "left" in the text, and will match and align the encoded text segment "banana", the text material "yellow, long strip", multiple images of bananas, and the control information "right"; the server will input the matched and aligned data streams into the content generation model.

[0158] Among them, the content generation model is pre-trained by the server based on the training corpus, and the training corpus includes plain text, text-image, text-image-control information. The training process of the content generation model includes: encoding the training corpus to obtain various types of features of the training corpus (text features, image features and control features); matching and aligning the features of each type, and inputting the matched and aligned features into the content generation model; the content generation model processes the features to obtain the generated content corresponding to the training corpus; determining the loss value of the content generation model, and iteratively training the content generation model based on the loss value until the loss value is less than the preset loss, so as to obtain a trained content generation model. It should be noted that the above description of the training process of the content generation model is only exemplary, and the embodiments of the present application do not limit the training method of the content generation model. In addition, the server for training the content generation model and the server providing the content generation service may be the same or different, and the embodiments of the present application do not limit this.

[0159] In the above step 208, the input text and the multimodal information obtained through retrieval are integrated, and the integrated information is input into the content generation model. That is, the input of the content generation model combines multimodal information such as text, images and interactive signals, and matches and aligns multiple data streams under the multimodal information, which enables the content generation model to complete content generation more accurately and with high quality, thereby ensuring the knowledge extensibility of the content generation model and the accuracy of the generation effect.

[0160] The above step 208 is the process of information integration and content generation. Figure 5 The process shown in the above step 208 is described with an example. Figure 5 is a schematic diagram of a process of information integration and content generation provided by an embodiment of the present application, such as Figure 5 As shown, the server integrates the text (which includes the content appeal description and control information) and the text material in the target material (determined based on the text recalled by the retrieval engine and / or the text generated by the large language model) to obtain multiple groups of text, and inputs the integrated multiple groups of text into the text encoder to obtain multiple groups of text features (the dimension of each group of text features is X, and there are N groups of text features in total, where X and N are positive integers greater than 0); the server inputs the multiple groups of image materials in the target material (determined based on the images recalled by the retrieval engine) into the image encoder to obtain multiple groups of image features; the server inputs the control information corresponding to the selection operation (click (point), box (box) and smear (mask)) into the control information encoder to obtain multiple groups of control features; the server matches and aligns the multiple groups of text features, the multiple groups of image features and the multiple groups of control features, and inputs the matched and aligned features into the multi-mode multi-stream generation model (that is, the content generation model), and the multi-mode multi-stream generation model processes the features to output the expected content.

[0161] 209. The server sends the expected content to the terminal.

[0162] 2010. The terminal displays the expected content.

[0163] In some embodiments, the terminal displays at least one expected content corresponding to the text, and the terminal displays a variant of the expected content (for example, changing the position of an element in the expected content) in response to a transformation instruction for any expected content. In some embodiments, the server modifies the control information based on the transformation instruction for any expected content, transforms the expected content based on the modified control information, obtains a variant of the expected content, and sends the variant of the expected content to the terminal. In other embodiments, the terminal displays the expected content after the details are expanded (for example, the expected content after the resolution is increased) in response to a detail expansion instruction for any expected content. In some embodiments, the server adds the control information corresponding to the detail expansion instruction to the control information based on the detail expansion instruction for any expected content, expands the details of the expected content based on the modified control information, obtains the expected content after the details are expanded, and sends the expected content after the details are expanded to the terminal.

[0164] Below through Figure 6 and Figure 7 , combined with the above Figures 3 to 5 , the process shown in the above steps 201 to 2010 is illustrated by way of example. Figure 6 is a system architecture diagram of a content generation system provided in an embodiment of the present application, such as Figure 6 As shown, the content generation system includes multiple terminals 601 and a cloud server 602. The terminal 601 includes a PC and a mobile terminal, and each terminal 601 includes a user interaction interface and hardware, and the user interaction interface is used to input text and display the expected content corresponding to the text; the cloud server 602 provides content generation services, and the content generation services include a multimodal recommendation completion module, a knowledge blind spot identification module, a knowledge recall module, a knowledge optimization module, an information integration module, and a multi-mode multi-stream content generation module. The cloud server 602 sends instructions to the hardware through an operating system (operating system, OS) to execute the functions of the above-mentioned modules, wherein the hardware in the cloud server 602 includes a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a memory and an object storage service (OBS), etc.

[0165] Figure 7 is a flow chart of a content generation method provided by an embodiment of the present application, such as Figure 7As shown, the multimodal recommendation and completion module in the server is used to: identify the input intention (structural description, style description, quality description, etc.) and recommend multimodal completion during the input process. The multimodal content includes pictures and texts to help users quickly construct the desired text. The recommended completion content in the form of images will be automatically converted into text fragments for filling after the user selects it (corresponding to the above steps 201 and 202); the terminal constructs a content generation instruction based on the input operation and the selection operation of the recommended material, and sends the content generation instruction to the server (corresponding to the above step 203); the knowledge blind spot identification module in the server is used to: identify the parts of the content generation instruction that are difficult for the model to understand and generate by querying the knowledge blind spot whitelist (that is, the preset text fragment set), and provide data support for the downstream (corresponding to the above step 204); the knowledge recall module in the server is used to: recall sufficient multimodal information for the knowledge blind spot, including text, images and videos (corresponding to the above step 205); the knowledge optimization module in the server is used to: select the recalled The server can extract the key information in the multimodal information returned, taking into account the diversity, coordination, aesthetics, etc. of the information, determine the selected information, and support users to click and select as well as partially select the selected information (corresponding to the above steps 206 and 207); the information integration module in the server is used to integrate the control information, original content generation instructions and expanded information generated by the user through interaction (i.e., complete the matching association between the information), and integrate the information into a state recognizable by the content generation model; the multi-mode multi-stream generation module in the server is used to use the integrated information as the input of the multi-mode multi-stream generation model, and generate the expected content through the multi-mode multi-stream generation model (corresponding to the above step 208); the server sends at least one expected content generated by the model to the terminal, and the terminal displays the at least one expected content. If any content meets the expectations, the user can export the content, otherwise, it can return to the text input stage to optimize the content generation instructions or select one of the displayed contents for variation and / or detail expansion (corresponding to the above steps 209 and 2010).

[0166] In the above method, when generating the expected content described by the text, the first material corresponding to the target text segment in the text is determined, and the target material in the first material is used as a supplementary description of the expected content, thereby generating the expected content based on the text and the target material. In the above method, since the target material is added on the basis of the text to describe the expected content, for difficult-to-understand text, the generated content can be made more in line with expectations, the accuracy of the generated content can be improved, and the number of debugging times for the text can be reduced, thereby saving the time cost and computing resources brought by debugging; further, in the acquisition stage of the text to be processed, based on the input text fragment, the next input intention is predicted, and the material is recommended to the user in the form of image and text. Compared with only text recommendation, the user can more intuitively feel the generated content corresponding to different text fragments, thereby helping the user to accurately align the future generation effect. In addition, text filling based on the material selected by the user can help the user quickly build the desired text and improve the text input efficiency; in addition, the input text and the multimodal information obtained through retrieval are integrated, and the integrated information is input into the content generation model, that is, the input of the content generation model combines multimodal information such as text, image and interactive signal at the same time, and multiple data streams under the multimodal information are matched and aligned, which can enable the content generation model to complete content generation more accurately and with high quality, and ensure the knowledge extensibility of the content generation model and the accuracy of the generation results.

[0167] Figure 8 80 is a schematic diagram of the structure of a content generation device provided in an embodiment of the present application. The device is applied to a server and includes an acquisition module 801, a first determination module 802, a second determination module 803 and a generation module 804.

[0168] The acquisition module 801 is used to acquire the text to be processed, where the text is used to describe the expected content;

[0169] The first determination module 802 is used to determine a target text segment in the text;

[0170] The second determination module 803 is used to determine at least one first material corresponding to the target text segment, where the first material is used to supplement the description of the expected content;

[0171] The generating module 804 is used to generate the expected content corresponding to the text based on the text and the target material in the at least one first material.

[0172] Optionally, the first determining module 802 is configured to perform at least one of the following:

[0173] In response to a marking operation performed by the terminal on any text segment in the text, determining the text segment corresponding to the marking operation as the target text segment;

[0174] Based on a preset text segment set, the text is detected, and a text segment in the text that belongs to the preset text segment set is determined as the target text segment. The preset text segment set is a knowledge blind spot whitelist of a content generation model used to generate the expected content.

[0175] Optionally, the device further comprises:

[0176] The sending module is used to send prompt information about the target text segment to the terminal, where the prompt information is used to prompt that the target text segment belongs to the preset text segment set and prompt to modify the target text segment.

[0177] Optionally, the at least one first material includes a text material and an image material, and the second determining module 803 includes:

[0178] A retrieval unit, used to search the target text segment to obtain at least one retrieval result of the target text segment, wherein the retrieval result includes a text retrieval result, an image retrieval result, and a video retrieval result;

[0179] A first determining unit, configured to determine a text material in the at least one first material based on the text search result, wherein the text material in the first material is used to supplement the text description of the expected content;

[0180] The second determination unit is used to determine the video key frame in the video retrieval result, and based on the image retrieval result and the video key frame, determine the image material in the at least one first material, wherein the image material in the first material is used to supplement the image description of the expected content.

[0181] Optionally, the retrieval unit is further used for:

[0182] Performing image retrieval on the text material in the first material to obtain image retrieval results corresponding to the text material;

[0183] The second determining unit is used for:

[0184] Based on the image retrieval result corresponding to the text material in the first material, the image retrieval result corresponding to the target text segment and the video key frame, the image material in the at least one first material is determined.

[0185] Optionally, the second determining module 803 is used for at least one of the following:

[0186] In response to a selection operation of at least one of the first materials by the terminal, determining the first material corresponding to the selection operation as the target material;

[0187] A first material in the at least one first material that is sorted in a front target position is determined as the target material.

[0188] Optionally, the target material includes text material and image material, and the generating module 804 includes:

[0189] A first encoding unit, used for encoding the text and the text material in the target material to obtain text features;

[0190] A second encoding unit, used for encoding the image material in the target material to obtain image features;

[0191] The processing unit is used to process the text feature and the image feature through a content generation model to obtain the expected content corresponding to the text.

[0192] Optionally, the acquisition module 801 further includes:

[0193] An acquisition unit, used for acquiring control information, where the control information is used for controlling the generation of the expected content;

[0194] The generating module 804 further includes a third encoding unit, which is used to encode the control information to obtain a control feature;

[0195] The processing unit is used to:

[0196] The content generation model is used to process the text feature and the image feature based on the control feature to obtain the expected content corresponding to the text.

[0197] Optionally, the acquisition unit is used for at least one of the following:

[0198] Recognize the text to obtain control information in the text, where the control information in the text is used to indicate the relationship between elements in the expected content;

[0199] In response to a selection operation of a target element in the target material based on a terminal, control information corresponding to the selection operation is acquired, where the control information corresponding to the selection operation indicates to use the target element when generating the expected content.

[0200] Optionally, the text is input based on a terminal, and the device further comprises:

[0201] A third determination module is used to obtain an input text segment and predict the input intention based on the input text segment;

[0202] a fourth determination module, configured to obtain at least one second material corresponding to the predicted input intention from the Internet and / or a database, wherein the at least one second material is used to complete the input text segment, and the database stores text and content generated based on the text;

[0203] The sending module is used to send the at least one second material to the terminal.

[0204] Optionally, the device further comprises:

[0205] a fifth determining module, configured to determine, in response to a selection operation of any second material by the terminal, a text segment corresponding to the second material, wherein the text segment corresponding to the second material is used to complete the input text segment;

[0206] The generating module is further used to send the text segment corresponding to the second material to the terminal.

[0207] Optionally, the at least one second material includes a text material and an image material, and the fifth determining module is used to:

[0208] If the second material is a text material, determining a text segment corresponding to the predicted input intention from the second material;

[0209] If the second material is an image material and the image material comes from the Internet, the second material is converted into text by using an image-to-text model, and a text segment corresponding to the predicted input intention is determined from the text converted from the second material;

[0210] If the second material is an image material and the image material comes from the database, the text corresponding to the second material is obtained from the database, and the text segment corresponding to the predicted input intention is determined from the text corresponding to the second material.

[0211] It should be noted that, in other embodiments, the steps that the above modules are responsible for implementing can be specified as needed, and all the functions of the above device can be realized by respectively implementing different steps in the above content generation method through the above modules. That is, the content generation device provided in the above embodiment only uses the division of the above functional modules as an example when implementing the content generation method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment belongs to the same concept as the corresponding method embodiment. The specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0212] Fig. 9 90 is a schematic diagram of the structure of a content generation device provided in an embodiment of the present application, which is applied to a terminal. The device includes a sending module 901 and a display module 902.

[0213] The sending module 901 is used to obtain the text to be processed and send the text to the server, where the text is used to describe the expected content;

[0214] The display module 902 is used to receive the expected content corresponding to the text returned by the server and display the expected content. The expected content is generated based on the text and target material. The target material is a material in at least one first material corresponding to a target text segment in the text. The first material is used to supplement the description of the expected content.

[0215] Optionally, the device further comprises:

[0216] The marking module is used for marking any text segment in the text as the target text segment in response to a marking operation on the text segment.

[0217] Optionally, the display module 902 is further used for:

[0218] Prompt information for the target text segment is displayed, where the prompt information is used to prompt that the target text segment belongs to a preset text segment set and prompt to modify the target text segment, where the preset text segment set is a knowledge blind spot whitelist of a content generation model used to generate the expected content.

[0219] Optionally, the display module 902 is further used for:

[0220] Displaying the at least one first material corresponding to the target text segment;

[0221] The device also includes a determination module for, in response to a selection operation on at least one of the first materials, determining the first material corresponding to the selection operation as the target material, and / or determining the first material sorted in the front target position among the at least one first material as the target material.

[0222] Optionally, the device further comprises:

[0223] A generating module, configured to generate control information corresponding to a selection operation on a target element in the target material in response to the selection operation on the target element, wherein the control information corresponding to the selection operation indicates to use the target element when generating the expected content;

[0224] The generating module is also used to send control information corresponding to the selection operation to the server.

[0225] Optionally, the display module 902 is further used for:

[0226] In response to the input operation, based on the input text segment, display at least one second material, the second material is used to complete the input text segment, the second material corresponds to the predicted input intention, and the predicted input intention is determined based on the input text segment;

[0227] In response to a selection operation on any of the second materials, a text segment corresponding to the second material is displayed after the input text segment.

[0228] It should be noted that, in other embodiments, the steps that the above modules are responsible for implementing can be specified as needed, and all the functions of the above device can be realized by respectively implementing different steps in the above content generation method through the above modules. That is, the content generation device provided in the above embodiment only uses the division of the above functional modules as an example when implementing the content generation method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment belongs to the same concept as the corresponding method embodiment. The specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0229] Among them, the acquisition module 801, the first determination module 802, the second determination module 803, the generation module 804, the sending module 901 and the display module 902 can all be implemented by software, or can be implemented by hardware. Exemplarily, the implementation of the acquisition module 801 is introduced below by taking the acquisition module 801 as an example. Similarly, the implementation of the first determination module 802, the second determination module 803, the generation module 804, the sending module 901 and the display module 902 can refer to the implementation of the acquisition module 801.

[0230] As an example of a software functional unit, the acquisition module 801 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the acquisition module 801 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.

[0231] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0232] As an example of a hardware functional unit, the acquisition module 801 may include at least one computing device, such as a server, etc. Alternatively, the acquisition module 801 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0233] The multiple computing devices included in the acquisition module 801 can be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module 801 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the acquisition module 801 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0234] The present application also provides a server 1000. Fig.10 is a schematic diagram of the structure of a server provided in an embodiment of the present application, such as Fig.10 As shown, the server 1000 includes: a bus 1001, a processor 1002, a memory 1003 and a communication interface 1004. The processor 1002, the memory 1003 and the communication interface 1004 communicate with each other through the bus 1001. The server 1000 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the server 1000.

[0235] The bus 1001 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.10 The bus 1001 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 1001 may include a path for transmitting information between various components of the server 1000 (for example, the memory 1003, the processor 1002, and the communication interface 1004).

[0236] The processor 1002 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0237] The memory 1003 may include a volatile memory, such as a random access memory (RAM). The memory 1003 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0238] The memory 1003 stores executable program codes, and the processor 1002 executes the executable program codes to respectively implement the functions of the aforementioned acquisition module 801, the first determination module 802, the second determination module 803, and the generation module 804, thereby implementing the content generation method. That is, the memory 1003 stores instructions for executing the content generation method.

[0239] Alternatively, the memory 1003 stores executable program codes, and the processor 1002 executes the executable program codes to respectively implement the functions of the aforementioned sending module 901 and display module 902, thereby implementing the content generation method. That is, the memory 1003 stores instructions for executing the content generation method.

[0240] The communication interface 1004 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the server 1000 and other devices or communication networks.

[0241] The embodiment of the present application also provides a server cluster. Fig.11 is a schematic diagram of a server cluster provided in an embodiment of the present application, such as Fig.11 As shown, the server cluster includes at least one server 1000. The memory 1003 in one or more servers 1000 in the server cluster may store the same instructions for executing the content generation method.

[0242] In some possible implementations, the memory 1003 of one or more servers 1000 in the server cluster may also store partial instructions for executing the content generation method. In other words, the combination of one or more servers 1000 may jointly execute instructions for executing the content generation method.

[0243] It should be noted that the memory 1003 in different servers 1000 in the server cluster can store different instructions, which are respectively used to execute part of the functions of the content generation device. That is, the instructions stored in the memory 1003 in different servers 1000 can implement the functions of one or more of the aforementioned acquisition module 801, the first determination module 802, the second determination module 803 and the generation module 804.

[0244] It should be understood that Fig.11 The functions of the server 1000 shown in the figure may also be performed by multiple servers 1000 .

[0245] In some possible implementations, one or more servers in the server cluster may be connected via a network, which may be a wide area network or a local area network, etc. Fig.12 A possible implementation is shown. Fig.12 is a schematic diagram of a possible implementation of a server cluster provided in an embodiment of the present application, such as Fig.12 As shown, two servers 1000 are connected via a network. Specifically, the two servers 1000 are connected to the network via a communication interface in each server.

[0246] The present application embodiment also provides another server cluster. The connection relationship between the servers in the server cluster can be similarly referred to as Fig.11 or Fig.12 The difference is that the memory 1003 in one or more servers 1000 in the server cluster may store the same instructions for executing the content generation method.

[0247] In some possible implementations, the memory 1003 of one or more servers 1000 in the server cluster may also store partial instructions for executing the content generation method. In other words, the combination of one or more servers 1000 may jointly execute instructions for executing the content generation method.

[0248] It should be noted that the memory 1003 in different servers 1000 in the server cluster can store different instructions, which are respectively used to execute part of the functions of the content generation device. That is, the instructions stored in the memory 1003 in different servers 1000 can implement the functions of one or more of the aforementioned acquisition module 801, the first determination module 802, the second determination module 803 and the generation module 804.

[0249] An embodiment of the present application provides a terminal, which includes a processor and a memory; the memory stores program code, and the processor is used to execute the program code stored in the memory, so that the terminal executes the steps performed by the terminal in the content generation method provided in the above embodiment.

[0250] The present application implements a computer program product, which may be software or a program product containing program code and capable of running on a server, a terminal, or being stored in any available medium. When the computer program product runs on at least one server, the at least one server executes the steps executed by the server in the content generation method provided in the above method embodiment; when the computer program product runs on a terminal, the terminal executes the steps executed by the terminal in the content generation method provided in the above embodiment.

[0251] The embodiment of the present application provides a computer-readable storage medium, which can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes a program code, and when the program code is executed by at least one server, the at least one server executes the steps performed by the server in the content generation method provided in the above method embodiment; when the program code is executed by a terminal, the terminal executes the steps performed by the terminal in the content generation method provided in the above method embodiment.

[0252] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the images and texts involved in this application are obtained with full authorization. In some embodiments, the embodiments of the present application provide a permission inquiry page, which is used to inquire whether to grant permission to obtain the above information. In the permission inquiry page, an authorization consent control and a authorization rejection control are displayed. When a trigger operation for the authorization consent control is detected, the content generation method provided in the embodiments of the present application is used to obtain the above information.

[0253] Those of ordinary skill in the art will appreciate that the various method steps and units described in the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the steps and components of each embodiment have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0254] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0255] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the unit is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.

[0256] The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiment of the present application.

[0257] In addition, each unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or software unit.

[0258] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computing device (which can be a personal computer, a server, or a computing device, etc.) to execute all or part of the steps of the method in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0259] In this application, the words such as "first" and "second" are used to distinguish the same or similar items with basically the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second" and "nth", nor is the quantity and execution order limited. It should also be understood that although the following description uses the terms first, second, etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first material can be referred to as the second material, and similarly, the second material can be referred to as the first material. The first material and the second material can both be materials, and in some cases, can be separate and different materials.

[0260] The term "at least one" in this application means one or more, the term "plurality" in this application means two or more, and the terms "system" and "network" are often used interchangeably in this document.

[0261] It should also be understood that the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting." Similarly, the phrase "if it is determined that ..." or "if [a stated condition or event] is detected" may be interpreted to mean "upon determining that ..." or "in response to determining that ..." or "upon detecting [a stated condition or event]" or "in response to detecting [a stated condition or event]," depending on the context.

[0262] The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0263] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer program instructions. When the computer program instructions are loaded and executed on a computer, the process or function in accordance with the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0264] The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital video disc (DVD), or a semiconductor medium (e.g., a solid-state drive)), etc.

[0265] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0266] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A content generation method, characterized in that: Applied to a server, the method comprises: Acquire text to be processed, where the text is used to describe expected content; Determine a target text segment in the text, and determine at least one first material corresponding to the target text segment, where the first material is used to supplement the description of the expected content; The expected content corresponding to the text is generated based on the text and a target material in the at least one first material.

2. The method according to claim 1, characterized in that: The determining of the target text segment in the text includes at least one of the following: In response to a marking operation performed by a terminal on any text segment in the text, determining the text segment corresponding to the marking operation as the target text segment; Based on a preset text segment set, the text is detected, and text segments in the text that belong to the preset text segment set are determined as the target text segments. The preset text segment set is a knowledge blind spot whitelist of a content generation model used to generate the expected content.

3. The method according to claim 2, characterized in that After determining the text segment in the text that belongs to the preset text segment set as the target text segment, the method further includes: Prompt information about the target text segment is sent to the terminal, where the prompt information is used to prompt that the target text segment belongs to the preset text segment set and to prompt to modify the target text segment.

4. The method according to any one of claims 1 to 3, characterized in that The at least one first material includes a text material and an image material, and the determining the at least one first material corresponding to the target text segment includes: Searching the target text segment to obtain at least one search result of the target text segment, wherein the search result includes a text search result, an image search result, and a video search result; Determining a text material in the at least one first material based on the text search result, wherein the text material in the first material is used to supplement the text description of the expected content; Determine a video key frame in the video retrieval result, and based on the image retrieval result and the video key frame, determine an image material in the at least one first material, wherein the image material in the first material is used to supplement the image description of the expected content.

5. The method according to claim 4, characterized in that The method further comprises: Performing image retrieval on the text material in the first material to obtain image retrieval results corresponding to the text material; The determining, based on the image retrieval result and the video key frame, an image material in the at least one first material comprises: Based on the image retrieval results corresponding to the text material in the first material, the image retrieval results corresponding to the target text segment and the video key frame, the image material in the at least one first material is determined.

6. The method according to any one of claims 1 to 5, characterized in that The process of determining the target material includes at least one of the following: In response to a selection operation of at least one of the first materials by a terminal, determining the first material corresponding to the selection operation as the target material; A first material in the at least one first material that is sorted in a front target position is determined as the target material.

7. The method according to any one of claims 1 to 6, characterized in that The target material includes a text material and an image material, and the generating the expected content corresponding to the text based on the text and the target material in the at least one first material includes: Encoding the text and the text material in the target material to obtain text features; Encoding the image material in the target material to obtain image features; The text features and the image features are processed through a content generation model to obtain the expected content corresponding to the text.

8. The method according to claim 7, characterized in that The method further comprises: Acquiring control information, where the control information is used to control the generation of the expected content; Encoding the control information to obtain a control feature; The processing of the text features and the image features by a content generation model to obtain the expected content corresponding to the text includes: The text feature and the image feature are processed by the content generation model based on the control feature to obtain the expected content corresponding to the text.

9. The method according to claim 8, characterized in that The obtaining of the control information includes at least one of the following: Recognize the text to obtain control information in the text, where the control information in the text is used to indicate the relationship between elements in the expected content; In response to a selection operation of a target element in the target material based on a terminal, control information corresponding to the selection operation is acquired, where the control information corresponding to the selection operation indicates to use the target element when generating the expected content.

10. The method according to any one of claims 1 to 9, characterized in that The text is input based on a terminal, and the method further comprises: Obtaining an input text segment, and predicting an input intention based on the input text segment; Acquire at least one second material corresponding to the predicted input intention from the Internet and / or a database, wherein the at least one second material is used to complete the input text segment, and the database stores text and content generated based on the text; The at least one second material is sent to the terminal.

11. The method according to claim 10, characterized in that The method further comprises: In response to a selection operation of any of the second materials by the terminal, determining a text segment corresponding to the second material, where the text segment corresponding to the second material is used to complete the input text segment; The text segment corresponding to the second material is sent to the terminal.

12. The method according to claim 11, characterized in that The at least one second material includes a text material and an image material, and determining a text segment corresponding to the second material includes: If the second material is a text material, determining a text segment corresponding to the predicted input intention from the second material; If the second material is an image material and the image material comes from the Internet, converting the second material into text through an image-to-text model, and determining a text segment corresponding to the predicted input intent from the text converted from the second material; If the second material is an image material and the image material comes from the database, the text corresponding to the second material is obtained from the database, and a text segment corresponding to the predicted input intention is determined from the text corresponding to the second material.

13. A content generation method, characterized in that: Applied to a terminal, the method comprises: Acquire text to be processed, and send the text to a server, where the text is used to describe expected content; Receive expected content corresponding to the text returned by the server, and display the expected content, wherein the expected content is generated based on the text and target material, wherein the target material is material in at least one first material corresponding to a target text segment in the text, and the first material is used to supplement the description of the expected content.

14. The method according to claim 13, characterized in that The method further comprises: In response to a marking operation on any text segment in the text, the text segment is marked as the target text segment.

15. The method according to claim 13, characterized in that The method further comprises: Prompt information for the target text segment is displayed, wherein the prompt information is used to prompt that the target text segment belongs to a preset text segment set and prompt to modify the target text segment, and the preset text segment set is a knowledge blind spot whitelist of a content generation model used to generate the expected content.

16. The method according to any one of claims 13 to 15, characterized in that The method further comprises: Displaying the at least one first material corresponding to the target text segment; In response to a selection operation on at least one of the first materials, the first material corresponding to the selection operation is determined as the target material, and / or the first material in the at least one first material that is sorted in a front target position is determined as the target material.

17. The method according to any one of claims 16, characterized in that The method further comprises: In response to a selection operation on a target element in the target material, generating control information corresponding to the selection operation, wherein the control information corresponding to the selection operation indicates that the target element is used when generating the expected content; Send control information corresponding to the selection operation to the server.

18. The method according to any one of claims 13 to 17, characterized in that The obtaining of the text to be processed includes: In response to the input operation, based on the input text segment, display at least one second material, the second material is used to complete the input text segment, the second material corresponds to the predicted input intention, and the predicted input intention is determined based on the input text segment; In response to a selection operation on any of the second materials, a text segment corresponding to the second material is displayed after the input text segment.

19. A content generating device, characterized in that: The device comprises: An acquisition module, used for acquiring text to be processed, wherein the text is used for describing expected content; A first determination module, used to determine a target text segment in the text; A second determination module is used to determine at least one first material corresponding to the target text segment, where the first material is used to supplement the description of the expected content; A generating module is used to generate the expected content corresponding to the text based on the text and the target material in the at least one first material.

20. A content generating device, characterized in that: The device comprises: A sending module, used for acquiring text to be processed and sending the text to a server, wherein the text is used for describing expected content; A display module is used to receive the expected content corresponding to the text returned by the server and display the expected content, wherein the expected content is generated based on the text and a target material, wherein the target material is a material in at least one first material corresponding to a target text segment in the text, and the first material is used to supplement the description of the expected content.

21. A server, characterized in that: The server includes a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the server executes the content generation method according to any one of claims 1 to 12.

22. A server cluster, characterized in that: comprising at least one server, each server comprising a processor and a memory; The processor of the at least one server is used to execute instructions stored in the memory of the at least one server, so that the server cluster executes the content generation method according to any one of claims 1 to 12.

23. A terminal, characterized in that: The terminal includes a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the terminal executes the content generating method according to any one of claims 13 to 18.

24. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a server, the server performs the content generation method according to any one of claims 1 to 12.

25. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a server cluster, the server cluster performs the content generation method according to any one of claims 1 to 12.

26. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a terminal, the terminal executes the content generation method according to any one of claims 13 to 18.