A content generation method, apparatus, storage medium, and electronic device
By filtering and adjusting contextual information, and using contextual samples to filter the model and user descriptions to adjust the model, the problem of generated content not meeting user expectations in AIGC technology has been solved, achieving high quality and style controllability of generated content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2023-08-29
- Publication Date
- 2026-04-10
AI Technical Summary
Existing AIGC technology struggles to accurately generate content that meets user expectations based on user descriptions, especially in terms of content layout and details.
By obtaining initial content from user input, generating descriptions and reference content, using contextual samples to filter the model and adjusting the model based on user descriptions, filtering out highly relevant recommendation context information, generating descriptions based on recommendation context information and initial content, determining higher-quality recommendation content generation descriptions, and using the target content generation model to generate content.
It achieves controllability in the quality, style, and content of the generated content, ensuring that the generated content fully meets user expectations, especially in terms of content layout and details, which are more in line with the reference content.
Smart Images

Figure CN117150130B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of computer technology, and particularly relates to a content generation method and device, a storage medium and an electronic device. BACKGROUND
[0002] With the rapid development of computer technology, the application of the way of generating content by artificial intelligence (AIGC) is more and more widely used. The AIGC can generate content in a short time, for example, the AIGC technology can accurately generate the content expected by the user according to the generation description (prompt) of the user. SUMMARY
[0003] The present specification provides a content generation method, device, storage medium and electronic device, and the technical solution is as follows:
[0004] In a first aspect, the present specification provides a content generation method, which comprises:
[0005] obtaining an initial content generation description and reference content input by a user;
[0006] determining recommended context information based on the initial content generation description and the reference content, and determining a recommended content generation description based on the recommended context information and the initial content generation description;
[0007] generating content based on the recommended content generation description and the reference image by using a target content generation model.
[0008] In a second aspect, the present specification provides a context sample screening model training method, which comprises:
[0009] creating an initial context sample screening model, and determining context candidate data for the initial context sample screening model;
[0010] obtaining an initial content generation description sample and a reference content sample, and performing at least one round of model training on the initial context sample screening model based on the initial content generation description sample, the reference content sample and the context candidate data to obtain a context sample screening model after model training.
[0011] In a third aspect, the present specification provides a user description adjustment model training method, which comprises:
[0012] creating an initial user description adjustment model;
[0013] obtain an initial content generation description sample and first context sample information, the first context sample information consisting of at least one first context sample output by a context sample screening model for the initial content generation description sample;
[0014] label the first context sample information with a context sample ranking label based on the initial content generation description sample;
[0015] perform at least one round of model training on the initial user description adjustment model based on the first context sample information, the initial content generation description sample, and the context sample ranking label, to obtain a model-trained user description adjustment model.
[0016] In a fourth aspect, the present specification provides a content generation device, the device comprising:
[0017] a content obtaining module configured to obtain an initial content generation description input by a user and reference content;
[0018] a description determining module configured to determine recommended context information based on the initial content generation description and the reference content, and determine a recommended content generation description based on the recommended context information and the initial content generation description;
[0019] a content generation module configured to perform content generation based on the recommended content generation description and the reference image using a target content generation model.
[0020] In a fifth aspect, the present specification provides a context sample screening model training device, the device comprising:
[0021] a model creating module configured to create an initial context sample screening model and determine context candidate data for the initial context sample screening model;
[0022] a model training module configured to obtain an initial content generation description sample and reference content sample, and perform at least one round of model training on the initial context sample screening model based on the initial content generation description sample, the reference content sample, and the context candidate data, to obtain a model-trained context sample screening model.
[0023] In a sixth aspect, the present specification provides a user description adjustment model training device, the device comprising:
[0024] a model creating module configured to create an initial user description adjustment model;
[0025] a data obtaining module configured to obtain an initial content generation description sample and first context sample information, the first context sample information consisting of at least one first context sample output by a context sample screening model for the initial content generation description sample;
[0026] The data acquisition module is further configured to label a context sample ranking label for the first context sample based on the initial content generation description sample.
[0027] The model training module is configured to perform at least one round of model training on the initial user description adjustment model based on the first context sample information, the initial content generation description sample, and the context sample ranking label, to obtain a user description adjustment model after model training.
[0028] In a seventh aspect, the present specification provides a computer storage medium storing at least one instruction adapted to be loaded by a processor and to execute the method steps of one or more embodiments of the present specification.
[0029] In an eighth aspect, the present specification provides a computer program product storing at least one instruction adapted to be loaded by a processor and to execute the method steps of one or more embodiments of the present specification.
[0030] In a ninth aspect, the present specification provides an electronic device, which can include a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the method steps of one or more embodiments of the present specification.
[0031] The technical solutions provided by some embodiments of the present specification have at least the following beneficial effects:
[0032] In one or more embodiments of the present specification, according to the initial content generation description input by the user, the initial content generation description and the reference content are used to first determine the recommended context information highly related to the current user's input, and then the recommended context information and the initial content generation description are used to determine the recommended content generation description with better quality, and finally the recommended content generation description and the reference content are used to generate content by using the target content generation model. Compared with the initial content generation description prompt and the reference content, the content generated in this way is more controllable in terms of quality, style, content, etc., and the generated content completely meets the user's expectations, especially in terms of content layout and details. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the present specification or the prior art, the following will briefly introduce the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present specification, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0034] Figure 1is a scenario schematic diagram of a content generation system provided in the specification;
[0035] Figure 2 is a flow schematic diagram of a content generation method provided in the specification;
[0036] Figure 3 is a flow schematic diagram of another content generation method provided in the specification;
[0037] Figure 4 is a flow schematic diagram of a context sample screening model training method provided in the specification;
[0038] Figure 5 is a flow schematic diagram of a user description adjustment model training method provided in the specification;
[0039] Figure 6 is a structural schematic diagram of a content generation apparatus provided in the specification;
[0040] Figure 7 is a structural schematic diagram of a model training apparatus provided in the specification;
[0041] Figure 8 is a structural schematic diagram of another model training apparatus provided in the specification;
[0042] Figure 9 is a structural schematic diagram of an electronic device provided in the specification. DETAILED DESCRIPTION
[0043] The technical solutions in the specification will be described clearly and completely in the specification in combination with the drawings in the specification. Obviously, the described embodiments are only part of the embodiments of the specification, rather than all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the specification.
[0044] In the description of the specification, it is understood that the terms "first", "second", etc. are only for the purpose of description and cannot be understood as indicating or implying relative importance. In the description of the specification, it should be noted that, unless otherwise expressly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally also include steps or units not listed, or can optionally also include other steps or units inherent to these processes, methods, products or devices. The specific meaning of the above terms in the specification can be understood by the person skilled in the art according to the specific circumstances. In addition, in the description of the specification, "multiple" means two or more, unless otherwise specified. The association relationship of the associated objects is described, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after are a "or" relationship.
[0045] In related technologies, AIGC technology has been applied in creative inspiration acquisition and work modification, etc., but it is still difficult to achieve controllable generation by using current AIGC. Specifically, AIGC technology is difficult to accurately generate the content expected by the user according to the user's description (prompt), especially in terms of content layout and details, which shows that the content generation of AIGC in related technologies has certain limitations.
[0046] The specification will be described in detail below in conjunction with specific embodiments.
[0047] Please refer to Figure 1 , a scene schematic diagram of a content generation system provided in the specification. As Figure 1 indicated, the content generation system can at least include a client cluster and a service platform 100.
[0048] The client cluster can include at least one client, such as Figure 1 indicated, specifically including a client 1 corresponding to a user 1, a client 2 corresponding to a user 2,..., and a client n corresponding to a user n, n is an integer greater than 0.
[0049] The clients in the client cluster can be electronic devices with communication functions, including but not limited to wearable devices, handheld devices, personal computers, tablet computers, vehicle-mounted devices, smart phones, computing devices, or other processing devices connected to wireless modems, etc. Electronic devices can be called different names in different networks, such as user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic device in 5G network or future evolution network, etc.
[0050] The service platform 100 can be a separate server device, such as a rack-mounted, blade, tower, or cabinet server device, or a hardware device with strong computing power, such as a workstation, mainframe computer, etc. It can also be a server cluster composed of multiple servers. Each server in the service cluster can be composed in a symmetrical manner, where each server is functionally and positionally equivalent in the transaction link, and each server can independently provide services externally. The independent service can be understood as not requiring the assistance of another server.
[0051] In one or more embodiments of the present specification, the service platform 100 can establish a communication connection with at least one client in the client cluster, and complete the interaction of data in the content generation process based on the communication connection, such as online transaction data interaction. For example, the service platform 100 can implement application deployment to several clients based on the target generation content model obtained by the content generation method of the present specification, and the clients can execute the content generation method of one or more embodiments of the present specification to generate content.
[0052] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0053] The content generation system embodiments provided in this specification and the content generation methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the content generation method involved in one or more embodiments of this specification can be the aforementioned service platform 100; the execution entity corresponding to the content generation method involved in one or more embodiments of this specification can also be the electronic device corresponding to the client, specifically determined based on the actual application environment. The implementation process of the content generation system embodiments can be detailed in the following method embodiments, and will not be repeated here.
[0054] based on Figure 1 The following is a detailed description of the content generation method provided by one or more embodiments of this specification, illustrated in the scene diagram.
[0055] Please see Figure 2 This document provides a flowchart illustrating a content generation method according to one or more embodiments of the present specification. This method can be implemented using a computer program and can run on a content generation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The content generation device can be an electronic device.
[0056] Specifically, the content generation method includes:
[0057] S102: Obtain an initial content generation description and reference content input by a user;
[0058] The initial content generation description can be description information of content expected to be generated by the user, such as content style description, content description, content specification description, etc. The initial content generation description is used to instruct the target content generation model to generate content that meets the user's expectations.
[0059] The initial content generation description is usually a descriptive text. A target content corresponding to the content generation description can be automatically generated based on the given content generation description and reference content. The content generation description can be a text input by the user, and of course, it can also be a text converted from the user's voice. The form of obtaining the content generation description and the content of the content generation description are not limited in the present specification.
[0060] It should be noted that the data types of the generated content and the reference content involved in one or more embodiments of the present specification are not limited. For example, the user can expect to generate an image, text, audio, video, etc. based on the initial content generation description. The data types of the reference content and the expected generated content can be the same.
[0061] The reference content is a content reference provided when the target content generation model generates the corresponding content based on the content generation description (for example, the user expects to generate an image content, and the reference content is a reference image content). The reference content is used to instruct the target content generation model to generate target content data that matches the content style based on the content generation description. In this way, the target content data and the reference content can be highly consistent.
[0062] Illustratively, the electronic device can provide an AIGC-based content generation service to the outside. The user can input an initial content generation description prompt and reference content to the content generation service. At this time, the electronic device can obtain the aforementioned data.
[0063] S104: Determine recommended context information based on the initial content generation description and the reference content, and determine a recommended content generation description based on the recommended context information and the initial content generation description;
[0064] The database containing a plurality of context candidate data is pre-configured, and the context candidate data is composed of a content generation description pair (the content generation description pair includes a sample initial content generation description and a sample high-quality description content screened out) determined by high-quality description content of a large number of sample content generation data of a sample user based on platform content generation history data, so as to select sample high-quality description content reference for other user content generation description. The content generation history data of the sample user in the entire content generation process generally includes a plurality of sample content generation descriptions input by the sample user and content generation data output by the model based on each round of sample content generation description. Generally, the last content generation data is the sample user's desired generated content. For the content generation link, each content generation process can be annotated with "sample high-quality content generation description" (such as automatically annotating the last round or last x rounds of sample content generation description as sample high-quality content generation description).
[0065] It can be understood that generally in the entire content generation process, the last round or last x rounds (such as last three rounds) of sample content generation description of the plurality of sample content generation descriptions can be considered as the "sample high-quality content generation description" most suitable for content generation,
[0066] For example, the sample content generation data is pre-processed and integrated, and expert end service can be called for description content processing and description content summary to annotate high-quality content generation description. It can also be automatically annotated as sample high-quality content generation description of the last round or last x rounds of sample content generation description. The content generation description data pair of "sample initial content generation description + sample high-quality content generation description" is obtained. A group of content generation description data pairs is a context candidate data. Based on this, a plurality of context candidate data can be obtained.
[0067] For example, generally, content generation based on a target generation content model involves input of a plurality of rounds of content generation description and output of generated content data based on each round of content generation description. The input data of the user for content generation can be represented as {X1, X2,...Xn} in order of input rounds. X1 represents the first round of input content generation description (not limited to text, image, video, etc. Data types), X2 represents the second round of input content generation description (not limited to text, image, video, etc. Data types), and Xn represents the nth round of input content generation description (not limited to text, image, video, etc. Data types). Wherein, i and n are positive integers.
[0068] Further, the sample content generation description of the first round or the first y1 rounds (the value of y1 can be customized based on actual conditions, such as the first 3 rounds, or the last 10% of the total number of rounds) is generally referred to as the "sample initial content generation description"; further, the sample content generation description of the last round or the last y2 rounds (the value of y2 can be customized based on actual conditions, such as the last 3 rounds, or the last 10% of the total number of rounds) is generally referred to as the "sample high-quality content generation description", or, the sample content generation data {X1, X2,... Xn} is pre-processed and integrated, and expert services can be called for description content processing and description content summary to obtain the sample high-quality content generation description.
[0069] A plurality of context candidate data, each context candidate data consisting of "sample initial content generation description + sample high-quality content generation description" input by a sample user, the "sample initial content generation description" to the "sample high-quality content generation description" is regarded as a description adjustment process, the "sample initial content generation description" is the initial state description before the description adjustment process (which can be regarded as the upper text data), and the "sample high-quality content generation description" is the target state description after the description adjustment process (which can be regarded as the lower text data), both of which constitute the context candidate data, and the context candidate data from the upper text data to the lower text data is a description adjustment process, and the description adjustment characteristics in the description adjustment process can be used for reference for the description adjustment of the "sample initial content generation description" input this time.
[0070] The high-quality content generation description has better description quality than the initial content generation description, and the combination of the two constitutes data pairs with high-quality description adjustment characteristics;
[0071] Illustratively, based on the "initial content generation description and reference content" input by the current user this time, one or more recommended context data that fit the content input by the current user this time (initial content generation description and reference content) are selected from the plurality of context candidate data in the database, and the one or more recommended context data can be referred to as first context information. These first context information have inherent characteristics that can be used for reference for the description adjustment of the "initial content generation description" to facilitate the adjustment of the "initial content generation description" to a recommended content generation description sample with better description quality, and further based on the description adjustment characteristics of these recommended context data, the initial content generation description input this time is adjusted to obtain a recommended content generation description. The recommended content generation description refers to the high-quality description adjustment characteristics (such as description information style adjustment characteristics, description information architecture adjustment characteristics, and description information syntax adjustment characteristics) of the previous context candidate data, and the initial content generation description is adjusted according to these referenceable high-quality description adjustment characteristics to obtain the recommended content generation description.
[0072] In an implementable embodiment, the context candidate data can be subjected to context matching processing based on the initial content generation description and the reference content, to obtain first context information;
[0073] Illustratively, the context matching processing of the context candidate data based on the initial content generation description and the reference content can be to filter out recommended context data (content generation description samples input by other sample users during content generation) that is suitable for the current user input and fits the current user's content generation, for example, the recommended context data that fits the current user's content generation can be selected based on at least one dimension of content generation type, content generation scene, content layout, content theme, etc. as input content description features, which can be understood as that the "sample initial content generation description" in the recommended context data is similar to the "initial content generation description" of the current user in terms of the input content description features of at least one dimension of "content generation type, content generation scene, content layout, content theme, etc.".
[0074] Optionally, the initial content generation description can be subjected to feature engineering to extract sub-description feature vectors of at least one dimension of content generation type, content generation scene, content layout, content theme, etc., and all the sub-description feature vectors constitute the description feature vector of the initial content generation description; similarly, the "sample initial content generation description" in the context candidate data can be subjected to feature engineering to extract candidate sub-description feature vectors of at least one dimension of content generation type, content generation scene, content layout, content theme, etc., and all the candidate sub-description feature vectors constitute the candidate description feature vector of the "sample initial content generation description" in the context candidate data; then the description feature vector and the candidate description feature vector are matched to determine one or more recommended description feature vectors corresponding to the "sample initial content generation description" in the recommended context data, for example, the vector distance (such as Euclidean vector distance) between the description feature vector and the candidate description feature vector can be calculated, and one or more recommended description feature vectors are selected based on the vector distance, to obtain the recommended context data "recommended sample initial content generation description + recommended sample high-quality content generation description" corresponding to the recommended description feature vector.
[0075] Then the reference context (i.e. the recommended context data) in the first context information is subjected to reordering processing, to obtain a second context queue, the recommended context information is determined from the second context queue, and the initial content generation description is adjusted based on the recommended context information to obtain a recommended content generation description.
[0076] Assuming that the reference contexts in the first context information can be represented as {W1, W2,... Wn}, W1, W2,... Wn respectively represent reference contexts, the reference contexts in the first context information {W1, W2,... Wn} are unordered, the reference contexts in {W1, W2,... Wn} are reordered to obtain a second context queue {Z1, Z2,... Zn}, Z1, Z2,... Zn represent the reference contexts after reordering, and then a target number of recommended contexts V are selected from the second context queue {Z1, Z2,... Zn} in the order of sorting, all recommended contexts V form recommended context information, and each recommended context V is a data pair of "sample initial content generation description (regarded as an initial content generation description sample) + sample high-quality content generation description (regarded as a recommended content generation description sample)";
[0077] Illustratively, the reference contexts in the first context information can determine their relative content description quality, that is, the reference contexts in the first context information are labeled with quality order from the description quality fitting dimension, and the second context queue is obtained based on the quality order. Based on this, a target number of recommended contexts can be selected to form recommended context information according to the sorting priority, so as to adjust the initial content generation description based on the recommended context information to obtain the recommended content generation description.
[0078] S106: generating content based on the recommended content generation description and the reference image using a target generation content model.
[0079] In this specification, the target generation content model can be trained based on a basic artificial intelligence generation model in AIGC. The basic artificial intelligence generation model may, for example, be trained based on a diffusion model network.
[0080] The input of the target generation content model is the recommended content generation description and the reference image, and the target generation content model is controlled to generate a target generation content that meets the recommended content generation description and is consistent with the reference image.
[0081] In a feasible implementation, considering the style matching of the generated content and the content compliance, the electronic device can perform at least one of the following steps after determining the recommended content generation description based on the recommended context information and the initial content generation description:
[0082] A2: The electronic device performs user description style matching processing on the recommended content generation description based on the initial content generation description, obtains a style matching detection result, performs style strengthening description processing on the recommended content generation description based on the style matching detection result, and obtains a recommended content generation description after style strengthening description processing;
[0083] Illustratively, the recommended content generation description combined with the recommendation context can have a description style difference from the user input description style. At this time, the recommended content generation description can be subjected to user description style matching processing based on the initial content generation description, so as to obtain a style matching detection result.
[0084] The style matching detection result can be a style matching type or a style non-matching type.
[0085] If the style matching detection result belongs to the style matching type, the recommended content generation description is subjected to style strengthening description processing, and a recommended content generation description after style strengthening description processing is obtained. The style strengthening description processing can be to generate a new recommended content generation description by re-executing S104.
[0086] Optionally, the electronic device performing "user description style matching processing on the recommended content generation description based on the initial content generation description, and obtaining a style matching detection result" can be implemented by using a description detection model based on a machine learning network. The description detection model can be used to perform user description style matching processing on the recommended content generation description, and obtain a style matching detection result.
[0087] A4: The recommended content generation description is subjected to content compliance detection processing, a content compliance detection result is obtained, and the recommended content generation description is subjected to content strengthening description processing based on the content compliance detection result, and a recommended content generation description after content strengthening description processing is obtained.
[0088] Illustratively, the recommended content generation description combined with the recommendation context can have an implicit content compliance problem, such as the existence of pornographic / political / violent content in the recommended content generation description.
[0089] The content compliance detection result can be a content non-compliance type or a content compliance type.
[0090] If the content compliance detection result is a content non-compliance type, the recommended content generation description is subjected to content strengthening description processing based on the content compliance detection result, and a recommended content generation description after content strengthening description processing is obtained. The strengthening description processing can be to generate a new recommended content generation description by re-executing S104.
[0091] Illustratively, the step of the electronic device performing content compliance detection processing on the generated description of the recommended content to obtain a content compliance detection result can be that the generated description of the recommended content is subjected to content compliance detection processing by using a description detection model to obtain a content compliance detection result.
[0092] In a feasible implementation, the training process of the description detection model can refer to the following manner:
[0093] The description detection model can be created based on a machine learning model and trained by using model sample data, specifically as follows:
[0094] Sample data collection: a large amount of model sample data is collected, and the model sample data is the generated description of the recommended content and the corresponding initial content generation description in the foregoing application process;
[0095] Sample data labeling: the model sample data is labeled with a style category label (i.e., a category label indicating whether the generated description of the recommended content is consistent with the corresponding initial content generation description) and a content compliance score label;
[0096] Model forward propagation training: based on the machine learning model, an initial description detection model is created, the model sample data is input into the initial description detection model to perform at least one round of model training, and a style matching detection result (i.e., an actual style prediction category) and a content compliance detection result (i.e., an actual content compliance score) corresponding to the model sample data are obtained;
[0097] Model backward propagation fine-tuning: in each round of model training, a first model loss is calculated based on an actual style prediction category and a style category label based on a model loss calculation formula, a second model loss is calculated based on an actual content compliance score and a content compliance score label based on the model loss calculation formula, the first model loss and the second model loss are used to adjust model coefficients of the initial description detection model, until the initial description detection model satisfies a model training end condition, and a trained description detection model is obtained;
[0098] Optionally, the first model loss and the second model loss can be calculated by using a set model loss function, such as an Euclidean distance loss function, a cross-entropy loss function, a hinge loss function, and the like.
[0099] For example, the first model loss can be calculated by using the following formula:
[0100] Loss 1 = CrossEntropy(pred, y)
[0101] Wherein, Loss 1 is the first model loss, CrossEntropy() represents the cross-entropy calculation process, pred is the actual style prediction category, and y is the style category label (whether it is in one style).
[0102] Optionally, the model end training condition of the model can include, for example, that the value of the loss function is less than or equal to a preset loss function threshold, the number of iterations reaches a preset number threshold, and the like. The specific model end training condition can be determined based on actual conditions, and is not specifically limited here.
[0103] It should be noted that the machine learning model involved in one or more embodiments of the present specification includes, but is not limited to, fitting of one or more of a convolutional neural network (CNN) model, a deep neural network (DNN) model, a recurrent neural network (RNN), a model, an embedding model, a gradient boosting decision tree (GBDT) model, a logistic regression (LR) model, and the like.
[0104] In one or more embodiments of the present specification, in order to overcome the content generation limitations of AIGC in the related art, an initial content generation description prompt is generated according to user input, highly relevant recommended context information is first screened out using the initial content generation description and reference content, a quality more optimal recommended content generation description is determined based on the recommended context information and the initial content generation description, and content generation is performed using the recommended content generation description and the reference content. Compared with the initial content generation description prompt and the reference content, the content generated in terms of quality, style, content, and the like is more controllable, and the content generated can fully meet the user's expectations, especially in terms of content layout and details, and can better match the reference content.
[0105] See Figure 3 , Figure 3 is a flow diagram of another embodiment of a content generation method according to one or more embodiments of the present specification. Specifically:
[0106] S202: input the initial content generation description and the reference content into a context sample screening model;
[0107] S204: Control the context sample screening model to match a plurality of reference contexts from the context candidate data based on the initial content generation description and the reference content, generate first context information composed of a plurality of reference contexts, and output the first context information.
[0108] The context sample screening model is a context sample screening model created and trained based on a machine learning model. The input of the context screening model is the initial content generation description and the reference content. The output of the context screening model is the first context information matched from the context candidate data that fits the initial content generation description and the reference content.
[0109] A plurality of context candidate data is pre-configured. The context candidate data is a pair of context description data of a sample initial content generation description sample and a recommended content generation description sample input by a sample user, that is, "sample initial content generation description (as an initial content generation description sample) + sample high-quality content generation description (as a recommended content generation description sample)". The context sample screening model is used to screen one or more reference contexts that fit the current user's content generation from the "sample initial content generation description" of the plurality of context candidate data based on the model input. The set of one or more reference contexts can be referred to as first context information. The reference contexts in these first context information have adjustment characteristics that can be described, that is, "sample initial content generation description (as an initial content generation description sample) modulation to sample high-quality content generation description (as a recommended content generation description sample)". Since the "sample initial content generation description" that fits the current user's content description is matched, the adjustment experience learning of the description adjustment process of these "sample initial content generation description (as an initial content generation description sample) + sample high-quality content generation description (as a recommended content generation description sample)" is used to adjust the description of the input initial content generation description, thereby generating a recommended content generation description. The recommended content generation description refers to the high-quality quality description characteristics (such as description information style, description information architecture, etc.) of the adjustment process of the past context candidate data, thereby facilitating the description adjustment of the initial content generation description.
[0110] S206: Input the first context information and the initial content generation description into the user description adjustment model, and the user description adjustment model includes a context ranking network and a large language model network.
[0111] The user description adjustment model is a user description adjustment model created and trained based on a machine learning model. In some embodiments, the user description adjustment model includes a model structure composed of a context ranking network and a large language model network (that is, Large Language Model).
[0112] S208: reordering the reference contexts in the first context information through the context ordering network to obtain a second context queue, and selecting a target number of recommended contexts from the second context queue to obtain recommended context information;
[0113] The processing object of the context ordering network in the user description adjustment model is the reference contexts in the first context information. The context ordering network is used for ordering the multiple reference contexts (i.e., the recommended context data) in the first context information to obtain a second context queue, selecting a target number of recommended contexts from the second context queue to form recommended context information, and outputting the recommended context information.
[0114] Suppose that the reference contexts in the first context information can be represented as {W1, W2,... Wn}, W1, W2,... Wn represent reference contexts respectively. The reference contexts in the first context information {W1, W2,... Wn} are unordered. The reference contexts in {W1, W2,... Wn} are reordered to obtain a second context queue {Z1, Z2,... Zn}, Z1, Z2,... Zn represent the reference contexts after reordering. Then, a target number of recommended contexts V are selected from the second context queue {Z1, Z2,... Zn} according to the ordering sequence. All the recommended contexts V form recommended context information, and each recommended context V is a data pair of "sample initial content generation description (regarded as an initial content generation description sample) + sample high-quality content generation description (regarded as a recommended content generation description sample)".
[0115] Illustratively, the context ordering network calculates the relative content description quality between the reference contexts in the first context information {W1, W2,... Wn} during processing, labels the quality ordering sequence for the reference contexts from the description quality fitting dimension, and orders based on the quality ordering sequence to obtain the second context queue.
[0116] In actual applications, directly recommending and fine-tuning the initial content generation description based on the first context information may have two phenomena. The first phenomenon is that the data volume is too large, which affects the overall processing speed and reduces system efficiency. The second phenomenon is that a large amount of context data may contain some noise, which reduces the quality of the final prompt. To solve this limitation, an LLM large language model fine-tuning method based on intelligent context learning and reordering is proposed. A small number of appropriate samples can be selected from a large number of first context information as a second context queue for context description adjustment learning and current description prompt prediction of high-quality context candidate data.
[0117] S210: performing description recommendation fine-tuning on the initial content generation description based on the recommended context information by the large language model network.
[0118] The large language model network in the user description adjustment model can be directly obtained by training a Large Language Model model in the related art. The processing object of the large language model network is the recommended context information and the initial content generation description. The large language model network controls the description recommendation fine-tuning based on the initial content generation description to obtain the recommended content generation description.
[0119] S212: performing content generation based on the recommended content generation description and the reference image by a target generation content model.
[0120] For details, refer to the method steps of other embodiments of the present specification
[0121] In one or more embodiments of the present specification, in order to overcome the content generation limitations of AIGC in the related art, according to the initial content generation description prompt input by the user, the initial content generation description and the reference content are used to filter out highly relevant recommended context information by using a context sample screening model and a user description adjustment model, and then the recommended content generation description with better quality is determined based on the recommended context information and the initial content generation description. The content generation is performed by using the recommended content generation description and the reference content, which can be more controllable in terms of quality, style, content, etc. than the initial content generation description prompt and the reference content, and the content completely generated according to the user's expectation can be obtained, especially in terms of content layout and details.
[0122] Please refer to Figure 4 A flowchart of a context sample screening model training method is provided for one or more embodiments of the present specification. The method can be implemented by relying on a computer program and can be run on a model training device based on the von Neumann architecture. The computer program can be integrated in an application or run as an independent tool application. The model training device can be an electronic device. The model training device can be the same device as the content generation device or can be a different device. Specifically:
[0123] S302: creating an initial context sample screening model and determining context candidate data for the initial context sample screening model.
[0124] Specifically, a machine learning model is used to create an initial context sample screening model based on the context sample screening task.
[0125] Illustratively, the key to context learning is the screening of context candidate data, and the selected context candidate data can have a significant effect on improving the effect of randomly selected contexts. The selected context refers to a context that is more accurate or more suitable for the initial prompt, and is more consistent with the scene of adjusting the initial prompt. In this link, a context sample screening model is trained from a series of context candidate data, so that the context sample screening model can select the most suitable context sample for the initial content generation description sample from the context candidate data. After training, the context sample screening model has the ability to select context samples based on multi-modal understanding, and can be used to select a high-quality context sample pool according to the initial prompt and reference image input by the user. The context sample pool is the first context information in some embodiments;
[0126] For the initial content generation description sample input by the user, the content generation description with better quality annotated from the database in advance can be used as the context candidate data. Generally, the context candidate data can be obtained by manual screening, that is, for the sample content generation of a sample user, the content generation description with better quality is selected from the content generation description prompt adopted by the sample user as the context candidate data.
[0127] S304: Obtain an initial content generation description sample and a reference content sample, and perform at least one round of model training on the initial context sample screening model based on the initial content generation description sample, the reference content sample, and the context candidate data to obtain the context sample screening model after model training.
[0128] The initial content generation description sample is the initial content generation description input by the sample user in the model training stage;
[0129] Optionally, the recommended content generation description sample can be annotated from the context candidate data based on the initial content generation description sample. In each round of model training process, the initial content generation description sample is input, and the expert end service is manually annotated to annotate the description matching classification label for the context candidate data in the database based on the initial content generation description sample. The description matching classification label can be a probability value of the description matching degree.
[0130] The model structure of the initial context sample screening model can be composed of at least four parts, and the model structure can include a description feature encoding module, a reference content feature encoding module, a context feature encoding module, and a fusion matching module.
[0131] In a feasible implementation, the model training process can refer to the following manner:
[0132] B2: inputting the initial content generation description sample and the reference content sample into an initial context sample screening model for at least one round of model training, and in the model training process, controlling the initial context sample screening model to determine a sample user description matching result corresponding to the context candidate data based on the initial content generation description sample and the reference content sample;
[0133] In the model forward propagation training process:
[0134] The input of the initial context sample screening model is the initial content generation description sample and the reference content sample. The initial content generation description sample and the reference content sample are input into the initial context sample screening model, and the internal processing process can refer to the following interpretation:
[0135] B2-2: determining the content generation description feature corresponding to the initial content generation description sample through a description feature encoding module, determining the reference content feature corresponding to the reference content sample through a reference content feature encoding module, and determining the context feature corresponding to the context candidate data through a context feature encoding module;
[0136] The input of the description feature encoding module is the initial content generation description sample, which is used for description feature encoding of the initial content generation description sample. The output is the content generation description feature;
[0137] The input of the reference content feature encoding module is the reference content sample, which is used for content feature encoding of the reference content sample. The output is the reference content feature;
[0138] The input of the context feature encoding module is the context candidate data, which is used for context feature encoding of the context candidate data. The output is the context feature.
[0139] B2-4: performing description matching classification for the sample user based on the content generation description feature, the reference content feature and the context feature through a fusion matching module, to obtain a sample user description matching result corresponding to the context candidate data;
[0140] The fusion matching module is the content generation description feature, the reference content feature and the context feature, which is used for description matching classification for the sample user. The output is the sample user description matching result corresponding to the context candidate data;
[0141] Generally, the sample user description matching result can represent whether a certain context candidate data is suitable as the context input by the current user. The sample user description matching result can be a description matching probability, based on which it can be determined whether it belongs to the user description matching type or the user description non-matching type.
[0142] For example, a probability threshold can be set. If the probability of a description matching is greater than or equal to the probability threshold, it is considered to belong to the user description matching type. Conversely, if the probability of a description matching is less than the probability threshold, it is considered to belong to the user description non-match type.
[0143] B2-6: If the sample user description matching result is a user description matching type, then output the first context sample corresponding to the user description matching type.
[0144] Understandably, multiple context candidate data sets are pre-maintained, and based on the above method, sample user description matching results corresponding to several context candidate data sets can be obtained. Therefore, the number of possible first context samples output is multiple.
[0145] If the user description matching result of a certain context candidate data is a user description mismatch type, then "a certain context candidate data" will not be output.
[0146] B4: Obtain the description matching classification labels of the context candidate data based on the description samples generated from the initial content, and adjust the model parameters of the initial context sample screening model based on the sample user description matching results and the description matching classification labels until the initial context sample screening model completes training, thereby obtaining the context sample screening model.
[0147] During the backpropagation training process: the model classification loss is calculated based on the sample user description matching results and the description matching classification labels. The model parameters of the initial context sample screening model are adjusted using the model classification loss until the initial context sample screening model meets the model end training conditions, and then the context sample screening model is obtained.
[0148] For illustrative purposes, the classification loss is represented by the Softmax Loss formula. The model classification loss can be calculated using the Softmax Loss formula, and can be expressed as follows:
[0149] Loss = Softmax Loss(pred, y)
[0150] Where pred is the predicted sample user description matching result, and y is the description matching classification label;
[0151] The model's training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific training termination conditions can be determined based on actual circumstances and are not specifically limited here.
[0152] In one or more embodiments of the present specification, a model training manner of a context sample screening model is shown. The key of context learning is the screening of context candidate data. In actual application, the context candidate data selected by the context sample screening model can have a significant effect improvement compared with the randomly selected context. The selected context refers to the context that is more accurate or more suitable for the initial prompt and more consistent with the scene of describing and adjusting the initial prompt. In this link, the context sample screening model is trained from a series of context candidate data, so that the context sample screening model can select the context sample most suitable for the initial content generation description sample from the context candidate data. After training, the context sample screening model has the context sample selection capability based on multi-modal understanding,
[0153] Please refer to Figure 5 A flowchart of a user description adjustment model training method is provided for one or more embodiments of the present specification. The method can be implemented by relying on a computer program and can run on a model training device based on the von Neumann system. The computer program can be integrated in an application or run as a standalone tool application. The model training device can be an electronic device. The model training device can be the same device as the content generation device or can be a different device. Specifically:
[0154] S402: Create an initial user description adjustment model;
[0155] Specifically, a machine learning model is used to create an initial user description adjustment model based on the user description adjustment task. In some embodiments, the initial user description adjustment model includes a model structure composed of a context ranking network and a large language model network (i.e., Large Language Model).
[0156] S404: Obtain an initial content generation description sample and first context sample information, wherein the first context sample information is composed of at least one first context sample output by the context sample screening model for the initial content generation description sample;
[0157] S406: Label the first context sample information with a context sample ranking label based on the initial content generation description sample;
[0158] Data labeling phase: After the first step of screening by the context sample screening model, the context subsequent data obtains the first context sample information. The expert end service is used to manually label the relative good and bad of each first context sample in the first context sample information, and the ranking is generated, that is, each first context sample is labeled with a context sample ranking label considering all first context samples;
[0159] S408: At least one round of model training is performed on the initial user description adjustment model based on the first context sample information, the initial content generation description sample, and the context sample ranking label, to obtain a user description adjustment model after model training.
[0160] The initial user description adjustment model can include a context ranking network and a large language model network. Optionally, only the last i layers of the large language model network are selected for training, and the weights of other parts of the model network remain unchanged.
[0161] C2: The first context sample information and the initial content generation description sample are input into the initial user description adjustment model.
[0162] C4: In each round of model training, the first context sample in the first context sample information is reordered by the user description adjustment model to obtain a second context sample queue, the recommended context sample information is determined from the second context sample queue, and the recommended content generation description sample is obtained by description recommendation fine-tuning of the initial content generation description sample based on the recommended context sample information.
[0163] Further, the first context sample in the first context sample information is reordered by the context ranking network to obtain a second context sample queue, and a target number of recommended context samples are selected from the second context sample queue to obtain recommended context sample information.
[0164] The input of the context ranking network is the first context information selected in the first step, and the output is the ranking of each first context from important to less important. Based on the ranking, the first context can obtain a second context sample queue. The context ranking network selects a target number of recommended context samples (for example, the top K (for example, K=3 or 5) context data) from the second context sample queue to obtain recommended context sample information.
[0165] Further, the initial content generation description sample is fine-tuned by the large language model network based on the recommended context sample information to obtain the recommended content generation description sample.
[0166] The input of the large language model network is the target number of recommended context samples and the initial content generation description sample, and the output is the recommended content generation description sample obtained by fine-tuning the initial content generation description sample.
[0167] C6: determining a model comprehensive loss based on the second context sample queue, the context sample ranking label, the recommended content generation description sample and the initial content generation description sample, and performing model parameter adjustment on the initial user description adjustment model based on the model comprehensive loss to obtain a user description adjustment model after model training.
[0168] Illustratively, a context ranking loss is determined based on the context sample ranking information corresponding to the second context sample queue and the context sample ranking label, a recommended description prediction loss is determined based on the recommended content generation description label corresponding to the recommended content generation description sample and the initial content generation description sample, and a model comprehensive loss is determined based on the context ranking loss and the recommended description prediction loss.
[0169] Optionally, the context ranking loss and the recommended description prediction loss are calculated using a set model loss function, which can be a Euclidean distance loss function, a cross-entropy loss function, a hinge loss function, etc.
[0170] Illustratively, the context ranking loss can also be referred to as ranking loss,
[0171] Loss a = Ranking Loss(pred, y)
[0172] wherein Loss a is the context ranking loss, Pred is the context sample ranking order, and y is the context sample ranking label.
[0173] Illustratively, the recommended description prediction loss can be a regression loss, and the model loss function can be a Euclidean distance loss function, as follows:
[0174] Loss b = L2(pred, y)
[0175] wherein Loss b is the recommended description prediction loss, Pred is the recommended content generation description sample, and y is the recommended content generation description label corresponding to the initial content generation description sample.
[0176] Optionally, the model end training condition of the model can include, for example, that the value of the loss function is less than or equal to a preset loss function threshold value, the number of iterations reaches a preset number threshold, etc. The specific model end training condition can be determined based on actual conditions, and is not limited here.
[0177] In one or more embodiments of the present specification, considering that directly based on the first context information in practical applications, the initial content generation description may be recommended to fine-tune two phenomena, the first phenomenon is that the data volume is too large, which affects the overall processing speed and reduces the system efficiency; The second phenomenon is that a large amount of context data may contain some noise, which will reduce the quality of the final prompt; In order to solve this limitation, the LLM large language model fine-tuning method based on intelligent context learning and reordering is proposed, and the user description adjustment model is trained based on this. Using the user description adjustment model can select a small amount of appropriate samples from a large amount of first context information as a second context queue, and perform context description adjustment learning and current description prediction of high-quality context candidate data.
[0178] The following will be combined Figure 6 The content generation device provided in the present specification will be described in detail. It should be noted that Figure 6 The content generation device shown in the present specification is used to execute the method of the present specification Figures 1-5 The method of the embodiment shown in the present specification is only shown in relation to the present specification for the purpose of illustration, and the specific technical details not disclosed are described with reference to the present specification Figures 1-5 The embodiment shown in the present specification.
[0179] Please refer to Figure 6 , which shows the structural schematic diagram of the content generation device of the present specification. The content generation device 1 can be realized by software, hardware or combination of the two to become all or part of the electronic device. According to some embodiments, the content generation device 1 includes a content acquisition module 11, a description determination module 12 and a content generation module 13, which are specifically used for:
[0180] The content acquisition module 11 is used to acquire the initial content generation description and reference content input by the user;
[0181] The description determination module 12 is used to determine the recommended context information based on the initial content generation description and the reference content, and determine the recommended content generation description based on the recommended context information and the initial content generation description;
[0182] The content generation module 13 is used to generate content based on the recommended content generation description and the reference image using a target generation content model.
[0183] Optionally, the description determination module 12 is used to:
[0184] Based on the initial content generation description and the reference content, the context candidate data is subjected to context matching processing to obtain first context information;
[0185] reordering the reference contexts in the first context information to obtain a second context queue, determining recommended context information from the second context queue, and adjusting the initial content generation description based on the recommended context information to obtain a recommended content generation description.
[0186] Optionally, the description determining module 12 is configured to:
[0187] inputting the initial content generation description and the reference content into a context sample screening model;
[0188] controlling the context sample screening model to match a plurality of reference contexts from context candidate data based on the initial content generation description and the reference content, to generate first context information composed of the plurality of reference contexts, and to output the first context information.
[0189] Optionally, the description determining module 12 is configured to:
[0190] inputting the first context information and the initial content generation description into a user description adjustment model, the user description adjustment model including a context ordering network and a large language model network;
[0191] reordering the reference contexts in the first context information through the context ordering network to obtain a second context queue, and selecting a target number of recommended contexts from the second context queue to obtain recommended context information;
[0192] performing description recommendation fine-tuning on the initial content generation description based on the recommended context information through the large language model network to obtain a recommended content generation description.
[0193] Optionally, the apparatus 1 is further configured to:
[0194] performing user description style matching processing on the recommended content generation description based on the initial content generation description to obtain a style matching detection result, performing style strengthening description processing on the recommended content generation description based on the style matching detection result to obtain a recommended content generation description after style strengthening description processing; and / or,
[0195] performing content compliance detection processing on the recommended content generation description to obtain a content compliance detection result, performing content strengthening description processing on the recommended content generation description based on the content compliance detection result to obtain a recommended content generation description after content strengthening description processing.
[0196] Optionally, the apparatus 1 is further configured to:
[0197] adopting the description detection model to perform user description style matching processing on the recommended content generation description, to obtain a style matching detection result; and / or
[0198] adopting the description detection model to perform content compliance detection processing on the recommended content generation description, to obtain a content compliance detection result.
[0199] It should be noted that the content generation apparatus provided in the above embodiments is only used as an example to illustrate the division of the above functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the content generation apparatus and the content generation method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be described here.
[0200] The serial number of the above specification is only for description, not representing the pros and cons of the embodiments.
[0201] Please refer to Figure 7 which shows a structural schematic diagram of the context sample screening model training apparatus of the present specification. The reference feature extraction model training apparatus 2 can be realized by software, hardware or a combination of both to become all or part of an electronic device. According to some embodiments, the context sample screening model training apparatus 2 includes a model creation module 21, a model training module 22, specifically for:
[0202] The model creation module 21 is configured to create an initial context sample screening model and determine context candidate data for the initial context sample screening model.
[0203] The model training module 22 is configured to obtain an initial content generation description sample and a reference content sample, and perform at least one round of model training on the initial context sample screening model based on the initial content generation description sample, the reference content sample and the context candidate data, to obtain a context sample screening model after model training.
[0204] Optionally, the model training module 22 is configured to:
[0205] input the initial content generation description sample and the reference content sample into the initial context sample screening model for at least one round of model training, and in the model training process, control the initial context sample screening model to determine a sample user description matching result corresponding to the context candidate data based on the initial content generation description sample and the reference content sample;
[0206] obtaining a description matching classification label annotated by the initial content generation description sample on the context candidate data, and adjusting model parameters of the initial context sample screening model based on the sample user description matching result and the description matching classification label until the initial context sample screening model is trained to obtain a context sample screening model.
[0207] Optionally, the initial context sample screening model comprises a description feature encoding module, a reference content feature encoding module, a context feature encoding module and a fusion matching module.
[0208] The model training module 22 is configured to:
[0209] The description feature encoding module is configured to determine content generation description features corresponding to the initial content generation description sample, the reference content feature encoding module is configured to determine reference content features corresponding to the reference content sample, and the context feature encoding module is configured to determine context features corresponding to the context candidate data.
[0210] The fusion matching module is configured to perform sample user description matching classification based on the content generation description features, the reference content features and the context features to obtain a sample user description matching result corresponding to the context candidate data.
[0211] If the sample user description matching result is a user description matching type, a first context sample corresponding to the user description matching type is output.
[0212] Optionally, the model training module 22 is configured to:
[0213] The model training module 22 is configured to calculate a model classification loss based on the sample user description matching result and the description matching classification label, and adjust model parameters of the initial context sample screening model based on the model classification loss.
[0214] It should be noted that the context sample screening model training device provided in the above embodiments is used to execute the context sample screening model training method, and only the division of the above functional modules is used as an example for illustration. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the context sample screening model training device and the context sample screening model training method provided in the above embodiments belong to the same concept, and the implementation process is described in detail in the method embodiments. Therefore, it is not repeated here.
[0215] Please refer to Figure 8FIG. 1 is a structural schematic diagram of a target generation content model training apparatus of the present specification. The user description adjustment model training apparatus 3 can be implemented by software, hardware, or a combination of both as all or part of an electronic device. According to some embodiments, the user description adjustment model training apparatus 3 includes a model creation module 31, a data acquisition module 32, and a model training module 33, and is specifically configured to:
[0216] The model creation module 31 is configured to create an initial user description adjustment model.
[0217] The data acquisition module 32 is configured to acquire initial content generation description samples and first context sample information, the first context sample information consisting of at least one first context sample output by a context sample screening model for the initial content generation description samples.
[0218] The data acquisition module 32 is further configured to label the first context sample information with a context sample order label based on the initial content generation description samples.
[0219] The model training module 33 is configured to perform at least one round of model training on the initial user description adjustment model based on the first context sample information, the initial content generation description samples, and the context sample order label, to obtain a model-trained user description adjustment model.
[0220] Optionally, the model training module 33 is configured to:
[0221] input the first context sample information and the initial content generation description samples into the initial user description adjustment model;
[0222] in each round of model training, reorder the first context samples in the first context sample information through the user description adjustment model to obtain a second context sample queue, determine recommended context sample information from the second context sample queue, and perform description recommendation fine-tuning on the initial content generation description samples based on the recommended context sample information to obtain recommended content generation description samples;
[0223] determine a model comprehensive loss based on the second context sample queue, the context sample order label, the recommended content generation description samples, and the initial content generation description samples, adjust model parameters of the initial user description adjustment model using the model comprehensive loss, and obtain a model-trained user description adjustment model.
[0224] Optionally, the initial user description adjustment model includes a context ordering network and a large language model network, and the model training module 33 is configured to:
[0225] The first context sample in the first context sample information is reordered by the context reordering network to obtain a second context sample queue, and a target number of recommended context samples are selected from the second context sample queue to obtain recommended context sample information.
[0226] The large language model network performs description recommendation fine-tuning on the initial content generation description sample based on the recommended context sample information to obtain a recommended content generation description sample.
[0227] Optionally, the model training module 33 is configured to:
[0228] The context reordering loss is determined based on the context sample reordering information and the context sample reordering label corresponding to the second context sample queue, the recommended description prediction loss is determined based on the recommended content generation description label corresponding to the recommended content generation description sample and the initial content generation description sample, and the model comprehensive loss is determined based on the context reordering loss and the recommended description prediction loss.
[0229] It should be noted that the user description adjustment model training apparatus provided in the above embodiments is only used as an example to illustrate the division of the above functional modules in the execution of the user description adjustment model training method. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the user description adjustment model training apparatus and the user description adjustment model training method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments. Therefore, it is not repeated here.
[0230] The present specification also provides a computer storage medium, which can store a plurality of instructions, the instructions being suitable for being loaded and executed by a processor to perform the content generation method of the above-described Figures 1-5 embodiments. The specific implementation process can be referred to the specific description of the above-described Figures 1-5 embodiments, which is not repeated here.
[0231] The present specification also provides a computer program product, which stores at least one instruction, the at least one instruction being loaded and executed by the processor to perform the content generation method of the above-described Figures 1-5 embodiments. The specific implementation process can be referred to the specific description of the above-described Figures 1-5 embodiments, which is not repeated here.
[0232] Please refer to Figure 9A structural block diagram of an electronic device is provided for an embodiment of the present specification. The electronic device in the present specification can include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 can be connected through the bus 150.
[0233] The processor 110 can include one or more processing cores. The processor 110 connects various parts within the terminal through various interfaces and lines, performs various functions of the terminal 100 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Alternatively, the processor 110 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 110 can integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes an operating system, a user interface, and an application program; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 110, but can be implemented by a separate communication chip.
[0234] The memory 120 can include a random access memory (RAM) and can also include a read-only memory (ROM). Alternatively, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, codes, code sets or instruction sets.
[0235] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In an embodiment of the present specification, the input device 130 can be a temperature sensor for obtaining the operating temperature of the terminal. The output device 140 can be a speaker for outputting an audio signal.
[0236] In addition, those skilled in the art will understand that the structure of the terminal shown in the above figures does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the terminal may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WIFI) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.
[0237] In the embodiments of this specification, the executing entity for each step can be the terminal described above. Optionally, the executing entity for each step is the terminal's operating system. The operating system can be Android, iOS, or other operating systems; this specification does not limit this.
[0238] exist Figure 9 In the electronic device, the processor 110 can be used to call a program stored in the memory 120 and execute it to implement at least one of the content generation method, the context sample filtering model training method, and the user description adjustment model training method as described in the various method embodiments of this specification.
[0239] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0240] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the initial content generation description, reference content, and user information involved in this specification were all obtained with full authorization.
[0241] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.
Claims
1. A content generation method, the method comprising: Get the initial content input by the user and generate descriptions and reference content; Recommendation context information is determined based on the initial content generation description and the reference content; and a recommendation content generation description is determined based on the recommendation context information and the initial content generation description. Based on the recommended content description and the reference image, content is generated using a target content generation model; The step of determining recommendation context information based on the initial content generation description and the reference content, and determining recommendation content generation description based on the recommendation context information and the initial content generation description, includes: The initial content is used to generate a description, and the reference content is input into the context sample filtering model; The context sample filtering model is controlled to generate a description based on the initial content and the reference content, match multiple reference contexts from the context candidate data, generate a first context information composed of multiple reference contexts, and output the first context information; The reference contexts in the first context information are reordered to obtain a second context queue. Recommended context information is determined from the second context queue, and the initial content generation description is adjusted based on the recommended context information to obtain a recommended content generation description.
2. The method according to claim 1, wherein the step of reordering the reference contexts in the first context information to obtain a second context queue, determining the recommendation context information from the second context queue, and adjusting the initial content generation description based on the recommendation context information to obtain the recommendation content generation description includes: The first context information and the initial content are used to generate a description which is then input into the user description adjustment model, which includes a context sorting network and a large language model network. The reference contexts in the first context information are reordered through the context sorting network to obtain a second context queue, and a target number of recommended contexts are selected from the second context queue to obtain recommended context information. The large language model network fine-tunes the initial content generation description based on the recommendation context information to obtain the recommended content generation description.
3. The method according to claim 1, further comprising, after determining the recommended content generation description based on the recommended context information and the initial content generation description: Based on the initial content generation description, the recommended content generation description is subjected to user description style matching processing to obtain style matching detection results. Based on the style matching detection results, the recommended content generation description is subjected to style enhancement description processing to obtain the style enhancement description of the recommended content generation description. And / or, The generated description of the recommended content is subjected to content compliance detection processing to obtain the content compliance detection result. Based on the content compliance detection result, the generated description of the recommended content is subjected to content enhancement processing to obtain the generated description of the recommended content after content enhancement processing.
4. The method according to claim 1, further comprising: A description detection model is used to perform user description style matching on the descriptions generated for the recommended content to obtain style matching detection results. And / or, A description detection model is used to perform content compliance detection on the generated descriptions of the recommended content, and the content compliance detection results are obtained.
5. A method for training a contextual sample selection model, the method comprising: Create an initial context sample filtering model and determine the context candidate data for the initial context sample filtering model; Obtain initial content to generate description samples and reference content samples. Based on the initial content to generate description samples, the reference content samples and the context candidate data, train the initial context sample selection model for at least one round to obtain the trained context sample selection model.
6. The method according to claim 5, wherein the step of training the initial context sample selection model for at least one round based on the initial content-generated description sample, the reference content sample, and the context candidate data to obtain the trained context sample selection model includes: The initial content generated description sample and the reference content sample are input into the initial context sample filtering model for at least one round of model training. During the model training process, the initial context sample filtering model is controlled based on the initial content generated description sample and the reference content sample to determine the sample user description matching result corresponding to the context candidate data. Obtain the description matching classification labels of the context candidate data based on the description samples generated from the initial content, and adjust the model parameters of the initial context sample screening model based on the sample user description matching results and the description matching classification labels until the initial context sample screening model completes training, thus obtaining the context sample screening model.
7. The method according to claim 6, wherein the initial context sample screening model comprises a descriptive feature encoding module, a reference content feature encoding module, a context feature encoding module, and a fusion matching module. The step of generating a description sample based on the initial content and the reference content sample to control the initial context sample filtering model to determine the sample user description matching result corresponding to the context candidate data includes: The description feature encoding module determines the content generation description feature corresponding to the initial content generation description sample, the reference content feature encoding module determines the reference content feature corresponding to the reference content sample, and the context feature encoding module determines the context feature corresponding to the context candidate data. The fusion matching module performs description matching and classification for sample users based on the content-generated description features, the reference content features, and the context features, and obtains the sample user description matching results corresponding to the context candidate data. If the sample user description matching result is a user description matching type, then the first context sample corresponding to the user description matching type is output.
8. The method according to claim 6, wherein adjusting the model parameters of the initial context sample screening model based on the sample user description matching result and the description matching classification label includes: Based on the sample user description matching results and the description matching classification labels, the model classification loss is calculated, and the model parameters of the initial context sample screening model are adjusted using the model classification loss.
9. A method for training a user description-adjusted model, the method comprising: Create an initial user description and adjust the model; Obtain initial content to generate description samples and first context sample information, wherein the first context sample information consists of at least one first context sample output by the context sample filtering model for generating description samples for the initial content; Based on the initial content, generate descriptive samples and label the first context sample information with context sample sorting tags; Based on the first context sample information, the initial content generated description sample, and the context sample sorting label, the initial user description adjustment model is trained for at least one round to obtain the trained user description adjustment model.
10. The method according to claim 9, wherein the step of training the initial user description adjustment model based on the first context sample information, the initial content generated description sample, and the context sample sorting label for at least one round to obtain the trained user description adjustment model includes: The first context sample information and the initial content are used to generate a description sample, which is then input into the initial user description adjustment model. In each round of model training, the model is adjusted by user description to reorder the first context samples in the first context sample information to obtain the second context sample queue, and the recommended context sample information is determined from the second context sample queue. Based on the recommended context sample information, the initial content generated description sample is fine-tuned to obtain the recommended content generated description sample. Based on the second context sample queue, the context sample sorting labels, the recommended content generated description samples, and the initial content generated description samples, the model comprehensive loss is determined. The model parameters of the initial user description adjustment model are then adjusted using the model comprehensive loss to obtain the user description adjustment model after model training.
11. The method according to claim 10, wherein the initial user description adjustment model comprises a context ranking network and a large language model network. The process of reordering the first context samples in the first context sample information through a user description adjustment model to obtain a second context sample queue, determining recommendation context sample information from the second context sample queue, and fine-tuning the initial content generation description sample based on the recommendation context sample information to obtain recommendation content generation description samples includes: The first context samples in the first context sample information are reordered through the context sorting network to obtain the second context sample queue, and a target number of recommended context samples are selected from the second context sample queue to obtain recommended context sample information; The large language model network uses the recommendation context sample information to generate description samples from the initial content, and then fine-tunes the description recommendations to obtain recommended content generation description samples.
12. The method according to claim 10, wherein determining the model comprehensive loss based on the second context sample queue, the context sample sorting labels, the recommended content generated description samples, and the initial content generated description samples includes: The context ranking loss is determined based on the context sample ranking information and context sample ranking labels corresponding to the second context sample queue. The recommendation description prediction loss is determined based on the recommendation content generated description samples and the recommendation content generated description labels corresponding to the initial content generated description samples. The model comprehensive loss is determined based on the context ranking loss and the recommendation description prediction loss.
13. A content generation apparatus, the apparatus comprising: The content acquisition module is used to acquire the initial content input by the user and generate descriptions and reference content. The description determination module is used to determine recommendation context information based on the initial content and the reference content, and to determine the recommended content and generate a description based on the recommendation context information and the initial content. The content generation module is used to generate content based on the recommended content generation description and the reference image using a target content generation model; The step of determining recommendation context information based on the initial content generation description and the reference content, and determining recommendation content generation description based on the recommendation context information and the initial content generation description, includes: The initial content is used to generate a description, and the reference content is input into the context sample filtering model; The context sample filtering model is controlled to generate a description based on the initial content and the reference content, match multiple reference contexts from the context candidate data, generate a first context information composed of multiple reference contexts, and output the first context information; The reference contexts in the first context information are reordered to obtain a second context queue. Recommended context information is determined from the second context queue, and the initial content generation description is adjusted based on the recommended context information to obtain a recommended content generation description.
14. A context sample selection model training device, the device comprising: The model creation module is used to create an initial context sample filtering model and determine the context candidate data for the initial context sample filtering model. The model training module is used to obtain initial content generation description samples and reference content samples, and to train the initial context sample screening model at least once based on the initial content generation description samples, the reference content samples and the context candidate data, so as to obtain the trained context sample screening model.
15. A user description adjustment model training apparatus, the apparatus comprising: The model creation module is used to create an initial user description and adjust the model. The data acquisition module is used to acquire initial content generation description samples and first context sample information. The first context sample information consists of at least one first context sample output by the context sample filtering model for the initial content generation description sample. The data acquisition module is also used to generate descriptive samples based on the initial content and to label the first context sample information with context sample sorting tags; The model training module is used to perform at least one round of model training on the initial user description adjustment model based on the first context sample information, the initial content generated description sample, and the context sample sorting label, so as to obtain the user description adjustment model after model training.
16. A computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the steps of the method as claimed in any one of claims 1-4, 5-8, or 9-12.
17. A computer program product storing at least one instruction, said at least one instruction being loaded by a processor and executing the steps of the method as claimed in any one of claims 1-4, 5-8 or 9-12.
18. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method as described in any one of claims 1-4, 5-8, or 9-12.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN116580408A