Material generation method and device for recommendation information, electronic equipment and storage medium

By obtaining the picture style of audience tags and historical recommendation information, and using large language models and literary and artistic graphics models to generate picture materials that meet the layout information, solving the problems of insufficient material diversity and high operation costs of the recommender, and achieving efficient and diversified material generation.

CN120259453APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410011053.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The material diversity of recommended information in the prior art is insufficient, resulting in low click-through rate, and the recommender needs to upload image materials to increase the cost of delivery.

Method used

By obtaining the picture style of audience tags and historical recommendation information, the pre-trained first language model generates description text, and combines the literary and graphic model to automatically create picture materials that match the layout information.

Benefits of technology

It improves the efficiency and diversity of image materials for recommendation information, reduces the operation process of the recommender, and the generated materials are easier to be followed and clicked by the audience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259453A_ABST
    Figure CN120259453A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a recommendation information material generation method and device, electronic equipment and a computer readable storage medium, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring an interest label, a picture style of historical recommendation information and typesetting information of a recommendation information carrier; inputting the audience labels and the picture styles into a first large language model to obtain a description text which is output by the first large language model and is used for describing picture materials; and inputting the description text and the typesetting information into the text graph model to obtain a picture material which is output by the text graph model and conforms to the typesetting information. According to the embodiment of the invention, the pushing information carriers of respective typesetting can be automatically adapted, the generation efficiency, the adaptation degree and the diversity of the picture materials of the recommendation information can be remarkably improved, and more importantly, the operation process of a recommendation party is greatly reduced because the recommendation party does not need to upload any initial picture material.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating materials for recommendation information. Background Art

[0002] When placing recommendation information, fixed materials created by the recommender are usually required (generally only one picture material). This makes the picture materials of the exposed recommendation information the same for each user. In this regard, the optimization method of related technologies is to recommend uploading multiple picture materials and splicing the materials according to interests during exposure.

[0003] Problems existing in related technologies:

[0004] 1. Insufficient diversity of materials: Due to the limited number of picture materials created by the recommender, it is difficult for the number of picture materials to cover thousands of interest tags of the population. Therefore, the styles of the finally exposed materials are still relatively single, resulting in a low click-through rate.

[0005] 2. Low efficiency: The recommender needs to upload picture materials and consider different resolution requirements for different recommendation information carriers, which increases the placement cost for the recommender. Summary of the Invention

[0006] Embodiments of this application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating materials for recommendation information, which can solve the above problems of the prior art. The technical solutions are as follows:

[0007] According to one aspect of the embodiments of this application, a method for generating materials for recommendation information is provided. The method includes:

[0008] Obtain audience tags, the picture style of historical recommendation information, and the layout information of the recommendation information carrier, where the audience tags include at least one interest tag;

[0009] Input the audience tags and the picture style into a first large language model to obtain a description text output by the first large language model for describing the picture material;

[0010] Input the description text and the layout information into a text-to-image model to obtain a picture material that conforms to the layout information output by the text-to-image model.

[0011] As an optional embodiment, the first large language model is trained in the following manner:

[0012] Construct a first training set, where the first training set includes multiple sample description texts, and the sample description texts are obtained by screening based on predetermined evaluation metrics. The evaluation metrics include metric values in multiple dimensions, and the metric values of each dimension of the sample description texts are all greater than a preset threshold;

[0013] Perform supervised fine-tuning training on a pre-trained initial large language model according to the first training set to obtain a first model;

[0014] Obtain a second training set according to the audience preference ranking between the same audience labels and different picture styles, and a preset set of description text templates;

[0015] Perform supervised fine-tuning training on the first model according to the second training set to obtain a reward model;

[0016] Train the first model under reinforcement learning based on human feedback according to the reward model to obtain the first large language model.

[0017] As an alternative embodiment, the audience label further includes at least one aversion label;

[0018] The step of inputting the audience label and the picture style into the first large language model to obtain a description text for describing the picture material output by the first large language model includes:

[0019] Input the at least one interest label and the picture style into the first large language model to obtain a first description text output by the first large language model, and input the at least one aversion label into the first large language model to obtain a second description text output by the first large language model;

[0020] The step of inputting the description text and the layout information into a text-to-image model to obtain a picture material that conforms to the layout information output by the text-to-image model includes:

[0021] Use the first description text as a positive prompt, use the second description text as a negative prompt, and input the positive prompt, the negative prompt, and the layout information into the text-to-image model to obtain the picture material output by the text-to-image model;

[0022] Wherein, the picture material contains the information described by the first description text and does not contain the information described by the second description text.

[0023] As an alternative embodiment, before inputting the description text and the layout information into the text-to-image model, it further includes:

[0024] Obtain a template picture, and the template picture includes a target area to be filled;

[0025] Inputting the description text and typesetting information into a text-to-image model to obtain picture materials that conform to the typesetting information output by the text-to-image model, including:

[0026] Inputting the description text, typesetting information, and template picture into a text-to-image model to obtain picture materials with a target picture filled in the target area of the template picture;

[0027] Wherein, the target picture is a picture drawn according to the description text.

[0028] As an alternative embodiment, inputting the description text and typesetting information into a text-to-image model to obtain picture materials that conform to the typesetting information output by the text-to-image model, including:

[0029] Based on the description text and typesetting information, by invoking the text-to-image model to perform the following operations to obtain the picture materials:

[0030] Extracting the text features of the description text and randomly generating an initial noise picture;

[0031] Taking the initial noise picture as the noise picture to be processed for the first image denoising operation, and performing the image denoising operation a preset number of times;

[0032] Adjusting the denoised picture obtained from the last image denoising operation according to the typesetting information to obtain the picture materials;

[0033] Wherein, the image denoising operation includes:

[0034] According to the text features and the noise picture to be processed, calling the denoising network of the stable diffusion model to denoise the noise picture to be processed to obtain a denoised picture, and taking this denoised picture as the noise picture to be processed for the next image denoising operation.

[0035] As an alternative embodiment, inputting the description text, typesetting information, and template picture into a text-to-image model to obtain picture materials with a target picture filled in the target area of the template picture, including:

[0036] Based on the description text and typesetting information, by invoking the text-to-image model to perform the following operations to obtain the picture materials:

[0037] Extracting the text features of the description text and randomly generating an initial noise picture;

[0038] Filling the initial noise picture into the target area of the template picture and taking it as the noise picture to be processed for the first image denoising operation, and performing the image denoising operation a preset number of times;

[0039] Adjust the denoised image obtained from the last image denoising operation according to the typesetting information to obtain the picture material;

[0040] Among them, the image denoising operation includes:

[0041] According to the text features and the to-be-processed noisy image, call the denoising network of the Stable Diffusion model to perform denoising processing on the target area in the to-be-processed noisy image to obtain a denoised image, and use this denoised image as the to-be-processed noisy image for the next image denoising operation.

[0042] As an alternative embodiment, the method further includes:

[0043] Call a pre-trained second large language model to extract keywords of at least one target part of speech that match the interest tag from a pre-established keyword library, and perform text supplementation according to the keywords to obtain copywriting material;

[0044] Generate recommendation information based on the copywriting material and the picture material;

[0045] Among them, the second large language model is trained with the interest tags of multiple sample audiences as training samples and the keywords of the copywriting material in the respective historical recommendation information of the sample audiences as training labels.

[0046] As an alternative embodiment, after obtaining the picture material that conforms to the typesetting information, it further includes:

[0047] Input the picture material into an image recognition model to obtain the target object included in the picture material output by the image recognition model;

[0048] If it is determined that the target object does not belong to the objects to be avoided related to the recommender, retain the picture material;

[0049] If it is determined that the target object belongs to the objects to be avoided, generate a new negative prompt word according to the objects to be avoided;

[0050] Input the positive prompt word, negative prompt word, new negative prompt word, and typesetting information into a text-to-image model to obtain a new picture material that conforms to the typesetting information output by the text-to-image model.

[0051] As an alternative embodiment, the keywords in the historical recommendation information of the sample audiences are determined by the following method:

[0052] Perform word segmentation on the copywriting material in the historical recommendation information;

[0053] Determine the term frequency-inverse document frequency (TF-IDF) and TextRank of each segmented word in each copywriting material, and the text rank;

[0054] Weight the TF-IDF and TextRank of each segmented word to obtain the importance of each segmented word;

[0055] Select a preset number of segmented words as keywords in descending order of importance.

[0056] According to another aspect of the embodiments of the present application, there is provided a device for generating materials for recommended information, the device includes:

[0057] A basic information acquisition module, configured to acquire an audience label, a picture style of historical recommended information, and typesetting information of a recommended information carrier, where the audience label includes at least one interest label;

[0058] A text supplement module, configured to input the audience label and the picture style into a first large language model, and obtain a description text output by the first large language model for describing picture materials;

[0059] A picture creation module, configured to input the description text and the typesetting information into a text-to-image model, and obtain picture materials that conform to the typesetting information output by the text-to-image model.

[0060] According to another aspect of the embodiments of the present application, there is provided an electronic device, the electronic device includes a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above-mentioned method for generating materials for recommended information.

[0061] According to still another aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned method for generating materials for recommended information are implemented.

[0062] According to one aspect of the embodiments of the present application, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for generating materials for recommended information are implemented.

[0063] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:

[0064] On the one hand, in order to efficiently generate picture materials, by calling a pre-trained text-to-image model, picture materials can be obtained through picture creation according to the prompt words. On the other hand, the prompt words of the present application include descriptive texts, which are generated by calling a first large language model. When generating the descriptive texts in the embodiments of the present application, not only the audience tags, especially the interest tags, but also the picture styles of the historical recommendation information corresponding to the audience are considered. The descriptive texts generated in this way can fully describe the content that conforms to the audience's interests and fits the user's perception. Furthermore, the picture materials generated using such descriptive texts are also likely to be noticed and clicked by the audience. On the other hand, in order to adapt to recommendation information carriers of different sizes and resolutions, when calling the text-to-image model, the layout information of the recommendation information carrier is also input, so that the present application can automatically adapt to the recommendation information carriers with their respective layouts, which can significantly improve the generation efficiency, adaptability and diversity of the picture materials of the recommendation information. More importantly, since the present application does not require the recommender to upload any initial picture materials, the operation process of the recommender is also greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description in the embodiments of the present application.

[0066] Figure 1 Structural schematic diagram of the information recommendation system provided by the embodiment of the present application;

[0067] Figure 2 Flow schematic diagram of a method for generating materials of recommendation information provided by the embodiment of the present application;

[0068] Figure 3 Interface schematic diagram of a recommender editing picture materials on a recommendation information platform provided by the embodiment of the present application;

[0069] Figure 4 Flow schematic diagram of a method for generating materials of recommendation information provided by the embodiment of the present application;

[0070] Figure 5 Schematic diagram of a template picture provided by the embodiment of the present application;

[0071] Figure 6 Interface schematic diagram of editing picture materials provided by another embodiment of the present application;

[0072] Figure 7 Schematic diagram of a way for a recommendation host to create recommendation information provided by the embodiment of the present application;

[0073] Figure 8 Processing flow schematic diagram of a text-to-image model provided by the embodiment of the present application;

[0074] Figure 9 Schematic diagram of the processing flow of the text-to-image model provided in another embodiment of the present application

[0075] Figure 10 Schematic diagram of the process for generating materials for recommended information provided in an embodiment of the present application;

[0076] Figure 11 Schematic diagram of the process for generating picture materials provided in an embodiment of the present application;

[0077] Figure 12 Device for generating materials for recommended information provided in an embodiment of the present application;

[0078] Figure 13 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0079] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0080] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an" and "the" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology, etc. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0081] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the drawings.

[0082] First, several terms related to the present application will be introduced and explained:

[0083] Recommendation information is information that is pushed to a terminal device for display in the form of pictures, text, videos, or any combination thereof. For example, the recommendation information can be any content such as advertisements, pushed news, etc., and the embodiments of the present application do not make specific limitations in this regard. In the following embodiments of the present application, advertisements will be used as an example of the recommendation information for description.

[0084] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0085] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0086] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0087] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0088] A large language model (LLM) refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. Large language models can handle various natural language tasks, such as text classification, question answering, dialogue, etc., and are an important approach to artificial intelligence.

[0089] Text-to-image generation is used to process the descriptive text into a digital expression that can be understood by a neural network model and added to the process of image denoising to generate an image. A text-to-image model is used to implement the text-to-image generation.

[0090] The method, device, electronic device, computer-readable storage medium, and computer program product for generating the material of recommended information provided in this application are intended to solve the above technical problems in the prior art.

[0091] The technical solutions of the embodiments of this application and the technical effects produced by the technical solutions of this application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0092] See Figure 1 , Figure 1 is a schematic structural diagram of the information recommendation system provided by the embodiments of this application. To support a news application, the method for generating the material of recommended information provided by the embodiments of this application can be realized through the cooperation between the server and the terminal. The terminal 100 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two. The terminal 100 sends the news browsing interface to the server 200. The server 200 generates a preview interface image of the news to be displayed in the news browsing interface, performs blank area recognition processing on the preview interface image to obtain multiple blank areas in the information browsing interface, obtains the interest tags of the user of the application program, the picture style of the user's historical recommended information, and the layout information of the blank area, calls the pre-trained first large language model, uses the interest tags and picture style as keywords for text supplementation to generate a descriptive text for describing the picture material, calls the pre-trained text-to-image model, performs picture creation according to the descriptive text and layout information to obtain a picture material that meets the audience's interests and the layout information, the server 200 sends the picture material to the terminal, and displays the picture material and the recommended content of the recommender in the blank area of the information browsing interface of the terminal.

[0093] In some embodiments, the method for generating the material of the recommended information provided by the embodiments of the present application can be implemented independently based on the terminal. The terminal performs blank area recognition processing on a news browsing interface including at least one piece of information to obtain multiple blank areas in the news browsing interface; by obtaining the interest tags of the user of the application, the picture style of the user's historical recommended information, and the layout information of the blank area, a pre-trained first large language model is called, and the interest tags and picture style are used as keywords for text supplementation to generate a description text describing the picture material. A pre-trained text-to-image model is called, and image creation is performed according to the description text and the layout information to obtain news picture materials that meet the audience's interests and the layout information, and in at least one blank area of the information browsing interface, the news picture materials corresponding to the blank area are displayed. It should be understood that the embodiments of the present application can not only be used to generate news picture materials, but also be used for generating picture materials such as advertisements, short video covers, self-media profiles (such as moments covers, blog covers), and self-media copywriting (such as moments, Weibo).

[0094] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 100 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and no limitation is made in the embodiments of the present application.

[0095] In some embodiments, the terminal or the server can implement the method for generating the material of the recommended information provided by the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application program (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a news APP or an e-commerce APP; it can also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module or plug-in.

[0096] In each embodiment of this application, when collecting and processing relevant data in practical applications, the informed consent or separate consent of the personal information subject should be obtained in strict accordance with the requirements of relevant national laws and regulations, and subsequent data use and processing behaviors should be carried out within the scope authorized by laws and regulations and the personal information subject.

[0097] An embodiment of this application provides a method for generating materials for recommendation information. As Figure 2 shown, this method includes:

[0098] S101. Obtain the audience tags, the picture style of historical recommendation information, and the layout information of the recommendation information carrier.

[0099] In an embodiment of this application, all audiences meeting the recommendation conditions can be determined according to the population targeting of the recommender. For example, if the recommender wants to recommend a pair of sneakers, all users in Beijing, aged 20 - 40, male, who meet the above - mentioned recommendation conditions within the past 7 days can be retrieved to form an audience set.

[0100] In some embodiments, when the number of audiences in the audience set is too large, each audience can be sorted according to a preset sorting rule, and a certain number of audiences with lower rankings can be deleted from the audience set. The embodiment of this application does not make specific limitations on the sorting rule. For example, it can be sorted in descending order of the usage duration of the application, or sorted in descending order of the usage times of the application, and so on.

[0101] For each audience in the audience set of the embodiment of this application, the audience tags of this audience are obtained from a pre - constructed tag library. The audience tags include at least one interest tag. The embodiment of this application does not make specific limitations on the construction method of the tag library. For example, when the recommendation information platform recommends recommendation information to the audience, it can judge whether the audience is interested in the recommendation information according to whether the audience clicks on the recommendation information after the recommendation. If the audience clicks on the recommendation information, it is determined that the audience is interested in the recommendation information. If the audience does not click on the recommendation information, it is determined that the audience is not interested in the recommendation information. By counting the recommendation information that the audience is interested in, the interest tags of the audience can be obtained. Similarly, by counting the recommendation information that the audience is not interested in, the dislike tags of the audience can be obtained.

[0102] It should be understood that the historical recommendation information in the embodiment of this application all refers to the historical recommendation information that the audience is interested in (i.e., there is a conversion behavior). By obtaining the picture style of the historical recommendation information that the audience is interested in, it lays a foundation for personalized generation of picture materials for recommendation information.

[0103] The picture style in the embodiment of this application refers to the painting style, color tone, theme, clarity, saturation, etc. of the picture of the recommendation information.

[0104] The recommended information carrier, which is the blank area on the web page for displaying recommended information, and the layout information, which can refer to the size and resolution of the blank area. Since there are differences in the size and resolution of the blank area displayed on different terminals and different web pages, when generating the picture materials of the recommended information in this application, it is also necessary to obtain the layout information of the recommended information carrier so that the finally generated picture materials can meet the requirements of the layout information.

[0105] S102. Invoke the pre-trained first large language model to supplement the text with the interest tags and picture styles as keywords, and generate the descriptive text for describing the picture materials.

[0106] The prompt words of the text-to-image model play a very important role in the finally generated pictures. Therefore, in the embodiments of this application, the large language model is first used to process and optimize the keywords, and then the text-to-image is used for picture creation. The embodiments of this application utilize the ability of the large language model to generate natural language texts, and generate the prompt words by invoking the pre-trained first large language model.

[0107] Through research, it is found that if only the interest tags of the audience are used as the prompt words of the text-to-image model, the effect of the finally generated picture materials is not good. Therefore, in this application, the text is supplemented with the interest tags and picture styles as keywords to generate the descriptive text for describing the picture materials, and then the picture materials generated with the descriptive text as the prompt words are more likely to attract the attention and love of the audience in terms of both style and content.

[0108] Taking the interest tags [cat, doll] and the picture style [watercolor Chinese style, tone: warm, color: pink, theme: spring] as an example, the descriptive text generated by the first large language model can be "a picture depicting an orange cat playing with Crayon Shin-chan under the cherry blossom tree in spring in watercolor painting style, with the tone mainly warm and light pink". Obviously, the descriptive text obtained after text supplementation by the large language model has more information and is more specific, which helps the text-to-image model generate picture materials with rich content and controllable content, and is especially suitable for the recommendation scenario.

[0109] S103. Input the descriptive text and the layout information into the text-to-image model, and obtain the picture materials that meet the layout information output by the text-to-image model.

[0110] The embodiments of this application can invoke the pre-trained text-to-image model to create pictures based on the prompt words. The prompt words of the embodiments of this application include the descriptive text, so as to obtain the picture materials that meet the description of the descriptive text. Moreover, by inputting the layout information into the text-to-image model, the size of the picture materials can be limited, realizing the limitation of picture creation in two dimensions of the content and size of the pictures, and finally obtaining the picture materials that meet the layout information and the preferences of the audience.

[0111] In one embodiment, the text-to-image model of the embodiments of the present application is the Stable Diffusion model. An important advantage of this model is its higher computational efficiency, which can be deployed on mobile terminals, facilitating the practical implementation of this solution and helping to reduce the pressure on the server side.

[0112] It should be emphasized that the embodiments of the present application generate descriptive text by using a large language model and then generate picture materials by the text-to-image model, rather than directly generating picture materials based on audience tags and picture styles, considering the following situations:

[0113] Specifically for image generation: The text-to-image model is specifically designed for image generation. It can generate high-quality and high-resolution pictures. Although the large language model can generate pictures through specific encoding methods, its main design goal is text generation. Therefore, the effect of image generation may be inferior to that of the text-to-image model.

[0114] High generation quality: The text-to-image model can generate pictures with good visual effects and realism. When the large language model generates pictures, it may be restricted by the encoding method, and the generated picture quality may be low.

[0115] Training stability: Compared with the large language model that requires a large amount of computing resources and training time, the text-to-image model shows better stability during the training process, reducing the difficulty and complexity of model training.

[0116] Controllability: The text-to-image model can control the diversity and quality of the generated pictures by adjusting the noise level during the diffusion process, which makes the model have better controllability during the picture creation process.

[0117] Applicable to multiple tasks: The text-to-image model can be applied to various picture generation tasks, such as image restoration, image denoising, image super-resolution, etc. The application of the large language model in picture generation tasks may be relatively limited.

[0118] Generally speaking, the text-to-image model has advantages in aspects such as specifically for image generation, generation quality, training stability, controllability, and applicability to multiple tasks, making it more advantageous in picture creation tasks.

[0119] The method for generating materials of recommended information provided by the embodiments of the present application, on the one hand, in order to efficiently generate picture materials, by calling a pre-trained text-to-image model, picture materials can be obtained through picture creation according to the prompt words. On the other hand, the prompt words of the present application include descriptive texts, which are generated by calling a pre-trained first large language model. When generating the descriptive texts in the embodiments of the present application, not only the audience tags, especially the interest tags, but also the picture styles of the historical recommended information corresponding to the audience are considered. The descriptive texts generated in this way can fully describe the content that meets the audience's interests and fits the user's perception. Furthermore, the picture materials generated using such descriptive texts are also easily noticed and clicked by the audience. On the other hand, in order to adapt to recommended information carriers of different sizes and resolutions, when calling the text-to-image model, the layout information of the recommended information carrier is also input, so that the present application can automatically adapt to the recommended information carriers with their respective layouts, which can significantly improve the generation efficiency, adaptability and diversity of the picture materials of the recommended information. More importantly, since the present application does not require the recommender to upload any initial picture materials, the operation process of the recommender is also greatly reduced.

[0120] Please refer to Figure 3 , which exemplarily shows a schematic diagram of the interface for the recommender in the embodiments of the present application to edit picture materials on the recommended information platform. As shown in the figure, the interface provides two simple controls: a material upload control 301 and an automatic creation control 302. When the recommender triggers a preset operation on the material upload control 301, an input box 303 is displayed. The recommender imports the pre-created picture materials into the input box 303, and the recommended information platform will subsequently generate recommended information based on the picture materials imported at this time. If the recommender triggers a preset operation on the automatic creation control 302, a prompt message 304 is displayed, prompting the recommender that no picture materials need to be uploaded independently in the future, but picture materials that meet the audience's interests and the layout information of the recommended information carrier are automatically generated according to the material generation method provided by the embodiments of the present application. The embodiments of the present application can greatly reduce the operation process of the recommender for placing recommended information.

[0121] It should be noted that the descriptive texts generated by the large language model may not generate appropriate descriptive texts in all cases. In some cases, the descriptive texts generated by the large language model may have phenomena such as outputting biased and ambiguous descriptive texts, or not following the instructions (i.e., the input audience tags and picture styles). Such incorrect information may lead to the audience's rejection of the generated picture material solutions, or even have inestimable negative impacts. To solve the above problems, in some embodiments, the first large language model of the embodiments of the present application can be trained in the following ways:

[0122] S11. Construct a first training set, where the first training set includes multiple sample description texts. The sample description texts are obtained by screening based on predetermined evaluation indicators. The evaluation indicators include indicator values ​​of multiple dimensions. The indicator values ​​of each dimension of the sample description texts are all greater than a preset threshold.

[0123] The embodiment of the present application does not limit the length of the sample description file, which can be determined based on internal testing, demand, etc. The embodiment of the present application does not specifically limit the dimensions of the evaluation index, which can include, for example, emotional color type (such as romance, warmth, vitality, indifference, nagging, depression, etc.), degree of exaggeration, etc.

[0124] The embodiment of the present application can set preset thresholds for indicators of different dimensions. When the indicator value of a sample description text in one dimension is higher than the preset threshold of the dimension, the sample is considered to meet the requirements in the dimension. Therefore, if the indicator values ​​of a sample description text in all dimensions meet the requirements, it will be included in the first training set.

[0125] S12 performs supervised fine-tuning training on the pre-trained initial large language model according to the first training set to obtain a first model.

[0126] The initial large language model in the embodiment of the present application can be a generative supervised training model. The initial large language model is fine-tuned in a supervised manner through the first training set, so that the initial large language model can output descriptive text in a picture style that is consistent with the audience label and historical recommendation information.

[0127] S13. Obtain a second training set based on the material preference ranking between the same audience label and different picture materials, and the first model.

[0128] Specifically, a second training set is constructed based on a material preference ranking between the same audience label and different picture materials and a preset description text template set.

[0129] For each audience, we construct a sample pair of the audience tag of the audience and each image material pushed to the audience in the historical period, and sort them according to the click-through rate of each image material. We further summarize the sorting results of all audiences with the same audience tag to obtain the material preference ranking between the same audience tag and different image materials.

[0130] After determining the material preference ranking, since the first model is trained based on the generative supervised training model, and the generative supervised training model itself has the ability to provide textual descriptions of images, the embodiment of the present application can input the top-ranked image materials into the first model, and the first model will provide textual descriptions of the image materials to obtain description texts of the image materials.

[0131] In some embodiments, the audience labels can be used as training samples, and the description texts of the picture materials ranked higher can be used as training labels to obtain a second training set.

[0132] In other embodiments, the audience labels can be used as training samples, and the description texts sorted according to the sorting of the material preferences of the audience labels can be used as training labels to obtain a second training set.

[0133] S14. Perform supervised fine-tuning training on the first model according to the second training set to obtain a reward model.

[0134] Specifically, perform supervised training on the first model using the second training set to obtain a credible reward model. A credible reward model (Trust Reward Model, TRM) refers to learning from the second training set how to assign different rewards (generally returned in the form of scores) to different description texts under the same input, so that the model learns in the direction of obtaining higher rewards, thereby being able to output more credible results that better meet the actual needs of users.

[0135] S15. According to the reward model, perform training on the first model under reinforcement learning based on human feedback to obtain the first large language model.

[0136] The reinforcement learning from human preferences (RLHF) follows the following steps: Initialize a new large language model S based on the parameters of the first model; Based on the new audience labels and picture styles (i.e., the prompt t), let the model S generate description texts for each prompt, and then input the description texts into the credible reward model (TRM); The TRM will calculate a score for each description text as a scalar reward, and the high or low score indicates the quality of the response; Adopt the RLHF method to continuously update its strategy based on the total reward score obtained by the model S until convergence. The model S trained at this time is the first large language model that meets the requirements, and this model has the ability to output description texts that better meet the needs of the audience.

[0137] Based on the above embodiments, as an optional embodiment, the audience labels further include at least one aversion label. It should be understood that the aversion label is determined based on the recommended information that the audience dislikes. When the recommendation information platform delivers recommendation information to the audience, feedback controls can be provided near the recommendation information so that the audience can indicate whether they like, are indifferent to, or dislike the recommendation information by triggering a preset operation on the feedback control.

[0138] In the embodiment of the present application, the audience tags and the picture style are input into the first large language model, and the description text for describing the picture material output by the first large language model is obtained, including:

[0139] The first large language model is called, and text supplementation is performed using at least one interest tag and the picture style as keywords to obtain a first description text, and text supplementation is performed using at least one aversion tag as a keyword to obtain a second description text.

[0140] In the embodiment of the present application, the first large language model is called. On the one hand, text supplementation is performed using the interest tag and the picture style as keywords to obtain a first description text, so that the first description text describes the interest tag of the audience and the picture style of the picture material in natural language. On the other hand, text supplementation is performed according to the aversion tag as a keyword to obtain a second description text, so that the second description text describes the aversion tag of the audience in natural language.

[0141] The description text and the layout information are input into the text-to-image model, and the picture material that conforms to the layout information output by the text-to-image model is obtained, including:

[0142] The text-to-image model is called, the first description text is used as a positive prompt, the second description text is used as a negative prompt, and picture creation is performed according to the positive prompt, the negative prompt, and the layout information to obtain picture material that conforms to the layout information.

[0143] The prompts required by the text-to-image model can include negative prompts in addition to the regular positive prompts. It should be understood that the positive prompts are used to describe the information that should be included in the picture, while the negative prompts are used to describe the information that should not be included in the picture. In the present application, by using the first description text as the positive prompt and the second description text as the negative prompt, the finally generated picture contains the information described by the first description text and does not contain the information described by the second description text, which can meet the interests of the audience while reducing the risk that the recommended information contains information that the audience dislikes.

[0144] Please refer to Figure 4, which exemplarily shows a schematic flowchart of a method for generating materials for recommended information provided by an embodiment of the present application. As shown in the figure, obtain the audience tags of the audience, where the audience tags include at least one interest tag [dog, doll] and at least one aversion tag [cat]. At the same time, also obtain the picture style [warm, pink] of the historical recommended information of this audience and the layout information [600*800] of the recommended information carrier. Call a pre-trained large language model to supplement text with the at least one interest tag and the picture style as keywords to obtain a first description text: "Describe a scene where a dog stands beside a doll under a tree, with the tone mainly warm and light pink", and supplement text with the at least one aversion tag as a keyword to obtain a second description text: "Describe a scene where a cat yawns"; call the text-to-image model, use the first description text as a positive prompt, use the second description text as a negative prompt, and perform picture creation according to the positive prompt, negative prompt, and layout information [600*800] to obtain a picture material with a size of 600*800. It can be seen that the generated picture material describes a scene where a dog stands beside a doll under a tree and there is no cat. Since the picture material contains information that the audience is interested in and does not contain information that the audience hates, there is a high probability that the audience will click on this recommended information.

[0145] Based on the above embodiments, as an alternative embodiment, before calling the pre-trained text-to-image model, it further includes:

[0146] Obtain a template picture, where the template picture includes a target area to be filled.

[0147] The template picture in the embodiment of the present application can be provided by the recommender. The purpose is to enable the picture material to be generated based on the template picture. The template picture includes a target area to be filled. It can be understood that before filling, the target area is blank. Please refer to Figure 5 , which exemplarily shows a schematic diagram of a template picture provided by an embodiment of the present application. As shown in the figure, the template picture includes a house, and the blank area 501 beside the house is the target area to be filled. Through the text-to-image model of the embodiment of the present application, the target picture can be filled in the area 501, and the target picture is also the picture drawn according to the description text.

[0148] The embodiment of the present application does not specifically limit the size of the template picture and the size of the target area in the target picture. When the layout information of the recommended information carrier does not match the size of the template picture, the size of the template picture needs to be aligned with the recommended information carrier.

[0149] Input the description text and the layout information into the text-to-image model to obtain the picture material output by the text-to-image model that conforms to the layout information, including:

[0150] Input the described text, typesetting information, and template image into a text-to-image model to obtain image material with the target image filled in the target area of the template image;

[0151] Wherein, the target image is an image drawn according to the described text.

[0152] In the embodiment of the present application, through a text-to-image model, a target image is obtained based on the described text, and the target image is filled into the target area of the template image. Then, the filled target image is adjusted according to the typesetting information to obtain image material. The embodiment of the present application can meet the requirement that the recommender independently sets the basic information of the image material (i.e., the information in the template image).

[0153] In Figure 3 Based on the embodiment shown, please refer to Figure 6 , which exemplarily shows a schematic diagram of an interface for editing image material provided by another embodiment of the present application. When the recommender imports an image into the input box 601, a process of automatically identifying whether the imported image contains a target area will be triggered. In some embodiments, by identifying the pixel values of each pixel point in the image, if the pixel values of multiple adjacent pixel points are all 0 or 255, then these pixel points of the vectors are marked as the target area. If no target area is identified, the area adjustment control 602 is displayed. When the recommender triggers a preset operation on the area adjustment control 602, the imported image and an indication box 603 for indicating the size of the target area to be filled will be displayed. The recommender can achieve independent setting of the target area by adjusting the size and position of the indication box 603. When the target area is set, the target area in the originally imported image will become a blank area to prompt the recommender.

[0154] Please refer to Figure 7, which exemplarily shows a schematic diagram of the way for the recommendation master to create recommendation information provided by the embodiments of the present application. As shown in the figure, the recommendation master first customizes the copywriting of the recommendation information, and then selects the generation method of the picture material. The present application provides three methods for the recommendation master to choose independently. Method 1: The recommendation master uploads the customized picture material. Subsequently, when the recommendation information is exposed, both the copywriting and the picture material of the recommendation information are the copywriting and picture material customized by the recommendation master. Method 2: The recommendation master uploads a template picture, and the template picture includes a target area. By obtaining the audience tags of the audience, the picture style of the historical recommendation information, and the layout information of the recommendation information carrier through the embodiments of the present application, the generated picture material is a picture in which the target area of the template picture is filled with the target picture, and the target picture is a picture drawn according to the description text. Method 2 has an obvious improvement compared with Method 1. The recommendation party only needs to provide a basic picture background. Method 3 has the highest degree of freedom. The recommendation party does not need to provide any pictures, and can create picture materials with different contents for different audiences through the embodiments of the present application.

[0155] Based on the above embodiments, as an optional embodiment, a pre-trained text-to-image model is called, and image creation is performed according to the description text and the layout information to obtain picture materials that meet the layout information, including:

[0156] Based on the description text and the layout information, the following operations are performed by calling the text-to-image model to obtain the picture materials:

[0157] S201. Extract the text features of the description text and randomly generate an initial noise picture;

[0158] S202. Use the initial noise picture as the noise picture to be processed for the first image denoising operation, and perform the image denoising operation a preset number of times;

[0159] S203. Adjust the denoised picture obtained from the last image denoising operation according to the layout information to obtain picture materials;

[0160] Among them, the image denoising operation includes:

[0161] According to the text features and the noise picture to be processed, call the denoising network of the StableDiffusion model to perform denoising processing on the noise picture to be processed to obtain a denoised picture;

[0162] Based on the denoised picture, obtain the noise picture to be processed for the next image denoising operation.

[0163] Please refer to Figure 8, which exemplarily shows a schematic diagram of the processing flow of the text-to-image model provided by the embodiments of the present application. As shown in the figure, the text-to-image model includes a text encoder, a StableDiffusion model, and an image decoder. The description text "a puppy playing football" is input into the text encoder, and the text encoder extracts features from the description text to obtain the text features of the description text. An initial noise image is randomly generated. In the figure, the size of the initial noise image is 64*64. The size of 64*64 is to ensure computational efficiency and perform calculations in a low-dimensional space. The initial noise image is used as the noise image to be processed based on which the first image denoising operation is performed. The noise image to be processed and the text features are called to the StableDiffusion model (Diffusion model) for 50 image denoising operations. The denoised image after 50 image denoising operations is decoded by the image decoder and adjusted according to the layout information to obtain the picture material. It can be seen that the final picture material depicts a scene of a puppy playing football.

[0164] In the process of performing 50 image denoising operations in the embodiments of the present application, for the first image denoising operation, the initial noise image is used as the noise image to be processed based on which the first image denoising operation is performed. For the 2nd - 50th image denoising operations, the denoised image obtained in the previous time is used as the noise image to be processed.

[0165] Based on the above embodiments, as an optional embodiment, a pre-trained text-to-image model is called to perform picture creation based on the description text, layout information, and template picture, and picture material with the target picture filled in the target area of the template picture is obtained, including:

[0166] Based on the description text and layout information, the following operations are performed by calling the text-to-image model to obtain the picture material:

[0167] S301. Extract the text features of the description text and randomly generate an initial noise image;

[0168] S302. Fill the initial noise image into the target area of the template picture and use it as the noise image to be processed based on which a preset number of image denoising operations are performed;

[0169] S303. Adjust the denoised image obtained from the last image denoising operation according to the layout information to obtain the picture material;

[0170] Among them, the image denoising operation includes:

[0171] According to the text feature and the to-be-processed noisy picture, call the denoising network of the Stable Diffusion model to perform denoising processing on the target area in the to-be-processed noisy picture, obtain a denoised picture, and use this denoised picture as the to-be-processed noisy picture for the next image denoising operation.

[0172] In the embodiment of the present application, after randomly generating an initial noisy picture, fill the initial noisy picture into the target area of the template picture, and use the template picture filled with the initial noisy picture as the to-be-processed noisy picture for the first image denoising operation. It should be noted that the image denoising operation in the embodiment of the present application only performs denoising operations on the target area in the to-be-processed noisy picture, which can allow the embodiment of the present application to modify the target area without affecting other areas of the picture.

[0173] Please refer to Figure 9 , which exemplarily shows a schematic diagram of the processing flow of the text-to-image model provided by another embodiment of the present application. As shown in the figure, the text-to-image model includes a text encoder, a Stable Diffusion model, and an image decoder. Input the description text "a puppy playing football" into the text encoder. The text encoder extracts features from the description text to obtain the text feature of the description text. Randomly generate an initial noisy picture. The size of the initial noisy picture in the figure is 64*64. The size of 64*64 is to ensure computational efficiency and calculate in a low-dimensional space. Fill the initial noisy picture into the target area of the template picture as the to-be-processed noisy picture for the first image denoising operation. Call the Stable Diffusion model (Diffusion model) to perform 50 image denoising operations on the to-be-processed noisy picture and the text feature. Decode the denoised picture after 50 denoising operations through the image decoder and adjust it according to the typesetting information to obtain picture materials. It can be seen that the final picture materials depict a scene of a puppy playing football.

[0174] During the process of performing 50 image denoising operations in the embodiment of the present application, for the first image denoising operation, use the template picture filled with the initial noisy picture in the target area as the to-be-processed noisy picture for the first image denoising operation, call the denoising network of the Stable Diffusion model to perform denoising processing on the target area in the to-be-processed noisy picture, and obtain the denoised picture obtained from the first image denoising operation. For the 2nd - 50th image denoising operations, use the denoised picture obtained in the previous time as the to-be-processed noisy picture.

[0175] On the basis of the above embodiments, as an optional embodiment, in order to overcome the problem that on the one hand, the copywriting materials in the recommended information lack diversity because they are the same for all audiences in the related art, and on the other hand, creating copywriting materials requires a large amount of manual work by the recommender and has low efficiency, the method for generating materials of the recommended information in the embodiment of the present application further includes:

[0176] Invoke the pre-selected and trained second large language model to extract keywords of at least one target part of speech that match the interest tags from a pre-established keyword library, and perform text supplementation based on the keywords to obtain copywriting materials;

[0177] Generate recommendation information based on the copywriting materials and the picture materials;

[0178] The second large language model of the embodiments of the present application is trained with the interest tags of multiple sample audiences as training samples and the keywords of the copywriting materials in each historical recommendation information of the sample audiences as training labels, so as to have the ability to extract keywords related to the interests of the audiences. Further, using the text creation ability of the large language model itself for scripts, the extracted keywords are used for text supplementation to obtain copywriting materials. The embodiments of the present application can greatly reduce the labor cost of the recommender for generating copywriting materials.

[0179] In order to enable the large language model to better understand the composition of these training data, it is necessary to first segment these historical copywriting materials with higher click-through rates (for example, through jieba segmentation), and further perform keyword extraction and part-of-speech tagging, and then input them into the second large language model for learning and understanding, which can further improve the ability of the second large language model to capture keywords.

[0180] It should be noted that due to the large differences in copywriting in different industries, the embodiments of the present application can be divided into multiple models for training according to industries (such as self-media, education, finance, etc.) to obtain second large language models for different industries.

[0181] The following combines a specific example to illustrate the process of generating text materials in the embodiments of the present application.

[0182] Training data 1: Male, aged 20 - 25 years old, living in a second-tier city, hobbies are football and running, has clicked on a self-media advertisement copy: "The cold and arrogant senior brother, in order to improve his martial arts skills, broke into the dragon's pool and tiger's den by force, and actually saw the funny and cute junior brother easily break through the level", the keywords and parts of speech are as follows: cold and arrogant (adjective), senior brother (noun), break into by force (verb), funny and cute (adjective), junior brother (noun), break through the level (verb);

[0183] Training data 2: Female, aged 30 - 35 years old, living in a third-tier city, hobbies are yoga and singing, has clicked on a self-media advertisement copy: "This is a growth story about urban girls and love, warm and touching. Please follow their footsteps, feel the power of love, and together in the bustling city, look for the flower of happiness that belongs to you", the keywords and parts of speech are as follows: they (pronoun), urban girls (noun), warm (adjective), love (noun), flower of happiness (noun);

[0184] 2. The second largest language model uses the Generative Pre-Trained Transformer (GPT) model.

[0185] When exposing recommended information, the audience's interest tags are input into the trained big model. The big model selects keywords as prompts based on the interest tags and creates copy.

[0186] The keyword selection process is as follows: when the second language model is trained, a corresponding keyword library is accumulated, which is mainly a noun library, a verb library, an adjective library, etc. By performing statistical analysis on the training data, the second language model is required to randomly select 2-3 nouns preferred by the audience from the keyword library, randomly select 1-2 verbs from the verb library, and select 0-2 adjectives from the adjective library according to the audience's interest tags, and create based on these keywords. In some embodiments, the recommender can also pre-specify keyword information that must appear in the copywriting material - such as brand names, recently popular hot words, etc., to improve the click-through rate after exposure.

[0187] The steps are as follows:

[0188] ① Input the interest tags of the audience to be exposed into the second language model: user gender is female, age is 35-40 years old, permanent residence is a third-tier city, hobbies are yoga and singing;

[0189] ② The second largest language model randomly selects 2 nouns [city, love], 2 verbs [struggle, promotion], and 1 adjective [warmth] from the noun library;

[0190] ③The second language model creates copy based on selected keywords:

[0191] In this modern city full of challenges and opportunities, they are a young couple with dreams, working hard for love and career. She is an ordinary working woman, eager to make a name for herself in the workplace, and also eager to have a warm and sincere love. He is a vigorous newcomer in the workplace, dreaming of promotion and salary increase, and willing to work hard with her for the future. Their love story is like a ray of warm sunshine, illuminating each other's way forward.

[0192] See also Figure 10 , which exemplarily shows a flow chart of a method for generating materials for recommendation information provided by another embodiment of the present application, as shown in the figure, including:

[0193] Obtain interest tags of multiple sample audiences and copywriting materials of multiple (interesting) historical recommendation information;

[0194] For each sample audience, extract keywords from the copywriting materials of each historical recommendation information of the sample audience, determine the part of speech of the keywords, and save the extracted keywords and their parts of speech to the keyword library;

[0195] Use the interest tags of the sample audience as training labels, and use the corresponding keywords of the sample audience as the training label input to the second large language model for learning and understanding until the second large language model can understand the keywords that the sample audience with different interest tags is interested in;

[0196] Obtain the audience tags of the audience, the picture style of the historical recommendation information, and the layout information of the recommendation information carrier;

[0197] Input the audience tags and the picture style as keywords into the first large language model to obtain the descriptive text output by the first large language model for describing the picture materials;

[0198] Input the descriptive text and the layout information into the text-to-image model to obtain the picture materials that meet the layout information output by the text-to-image model;

[0199] Call the pre-trained second large language model, extract keywords with at least one target part of speech that match the interest tags from the keyword library, and supplement the text according to the keywords to obtain the copywriting materials;

[0200] Generate recommendation information according to the copywriting materials and the picture materials.

[0201] In the embodiments of the present application, by pre-obtaining the interest tags of the sample audience and the historical recommendation information, and training the second large language model to have the ability to understand the keywords that the sample audience with different interest tags is interested in, so that when generating recommendation information is needed, only the interest tags of the audience need to be input into the second large language model, and personalized copywriting materials can be obtained based on the keywords that the audience is interested in, without the need for the recommender to input any text information. Moreover, when obtaining the picture materials in the present application, it is also not necessary for the recommender to input any picture information. Only by obtaining the audience tags of the audience and the picture style of the historical recommendation information, expanding the text through the first large language model to obtain the descriptive text, and then using the layout information of the recommendation information carrier and the descriptive text, calling the text-to-image model, the picture materials that meet the layout information can be obtained. Finally, by combining the picture materials and the text materials, the recommendation information can be obtained. The present application significantly reduces the workload of the recommender, has almost no technical threshold and labor cost for putting into use, and can greatly improve the putting-into-use experience of the recommender.

[0202] Based on the above embodiments, as an alternative embodiment, the embodiments of the present application further provide relevant steps for text material quality detection. The copywriting materials output by the second large language model may have problems such as improper word order and missing information. Therefore, it is necessary to introduce a copywriting quality judgment logic, which is implemented based on training with the GPT model. The training content is text with quality problems output by humans, mainly divided into the following situations:

[0203] ① Information missing category: "The aloof senior brother, in order to improve his martial arts skills, actually saw the funny and cute junior brother";

[0204] ② Typo category: "The aloof senior brother, in order to improve his martial arts skills, broke into the dragon's pool and tiger's den by force, and actually saw the highly amusing and cute junior brother easily passing the level";

[0205] ③ Repetition category: "The aloof senior brother, in order to improve his martial arts skills martial arts, broke into the dragon's pool and tiger's den by force, and actually saw the funny and cute junior brother easily passing the level";

[0206] ④ Word order inversion category: "The aloof senior brother, in order to improve his martial arts skills, actually saw the funny and cute junior brother easily passing the level and breaking into the dragon's pool and tiger's den by force."

[0207] Input the sentence error data as training data into the second large language model for learning. Before the advertisement is exposed, review the copywriting materials generated by the second large language model. If it is determined that there is a text error, regenerate the copywriting materials through the second large language model until there is no more error reporting.

[0208] In actual application, there are often some objects that need to be avoided in the recommended information of some industries. For example, in the recommended information of the beauty industry, hair is generally avoided. In the recommended information of the pharmaceutical industry, diseased organs are generally avoided. More commonly, any recommender does not like the information of competitors to appear in their recommended information. Therefore, based on the above embodiments, as an alternative embodiment, after obtaining the picture material that conforms to the typesetting information, the following further includes:

[0209] S401. Input the picture material into an image recognition model to obtain the target object included in the picture material output by the image recognition model;

[0210] S402. If it is determined that the target object does not belong to the objects to be avoided related to the recommender, retain the picture material;

[0211] S402'. If it is determined that the target object belongs to the objects to be avoided, generate a new negative prompt word according to the objects to be avoided;

[0212] S403. Input the positive prompt, negative prompt, new negative prompt, and layout information into the text-to-image model to obtain new image materials that meet the layout information and are output by the text-to-image model.

[0213] In the embodiments of the present application, after the text-to-image model generates image materials, a pre-trained image recognition model is called to perform image recognition on the image materials to determine the target objects included in the image materials, and sensitive information of the target objects is detected. If it is determined that the target objects do not belong to the objects to be avoided related to the recommender, the image materials are retained; if it is determined that the target objects belong to the objects to be avoided, new negative prompts are generated according to the objects to be avoided, and then the text-to-image model is reselected and called, and image creation is performed using the positive prompt, negative prompt, new negative prompt, and layout information to obtain new image materials that meet the layout information.

[0214] Please refer to Figure 11 , which exemplarily shows a schematic flow chart of generating image materials provided by the embodiments of the present application. As shown in the figure, it includes:

[0215] S501. Obtain the audience tags of the audience, the picture style of the historical recommendation information, and the layout information of the recommendation information carrier. The audience tags include at least one interest tag and at least one aversion tag.

[0216] S502. Input the at least one interest tag and the picture style into the first large language model to obtain a first description text output by the first large language model. Input the at least one aversion tag into the first large language model to obtain a second description text output by the first large language model.

[0217] S503. Use the first description text as the positive prompt, use the second description text as the negative prompt, input the positive prompt, negative prompt, and layout information into the text-to-image model to obtain the image materials output by the text-to-image model.

[0218] S504. Input the image materials into the image recognition model to obtain the target objects included in the image materials output by the image recognition model.

[0219] S505. Determine whether the target objects belong to the objects to be avoided related to the recommender. If so, execute S506; if not, end the process.

[0220] S506. Generate new negative prompts according to the objects to be avoided.

[0221] S507. Input the positive prompt, negative prompt, new negative prompt, and layout information into the text-to-image model to obtain new image materials output by the text-to-image model, and return to execute S504.

[0222] Based on the above embodiments, as an optional embodiment, the keywords in the historical recommendation information of the sample audience are determined by the following method:

[0223] Segment the copywriting materials in the historical recommendation information;

[0224] Determine the term frequency-inverse document frequency (TF-IDF) and TextRank of each segmented word in each copywriting material;

[0225] Weight the TF-IDF and TextRank of each segmented word to obtain the importance of each segmented word;

[0226] Select a preset number of segmented words as keywords in descending order of importance.

[0227] It should be noted that in the keyword extraction scheme, the related technology provides two algorithms, namely TextRank and term frequency-inverse document frequency (TF-IDF). If only one of the keyword extraction methods is used, the final keyword result will perform poorly. Because TF-IDF is a keyword extraction method based on term frequency and inverse document frequency, it can quickly find the words that appear frequently in a specific text but rarely appear in the entire corpus. However, TF-IDF cannot capture the relationships between words and cannot handle synonyms and polysemous words. TextRank is a graph-based ranking algorithm that can capture the relationships between words and can handle synonyms and polysemous words. However, TextRank may ignore some words that appear frequently in a specific text but rarely appear in the entire corpus.

[0228] Therefore, in the embodiments of the present application, by preprocessing the copywriting materials, such as removing stop words, punctuation marks, etc., and performing word segmentation, on the one hand, TF-IDF is used to extract candidate keywords: calculate the weight of each word using TF-IDF, and select the m words with the highest weights as candidate keywords. On the other hand, TextRank is used to calculate the relationship between words: perform TextRank ranking on all candidate keywords. The TextRank algorithm is a graph-based ranking algorithm that considers the relationship between words and gives the ranking of each word. Combine the TF-IDF weight and TextRank ranking: for each candidate keyword, its final weight can be jointly determined by the TF-IDF weight and TextRank ranking. Normalize the TF-IDF weight and TextRank ranking, map the values to the range of [0, 1], and then perform weighted calculation to obtain the importance score of each candidate keyword finally.

[0229] It should be noted that the embodiments of the present application can set different weights for different industries. For example, in the self-media industry, the weight of TextRank can be made higher. Therefore, the weight of TF-IDF is specified as 0.4 and the weight of TextRank is specified as 0.6.

[0230] The embodiments of the present application provide a device for generating materials of recommended information, as Figure 12 shown. The device may include: a basic information acquisition module 121, a text supplement module 122, and an image creation module 123. Among them,

[0231] The basic information acquisition module 121 is configured to acquire audience tags, the picture style of historical recommended information, and the layout information of the recommended information carrier. The audience tags include at least one interest tag;

[0232] The text supplement module 122 is configured to input the audience tags and the picture style into the first large language model to obtain a description text output by the first large language model for describing picture materials;

[0233] The image creation module 123 is configured to input the description text and the layout information into a text-to-image model to obtain a picture material that conforms to the layout information output by the text-to-image model.

[0234] The device of the embodiments of the present application can execute the method provided by the embodiments of the present application, and its implementation principle is similar. The actions performed by each module in the device of the embodiments of the present application correspond to the steps in the method of the embodiments of the present application. For the detailed function descriptions of each module of the device, reference may specifically be made to the descriptions in the corresponding method shown above, and details are not described herein again.

[0235] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory. The processor executes the above computer program to implement the steps of a method for generating materials for recommended information. Compared with the related art, the following can be achieved: on the one hand, in order to efficiently generate picture materials, by calling a pre-trained text-to-image model, picture materials can be obtained through picture creation according to the prompt words. On the other hand, the prompt words of the present application include a description text, which is generated by calling a pre-trained first large language model. When generating the description text in the embodiment of the present application, not only the audience tags, especially the interest tags, are considered, but also the picture style of the historical recommended information corresponding to the audience is considered. The description text generated in this way can fully describe the content that conforms to the audience's interests and fits the user's perception. Furthermore, the picture materials generated using such a description text are also likely to be noticed and clicked by the audience. On the other hand, in order to adapt to recommended information carriers of different sizes and resolutions, in the embodiment of the present application, when calling the text-to-image model, the layout information of the recommended information carrier is also input, so that the present application can automatically adapt to the recommended information carriers with their respective layouts, which can significantly improve the generation efficiency, adaptability, and diversity of the picture materials of the recommended information. More importantly, since the present application does not require the recommender to upload any initial picture materials, the operation process of the recommender is also greatly reduced.

[0236] In an alternative embodiment, an electronic device is provided, as Figure 12 shown. Figure 12 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.

[0237] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0238] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 only a thick line is used to represent it herein, but it does not mean that there is only one bus or one type of bus.

[0239] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0240] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0241] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0242] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0243] Terms such as "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or textually described order.

[0244] It should be understood that although the flowcharts in the embodiments of the present application indicate each operation step by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this document, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0245] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, using other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A method for generating materials of recommended information, characterized in that, Including: Obtain the audience tags, the picture style of historical recommendation information, and the layout information of the recommendation information carrier, where the audience tags include at least one interest tag; Input the audience tags and the picture style into a first large language model to obtain a description text output by the first large language model for describing picture materials; Input the description text and the layout information into a text-to-image model to obtain picture materials that conform to the layout information output by the text-to-image model.

2. The method according to claim 1, wherein The first large language model is trained in the following way: Construct a first training set, which includes a plurality of first training samples and corresponding first training labels. The first training samples are the audience tags and the picture style of historical recommendation information, and the first training labels are sample description texts. The sample description texts are obtained by screening based on a pre-determined evaluation index. The evaluation index includes index values in multiple dimensions, and the index values of each dimension of the sample description text are all greater than a preset threshold; Perform supervised fine-tuning training on a pre-trained initial large language model according to the first training set to obtain a first model; According to the material preference ranking between the same audience tags and different picture materials, and the first model, obtain a second training set; Perform supervised fine-tuning training on the first model according to the second training set to obtain a reward model; According to the reward model, train the first model under reinforcement learning based on human feedback to obtain the first large language model.

3. The method according to claim 1, wherein The audience tags further include at least one dislike tag; The step of inputting the audience tags and the picture style into the first large language model to obtain a description text output by the first large language model for describing picture materials includes: Input the at least one interest tag and the picture style into the first large language model to obtain a first description text output by the first large language model, and input the at least one dislike tag into the first large language model to obtain a second description text output by the first large language model; The step of inputting the description text and the layout information into the text-to-image model to obtain picture materials that conform to the layout information output by the text-to-image model includes: Use the first description text as a positive prompt, use the second description text as a negative prompt, and input the positive prompt, the negative prompt, and the layout information into the text-to-image model to obtain the picture materials output by the text-to-image model; Wherein, the picture materials contain the information described by the first description text and do not contain the information described by the second description text.

4. The method according to any one of claims 1-3, characterized in that Before inputting the description text and the layout information into the text-to-image model, it further includes: Obtain a template picture, which includes a target area to be filled; The step of inputting the description text and the layout information into the text-to-image model to obtain picture materials that conform to the layout information output by the text-to-image model includes: Input the description text, the layout information, and the template picture into the text-to-image model to obtain picture materials with the target picture filled in the target area of the template picture; Wherein, the target picture is a picture drawn according to the description text.

5. The method according to claim 1, wherein Inputting the description text and typesetting information into the text-to-image model to obtain picture materials that conform to the typesetting information, including: Based on the description text and typesetting information, perform the following operations by calling the text-to-image model to obtain the picture materials: Extract the text features of the description text and randomly generate an initial noise picture; Use the initial noise picture as the noise picture to be processed for the first image denoising operation, and perform the image denoising operation a preset number of times; Adjust the denoised picture obtained from the last image denoising operation according to the typesetting information to obtain the picture materials; Among them, the image denoising operation includes: According to the text features and the noise picture to be processed, call the denoising network of the StableDiffusion model to denoise the noise picture to be processed to obtain a denoised picture, and use this denoised picture as the noise picture to be processed for the next image denoising operation.

6. The method according to claim 4, characterized in that, Inputting the description text, typesetting information, and template picture into the text-to-image model to obtain picture materials with the target picture filled in the target area of the template picture, including: Based on the description text and typesetting information, perform the following operations by calling the text-to-image model to obtain the picture materials: Extract the text features of the description text and randomly generate an initial noise picture; Fill the initial noise picture into the target area of the template picture and use it as the noise picture to be processed for the first image denoising operation, and perform the image denoising operation a preset number of times; Adjust the denoised picture obtained from the last image denoising operation according to the typesetting information to obtain the picture materials; Among them, the image denoising operation includes: According to the text features and the noise picture to be processed, call the denoising network of the StableDiffusion model to denoise the target area in the noise picture to be processed to obtain a denoised picture, and use this denoised picture as the noise picture to be processed for the next image denoising operation.

7. The method according to claim 1, characterized in that, It also includes: Call a pre-trained second large language model to extract keywords of at least one target part of speech that match the interest tags from a pre-established keyword library, and perform text supplementation according to the keywords to obtain copywriting materials; Generate recommendation information according to the copywriting materials and the picture materials; Among them, the second large language model is trained with the interest tags of multiple sample audiences as training samples and the keywords of the copywriting materials in each historical recommendation information of the sample audiences as training labels.

8. The method according to claim 3, wherein After obtaining the picture materials that conform to the typesetting information, it also includes: Input the picture materials into an image recognition model to obtain the target objects included in the picture materials output by the image recognition model; If it is determined that the target object does not belong to the objects to be avoided related to the recommender, retain the picture materials; If it is determined that the target object belongs to the objects to be avoided, generate new negative prompt words according to the objects to be avoided; Input the positive prompt, negative prompt, new negative prompt, and layout information into the text-to-image model to obtain new image materials that meet the layout information output by the text-to-image model.

9. The method according to claim 7, characterized in that, The keywords in the historical recommendation information of the sample audience are determined by the following method: Segment the text materials in the historical recommendation information; Determine the term frequency-inverse document frequency index TF-IDF and TextRank of each segmented word in each text material; Weight the TF-IDF and TextRank of each segmented word to obtain the importance of each segmented word; Select a preset number of segmented words as keywords in descending order of importance.

10. A material generation device for recommendation information, characterized in that, Including: A basic information acquisition module for acquiring the audience label, the image style of the historical recommendation information, and the layout information of the recommendation information carrier, where the audience label includes at least one interest label; A text supplement module for inputting the audience label and image style into the first large language model to obtain a description text for describing the image material output by the first large language model; An image creation module for inputting the description text and layout information into the text-to-image model to obtain new image materials that meet the layout information output by the text-to-image model.

11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method for generating materials for the recommendation information according to any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for generating materials for the recommendation information according to any one of claims 1-9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for generating materials for the recommendation information according to any one of claims 1-8.

Citation Information

Cited By

  • Recommendation information generation method and device, equipment, medium and program product

    CN121365168A

  • Method and device for generating picture-text content, storage medium and electronic device

    CN122550751A