Personalized image generation method, model training method, equipment, medium and product

By obtaining interactive scenarios and user information and generating personalized copy and images, the problem that a single marketing style is difficult to meet the needs of diverse users is solved, and the user's willingness to interact and information transmission effect is improved.

CN120147455APending Publication Date: 2025-06-13ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510225349.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Among the diverse user groups, a single marketing style is difficult to meet the needs of all users, resulting in poor information communication and insufficient user willingness to interact.

Method used

By obtaining the scene description information and image generation requirements of the interactive scene, as well as the user information and image preference information of the target user, personalized copy and images are generated to ensure that the content is highly relevant to the interactive scene and meets user preferences.

Benefits of technology

It achieves a high degree of relevance and user acceptance of personalized copywriting and images, and improves the user's interaction willingness and the effectiveness and fun of information transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147455A_ABST
    Figure CN120147455A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a personalized image generation method, a model training method, equipment, a medium and a product. The personalized image generation method comprises the following steps: acquiring scene description information and an image generation demand of an interaction scene, and acquiring user information and image preference information of a target user; generating a first cue word based on the scene description information and the user information, and inputting the first cue word into a trained text generation model to generate a personalized copywriting which conforms to the interaction scene and faces the target user by the text generation model; generating a second cue word based on the image generation demand, the image preference information and the personalized copywriting, and inputting the second cue word into a trained image generation model, and generating a personalized image which meets the image generation requirement and the image preference information and comprises the personalized copywriting by the image generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of data processing technologies, and in particular, to a personalized image generation method, a training method for a text generation model, an electronic device, a computer-readable medium, and a computer program product. Background Art

[0002] In various interactive scenarios that need to convey information to users, such as product marketing, advertising promotion, e-commerce recommendation, etc., in order to improve the conveying effect and the user's interaction willingness, the activity initiator usually conveys relevant information (such as marketing activity information, product information, advertising content, etc.) to users in a combination of text and images. This way, through the combination of vision and text, it can attract the user's attention and enhance the readability and attractiveness of the information. However, in the face of a large and diverse user group, a single marketing style often cannot meet the needs of all users. Summary of the Invention

[0003] In view of this, one or more embodiments of this specification provide a personalized image generation method, a training method for a text generation model, an electronic device, a computer-readable medium, and a computer program product.

[0004] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:

[0005] According to a first aspect of one or more embodiments of this specification, a personalized image generation method is proposed, including:

[0006] Obtain the scene description information and image generation requirements of the interactive scenario, and obtain the user information and image preference information of the target user;

[0007] Generate a first prompt word based on the scene description information and the user information, and input the first prompt word into a trained text generation model, so that the text generation model generates a personalized copywriting that conforms to the interactive scenario and is targeted at the target user;

[0008] Generate a second prompt word based on the image generation requirements, the image preference information, and the personalized copywriting, and input the second prompt word into a trained image generation model, so that the image generation model generates a personalized image that conforms to the image generation requirements and the image preference information and includes the personalized copywriting.

[0009] According to a second aspect of the embodiments of this specification, a training method for a text generation model is provided, including:

[0010] Obtain a plurality of training samples, each of the training samples including input information and a label, the input information including scene description information of an interaction scene and user information of an interacted user, the label including a copywriting in an image corresponding to the interaction scene, and the interacted user having performed a specified interaction behavior on the image corresponding to the interaction scene;

[0011] Input a prompt word generated from the input information in the training sample into a text generation model to be trained, so that the text generation model to be trained generates a predicted copywriting that conforms to the interaction scene and is targeted at the interacted user based on the prompt word, and train the text generation model to be trained with the optimization objective of minimizing the difference between the predicted copywriting and the label;

[0012] Among them, the trained text generation model is applied to the method described in the first aspect.

[0013] According to a third aspect of the embodiments of the present specification, there is provided an electronic device, including:

[0014] A processor;

[0015] A memory for storing instructions executable by the processor;

[0016] Among them, when the processor executes the executable instructions, it is used to implement the method described in the first aspect.

[0017] According to a fourth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in the first aspect.

[0018] According to a fifth aspect of the embodiments of the present specification, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.

[0019] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects:

[0020] In the embodiments of this specification, the scene description information and image generation requirements of the interaction scene are obtained, and the user information and image preference information of the target user are obtained. Then, a first prompt word is generated based on the scene description information and the user information, and the first prompt word is input into a trained text generation model to generate a personalized copywriting that conforms to the interaction scene and is targeted at the target user. On the one hand, it ensures that the personalized copywriting is highly relevant to the interaction scene, avoiding irrelevant promotions or information displays. On the other hand, it ensures that the personalized copywriting conforms to the user's preferences, which can increase the user's acceptance of the copywriting and effectively improve the user's interaction willingness. Further, a second prompt word is generated based on the image generation requirements, the image preference information, and the personalized copywriting, and the second prompt word is input into a trained image generation model to generate a personalized image that conforms to the image generation requirements and the image preference information and includes the personalized copywriting. The image content of the generated personalized image is both highly relevant to the interaction scene and meets the user's visual needs. The copywriting and image content in the personalized image both conform to the user's preferences, which can achieve the effect of text-image collaboration, form a dual attraction of vision and language, and enhance the effectiveness and interest of information transmission.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. Brief Description of the Drawings

[0022] Figure 1 is a schematic structural diagram of a personalized image generation service system provided by an exemplary embodiment.

[0023] Figure 2 is a training schematic diagram of a text generation model provided by an exemplary embodiment.

[0024] Figure 3 is a schematic structural diagram of a text generation model provided by an exemplary embodiment.

[0025] Figure 4 is a schematic structural diagram of another text generation model provided by an exemplary embodiment.

[0026] Figure 5 is a flowchart of a personalized image generation method provided by an exemplary embodiment.

[0027] Figure 6 is a flowchart of another personalized image generation method provided by an exemplary embodiment.

[0028] Figure 7 is a schematic structural diagram of an electronic device provided by an exemplary embodiment. Detailed Description of the Embodiments

[0029] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0030] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0031] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0032] The relevant terms that appear in the embodiments of this application are explained herein:

[0033] 1. Transformer: A deep learning model architecture, a model based on the attention mechanism, specifically designed for processing sequence data, and has achieved remarkable success especially in the field of natural language processing.

[0034] 2. Attention mechanism: The core component in Transformer, which is a mechanism for processing sequence data, especially suitable for natural language processing tasks. The attention mechanism allows the model to consider the information of all other elements in the sequence when processing each element in the sequence, so as to be able to capture long-distance dependencies and better understand the context.

[0035] Exemplarily, the implementation of the attention mechanism includes the following steps:

[0036] First, it involves the conversion of Query, Key, and Value. For each input position, three vectors are obtained through linear transformation: the Query vector (q), the Key vector (k), and the Value vector (v). These vectors are used to calculate attention scores and generate the output. By calculating the similarity between the Query vector and each Key vector (usually using the dot product or other similarity functions), the attention scores of each position to all other positions are obtained. Finally, the attention scores are used to weight and sum all the Value vectors to generate the final output.

[0037] It can be expressed by the following formula:

[0038] q = XW Q ;

[0039] k = XW K ;

[0040] v = XW V ;

[0041]

[0042] out = attnW o .

[0043] Among them, X represents the input vector, W Q represents the query matrix learned during the model training process, W K represents the key matrix learned during the model training process, W V represents the value matrix learned during the model training process, q represents the Query vector, k represents the Key vector, v represents the Value vector, softmax() is the scaling and normalization function, is the scaling factor (used to scale the dot product result to a moderate range to alleviate the problem of numerical instability), d is the dimension of the Query vector, k T is the transpose of the Key vector, W o represents the output weight learned during the model training process, and out represents the final output.

[0044] 3. Self-attention mechanism allows each element in the input sequence to attend to all elements in the sequence (including itself) in order to calculate its output representation. This process helps to capture long-range dependencies in the sequence. The implementation of Self-attention usually includes the following steps:

[0045] (1) Linear transformation: For each token in the input sequence, through three learnable weight matrices W Q , W V and WK They are respectively mapped to the query (Q), key (K), and value (V) vector spaces.

[0046] (2) Calculate attention scores: Use the dot product or other similarity functions to calculate the matching degree or correlation between each Query and all other Keys. The result of this step is a score matrix, indicating the association strength between different elements.

[0047] (3) Softmax normalization: Perform the softmax operation on the attention scores to ensure that the scores in each row add up to 1, obtaining a probability distribution that reflects the importance of each element in the current context.

[0048] (4) Weighted summation: Use the probability distribution obtained in the previous step as weights to perform weighted summation on the Values, generating the final output representation for each position.

[0049] Among them, in order to increase the expressive power of the model, the self-attention mechanism usually adopts the multi-head method, that is, multiple above processes are executed simultaneously, and each process uses a different parameter set. Finally, the outputs of each head are concatenated and passed through an additional linear layer to obtain the final output.

[0050] 4. The cross-attention mechanism is a mechanism for calculating attention between two different input sequences. Specifically, when processing one sequence (such as the target sequence), it depends on another sequence (such as the source sequence) to generate a better representation. In cross-attention, one sequence serves as Key and Value, and the other sequence serves as Query.

[0051] Suppose there are two input sequences: (1) Sequence A (Query sequence): used to generate the Query vector (q); (2) Sequence B (Key-Value sequence): used to generate the Key vector (k) and Value vector (v).

[0052] First, calculate the Query vector (q), Key vector (k), and Value vector (v): q = X A W Q ; k = X B W K ; v = X B W V ; X A is the representation of sequence A, X B is the representation of sequence B, and W Q represents the query matrix learned during the model training process, and W Kdenotes the key matrix learned during model training, W V denotes the value matrix learned during model training, q denotes the Query vector, k denotes the Key vector, and v denotes the Value vector.

[0053] Next, calculate the attention score through the following formula: attn = softmax(qk T / √d)v; Finally, perform a weighted sum on the value vectors through the attention scores to obtain an enhanced representation of sequence A in the context of sequence B.

[0054] In some embodiments, Figure 1 is a schematic diagram of the architecture of a personalized image generation service system provided by an exemplary embodiment. As Figure 1 shown, the system may include a server 11, a network 12, and several user terminals, such as a PC (Personal Computer) 13, a mobile phone 14, etc.

[0055] The server 11 may be a physical server including an independent host, or the server 11 may be a virtual server hosted by a host cluster. During operation, the server 11 may run the server-side program of the personalized image generation application to implement the corresponding personalized image generation platform.

[0056] The PC 13 and the mobile phone 14 are only some types of user terminals that users can use. In fact, users can obviously also use user terminals of the following types: tablet devices, laptop computers, personal digital assistants (PDAs), wearable devices (such as smart glasses, smart watches, etc.). One or more embodiments of this specification do not limit this. During operation, the user terminal may run the client-side program of the personalized image generation application to implement the client of the personalized image generation service. Among them, the application program of the client of the above-mentioned personalized image generation service can be started and run on the user terminal. The client-side program may be a native application installed on the user terminal, or the client-side program may be a small program, a fast application, or other similar forms. Of course, when using web technologies such as HTML5 or similar, relevant functions can be implemented through the page displayed by the browser. Here, the browser may be an independent browser application or a browser module embedded in some applications.

[0057] For the network 12 through which user terminals such as the PC 13 and the mobile phone 14 interact with the server 11, communication can be specifically implemented using a wired or wireless network based on the communication methods supported by the corresponding user terminals. This specification does not limit this. For example, if the PC 13 supports both wired and wireless communication, communication can be implemented using a wired or wireless network according to needs, while the mobile phone 14 usually only supports wireless communication, so wireless network can be used to implement communication.

[0058] Based on the problems in the related art, the server 11 can deploy relevant pre-trained models, such as text generation models, image generation models, image style description models, and image detection models, etc., to execute the personalized image generation method provided in the embodiments of this specification, generate personalized images for different users, and meet the different needs of a large and diverse user group, thereby enhancing the user's interaction willingness. Of course, it can also be executed by other electronic devices docked with the server 11, and the generated personalized images are transmitted to the server 11. This embodiment does not make any restrictions on this.

[0059] (1) Regarding the text generation model used in the embodiments of this specification.

[0060] In a possible implementation, a pre-trained language model can be directly adopted, such as a large language model (LLM). The pre-trained language model is based on deep learning technology, especially an artificial intelligence model trained using a large corpus, aiming to understand and generate text similar to human language, and has powerful natural language understanding and generation capabilities. The goal of the LLM is to achieve various applications through natural language processing capabilities, such as text generation, translation, summarization, question answering, and dialogue systems, so as to help improve the efficiency of human-computer interaction and the degree of automation.

[0061] In another possible implementation, in order to improve the generation effect of personalized copywriting, the pre-trained language model can be fine-tuned. The fine-tuning process can be executed by any electronic device with computing capabilities. This embodiment does not make any restrictions on this. As shown in Figure 2 it is possible to use the pre-trained language model as the text generation model to be trained and perform the following training process:

[0062] First, the electronic device can obtain a plurality of pre-prepared training samples. Each training sample includes input information and a label. The input information includes the scene description information of the interaction scene and the user information of the interacted user. The label includes the copywriting in the image corresponding to the interaction scene, and the interacted user has performed a specified interaction behavior on the image corresponding to the interaction scene. The training samples can clearly show the association between the input information and the label, providing high-quality labeled data for model learning.

[0063] Exemplarily, the interaction scenario is a marketing scenario for a certain product, the scenario description information is the relevant introduction information of the product, the product marketing scenario produces images for marketing, and the interacted users refer to those who have performed interaction behaviors such as clicking, favoriting, and liking on the images produced by the product marketing scenario. The user information of the interacted users includes, but is not limited to, the user's behavior data, interest data, basic attributes, etc.

[0064] In one example, assume that the interaction scenario is a marketing scenario for beauty products. A training sample includes: (1) The scenario description information includes: {Product type: Lipstick; Core selling points: Matte texture, long-lasting and non-smudging, 20 color numbers; Target population: Working women aged 25 - 35. Marketing theme: Workplace confident makeup look}. (2) The user information of the interacted users includes, but is not limited to: {User ID: U2023_Beauty_001, Consumption habit: {Monthly average beauty consumption: 800 yuan, Preferred categories: [Lipstick, Foundation], Purchase channel: E-commerce live streaming}, Basic attributes: {Age: 28, Occupation: Financial analyst}, Behavior characteristics: {Recently viewed: [Workplace dressing tutorials, Makeup skill videos]}}. (3) Label:

Essential for the workplace

[0065] In another example, assume that the interaction scenario is a game promotion scenario. A training sample includes: (1) The scenario description information includes: {Product type: Mobile game, Core selling points: [Open-world exploration, Multiplayer dungeons, Character cultivation], Target population: Student group aged 18 - 25, Marketing theme: New user registration}. (2) The user information of the interacted users includes, but is not limited to: {User ID: U2023_Game_045, Consumption habit: {Monthly average game consumption: 300 yuan, Preferred categories: [MMORPG, Competitive games], Purchase channel: In-game mall}, Basic attributes: {Age: 22, Occupation: College student}, Behavior characteristics: {Recently viewed: [Game strategy videos, E-sports competition live broadcasts]}}. (3) Label:

Epic open world

[0066] After obtaining multiple training samples, the electronic device can generate a prompt word based on the input information in the training samples, and then input the prompt word generated from the input information in the training samples into the text generation model to be trained, so that the text generation model to be trained generates a prediction copywriting that conforms to the interaction scenario and is targeted at the interacted users based on the prompt word, and takes minimizing the difference between the prediction copywriting and the label as the optimization goal to train the text generation model to be trained.

[0067] Among them, the role of the prompt is to transform the scene description information of the interaction scene and the user information of the interacted user into a structured input, which helps the text generation model understand the task objective and generate personalized copy that meets the user's needs. The prompt integrates the input information (including the description of the interaction scene and user information), abstracts it into a concise and structured prompt, enabling the model to extract important features from this information. Through these prompts, the model can better understand the relationship between the interaction scene and the user's needs. The prompt is not only the collation of information, but also clarifies the objectives and constraints of the generation task. It instructs the model to generate copy that meets the needs of the target user based on certain inputs and ensures that the generated copy can match the specific interaction scene and user profile.

[0068] In some embodiments, the user information includes user personalization information represented in text form and a first embedding vector for characterizing the user profile of the interacted user. The user personalization information represented in text form and the first embedding vector for characterizing the user profile of the interacted user each play different but complementary roles. They respectively provide the model with user information at different levels and in different forms, thus making the generated copy more personalized and accurate.

[0069] The user personalization information represented in text form can be determined specifically according to the characteristics of the interaction scene and contains detailed descriptions reflecting the user's preferences, needs, and behaviors. These information help the model more intuitively understand the basic characteristics and needs of the user and can be flexibly adjusted according to the interaction scene. It has the following functions:

[0070] (1) Strong scene relevance: The information in text form is usually customized according to the specific interaction scene and can accurately describe the user's behaviors and needs in this scene. For example, in the e-commerce product marketing scene, the user's personalization information includes historical purchase records, browsing history, shopping cart data, basic attribute information, etc., which can directly reflect the user's behavior patterns, preferences, and sensitivities in shopping; in the travel / travel scene, the user's personalization information includes travel history, destination preferences, budget, travel frequency, travel mode, basic attribute information, etc., which can show the user's travel interests and needs. This enables the text information to help the model accurately capture the user's immediate needs in a specific scene.

[0071] (2) Strong interpretability: The personalized information represented in text form is intuitive and easy to understand, and can provide clearer user characteristics. For example, information such as the user's preference for a certain brand and sensitivity to price can be clearly described, which helps the model adjust the output to match the user's specific needs.

[0072] The first embedding vector used to represent the user profile of the interacted user is obtained by comprehensively processing multiple relevant pieces of information of the user and converting them into a dense vector representation, which is usually composed of the global information of historical data and behavior patterns. It has the following functions:

[0073] (1) Comprehensive user characteristics: The first embedding vector is not just a single piece of user behavior or attribute information, but is created by comprehensively analyzing various types of user data (such as behavioral data, interests, historical interaction records, etc.). This enables it to capture the overall picture and potential needs of the user. For example, in an e-commerce scenario, the first embedding vector may integrate information such as the user's purchase habits, price sensitivity, brand preferences, etc., providing a more comprehensive reference for generating personalized recommendations or advertisements.

[0074] (2) Deep pattern learning: The first embedding vector can reflect the deep patterns and preference trends in user behavior, which may not be intuitively expressed by pure text information. For example, the user's repeated browsing or purchase behavior of a certain type of product can be reflected by the embedding vector, and this information may not be fully presented by simple text descriptions.

[0075] In summary, user personalized information represented in text form focuses more on describing the user's needs and behaviors in specific interaction scenarios through specific and interpretable information. It can provide intuitive and highly scenario-related input for the model, supporting precise content generation and recommendation. The first embedding vector, on the other hand, is obtained by comprehensively analyzing all relevant information of the user and converted into a numerical representation that can capture the user's potential interests, behavior patterns, and global characteristics. This vectorized representation can provide deep personalized insights for the model, supporting efficient personalized recommendation and content generation. The combination of the two enables the model to have both clearly interpretable personalized information and utilize the embedding vector for deep learning when dealing with complex user needs, thus generating content that better meets the user's interests and needs.

[0076] To adapt to the above user information, for each training sample, the electronic device can generate a text-based prompt word based on the user personalized information, scenario description information, and specified placeholder in each training sample, and then input the prompt word and the first embedding vector into the text generation model to be trained. The text generation model to be trained is then used to generate a prediction copy that conforms to the interaction scenario and is targeted at the interacted user based on the prompt word and the first embedding vector. The text generation model to be trained is trained with the optimization goal of minimizing the difference between the prediction copy and the label.

[0077] Please refer to Figure 3 , the text generation model to be trained includes a linear processing layer, an embedding layer, and multiple Transformer structures.

[0078] An embedding layer for mapping the text in the prompt other than the specified placeholder to an embedding vector based on a learnable embedding vector mapping relationship.

[0079] Among them, the text generation model to be trained is a pre-trained neural network model. In the pre-training stage, the text generation model establishes an initial embedding vector mapping relationship by learning existing large-scale text data. These initial embedding vector mapping relationships are usually learned based on common language features and patterns (such as the semantic meaning and grammatical structure of words and sentences, etc.). For example, the model can map different words, phrases, or grammatical structures into a vector space, so that words or sentences with similar semantics have similar vector representations in the vector space. It can be understood that in addition to establishing the initial embedding vector mapping relationship, the text generation model also accumulates certain general knowledge in the pre-training stage, such as common facts, historical events, cultural backgrounds, common sense of life, etc. This embodiment does not impose any restrictions on this.

[0080] Then, this learnable embedding vector mapping relationship can be further adjusted and optimized during this training process with the help of the personalized information in the training samples.

[0081] During this training process, the personalized information of each interacting user (such as purchase history, interest preferences, behavioral data, etc.) will be mapped to a preliminary embedding vector representation. The model will update the learnable parameters of the embedding layer through the backpropagation optimization process according to the error between the predicted copywriting and the labeled copywriting. In this way, the model can gradually learn the mapping relationship between the personalized information of the user and the embedding vector. As the training progresses, the embedding layer of the model will gradually increase the mapping ability for the personalized information of the interacting users. That is to say, the embedding vector mapping relationship included in the embedding layer further increases the mapping relationship between the personalized information of the interacting users and the embedding vector, thereby increasing the control ability of the personalized information over the generation result of the model, making the copywriting generated by the model more in line with the needs of individual users.

[0082] A linear processing layer for processing the vector dimension of the first embedding vector so that the processed first embedding vector is aligned with the embedding vector mapped by the embedding layer in the same vector space. In this embodiment, since the first embedding vector and the embedding vector generated by the embedding layer may be in different vector spaces. The linear processing layer converts the first embedding vector into a new representation by learning appropriate weights, so that this vector can be aligned with other input embedding vectors in a unified vector space, thereby ensuring that different input vectors can be effectively fused and processed in the same feature space. It can be understood that if the first embedding vector itself can already be aligned with the embedding vector mapped by the embedding layer in the same vector space, there is no need for the linear processing layer to further process the first embedding vector.

[0083] The embedding layer is also used to replace the specified placeholder in the prompt with the processed first embedding vector, so as to introduce the first embedding vector, effectively combine the first embedding vector with other parts of the prompt, obtain a second embedding vector that can synthesize the first embedding vector and the text-based user personalized information, and input the second embedding vector into multiple subsequent Transformer structures for further processing.

[0084] It can be understood that the relevant parameters of the linear processing layer, embedding layer and multiple Transformer structures included in the text generation model can be updated based on the above optimization objectives during this training process, so that the fine-tuned text generation model is more suitable for the personalized copywriting generation scenario. For the specific parameter update method, the parameter update algorithm in the relevant technology can be used for update, and this embodiment does not make any restrictions on this.

[0085] In a possible implementation manner, if only the second embedding vector needs to be input into multiple subsequent Transformer structures for processing, the multiple Transformer structures can refer to the model architecture in the relevant technology, and this embodiment does not make any restrictions on this.

[0086] In a possible implementation manner, in order to further enhance the control ability of the first embedding vector on the model generation result, the first embedding vector can be introduced again in the last Transformer structure among the multiple Transformer structures, and the second embedding vector is used as the input of the first Transformer structure among the multiple Transformer structures; the first embedding vector is used as a part of the input of the last Transformer structure among the multiple Transformer structures.

[0087] Exemplarily, the last Transformer structure among multiple Transformer structures includes a self-attention structure and a cross-attention structure; the input of the self-attention structure is the output of the previous Transformer structure; the input of the cross-attention structure includes the output of the self-attention structure and a first embedding vector. The output of the self-attention structure is used to calculate the query vector in the cross-attention structure, and the first embedding vector is used to calculate the key vector and value vector in the cross-attention structure. In this embodiment, by introducing the first embedding vector only in the last Transformer structure, the model can first learn general context and patterns in the previous Transformer structures, thereby constructing a relatively general representation. Reintroducing the first embedding vector in the last Transformer structure can enable the model to make personalized fine-tuning during the generation process. At this time, the model already has basic semantic and context information. By reintroducing the first embedding vector, it can ensure that the generated content is more in line with the interests, needs, or behavioral characteristics of specific users.

[0088] The other Transformer structures except the last one among multiple Transformer structures can refer to the model architecture in the related art, and this embodiment does not impose any restrictions on this.

[0089] Exemplarily, please refer to Figure 4 , Figure 4 Taking N Transformer structures as an example, where N is an integer greater than 0, the last Transformer structure among multiple Transformer structures includes a self-attention structure, a normalization layer, a cross-attention structure, a fusion layer, and a feed-forward neural network.

[0090] The self-attention structure is used to calculate based on the output of the previous Transformer structure. The specific calculation process can refer to the above description and will not be elaborated here.

[0091] The normalization layer is used to normalize the output of the self-attention structure.

[0092] The input of the cross-attention structure includes the output of the normalization layer and a first embedding vector. The output of the normalization layer is used to calculate the query vector in the cross-attention structure, and the first embedding vector is used to calculate the key vector and value vector in the cross-attention structure. The specific calculation process of the cross-attention structure can refer to the above description and will not be elaborated here.

[0093] The fusion layer is used to fuse the outputs of the self-attention structure and the cross-attention structure; for example, it can be fused based on bitwise addition, and the fused result is further normalized.

[0094] The feed-forward neural network is used to process the output of the fusion layer.

[0095] Of course, in addition to the linear processing layer, the embedding layer, and multiple Transformer structures, the text generation model to be trained can also include other structures. For example, it can include an LM head (Language Model Head) structure associated with multiple Transformer structures. The LM head structure can convert the output of the last Transformer structure into a probability distribution for predicting the next word.

[0096] In another possible implementation, the user information of the interacted user can also only include user personalization information represented in text form. Then, the prompt word does not need to include a specified placeholder. It only needs the embedding layer to convert the prompt word into an embedding vector, and then input the embedding vector into multiple Transformer structures for further processing.

[0097] (2) Regarding the image generation model used in the embodiments of this specification, the image generation model can generate another image that meets the requirements based on at least one of the input text and image.

[0098] In one possible implementation, a pre-trained image generation model in related technologies can be directly adopted, such as a diffusion model from text to image, a generative adversarial network, etc.

[0099] In another possible implementation, in order to improve the generation effect of personalized images, a pre-trained image generation model in related technologies can also be fine-tuned using relevant training samples. The fine-tuning process can be executed by any electronic device with computing power, and this embodiment does not make any restrictions on this. The training samples include input information and labels. The output information includes the image generation requirements of the interaction scenario and the image preference information of the interacted user. The labels include the images corresponding to the interaction scenarios.

[0100] (3) Regarding the image style description model used in the embodiments of this specification, the image style description model is used to analyze the input image to generate corresponding image style description information.

[0101] In a possible implementation, a pre-trained image style description model in related technologies can be directly adopted, such as the BLIP (Bootstrapping Language-Image Pretraining) model, etc.

[0102] In another possible implementation, in order to improve the generation effect of personalized images, a pre-trained image style description model in related technologies can also be fine-tuned using relevant training samples. The fine-tuning process can be executed by any electronic device with computing capabilities, and this embodiment places no restrictions on this. The training samples include input information and labels. The output information includes images in which the user has performed specified interaction behaviors, and the label is the style description information of the image.

[0103] (4) Regarding the image detection model used in the embodiments of this specification, this image detection model is used to detect the content of a specified dimension in the input image to determine whether it meets the requirements.

[0104] In a possible implementation, a pre-trained image detection model in related technologies can be directly adopted.

[0105] In another possible implementation, in order to improve the detection effect of image quality, a pre-trained image detection model in related technologies can also be fine-tuned using relevant training samples. The fine-tuning process can be executed by any electronic device with computing capabilities, and this embodiment places no restrictions on this. The training samples include input information and labels. The input information of positive samples includes qualified images, and the label is an identifier indicating passing the detection; the input information of negative samples includes unqualified images, and the label is an identifier indicating failing the detection.

[0106] In some embodiments, the following provides an exemplary description of a personalized image generation method provided by the embodiments of this specification, which can be executed by Figure 1 the server 11 in Figure 5 . The method includes:

[0107] In S501, obtain the scene description information and image generation requirements of the interaction scene, and obtain the user information and image preference information of the target user.

[0108] The scene description information of the interaction scene involves the basic background and requirements of the interaction scene. For example, in an advertising and marketing scene, the scene description may include information such as product category, core selling points, and target audience. If it is a travel recommendation, the scene description may involve the destination, season, activity type, etc.

[0109] The image generation requirements include the basic generation requirements of the image, such as the style, size, content, quality requirements, etc. of the image. The image generation requirements clarify the general outline of the image related to the interaction scenario.

[0110] The user information of the target user includes the user's historical data, behaviors, basic attributes, etc. For example, the user's age, gender, hobbies, purchase history, etc., which help the model understand the user's personalized needs. Exemplarily, the user information includes: the user's personalized information represented in text form, and, the first embedding vector used to characterize the user portrait of the target user.

[0111] The image preference information refers to the specific requirements of the user for aspects such as image style, color, content preference, etc. It helps to further refine the generated image and make the image more in line with the user's aesthetics or needs. For example, the user prefers a simple and modern style or has a preference for a certain specific color.

[0112] Exemplarily, if the target user is not a cold start user, since the target user already has a certain interaction history, the image preference information of the target user includes at least one of the following: the historical images related to the specified interaction behaviors of the target user, and, the image style preference information of the target user; wherein, the image style preference information of the target user is obtained by using the trained image style description model to identify the historical images related to the specified interaction behaviors of the target user. The specified interaction behaviors include, but are not limited to, behaviors such as clicking, collecting, liking, forwarding, etc. Since non-cold start users have accumulated a certain amount of historical data, the server can infer their preferences based on the user's historical interaction data and make more accurate personalized recommendations.

[0113] Exemplarily, if the target user is a cold start user and the server cannot directly infer the user's image preferences based on the target user's historical interaction behaviors, it can find reference users similar to the target user based on the first embedding vector of the target user. The similarity between the first embedding vector corresponding to the reference user and the first embedding vector corresponding to the target user meets the preset similarity condition. For example, the preset similarity condition is that the similarity between the first embedding vector corresponding to the reference user and the first embedding vector corresponding to the target user is greater than the preset threshold, or, the similarity between the first embedding vector corresponding to the reference user and the first embedding vector corresponding to the target user is the highest.

[0114] The image preference information of the reference user can be used to infer the preferences of cold start users. Then, the image preference information of the target user includes at least one of the following: historical images related to the specified interaction behavior of the reference user, and, the image style preference information of the reference user; wherein, the image style preference information of the reference user is obtained by using a trained image style description model to identify the historical images related to the specified interaction behavior of the reference user. By leveraging the information of a reference user similar to the target user, effective personalized recommendations can be made in the absence of direct data of the target user, which provides reasonable initial recommendations for cold start users.

[0115] In S502, a first prompt is generated based on the scene description information and user information, and the first prompt is input into a trained text generation model, so that the text generation model generates personalized copywriting that conforms to the interaction scene and is targeted at the target user.

[0116] In this step, the generated personalized copywriting includes promotional language, advertising copy, etc. for products or activities, and will be adjusted according to the user's interests and needs to increase the user's acceptance of the copywriting, so as to better attract the target user.

[0117] Exemplarily, please refer to Figure 6 , in the case where the user information includes: user personalized information represented in text form, and, a first embedding vector for characterizing the user portrait of the target user, the server can generate a first prompt based on the user personalized information, the scene description information, and a specified placeholder; input the first prompt and the first embedding vector into the text generation model, so that the text generation model converts the other text in the first prompt except the specified placeholder into an embedding vector, and replaces the specified placeholder with the first embedding vector to obtain a second embedding vector corresponding to the first prompt, and generates personalized copywriting that conforms to the interaction scene and is targeted at the target user at least based on the second embedding vector. Among them, the combination of the user personalized information represented in text form and the first embedding vector for characterizing the user portrait of the target user enables the text generation model to have both clearly interpretable personalized information and be able to utilize the embedding vector for deep learning when processing complex user needs, thereby generating content that better conforms to the user's interests and needs.

[0118] Among them, the text generation model includes a linear processing layer and an embedding layer. When the first embedding vector is introduced for the first time, the embedding layer is used to map the text other than the specified placeholder in the first prompt word into an embedding vector according to the embedding vector mapping relationship learned during the pre-training process and fine-tuning process of the text generation model; the linear processing layer is used to process the first embedding vector so that the processed first embedding vector and the embedding vector mapped by the embedding layer are aligned to the same vector space; the embedding layer is further used to replace the specified placeholder in the first prompt word with the processed first embedding vector, effectively combining the first embedding vector with other parts of the prompt word, so as to obtain a second embedding vector corresponding to the first prompt word, realizing the integration of the first embedding vector and the text-based user personal information.

[0119] To enhance the control ability of the first embedding vector over the model generation result, the first embedding vector can be introduced again. The text generation model is used to generate a personalized copy that conforms to the interaction scenario and is targeted at the target user according to the first embedding vector and the second embedding vector; among them, the text generation model includes multiple Transformer structures; the second embedding vector is used as the input of the first Transformer structure among the multiple Transformer structures; the first embedding vector is used as a part of the input of the last Transformer structure among the multiple Transformer structures.

[0120] The last Transformer structure among the multiple Transformer structures includes a self-attention structure and a cross-attention structure; the input of the self-attention structure is the output of the previous Transformer structure; the input of the cross-attention structure includes the output of the self-attention structure and the first embedding vector. The output of the self-attention structure is used to calculate the query vector in the cross-attention structure, and the first embedding vector is used to calculate the key vector and value vector in the cross-attention structure.

[0121] Please refer to Figure 4 , the last Transformer structure among the multiple Transformer structures further includes a normalization layer, a fusion layer, and a feed-forward neural network; the normalization layer is used to perform normalization processing on the output of the self-attention structure; the input of the cross-attention structure includes the output of the normalization layer and the first embedding vector, and the output of the normalization layer is used to calculate the query vector in the cross-attention structure; the fusion layer is used to perform fusion processing on the output of the self-attention structure and the output of the cross-attention structure; the feed-forward neural network is used to process the output of the fusion layer.

[0122] In S503, a second prompt is generated based on the image generation requirement, image preference information, and personalized copywriting, and the second prompt is input into a pre-trained image generation model, so that the image generation model generates a personalized image that meets the image generation requirement and image preference information and includes the personalized copywriting.

[0123] In this step, the server generates a second prompt based on the generated personalized copywriting, image generation requirement, and image preference information, and then inputs the second prompt into a pre-trained image generation model. The image generation model can generate an image that not only meets the image requirement but also expresses the content of the personalized copywriting according to the user's image preferences (such as color, style, composition, etc.). For example, in an e-commerce scenario, the generated image may contain a high-quality display of a certain product, and at the same time, a slogan matching the copywriting content is embedded in the image, enhancing the marketing effect of the image.

[0124] In this embodiment, a language generation model can be used to generate personalized copywriting that meets the interaction scenario and is targeted at the target user. On the one hand, it ensures that the personalized copywriting is highly relevant to the interaction scenario, avoiding irrelevant promotions or information displays. On the other hand, it ensures that the personalized copywriting meets the user's preferences, which can increase the user's acceptance of the copywriting and thus effectively improve the user's interaction willingness. Further, an image generation model is used to generate a personalized image that meets the image generation requirement and image preference information and includes the personalized copywriting. The image content of the generated personalized image is both highly relevant to the interaction scenario and meets the user's visual needs. The copywriting and image content in the personalized image both meet the user's preferences, which can achieve the effect of graphic and text collaboration, form a double attraction of vision and language, and enhance the effectiveness and interest of information transmission.

[0125] Exemplarily, please refer to Figure 6 , the second prompt also includes the user information of the target user, so that the personalized image generated by the image generation model conforms to the user information of the target user. By integrating the user information of the target user (such as age, gender, interests, behavior data, etc.) into the second prompt, the image generation model can generate an image that more precisely conforms to the user's characteristics. For example, for a young female user, the generated image may be more in line with her preferences in terms of style, color tone, and composition. Similarly, for a male user, an image that is more in line with his aesthetics or preferences may be generated.

[0126] In some embodiments, after obtaining the personalized image, the server can also use the trained image detection model to detect the content of a specified dimension in the personalized image to obtain a detection result; if the detection result indicates that the personalized image passes the detection, the personalized image is output to the user terminal of the target user. In this embodiment, by detecting the content of the specified dimension in the image, it can be ensured that the generated personalized image not only meets the user's preferences, needs, and scenario requirements, but also has high quality and adaptability. This detection can help ensure the quality and suitability of the image before output, thereby enhancing the user experience and the effectiveness of the image, and avoiding inappropriate or low-quality images from being sent to the target user.

[0127] The content of the specified dimension can be specifically set according to the actual application scenario, and this embodiment does not impose any restrictions on this.

[0128] For example, if the personalized image is a product marketing image, the content of the specified dimension can include whether the product appears in the image and whether the core selling points of the product (such as size, color, function, etc.) are clearly displayed.

[0129] Or, detect whether the image meets the expectations of the interaction scenario. For example, in a travel advertisement image, the specified dimension can be to check whether the image shows the travel scenarios preferred by the target user (such as beaches, cities, mountains, etc.).

[0130] Or, detect whether the resolution of the image meets the target output requirements to ensure that the image quality is high enough, clear, and without obvious noise or distortion. Or, check whether the brightness and contrast of the image meet the standards to ensure the visual effect of the image on different devices.

[0131] The various technical features in the above embodiments can be combined arbitrarily as long as there is no conflict or contradiction between the features. However, due to space limitations, they are not described one by one. Therefore, any combination of the various technical features in the above embodiments also belongs to the scope disclosed in this specification.

[0132] In some embodiments, the embodiments of this specification also provide an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor realizes the method described in any one of the above by running the executable instructions.

[0133] Figure 7 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 7, at the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710. Of course, it may also include other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 702 reads the corresponding computer program from the non-volatile memory 710 into the memory 708 and then runs it. Of course, in addition to the software implementation method, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0134] In some embodiments, the personalized image generation device can be applied to a device as Figure 7 shown to implement the technical solutions of this specification. Among them, the personalized image generation device may include:

[0135] An acquisition module, configured to acquire the scene description information and image generation requirements of the interaction scene, and, acquire the user information and image preference information of the target user.

[0136] A personalized copywriting generation module, configured to generate a first prompt word based on the scene description information and the user information, and input the first prompt word into a trained text generation model, so that the text generation model generates personalized copywriting that conforms to the interaction scene and is targeted at the target user.

[0137] A personalized image generation module, configured to generate a second prompt word based on the image generation requirements, the image preference information, and the personalized copywriting, and input the second prompt word into a trained image generation model, so that the image generation model generates a personalized image that conforms to the image generation requirements and the image preference information and includes the personalized copywriting.

[0138] Exemplarily, the user information includes: user personalized information represented in text form, and a first embedding vector used to characterize the user portrait of the target user.

[0139] The personalized copywriting generation module is specifically configured to generate a first prompt word based on the user personalized information, the scene description information, and a specified placeholder; input the first prompt word and the first embedding vector into the text generation model, so that the text generation model converts the other text in the first prompt word except the specified placeholder into an embedding vector, and replaces the specified placeholder with the first embedding vector to obtain a second embedding vector corresponding to the first prompt word, and at least generate personalized copywriting that conforms to the interaction scene and is targeted at the target user based on the second embedding vector.

[0140] Exemplarily, the text generation model includes a linear processing layer and an embedding layer; the embedding layer is configured to map the text other than the specified placeholder in the first prompt word to an embedding vector according to the embedding vector mapping relationship learned during the pre-training process and the fine-tuning process of the text generation model; the linear processing layer is configured to process the first embedding vector so that the processed first embedding vector and the embedding vector mapped by the embedding layer are aligned to the same vector space; the embedding layer is further configured to replace the specified placeholder in the first prompt word with the processed first embedding vector.

[0141] Exemplarily, the text generation model is configured to generate a personalized copy that conforms to the interaction scenario and is targeted at the target user according to the first embedding vector and the second embedding vector; wherein, the text generation model includes a plurality of Transformer structures; the second embedding vector serves as the input of the first Transformer structure among the plurality of Transformer structures; the first embedding vector serves as a part of the input of the last Transformer structure among the plurality of Transformer structures.

[0142] Exemplarily, the last Transformer structure among the plurality of Transformer structures includes a self-attention structure and a cross-attention structure; the input of the self-attention structure is the output of the previous Transformer structure; the input of the cross-attention structure includes the output of the self-attention structure and the first embedding vector, the output of the self-attention structure is used to calculate the query vector in the cross-attention structure, and the first embedding vector is used to calculate the key vector and the value vector in the cross-attention structure.

[0143] Exemplarily, the last Transformer structure among the plurality of Transformer structures further includes a normalization layer, a fusion layer, and a feed-forward neural network; the normalization layer is configured to perform normalization processing on the output of the self-attention structure; the input of the cross-attention structure includes the output of the normalization layer and the first embedding vector, and the output of the normalization layer is used to calculate the query vector in the cross-attention structure; the fusion layer is configured to perform fusion processing on the output of the self-attention structure and the output of the cross-attention structure; the feed-forward neural network is configured to process the output of the fusion layer.

[0144] Exemplarily, if the target user is not a cold start user, the image preference information of the target user includes: historical images related to the specified interaction behavior of the target user, and / or, the image style preference information of the target user; wherein, the image style preference information of the target user is obtained by using a trained image style description model to identify historical images related to the specified interaction behavior of the target user.

[0145] Exemplarily, if the target user is a cold start user, the image preference information of the target user includes: historical images related to the specified interaction behavior of a reference user, and / or, the image style preference information of the reference user; wherein, the image style preference information of the reference user is obtained by using a trained image style description model to identify historical images related to the specified interaction behavior of the reference user; the similarity between the first embedding vector corresponding to the reference user and the first embedding vector corresponding to the target user satisfies a preset similarity condition.

[0146] Exemplarily, the second prompt also includes the user information of the target user, so that the personalized image generated by the image generation model conforms to the user information of the target user.

[0147] Exemplarily, the device further includes an image detection module, configured to use a trained image detection model to detect the content of a specified dimension in the personalized image to obtain a detection result; if the detection result indicates that the personalized image passes the detection, output the personalized image to the user terminal of the target user.

[0148] In some embodiments, the training device of the text generation model can be applied to a device as shown in Figure 7 to implement the technical solutions of this specification. Among them, the training device of the text generation model may include:

[0149] A sample acquisition module, configured to acquire a plurality of training samples, each of the training samples including input information and a label, the input information including scene description information of an interaction scene and user information of an interacted user, the label including the copywriting in the image corresponding to the interaction scene, and the interacted user has performed a specified interaction behavior on the image corresponding to the interaction scene;

[0150] A training module, configured to input a prompt word generated from the input information in the training samples into a text generation model to be trained, so that the text generation model to be trained generates a predicted copywriting that conforms to the interaction scenario and is targeted at the interacted user based on the prompt word, and trains the text generation model to be trained with the optimization goal of minimizing the difference between the predicted copywriting and the label; wherein, the trained text generation model is applied to the above personalized image generation method.

[0151] Exemplarily, the user information includes: user personalized information represented in text form, and a first embedding vector for characterizing the user portrait of the interacted user.

[0152] The training module is specifically configured to input the prompt word generated from the user personalized information, the scenario description information and the specified placeholder, and the first embedding vector into the text generation model to be trained.

[0153] Wherein, the text generation model to be trained includes a linear processing layer, an embedding layer and multiple Transformer structures.

[0154] The embedding layer is configured to map other texts in the prompt word except the specified placeholder into embedding vectors based on a learnable embedding vector mapping relationship; wherein, the learnable embedding vector mapping relationship further increases the mapping relationship between the personalized information of the interacted user and the embedding vectors during this training process.

[0155] The linear processing layer is configured to process the first embedding vector so that the processed first embedding vector and the embedding vectors mapped by the embedding layer are aligned to the same vector space.

[0156] The embedding layer is further configured to replace the specified placeholder in the prompt word with the processed first embedding vector.

[0157] The second embedding vector serves as the input of the first Transformer structure among the multiple Transformer structures; the first embedding vector serves as a part of the input of the last Transformer structure among the multiple Transformer structures.

[0158] The implementation processes of the functions and roles of each module in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0159] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor realizes the steps of the method as described in any of the above embodiments by running the executable instructions.

[0160] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any of the above embodiments are realized.

[0161] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0162] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method as described in any of the above embodiments are realized.

[0163] The above description is only the preferred embodiment of one or more embodiments of this specification, and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the protection scope of one or more embodiments of this specification.

Claims

1. A personalized image generation method, comprising: Obtaining scene description information and image generation requirements of the interactive scene, as well as obtaining user information and image preference information of the target user; Generate a first prompt word based on the scene description information and the user information, and input the first prompt word into a trained text generation model, so that the text generation model generates a personalized copy that conforms to the interaction scene and is oriented to the target user; A second prompt word is generated based on the image generation requirement, the image preference information and the personalized text, and the second prompt word is input into a trained image generation model so that the image generation model generates a personalized image that meets the image generation requirement and the image preference information and includes the personalized text.

2. The method according to claim 1, wherein the user information comprises: User personalized information represented in text form, and a first embedding vector for characterizing a user portrait of the target user; The step of generating a first prompt word based on the scene description information and the user information, and inputting the first prompt word into a trained text generation model so that the text generation model generates a personalized copy that conforms to the interaction scene and is oriented to the target user, includes: Generate a first prompt word based on the user personalized information, the scene description information and a designated placeholder; The first prompt word and the first embedding vector are input into the text generation model, so that the text generation model converts other text in the first prompt word except the designated placeholder into an embedding vector, and replaces the designated placeholder with the first embedding vector to obtain a second embedding vector corresponding to the first prompt word, and generates personalized copy that conforms to the interaction scenario and is oriented to the target user based on at least the second embedding vector.

3. The method according to claim 2, wherein the text generation model comprises a linear processing layer and an embedding layer; The embedding layer is used to map other texts in the first prompt word except the designated placeholder into an embedding vector according to the embedding vector mapping relationship learned in the pre-training process and the fine-tuning process of the text generation model; The linear processing layer is used to process the first embedding vector so that the processed first embedding vector and the embedding vector mapped by the embedding layer are aligned to the same vector space; The embedding layer is further used to replace the designated placeholder in the first prompt word with the processed first embedding vector.

4. According to the method of claim 2 or 3, the text generation model is used to generate personalized copy that conforms to the interaction scenario and is oriented to the target user according to the first embedding vector and the second embedding vector; in, The text generation model includes multiple Transformer structures; The second embedding vector is used as an input of a first Transformer structure among the multiple Transformer structures; The first embedding vector is input as part of a last Transformer structure in the plurality of Transformer structures.

5. According to the method of claim 4, the last Transformer structure in the plurality of Transformer structures comprises a self-attention structure and a cross-attention structure; The input of the self-attention structure is the output of the previous Transformer structure; The input of the cross-attention structure includes the output of the self-attention structure and the first embedding vector, the output of the self-attention structure is used to calculate the query vector in the cross-attention structure, and the first embedding vector is used to calculate the key vector and the value vector in the cross-attention structure.

6. The method according to claim 5, wherein the last Transformer structure in the plurality of Transformer structures further comprises a normalization layer, a fusion layer, and a feed-forward neural network; The normalization layer is used to normalize the output of the self-attention structure; The input of the cross-attention structure includes the output of the normalization layer and the first embedding vector, and the output of the normalization layer is used to calculate the query vector in the cross-attention structure; The fusion layer is used to fuse the output of the self-attention structure and the output of the cross-attention structure; The feedforward neural network is used to process the output of the fusion layer.

7. The method according to claim 1, If the target user is not a cold start user, the image preference information of the target user includes: Historical images related to the specified interactive behavior of the target user, and / or image style preference information of the target user; wherein the image style preference information of the target user is obtained by identifying historical images related to the specified interactive behavior of the target user using a trained image style description model; If the target user is a cold start user, the image preference information of the target user includes: historical images related to the specified interaction behavior of the reference user, and / or image style preference information of the reference user; wherein the image style preference information of the reference user is obtained by identifying the historical images related to the specified interaction behavior of the reference user using a trained image style description model; and the similarity between the first embedding vector corresponding to the reference user and the first embedding vector corresponding to the target user satisfies a preset similarity condition. 8 . The method according to claim 1 , wherein the second prompt word further includes user information of the target user, so that the personalized image generated by the image generation model conforms to the user information of the target user.

9. The method according to claim 1, further comprising: Using the trained image detection model to detect the content of the specified dimension in the personalized image, and obtain a detection result; If the detection result indicates that the personalized image has passed the detection, the personalized image is output to the user terminal of the target user.

10. A method for training a text generation model, comprising: Acquire multiple training samples, each of which includes input information and a label, wherein the input information includes scene description information of the interaction scene and user information of the interacted user, the label includes text in the image corresponding to the interaction scene, and the interacted user performs a specified interaction behavior on the image corresponding to the interaction scene; Inputting the prompt words generated by the input information in the training sample into the text generation model to be trained, so that the text generation model to be trained generates a predicted text that conforms to the interaction scenario and is oriented to the interacted user based on the prompt words, and training the text generation model to be trained with minimizing the difference between the predicted text and the label as the optimization goal; Among them, the trained text generation model is applied to the method described in any one of claims 1 to 9.

11. The method according to claim 10, wherein the user information comprises: User personalized information represented in text form, and a first embedding vector for characterizing a user portrait of the interacted user; Inputting the prompt words generated by the input information in the training sample into the text generation model to be trained includes: Inputting the prompt word generated by the user personalized information, the scene description information and the designated placeholder and the first embedding vector into a text generation model to be trained; The text generation model to be trained includes a linear processing layer, an embedding layer and multiple Transformer structures; The embedding layer is used to map other texts in the prompt word except the designated placeholder into an embedding vector based on a learnable embedding vector mapping relationship; wherein the learnable embedding vector mapping relationship further increases the mapping relationship between the personalized information of the interacted user and the embedding vector during this training process; The linear processing layer is used to process the first embedding vector so that the processed first embedding vector and the embedding vector mapped by the embedding layer are aligned to the same vector space; The embedding layer is further used to replace the designated placeholder in the prompt word with the processed first embedding vector; The second embedding vector is used as an input of a first Transformer structure among the multiple Transformer structures; The first embedding vector is input as part of a last Transformer structure in the plurality of Transformer structures.

12. An electronic device comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 11 by executing the executable instructions.

13. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.

14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Image generation method and reward model training method

    CN122391412A

  • Personalized image generation method and model training method

    WO2026179202A1