Text generation model training method and device, text generation method and device and electronic equipment

By acquiring features from sample images and recommended text, and using a text interaction prediction model to filter target recommended text, the problem of inconsistency when large models generate recommended text is solved, thus improving the training and delivery performance of the text generation model.

CN121808092APending Publication Date: 2026-04-07BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411383601.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-07

Smart Images

  • Figure CN121808092A_ABST
    Figure CN121808092A_ABST
Patent Text Reader

Abstract

The invention discloses a text generation model training method and device, a text generation method and device and electronic equipment, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a plurality of first sample images and a plurality of first recommendation texts corresponding to each first sample image; determining an image-text matching index of each first recommended text corresponding to each first sample image; inputting the first image feature of each first sample image and the first text feature of each first recommendation text into a text interaction prediction model, and performing text interaction prediction on each first recommendation text to obtain a text interaction index of each first recommendation text; based on the image-text matching index and the text interaction index, screening a target recommendation text of each first sample image; and based on the plurality of first sample images and the respective target recommendation texts of the plurality of first sample images, performing text generation training on the to-be-trained text generation model to obtain the text generation model. According to the scheme, the training effect of the text generation model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a text generation model training, text generation method, device, and electronic device. Background Technology

[0002] With the development of artificial intelligence technology, AI models have demonstrated excellent text generation capabilities. In resource promotion scenarios, large models are often used to generate text based on the material images corresponding to the resources to be promoted, thereby obtaining the recommended copy for the resources to be promoted.

[0003] However, existing large models generate copy based on general world knowledge learned from large-scale datasets. When generating recommendation copy for resources to be promoted, there are sometimes discrepancies between the source images and the content of the recommendation copy. For example, the copy may contain exaggerated or false descriptions that do not match the source images. At the same time, the copy content may not match the attention needs of the target audience, resulting in poor delivery performance. Summary of the Invention

[0004] This application provides a text generation model training, text generation method, apparatus, and electronic device, which can filter out sample texts that semantically match sample images and have good delivery effects (interaction effects), effectively improving the construction quality of the model training dataset, thereby improving the training effect of the text generation model, and further improving the image-text semantic matching degree and delivery effect (interaction effect) of the recommended text generated by the text generation model. The technical solution of this application is as follows:

[0005] On the one hand, a method for training a text generation model is provided, the method comprising:

[0006] Obtain multiple first sample images and multiple first recommendation texts corresponding to each first sample image;

[0007] Determine the first image feature of each first sample image, the first text feature of each first recommended text corresponding to each first sample image, and the image-text matching index corresponding to each first recommended text;

[0008] The first image feature and the first text feature of each first recommended text are input into the text interaction prediction model. Based on the first image-text fusion feature between the first text feature of each first recommended text and the first image feature, text interaction prediction is performed on each first recommended text to obtain the text interaction index corresponding to each first recommended text.

[0009] Based on the image-text matching index and the text interaction index, the target recommended text corresponding to each first sample image is selected from the plurality of first recommended texts;

[0010] Based on the plurality of first sample images and the target recommendation text corresponding to each of the plurality of first sample images, the text generation model to be trained is trained to obtain the text generation model.

[0011] On the other hand, a text generation method is provided, the method comprising:

[0012] Obtain the target recommended image;

[0013] The target recommendation image is input into a text generation model to generate text, thereby obtaining the target recommendation text corresponding to the target recommendation image;

[0014] The text generation model is trained based on the text generation model training method described above.

[0015] On the other hand, a text generation model training device is provided, the device comprising:

[0016] The first sample data acquisition module is used to acquire multiple first sample images and multiple first recommendation texts corresponding to each first sample image;

[0017] The image-text matching index determination module is used to determine the first image feature of each first sample image, the first text feature of each first recommended text corresponding to each first sample image, and the image-text matching index corresponding to each first recommended text;

[0018] The text interaction index determination module is used to input the first image feature and the first text feature of each first recommended text into the text interaction prediction model, and perform text interaction prediction on each first recommended text based on the first image-text fusion feature between the first text feature of each first recommended text and the first image feature to obtain the text interaction index corresponding to each first recommended text.

[0019] The text filtering module is used to filter the target recommended text corresponding to each first sample image from the plurality of first recommended texts based on the image-text matching index and the text interaction index;

[0020] The first model training module is used to train the text generation model to be trained based on the plurality of first sample images and the target recommendation text corresponding to each of the plurality of first sample images, so as to obtain the text generation model.

[0021] On the other hand, a text generation apparatus is provided, the apparatus comprising:

[0022] The target recommendation image acquisition module is used to acquire target recommendation images;

[0023] The text generation module is used to input the target recommendation image into the text generation model to generate text, thereby obtaining the target recommendation text corresponding to the target recommendation image;

[0024] The text generation model is trained using the text generation model training device described above.

[0025] On the other hand, an electronic device is provided, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the text generation model training method or text generation method as described above.

[0026] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor to implement the text generation model training method or text generation method as described above.

[0027] On the other hand, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text generation model training method or text generation method as described above.

[0028] This application provides a text generation model training method, text generation apparatus, and electronic device, which have the following technical effects:

[0029] In the application scenario of resource promotion, this application acquires multiple first sample images for resource promotion and multiple first recommended texts corresponding to each first sample image. It determines the first image features of each first sample image, the first text features of each first recommended text corresponding to each first sample image, and the image-text matching index corresponding to each first recommended text. Then, based on the first image-text fusion feature between the first text features of each first recommended text and the first image features of the corresponding first sample image, a text interaction prediction model is used to perform text interaction prediction for each first recommended text, obtaining the text interaction index corresponding to each first recommended text. This can effectively improve the accuracy of text interaction index prediction. Based on the image-text matching metrics and text interaction metrics corresponding to each first recommended text, target recommended texts with good semantic matching with the first sample image and good delivery effect (interaction effect) are selected from multiple first recommended texts. This effectively improves the accuracy of sample recommended text selection and the construction quality of the model training dataset. Then, based on multiple first sample images and their corresponding target recommended texts, the text generation model is trained to obtain a text generation model, which improves the text generation training effect of the text generation model, thereby improving the text generation quality of the text generation model and enhancing the image-text semantic matching degree and delivery effect (interaction effect) of the generated text. Attached Figure Description

[0030] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a schematic diagram of an application environment provided in an embodiment of this application;

[0032] Figure 2 This is a flowchart illustrating a text generation model training method provided in an embodiment of this application;

[0033] Figure 3 This is a flowchart illustrating the training process of a text interaction prediction model provided in an embodiment of this application;

[0034] Figure 4 This is a schematic diagram of the structure of a text quality scoring model provided in an embodiment of this application;

[0035] Figure 5This is a schematic diagram of a process provided in this application embodiment of training a text generation model based on multiple first sample images and the target recommendation text corresponding to each of the multiple first sample images to obtain a text generation model.

[0036] Figure 6 This is a schematic diagram of a text generation model training method provided in an embodiment of this application;

[0037] Figure 7 This is a flowchart illustrating a text generation method provided in an embodiment of this application;

[0038] Figure 8 This is a block diagram of a text generation model training device provided in an embodiment of this application;

[0039] Figure 9 This is a block diagram of a text generation apparatus provided in an embodiment of this application;

[0040] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0042] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification, claims and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products or devices.

[0043] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0044] To facilitate understanding of the embodiments of this application, the technical terms involved in this application will be briefly introduced below:

[0045] Fine-tuning refers to adjusting an existing model. Fine-tuning can save computational resources and time, and improve computational efficiency.

[0046] The solutions provided in this application involve technologies such as natural language processing and pre-trained models in artificial intelligence, which are specifically illustrated through the following embodiments:

[0047] The text generation method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, the environment may include a client 10 and a server 20, which can be indirectly connected via wireless communication. A relevant object (such as a user) can send a text generation request carrying a target recommendation image to the server 20 through the client 10. The server 20 responds to the text generation request by acquiring the target recommendation image, inputting the target recommendation image into a text generation model to generate text, obtaining the target recommendation text corresponding to the target recommendation image, and feeding back the target recommendation text to the client 10. The text generation model is obtained by training the text generation model on multiple first sample images and their corresponding target recommendation texts. The target recommendation text corresponding to each first sample image is selected from multiple first recommendation texts based on their respective image-text matching metrics and text interaction metrics. The text interaction metric for each first recommendation text is obtained by inputting the first image features of the first sample image corresponding to each first recommendation text and the first text features of each first recommendation text into a text interaction prediction model, and then performing text interaction prediction on each first recommendation text based on the first image-text fusion feature between the first text features and the first image features. It should be noted that, Figure 1 This is just one example.

[0048] The client can be a physical device such as a smartphone, computer (e.g., desktop computer, tablet computer, laptop computer), digital assistant, smart voice interaction device (e.g., smart speaker), smart wearable device, in-vehicle terminal, etc., or it can be software running on the physical device, such as a computer program. The operating system corresponding to the first client can be Android, iOS (a mobile operating system developed by Apple), Linux, Microsoft Windows, etc.

[0049] The server side can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server may include network communication units, processors, and memory, etc. The server side can provide backend services to the corresponding clients.

[0050] The aforementioned client 10 and server 20 can be used to build a system for text generation, which can be a distributed system.

[0051] It should be noted that the text generation model training method and text generation method provided in this application can be applied to both the client and the server, and are not limited to the embodiments described above.

[0052] The following describes a specific embodiment of a text generation model training method provided in this application. Figure 2 This is a flowchart illustrating a text generation model training method provided in an embodiment of this application. This application provides the operational steps described in the embodiments or flowchart, but based on conventional or non-inventive methods, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual systems or products, the methods can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment) as shown in the embodiments or drawings. Specifically, as... Figure 2 As shown, the method may include:

[0053] S201, Obtain multiple first sample images and multiple first recommendation texts corresponding to each first sample image.

[0054] In one specific embodiment, the first sample image can be a sample image containing sample recommendation elements. Specifically, the sample recommendation elements can be multimedia elements used to introduce the resource to be promoted. The multimedia elements can include, but are not limited to, text elements, image elements, etc.; the resource to be promoted can include, but are not limited to, products, applications, novels, and other resources that need to be promoted; each of the multiple first recommendation texts corresponding to the first sample image can be a descriptive text used to introduce the resource to be promoted corresponding to the first sample image.

[0055] In one optional embodiment, the first sample image can be input into a pre-trained large language model for text generation processing to obtain multiple first recommended texts corresponding to the first sample image; alternatively, multiple first sample images and multiple first recommended texts corresponding to each first sample image can be collected from the historical data of the resource recommendation system. This application does not impose any particular limitation on this.

[0056] S202, determine the first image feature of each first sample image, the first text feature of each first recommended text corresponding to each first sample image, and the image-text matching index corresponding to each first recommended text.

[0057] In one specific embodiment, the first image feature of each first sample image is the representation information of the corresponding first sample image, i.e., the feature vector; specifically, the first image feature can represent the element features of multimedia elements in the first sample image. Optionally, the first sample image can be input into a pre-trained sample image feature extraction model for image feature extraction processing to obtain the first image feature. Specifically, the model structure of the sample image feature extraction model can be set according to the actual application, such as a vision transformer (visual encoder).

[0058] In a specific embodiment, the first text feature of each first recommended text can be the representation information of the corresponding first recommended text, i.e., the feature vector; specifically, the first text feature can represent the promotional resource features described by the corresponding first recommended text; optionally, the first recommended text can be input into a pre-trained text feature extraction model for text feature extraction processing to obtain the first text feature; specifically, the model structure of the text feature extraction model can be set according to the actual application, such as RoBERTa (Robustly Optimized BERT Approach, a pre-trained natural language processing model).

[0059] In a specific embodiment, the sample image feature extraction model and the text feature extraction model can also be pre-trained for image-text matching, thereby mapping the image and text to a common feature representation space. During the training process, the model parameters are adjusted by the contrastive loss function so that similar images and texts are closer in the feature representation space, and dissimilar images and texts are farther apart in the feature representation space.

[0060] In one specific embodiment, the image-text matching index corresponding to each first recommended text can be used to characterize the semantic matching degree between the first recommended text and the first sample image corresponding to the first recommended text. Specifically, the first text features of the first recommended text and the first image features of the first sample image corresponding to the first recommended text can be input into the image-text matching prediction model to perform image-text matching prediction on the first recommended text and the first sample image corresponding to the first recommended text, thereby obtaining the image-text matching index corresponding to each first recommended text. Specifically, the image-text matching prediction model can determine the feature distance between the first text features of the first recommended text and the first image features of the first sample image corresponding to the first recommended text, and based on this feature distance, determine the image-text matching index corresponding to each first recommended text. In an optional embodiment, the first text features and the first image features can be represented as feature vectors, and the corresponding feature distance can include, but is not limited to, cosine distance (i.e., cosine similarity), Euclidean distance, and Manhattan distance.

[0061] S203, input the first image features and the first text features of each first recommended text into the text interaction prediction model, and perform text interaction prediction for each first recommended text based on the first image-text fusion feature between the first text features and the first image features of each first recommended text to obtain the text interaction index corresponding to each first recommended text.

[0062] In a specific embodiment, the text interaction metrics corresponding to each first recommended text are used to predict the interaction effect of the text delivery target on the first recommended text. Specifically, the text interaction metrics may include: text click metrics and text conversion metrics. For illustration, the text delivery target may be a user account. The text click metrics may include, but are not limited to: UV click-through rate, PV click-through rate, etc., and the text conversion metrics may include, but are not limited to: purchase rate, download rate, subscription rate, etc.

[0063] In a specific embodiment, the above-mentioned text interaction prediction model may include: a text-image feature fusion layer and an interaction prediction layer. The first image features and the first text features of each first recommended text are input into the text interaction prediction model. Based on the first text-image fusion feature between the first text features of each first recommended text and the first image features, text interaction prediction is performed on each first recommended text to obtain the text interaction index corresponding to each first recommended text, which may include:

[0064] S2031, input the first image features and the first text features of each first recommended text into the image-text feature fusion layer, perform cross-attention learning on the first image features and the first text features of each first recommended text to obtain the first image-text fusion feature corresponding to each first recommended text.

[0065] In a specific embodiment, the first image-text fusion feature corresponding to each first recommended text is a fusion feature, i.e., a feature vector, obtained by cross-attention learning of the first text feature of each first recommended text and the first image feature of the first sample image corresponding to each first recommended text. Specifically, the image-text feature fusion layer can be used to perform cross-attention learning on the first text feature of each first recommended text and the first image feature of the first sample image corresponding to each first recommended text. The model structure of the image-text feature fusion layer can be set according to the actual application. In an optional embodiment, during the cross-attention learning process on the first text feature of each first recommended text and the first image feature of the first sample image corresponding to each first recommended text, the first text feature can be used as a query vector, and the first image feature can be used as a key vector and a value vector. Attention is paid to the information in the first image feature that is semantically related to the first text feature, so that the model can learn which key information in the text is most valuable for predicting the interaction effect, and help the model accurately judge the interaction effect of the recommended copy.

[0066] S2032, input the first image-text fusion feature corresponding to each first recommended text into the interaction prediction layer to perform text interaction prediction for each first recommended text, and obtain the text interaction index corresponding to each first recommended text.

[0067] In a specific embodiment, the interaction prediction layer can perform text interaction prediction for each first recommended text based on the first image-text fusion feature corresponding to each first recommended text; specifically, the model structure of the interaction prediction layer can be set according to the actual application, such as a feedforward neural network.

[0068] As can be seen from the above embodiments, the first image features and the first text features of each first recommended text are input into the image-text feature fusion layer. Cross-attention learning is performed on the first image features and the first text features of each first recommended text to obtain the first image-text fusion feature corresponding to each first recommended text. Then, the first image-text fusion feature corresponding to each first recommended text is input into the interaction prediction layer to perform text interaction prediction for each first recommended text to obtain the text interaction index corresponding to each first recommended text. By focusing on the information in the first image features that is semantically related to the first text features, the accuracy of the first image-text fusion feature in representing the dependency relationship between image and text modalities is improved, thereby helping the model to accurately judge the interaction effect of the recommended text and improve the accuracy of text interaction prediction.

[0069] In a specific embodiment, such as Figure 3 As shown, the above text interaction prediction model is trained in the following way:

[0070] S301, acquire multiple second sample images, multiple second recommended texts corresponding to each second sample image, and preset interaction metrics corresponding to each of the multiple second recommended texts.

[0071] In one specific embodiment, the second sample image can be a sample image containing sample recommendation elements. Specifically, the sample recommendation elements can be multimedia elements used to introduce the resource to be promoted; each of the multiple second recommendation texts corresponding to the second sample image can be descriptive text used to introduce the resource to be promoted corresponding to the second sample image.

[0072] In one specific embodiment, the preset interaction metrics corresponding to each second recommended text can be pre-labeled metrics that reflect the interaction effect of the text delivery object on the second recommended text. Optionally, the preset interaction metrics may include: preset click metrics and preset conversion metrics.

[0073] In one specific embodiment, multiple second sample images, multiple second recommended texts corresponding to each second sample image, and historical interaction metrics corresponding to each of the multiple second recommended texts can be collected from the historical data of the resource recommendation system. The historical interaction metrics corresponding to each second recommended text are then used as preset interaction metrics for that second recommended text. Specifically, the historical interaction metrics for each second recommended text can be indicator information that reflects the interaction effect of the text delivery object on the second recommended text within a historical time period. Optionally, historical interaction metrics may include: historical click metrics and historical conversion metrics.

[0074] S302, determine the current sample image from multiple second sample images, and use the multiple second recommended texts corresponding to the current sample image as multiple current recommended texts.

[0075] In one specific embodiment, the current sample image can be the second sample image corresponding to the current training round; specifically, the current sample image can include at least one second sample image.

[0076] S303, determine the second image features of the current sample image and the second text features of each current recommended text.

[0077] Specifically, the second image feature of the current sample image can be the representation information of the current sample image (the element features of multimedia elements in the current sample image); the second text feature of each current recommended text can be the representation information of each current recommended text (the promotional resource features described by the first recommended text).

[0078] In a specific embodiment, the detailed determination of the second image features of the current sample image and the second text features of each current recommended text can be found in the detailed determination of the first image features of each first sample image and the first text features of each first recommended text corresponding to each first sample image, which will not be repeated here.

[0079] S304, input the second image features and the second text features of each current recommended text into the text interaction prediction model to be trained, and perform text interaction prediction on each current recommended text based on the second image fusion features between the second text features and the second image features of each current recommended text to obtain the predicted interaction index corresponding to each current recommended text.

[0080] In a specific embodiment, the second image features and the second text features of each current recommended text are input into the text interaction prediction model to be trained. Based on the second image-text fusion feature between the second text features and the second image features of each current recommended text, text interaction prediction is performed on each current recommended text to obtain a detailed breakdown of the predicted interaction index corresponding to each current recommended text. This can be referred to in the above-described method of inputting the first image features and the first text features of each first recommended text into the text interaction prediction model, and based on the first image-text fusion feature between the first text features and the first image features of each first recommended text, text interaction prediction is performed on each first recommended text to obtain a detailed breakdown of the text interaction index corresponding to each first recommended text. This will not be repeated here.

[0081] In a specific embodiment, the predicted interaction metric corresponding to each current recommended text can characterize the predicted interaction effect of the text delivery object with respect to the current recommended text, that is, the probability that the current recommended text is predicted to be a text with high interaction effect.

[0082] S305, Based on the predicted interaction index corresponding to each current recommended text and the preset interaction index corresponding to each current recommended text, determine the interaction classification loss information.

[0083] In one specific embodiment, the interaction classification loss information is used to characterize the difference between the predicted interaction metric and the preset interaction metric for each current recommended text. Specifically, the smaller the difference between the predicted interaction metric and the preset interaction metric for each current recommended text, the smaller the value of the interaction classification loss information; conversely, the larger the difference between the predicted interaction metric and the preset interaction metric for each current recommended text, the larger the value of the interaction classification loss information. Specifically, the interaction classification loss information between the predicted interaction metric and the preset interaction metric for each current recommended text can be calculated by combining a first preset loss function. Illustratively, the first preset loss function may include, but is not limited to, the cross-entropy loss function.

[0084] In one specific embodiment, an interaction index threshold can be preset. The current recommended text with an interaction index greater than the preset threshold is taken as a positive sample (i.e., text with high interaction effect), and the current recommended text with an interaction index less than or equal to the preset threshold is taken as a negative sample (i.e., text with low interaction effect), so that the text interaction prediction model to be trained can learn to distinguish between text with high interaction effect and text with low interaction effect.

[0085] S306, Based on the index ranking information corresponding to the preset interaction indexes of multiple currently recommended texts, perform ranking analysis on the predicted interaction indexes of multiple currently recommended texts to determine the interaction ranking loss information.

[0086] In one specific embodiment, the interaction ranking loss information is used to characterize the ranking difference between the predicted interaction metrics of multiple currently recommended texts and the preset interaction metrics. Specifically, the smaller the difference between the ranking relationship of the predicted interaction metrics of multiple currently recommended texts and the ranking relationship of the preset interaction metrics, the smaller the value of the interaction ranking loss information; conversely, the greater the difference between the ranking relationship of the predicted interaction metrics of multiple currently recommended texts and the ranking relationship of the preset interaction metrics, the larger the value of the interaction ranking loss information. Specifically, the interaction ranking loss information between the ranking relationship of the predicted interaction metrics of multiple currently recommended texts and the ranking relationship of the preset interaction metrics can be calculated by combining a second preset loss function, which can be set according to the actual application.

[0087] In a specific embodiment, when the number of multiple currently recommended texts is greater than 2, the multiple currently recommended texts can be combined in pairs to obtain multiple current recommended text groups. According to the index size of a preset interaction index, the two current recommended texts in each current recommended text group are sorted to obtain the first index sorting information of each current recommended text group. And according to the index size of the predicted interaction index, the two current recommended texts in each current recommended text group are sorted to obtain the second index sorting information of each current recommended text group. Based on a second preset loss function, the sorting loss information between the first index sorting information and the second index sorting information of each current recommended text group is determined. Then, the sorting loss information of multiple current recommended text groups is fused to obtain the interaction sorting loss information.

[0088] In an optional embodiment, the ranking loss information between the first indicator ranking information and the second indicator ranking information of each current recommended text group, based on the second preset loss function, can be expressed as the following formula:

[0089] L Rank =max(0,1-(P) a -Pb ))

[0090] Among them, P a P represents the predicted interaction metric of the current recommended text with the better preset interaction metric among the two current recommended texts in each current recommended text group. b This represents the predicted interaction metric of the current recommended text that has a worse corresponding preset interaction metric among two current recommended texts.

[0091] In an optional embodiment, the fusion processing of the ranking loss information of multiple current recommended text groups can be performed by mean averaging or other fusion methods, and this application does not impose any particular limitation on this.

[0092] S307. Based on the interaction classification loss information and interaction ranking loss information, the text interaction prediction model to be trained is trained to obtain the text interaction prediction model.

[0093] In a specific embodiment, training the text interaction prediction model to be trained based on interaction classification loss information and interaction ranking loss information to obtain the text interaction prediction model may include: updating the model parameters of the text interaction prediction model to be trained based on interaction classification loss information and interaction ranking loss information; performing the next round of iterative training based on the updated text interaction prediction model to be trained (i.e., repeatedly determining the current sample image from multiple second sample images to updating the model parameters of the text interaction prediction model to be trained based on interaction classification loss information and interaction ranking loss information), until a preset convergence condition is met, and the text interaction prediction model to be trained corresponding to the preset convergence condition is taken as the text interaction prediction model.

[0094] In a specific embodiment, the text interaction prediction model to be trained may include: a text-image feature fusion layer to be trained and an interaction prediction layer to be trained. Accordingly, the above-mentioned updating of the model parameters of the text interaction prediction model to be trained based on interaction classification loss information and interaction ranking loss information includes: updating the model parameters of the text-image feature fusion layer and the interaction prediction layer to be trained based on interaction classification loss information and interaction ranking loss information. Optionally, the model parameters can be updated by combining gradient descent method.

[0095] In a specific embodiment, the preset convergence conditions can be set according to the actual application, such as the number of iterations of training reaching a preset number, the interaction classification loss information and / or interaction ranking loss information being less than a specified threshold, etc., which can be set according to the training speed and model accuracy requirements.

[0096] As can be seen from the above embodiments, a text interaction ranking task is added on the basis of the text interaction classification task. Based on the interaction classification loss information corresponding to the text interaction classification task and the interaction ranking loss information corresponding to the text interaction ranking task, the text interaction prediction model to be trained is trained. This allows the text interaction prediction model to learn to distinguish between texts with high interaction effect and texts with low interaction effect, while ensuring that the predicted interaction index output by the model has strict ranking. This ensures that the predicted interaction index output by the model can be used for comparison during the inference stage, thereby improving the accuracy and interpretability of subsequent text selection.

[0097] In a specific embodiment, the text interaction prediction model to be trained based on the interaction classification loss information and interaction ranking loss information can be used to obtain the text interaction prediction model, which may include:

[0098] S3071, determine the first loss change index corresponding to the interactive classification loss information and the second loss change index corresponding to the interactive ranking loss information.

[0099] In a specific embodiment, the first loss change index can characterize the change in interaction classification loss information corresponding to the interaction result classification task. Optionally, the first loss change index may include: the first loss change rate of interaction classification loss information and the first loss change amount of interaction classification loss information. The second loss change index can characterize the change in interaction ranking loss information corresponding to the interaction result sorting task. Optionally, the second loss change index may include: the second loss change rate of interaction ranking loss information and the second loss change amount of interaction ranking loss information.

[0100] In one specific embodiment, the current sample image can be the second sample image corresponding to the current training round, and correspondingly, the interaction classification loss information can be the interaction classification loss information corresponding to the current training round, and the interaction ranking loss information can be the interaction ranking loss information corresponding to the current training round. In an optional embodiment, a first loss change rate can be determined based on the ratio of the interaction classification loss information corresponding to the previous round of the current training round to the interaction classification loss information corresponding to the two rounds before the current training round, and the first loss change rate can be used as a first loss change index. A second loss change rate can be determined based on the ratio of the interaction ranking loss information corresponding to the previous round of the current training round to the interaction ranking loss information corresponding to the two rounds before the current training round, and the second loss change rate can be used as a second loss change index.

[0101] S3072, normalize the first loss change index and the second loss change index to obtain the first loss weight corresponding to the first loss change index and the second loss weight corresponding to the second loss change index.

[0102] In a specific embodiment, the first loss weight is used to adjust the contribution of interaction classification loss information to model training, and the second loss weight is used to adjust the contribution of interaction ranking loss information to model training. Specifically, a preset normalization function can be used to normalize the first loss change index and the second loss change index. The preset normalization function can be set according to the actual application.

[0103] In an optional embodiment, the preset normalization function can be a normalization exponential function. Optionally, during the normalization process of the first loss change index and the second loss change index based on the normalization exponential function, a preset temperature coefficient can be used to control the sensitivity of weight adjustment. The larger the preset temperature coefficient, the lower the sensitivity of weight adjustment. Specifically, the preset temperature coefficient can be set according to the weight adjustment requirements in actual applications. (Illustratively, the first loss weight...) Second loss weight This can be expressed as the following formula:

[0104]

[0105] Where k is used to identify the current training round. This indicates the first loss change metric corresponding to the current training round. This represents the second loss change index corresponding to the current training round, and τ represents the preset temperature coefficient.

[0106] S3073, based on the first loss weight and the second loss weight, the interaction classification loss information and the interaction ranking loss information are weighted and fused to obtain the target loss information.

[0107] In a specific embodiment, the target loss information can reflect the prediction performance of the text interaction prediction model to be trained for the text interaction effect. The target loss information is the loss information after weighted fusion of interaction classification loss information and interaction ranking loss information. Specifically, the interaction classification loss information is weighted based on the first loss weight to obtain the weighted interaction classification loss information, and the interaction ranking loss information is weighted based on the second loss weight to obtain the weighted interaction ranking loss information. The weighted interaction classification loss information and the weighted interaction ranking loss information are added together to obtain the target loss information.

[0108] S3074, based on the target loss information, train the text interaction prediction model to be trained to obtain the text interaction prediction model.

[0109] As can be seen from the above embodiments, during the training process of the text interaction prediction model, the weights of the two loss information are dynamically adjusted according to the interaction classification loss information corresponding to the text interaction classification task and the interaction ranking loss information corresponding to the text interaction ranking task. This dynamically balances the learning progress of the two training tasks, avoids over-optimization or insufficient learning of one task, and improves the training effect of the text interaction prediction model. This ensures that while the text interaction prediction model learns to distinguish between texts with high interaction effect and texts with low interaction effect, the predicted interaction index output by the model has strict ranking.

[0110] S204, based on image-text matching metrics and text interaction metrics, select the target recommended text corresponding to each first sample image from multiple first recommended texts.

[0111] In a specific embodiment, the target recommended text corresponding to each first sample image can be the first recommended text that semantically matches the corresponding first sample image and has a good interactive effect.

[0112] In a specific embodiment, the above-mentioned selection of the target recommended text corresponding to each first sample image from multiple first recommended texts based on image-text matching indicators and text interaction indicators may include:

[0113] S2041, perform weighted fusion processing on the image-text matching index and the text interaction index corresponding to each first recommended text to obtain the text quality index corresponding to each first recommended text.

[0114] In a specific embodiment, the text quality index corresponding to each first recommended text can be used to measure the text quality of the corresponding first recommended text. Specifically, the larger the value of the text quality index, the higher the text quality of the first recommended text; the smaller the value of the text quality index, the lower the text quality of the first recommended text.

[0115] In a specific embodiment, the text quality metric score corresponding to each first recommended text can be expressed as the following formula:

[0116] score=α×S1+(1-α)×S2

[0117] Where S1 represents the image-text matching index corresponding to each first recommended text, S2 represents the text interaction index corresponding to each first recommended text, and α represents the weighting coefficient of the quality index.

[0118] Specifically, the quality index weighting coefficient is a hyperparameter that can be adjusted according to the text filtering bias in actual applications. For example, when more attention is paid to the semantic matching degree between recommended text and images during the text filtering process, the quality index weighting coefficient can be increased, while when more attention is paid to the text interaction effect during the text filtering process, the quality index weighting coefficient can be decreased.

[0119] S2042, Based on the text quality indicators corresponding to each of the multiple first recommended texts, determine the target recommended text among the multiple first recommended texts.

[0120] Optionally, the top recommended texts with a corresponding text quality index greater than the quality index threshold can be selected as target recommended texts. Alternatively, the top N top recommended texts with the highest corresponding text quality index can be selected as target recommended texts. Specifically, the quality index threshold and the number N can be set according to the text filtering accuracy requirements in the actual application.

[0121] As can be seen from the above embodiments, by performing weighted fusion processing on the image-text matching index and the text interaction index corresponding to each first recommended text, a text quality index corresponding to each first recommended text is obtained. Based on the text quality index corresponding to each of the multiple first recommended texts, the target recommended text among the multiple first recommended texts is determined. The combination of image-text semantic matching degree and text interaction effect can be used to screen sample texts, which can improve the accuracy and rationality of sample text screening, thereby improving the construction quality of the training dataset.

[0122] In a specific embodiment, a text quality scoring model can also be constructed that includes the above-mentioned sample image feature extraction model, the above-mentioned text feature extraction model, the above-mentioned image-text matching prediction model, the above-mentioned text interaction prediction model, and the index fusion model, such as... Figure 4 As shown, each first sample image and the corresponding first recommended text are input into the text quality scoring model to predict the text quality of the first recommended text, thereby obtaining the text quality index of the first recommended text.

[0123] S205, based on multiple first sample images and the target recommendation text corresponding to each of the multiple first sample images, perform text generation training on the text generation model to be trained, and obtain the text generation model.

[0124] In one specific embodiment, the text generation model can generate text based on a recommended image (an image containing recommended elements, which can be multimedia elements used to introduce the resource to be promoted), resulting in recommended text that semantically matches the recommended image and has good interactive effects. Specifically, the text generation model can be an artificial intelligence model obtained by training the text generation model to be trained based on multiple first sample images and their corresponding target recommended texts. The model structure of the text generation model to be trained can be set according to the actual application; illustratively, the text generation model to be trained can adopt a multimodal large model structure.

[0125] In a specific embodiment, such as Figure 5 As shown, the text generation model to be trained may include: a preset image feature extraction layer, a preset image semantic mapping layer, and a preset text generation layer. Based on multiple first sample images and their corresponding target recommendation texts, the text generation model to be trained is trained to generate text, resulting in a text generation model that may include:

[0126] S501, determine the current training image from multiple first sample images;

[0127] S502, Input the current training image into the preset image feature extraction layer to extract image features and obtain the first sample image features of the current training image;

[0128] S503, input the features of the first sample image into a preset image semantic mapping layer for image semantic mapping processing to obtain the features of the second sample image;

[0129] S504, Input the features of the second sample image into the preset text generation layer for text generation processing to obtain the predicted recommendation text corresponding to the current training image;

[0130] S505, Based on the target recommended text corresponding to the current training image and the predicted recommended text corresponding to the current training image, determine the text loss information;

[0131] S506, based on text loss information, fine-tunes the text generation model to be trained, and obtains the text generation model.

[0132] In one specific embodiment, the current training image can be the first sample image corresponding to the current training round; specifically, the current training image may include at least one first sample image.

[0133] In one specific embodiment, the first sample image feature of the current training image is the representation information of the current training image, i.e., the feature vector; specifically, the first sample image feature can represent the element features of multimedia elements in the current training image. Specifically, the first sample image feature can be the feature obtained after inputting the current training image into a preset image feature extraction layer for image feature extraction. The model structure of the preset image feature extraction layer can be set according to the actual application, such as a vision transformer (visual encoder).

[0134] In a specific embodiment, the second sample image features can be the image semantic representation information after image semantic mapping of the first sample image features, i.e., feature vectors. Specifically, the preset image semantic mapping layer is used to semantically align the semantic spaces of the preset image feature extraction layer and the preset text generation layer. The model structure of the preset image semantic mapping layer can be set according to the actual application. In a specific embodiment, the preset image semantic mapping layer can be a network layer based on the transformer structure, such as Query Transformer (Q-Former, a deep learning model structure based on cross-attention mechanism).

[0135] In a specific embodiment, the predicted recommendation text can be the predicted text generated by the preset text generation layer based on the second sample image features of the current training image; specifically, the preset text generation layer can be used to generate text based on the second sample image features of the current training image, and the model structure of the preset text generation layer can be set according to the actual application.

[0136] In a specific embodiment, the preset text generation layer can employ a pre-trained large language model. The input to the preset text generation layer further includes a preset prompt text, which is used to instruct the generation of recommended text that semantically matches the input image and has good interactive effects. Correspondingly, the above-mentioned inputting the second sample image features into the preset text generation layer for text generation processing to obtain the predicted recommended text corresponding to the current training image can include: inputting the second sample image features and the preset prompt text into the preset text generation layer for text generation processing to obtain the predicted recommended text corresponding to the current training image.

[0137] In a specific embodiment, text loss information can reflect the text generation performance of the text generation model to be trained. Specifically, the smaller the difference between the target recommended text and the predicted recommended text corresponding to the current training image, the smaller the text loss information; conversely, the larger the difference between the target recommended text and the predicted recommended text corresponding to the current training image, the larger the text loss information. Specifically, the image generation loss can be determined by combining a third preset loss function that can calculate the difference between the target recommended text and the predicted recommended text corresponding to the current training image. Specifically, the third preset loss function can be set according to the actual application. Illustratively, the third preset loss function can adopt cross-entropy classification loss to maximize the probability of generating the t-th text word given the input first sample image and the first t-1 text words.

[0138] In a specific embodiment, the above-mentioned fine-tuning of the text generation model to be trained based on text loss information to obtain a text generation model may include: updating the model parameters of the text generation model to be trained based on text loss information; performing the next round of iterative training based on the updated text generation model to be trained (i.e., repeatedly determining the current training image from multiple first sample images to updating the model parameters of the text generation model to be trained based on text loss information) until a preset convergence condition is met, and using the text generation model to be trained when the preset convergence condition is met as the text generation model.

[0139] In one specific embodiment, gradient descent can be combined with the process of updating the model parameters of the text generation model to be trained based on text loss information.

[0140] In an optional embodiment, updating the model parameters of the text generation model to be trained based on text loss information may include: freezing the model parameters of a preset text generation layer, and updating the model parameters of a preset image feature extraction layer and a preset image semantic mapping layer based on text loss information; or freezing the model parameters of the preset image feature extraction layer and the preset image semantic mapping layer, and updating the model parameters of the preset text generation layer based on text loss information; or updating the model parameters of the preset image feature extraction layer, the preset image semantic mapping layer, and the preset text generation layer based on text loss information.

[0141] In a specific embodiment, the preset convergence conditions can be set according to the actual application, such as the number of iterations of training reaching a preset number, the text loss information being less than a specified threshold, etc., which can be set according to the training speed and model accuracy requirements.

[0142] As can be seen from the above embodiments, based on the text loss information between the target recommended text and the predicted recommended text corresponding to the current training image, the parameters of the preset image feature extraction layer, preset image semantic mapping layer and preset text generation layer in the text generation model to be trained are fine-tuned so that the generation capability of the text generation model is aligned with the business characteristics of the resource recommendation scenario, and generates recommended text that is beautifully written, highly matched with the promotional material image and attractive to users.

[0143] To illustrate, taking the first sample image as the advertising image and the corresponding first recommended text as the advertising copy as an example, the aforementioned text interaction prediction model can be a copywriting interaction prediction model, and the aforementioned text generation model can be a copywriting generation model. Figure 6 This is a schematic diagram of a text generation model training method provided in an embodiment of this application. Specifically, as shown... Figure 6 As shown, the method may include:

[0144] S0: Collect historical ad images and corresponding high-quality (highly matched with historical ad images and attractive to users) historical texts from the ad system log data. Train the text interaction prediction model based on the historical ad images and high-quality historical texts to obtain the trained text interaction prediction model.

[0145] S1: Obtain sample ad images and preset prompt text, and input the sample ad images and preset prompt text into the copy generation model to generate copy, and obtain the sample copy set corresponding to the sample ad images;

[0146] S2: Input the sample ad images and sample copy sets into the trained copy interaction prediction model, perform copy quality prediction on each sample copy in the sample copy set, obtain the copy quality index of each sample copy, and select high-quality sample copy in the sample copy set based on the copy quality index.

[0147] S3: Based on the sample ad images and the corresponding high-quality sample copy, the above copy generation model is fine-tuned to obtain the fine-tuned copy generation model. The fine-tuned copy generation model is then put online and applied to generate high-quality ad copy.

[0148] As can be seen from the technical solutions provided in the above embodiments of this application, in the application scenario of resource promotion, multiple first sample images for resource promotion and multiple first recommended texts corresponding to each first sample image are obtained. The first image features of each first sample image, the first text features of each first recommended text corresponding to each first sample image, and the image-text matching index corresponding to each first recommended text are determined. Then, based on the first image-text fusion feature between the first text features of each first recommended text and the first image features of the corresponding first sample image, a text interaction prediction model is used to perform text interaction prediction on each first recommended text, thereby obtaining the text interaction index corresponding to each first recommended text. This can effectively improve the text interaction index. To improve the accuracy of target prediction, and based on the image-text matching index and text interaction index corresponding to each first recommended text, target recommended texts that semantically match the first sample image and have good delivery effect (interaction effect) are selected from multiple first recommended texts. This effectively improves the accuracy of sample recommended text selection and the construction quality of the model training dataset. Then, based on multiple first sample images and their corresponding target recommended texts, the text generation model is trained to obtain a text generation model, improving the text generation training effect of the text generation model, thereby improving the text generation quality of the text generation model and enhancing the image-text semantic matching degree and delivery effect (interaction effect) of the generated text.

[0149] This application also provides a text generation method, such as... Figure 7 As shown, the text generation method may include:

[0150] S701, Obtain the target recommended image.

[0151] S702, Input the target recommendation image into the text generation model to generate text, and obtain the target recommendation text corresponding to the target recommendation image; wherein, the above text generation model is obtained after training based on the text generation model training method described above.

[0152] In a specific embodiment, the target recommendation image can be a recommendation image that needs to generate corresponding recommendation text. The target recommendation image can be an image containing recommendation elements. Specifically, the recommendation elements can be multimedia elements used to introduce the resource to be promoted. Multimedia elements can include, but are not limited to, text elements, image elements, etc.; the resource to be promoted can include, but are not limited to, products, applications, novels, and other resources that need to be promoted; the target recommendation text corresponding to the target recommendation image can be descriptive text used to introduce the resource to be promoted corresponding to the target recommendation image.

[0153] In a specific embodiment, the text generation model can generate target recommended text based on the target recommended image. Specifically, the text generation model is obtained by training the text generation model to be trained based on multiple first sample images and their corresponding target recommended texts. The target recommended text corresponding to each first sample image is selected from multiple first recommended texts based on their respective image-text matching indices and text interaction indices. The text interaction indices of each first recommended text are obtained by inputting the first image features of the first sample image corresponding to each first recommended text and the first text features of each first recommended text into the text interaction prediction model, and then performing text interaction prediction on each first recommended text based on the first image-text fusion features between the first text features and the first image features.

[0154] In a specific embodiment, the input to the text generation model may further include: a preset prompt text, which is used to instruct the generation of recommended text that semantically matches the input image and has good interactive effects. Accordingly, the above-mentioned inputting the target recommended image into the text generation model to generate text and obtain the target recommended text corresponding to the target recommended image may include: inputting the target recommended image and the preset prompt text into the text generation model to perform text generation processing to obtain the target recommended text.

[0155] As can be seen from the above embodiments, the text generation model trained based on the above-described text generation model training method generates text for the target recommendation image, thereby improving the text generation training effect of the text generation model and thus improving the text generation quality of the text generation model, as well as the image-text semantic matching degree and delivery effect (interaction effect) of the generated text.

[0156] In a specific embodiment, the text generation model described above may include: an image feature extraction layer, an image semantic mapping layer, and a text generation layer. The process of inputting the target recommendation image into the text generation model to generate the target recommendation text corresponding to the target recommendation image may include:

[0157] S7021, Input the target recommendation image into the image feature extraction layer to extract image features and obtain the first target image feature of the target recommendation image.

[0158] In a specific embodiment, the target recommendation image is input into the image feature extraction layer for image feature extraction to obtain a detailed refinement of the first target image feature of the target recommendation image. This can be referred to in the above description of inputting the current training image into the preset image feature extraction layer for image feature extraction to obtain a detailed refinement of the first sample image feature of the current training image, which will not be repeated here.

[0159] S7022, the first target image features are input into the image semantic mapping layer for image semantic mapping processing to obtain the second target image features.

[0160] In one specific embodiment, the first target image features are input into the image semantic mapping layer for image semantic mapping processing to obtain a more detailed version of the second target image features. This can be referred to in the above description of inputting the first sample image features into the preset image semantic mapping layer for image semantic mapping processing to obtain a more detailed version of the second sample image features, which will not be repeated here.

[0161] S7023, input the features of the second target image into the text generation layer for text generation processing to obtain the target recommendation text.

[0162] In one specific embodiment, the features of the second target image are input into the text generation layer for text generation processing to obtain a detailed version of the target recommended text. This can be referred to in the above description of inputting the features of the second sample image into the preset text generation layer for text generation processing to obtain a detailed version of the predicted recommended text corresponding to the current training image, which will not be repeated here.

[0163] As can be seen from the above embodiments, the target recommendation image is input into the image feature extraction layer for image feature extraction to obtain the first target image feature of the target recommendation image. Then, the first target image feature is input into the image semantic mapping layer for image semantic mapping processing to obtain the second target image feature. The second target image feature is then input into the text generation layer for text generation processing to obtain the target recommendation text. By aligning the semantic space of the image feature extraction layer and the semantic space of the text generation layer, the text generation layer can understand the cross-modal second target image feature, thereby improving the text generation quality of the text generation layer.

[0164] This application also provides a text generation model training device, such as... Figure 8 As shown, the text generation model training device may include:

[0165] The first sample data acquisition module 810 is used to acquire multiple first sample images and multiple first recommendation texts corresponding to each first sample image;

[0166] The image-text matching index determination module 820 is used to determine the first image feature of each first sample image, the first text feature of each first recommended text corresponding to each first sample image, and the image-text matching index corresponding to each first recommended text.

[0167] The text interaction index determination module 830 is used to input the first image features and the first text features of each first recommended text into the text interaction prediction model, and based on the first image-text fusion feature between the first text features and the first image features of each first recommended text, perform text interaction prediction for each first recommended text to obtain the text interaction index corresponding to each first recommended text.

[0168] The text filtering module 840 is used to filter the target recommended text corresponding to each first sample image from multiple first recommended texts based on image-text matching indicators and text interaction indicators.

[0169] The first model training module 850 is used to train the text generation model to be trained based on multiple first sample images and the target recommendation text corresponding to each of the multiple first sample images, so as to obtain the text generation model.

[0170] In a specific embodiment, the above-mentioned text interaction prediction model may include: a text-image feature fusion layer and an interaction prediction layer, and the above-mentioned text interaction index determination module 830 may include:

[0171] The image-text feature fusion unit is used to input the first image features and the first text features of each first recommended text into the image-text feature fusion layer, and to perform cross-attention learning on the first image features and the first text features of each first recommended text to obtain the first image-text fusion feature corresponding to each first recommended text.

[0172] The text interaction prediction unit is used to input the first image-text fusion feature corresponding to each first recommended text into the interaction prediction layer to perform text interaction prediction for each first recommended text, and obtain the text interaction index corresponding to each first recommended text.

[0173] In one specific embodiment, the above-mentioned text interaction prediction model is trained using the following device:

[0174] The second sample data acquisition module is used to acquire multiple second sample images, multiple second recommended texts corresponding to each second sample image, and preset interaction indicators corresponding to each of the multiple second recommended texts.

[0175] The current sample image determination module is used to determine the current sample image from multiple second sample images, and to use the multiple second recommended texts corresponding to the current sample image as multiple current recommended texts;

[0176] The feature determination module is used to determine the second image features of the current sample image and the second text features of each current recommended text;

[0177] The prediction interaction index determination module is used to input the second image features and the second text features of each current recommended text into the text interaction prediction model to be trained. Based on the second image-text fusion feature between the second text features and the second image features of each current recommended text, the module performs text interaction prediction for each current recommended text to obtain the prediction interaction index corresponding to each current recommended text.

[0178] The interaction classification loss information determination module is used to determine the interaction classification loss information based on the predicted interaction index corresponding to each current recommended text and the preset interaction index corresponding to each current recommended text.

[0179] The interaction ranking loss information determination module is used to perform ranking analysis on the predicted interaction indicators of multiple current recommended texts based on the indicator ranking information corresponding to the preset interaction indicators of multiple current recommended texts, and determine the interaction ranking loss information.

[0180] The second model training module is used to train the text interaction prediction model based on interaction classification loss information and interaction ranking loss information, so as to obtain the text interaction prediction model.

[0181] In one specific embodiment, the second model training module described above may include:

[0182] The loss change index determination unit is used to determine the first loss change index corresponding to the interactive classification loss information and the second loss change index corresponding to the interactive ranking loss information.

[0183] The index normalization unit is used to normalize the first loss change index and the second loss change index to obtain the first loss weight corresponding to the first loss change index and the second loss weight corresponding to the second loss change index.

[0184] The loss fusion unit is used to perform weighted fusion of interaction classification loss information and interaction ranking loss information based on the first loss weight and the second loss weight to obtain the target loss information.

[0185] The text interaction prediction model training unit is used to train the text interaction prediction model to be trained based on the target loss information, so as to obtain the text interaction prediction model.

[0186] In one specific embodiment, the text filtering module 840 described above may include:

[0187] The indicator fusion unit is used to perform weighted fusion processing on the image-text matching indicator and the text interaction indicator corresponding to each first recommended text to obtain the text quality indicator corresponding to each first recommended text.

[0188] The target recommendation text determination unit is used to determine the target recommendation text among the multiple first recommendation texts based on the text quality index corresponding to each of the multiple first recommendation texts.

[0189] In one specific embodiment, the text generation model to be trained may include: a preset image feature extraction layer, a preset image semantic mapping layer, and a preset text generation layer; the first model training module 850 may include:

[0190] The current training image determination unit is used to determine the current training image from multiple first sample images;

[0191] The first sample image feature unit is used to input the current training image into the preset image feature extraction layer to extract image features and obtain the first sample image features of the current training image.

[0192] The second sample image feature unit is used to input the first sample image features into a preset image semantic mapping layer for image semantic mapping processing to obtain the second sample image features.

[0193] The predictive recommendation text unit is used to input the features of the second sample image into the preset text generation layer for text generation processing, and obtain the predicted recommendation text corresponding to the current training image.

[0194] The text loss information determination unit is used to input the current training image into a preset image feature extraction layer to extract image features and obtain the first sample image features of the current training image.

[0195] The text generation model training unit is used to fine-tune the text generation model to be trained based on text loss information, so as to obtain the text generation model.

[0196] It should be noted that the apparatus and method embodiments described above are based on the same inventive concept.

[0197] This application also provides a text generation device, such as... Figure 9 As shown, the text generation device may include:

[0198] The target recommendation image acquisition module 910 is used to acquire the target recommendation image;

[0199] The text generation module 920 is used to input the target recommendation image into the text generation model to generate text and obtain the target recommendation text corresponding to the target recommendation image.

[0200] The text generation model described above is obtained after training using the text generation model training device described above.

[0201] In one specific embodiment, the above-mentioned text generation model may include: an image feature extraction layer, an image semantic mapping layer, and a text generation layer, and the above-mentioned text generation module 920 may include:

[0202] The image feature extraction unit is used to input the target recommendation image into the image feature extraction layer to extract image features and obtain the first target image features of the target recommendation image.

[0203] The image semantic mapping unit is used to input the first target image features into the image semantic mapping layer for image semantic mapping processing to obtain the second target image features;

[0204] The text output unit is used to input the features of the second target image into the text generation layer for text generation processing to obtain the target recommendation text.

[0205] It should be noted that the apparatus and method embodiments described above are based on the same inventive concept.

[0206] This application provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction or at least one program segment. The at least one instruction or at least one program segment is loaded and executed by the processor to implement the text generation model training method or text generation method provided in the above method embodiments.

[0207] Furthermore, Figure 10 A schematic diagram of the hardware structure of an electronic device for implementing the text generation model training method or text generation method provided in the embodiments of this application is shown. The electronic device may participate in or include the text generation model training device or text generation device provided in the embodiments of this application. Figure 10 As shown, the electronic device 100 may include one or more processors 1002 (shown as 1002a, 1002b, ..., 1002n in the figure) (processor 1002 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 10 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, the electronic device 100 may also include... Figure 10The more or fewer components shown, or having the same Figure 10 The different configurations shown.

[0208] It should be noted that the aforementioned one or more processors 1002 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be wholly or partially integrated into any other element within the electronic device 100 (or mobile device). As involved in the embodiments of this application, the data processing circuit serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0209] The memory 1004 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the text generation model training method or the program instructions / data storage device corresponding to the text generation method described in the embodiments of this application. The processor 1002 executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, thereby implementing the aforementioned text generation model training method or text generation method. The memory 1004 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1004 may further include memory remotely located relative to the processor 1002, and these remote memories can be connected to the electronic device 100 via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0210] The transmission device 1006 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 100. In one example, the transmission device 1006 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In one embodiment, the transmission device 1006 may be a radio frequency (RF) module for wireless communication with the Internet.

[0211] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows a user to interact with the user interface of the electronic device 100 (or mobile device).

[0212] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program for implementing at least one instruction or at least one program, which is loaded and executed by the processor to implement the text generation model training method or text generation method provided in the above method embodiments.

[0213] Optionally, in this embodiment, the storage medium may be located in at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0214] Embodiments of this application also provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text generation model training method or text generation method as provided in the method embodiments.

[0215] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0216] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.

[0217] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and apparatus embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0218] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0219] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for training a text generation model, characterized in that, The method includes: Obtain multiple first sample images and multiple first recommendation texts corresponding to each first sample image; Determine the first image feature of each first sample image, the first text feature of each first recommended text corresponding to each first sample image, and the image-text matching index corresponding to each first recommended text; The first image feature and the first text feature of each first recommended text are input into the text interaction prediction model. Based on the first image-text fusion feature between the first text feature of each first recommended text and the first image feature, text interaction prediction is performed on each first recommended text to obtain the text interaction index corresponding to each first recommended text. Based on the image-text matching index and the text interaction index, the target recommended text corresponding to each first sample image is selected from the plurality of first recommended texts; Based on the plurality of first sample images and the target recommendation text corresponding to each of the plurality of first sample images, the text generation model to be trained is trained to obtain the text generation model.

2. The method according to claim 1, characterized in that, The text interaction prediction model was trained in the following way: Acquire multiple second sample images, multiple second recommended texts corresponding to each second sample image, and preset interaction metrics corresponding to each of the multiple second recommended texts; From the plurality of second sample images, determine the current sample image, and take the plurality of second recommended texts corresponding to the current sample image as the plurality of current recommended texts; Determine the second image features of the current sample image and the second text features of each current recommended text; The second image feature and the second text feature of each current recommended text are input into the text interaction prediction model to be trained. Based on the second image-text fusion feature between the second text feature of each current recommended text and the second image feature, text interaction prediction is performed on each current recommended text to obtain the predicted interaction index corresponding to each current recommended text. Based on the predicted interaction index corresponding to each current recommended text and the preset interaction index corresponding to each current recommended text, the interaction classification loss information is determined. Based on the index ranking information corresponding to the preset interaction indexes of the multiple currently recommended texts, the predicted interaction indexes of the multiple currently recommended texts are analyzed to determine the interaction ranking loss information. Based on the interaction classification loss information and the interaction ranking loss information, the text interaction prediction model to be trained is trained to obtain the text interaction prediction model.

3. The method according to claim 2, characterized in that, The step of training the text interaction prediction model based on the interaction classification loss information and the interaction ranking loss information to obtain the text interaction prediction model includes: Determine the first loss change index corresponding to the interaction classification loss information and the second loss change index corresponding to the interaction ranking loss information. The first loss change index and the second loss change index are normalized to obtain the first loss weight corresponding to the first loss change index and the second loss weight corresponding to the second loss change index. Based on the first loss weight and the second loss weight, the interaction classification loss information and the interaction ranking loss information are weighted and fused to obtain the target loss information; Based on the target loss information, the text interaction prediction model to be trained is trained to obtain the text interaction prediction model.

4. The method according to claim 1, characterized in that, The text interaction prediction model includes an image-text feature fusion layer and an interaction prediction layer. The first image features and the first text features of each first recommended text are input into the text interaction prediction model. Based on the first image-text fusion feature between the first text features of each first recommended text and the first image features, text interaction prediction is performed on each first recommended text to obtain the text interaction index corresponding to each first recommended text, including: The first image feature and the first text feature of each first recommended text are input into the image-text feature fusion layer, and cross-attention learning is performed on the first image feature and the first text feature of each first recommended text to obtain the first image-text fusion feature corresponding to each first recommended text. The first image-text fusion feature corresponding to each first recommended text is input into the interaction prediction layer to perform text interaction prediction on each first recommended text, thereby obtaining the text interaction index corresponding to each first recommended text.

5. The method according to claim 3, characterized in that, The step of filtering the target recommended text corresponding to each first sample image from the plurality of first recommended texts based on the image-text matching index and the text interaction index includes: The image-text matching index and the text interaction index corresponding to each first recommended text are weighted and fused to obtain the text quality index corresponding to each first recommended text. Based on the text quality index corresponding to each of the plurality of first recommended texts, the target recommended text is determined among the plurality of first recommended texts.

6. The method according to any one of claims 1 to 5, characterized in that, The text generation model to be trained includes: a preset image feature extraction layer, a preset image semantic mapping layer, and a preset text generation layer. The text generation model is trained by performing text generation training based on the plurality of first sample images and the target recommendation text corresponding to each of the plurality of first sample images, resulting in the following: The current training image is determined from the plurality of first sample images; The current training image is input into the preset image feature extraction layer to extract image features, thereby obtaining the first sample image features of the current training image; The first sample image features are input into the preset image semantic mapping layer for image semantic mapping processing to obtain the second sample image features; The features of the second sample image are input into the preset text generation layer for text generation processing to obtain the predicted recommended text corresponding to the current training image; Based on the target recommended text corresponding to the current training image and the predicted recommended text corresponding to the current training image, determine the text loss information; Based on the text loss information, the text generation model to be trained is fine-tuned to obtain the text generation model.

7. A text generation method, characterized in that, The method includes: Obtain the target recommended image; The target recommendation image is input into a text generation model to generate text, thereby obtaining the target recommendation text corresponding to the target recommendation image; The text generation model is trained based on the text generation model training method described in any one of claims 1 to 6.

8. The method according to claim 7, characterized in that, The text generation model includes: an image feature extraction layer, an image semantic mapping layer, and a text generation layer. The step of inputting the target recommendation image into the text generation model to generate the target recommendation text corresponding to the target recommendation image includes: The target recommended image is input into the image feature extraction layer for image feature extraction to obtain the first target image feature of the target recommended image; The first target image features are input into the image semantic mapping layer for image semantic mapping processing to obtain the second target image features; The second target image features are input into the text generation layer for text generation processing to obtain the target recommendation text.

9. A text generation model training device, characterized in that, The device includes: The first sample data acquisition module is used to acquire multiple first sample images and multiple first recommendation texts corresponding to each first sample image; The image-text matching index determination module is used to determine the first image feature of each first sample image, the first text feature of each first recommended text corresponding to each first sample image, and the image-text matching index corresponding to each first recommended text; The text interaction index determination module is used to input the first image feature and the first text feature of each first recommended text into the text interaction prediction model, and perform text interaction prediction on each first recommended text based on the first image-text fusion feature between the first text feature of each first recommended text and the first image feature to obtain the text interaction index corresponding to each first recommended text. The text filtering module is used to filter the target recommended text corresponding to each first sample image from the plurality of first recommended texts based on the image-text matching index and the text interaction index; The first model training module is used to train the text generation model to be trained based on the plurality of first sample images and the target recommendation text corresponding to each of the plurality of first sample images, so as to obtain the text generation model.

10. A text generation device, characterized in that, The device includes: The target recommendation image acquisition module is used to acquire target recommendation images; The text generation module is used to input the target recommendation image into the text generation model to generate text, thereby obtaining the target recommendation text corresponding to the target recommendation image; The text generation model is trained using the text generation model training device described in claim 9.

11. An electronic device, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the text generation model training method as described in any one of claims 1 to 6 or the text generation method as described in any one of claims 7 to 8.

12. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the text generation model training method as described in any one of claims 1 to 6 or the text generation method as described in any one of claims 7 to 8.

13. A computer program product, characterized in that, The computer program product includes at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the text generation model training method as described in any one of claims 1 to 6 or the text generation method as described in any one of claims 7 to 8.