Method and device for generating push information

By extracting the characteristics of users' long-term and short-term behavior information, and combining scene characteristics, using a large language model to generate personalized push information, the problem of inadequate personalized information push in the existing technology is solved, and the efficiency and user experience of information push are improved.

CN120050325APending Publication Date: 2025-05-27ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125360.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology is difficult to provide more personalized and targeted information push, resulting in poor user experience and low information push efficiency.

Method used

By obtaining the user's long-term behavior information and short-term behavior information, long-term and short-term characteristics are coded and extracted respectively, and personalized push information is generated through a large language model based on scene characteristics, and network model parameters are adjusted to maximize the evaluation score of push information.

Benefits of technology

It realizes in-depth exploration of user personalized characteristics, generates more targeted push information, and improves the effectiveness and user experience of information push.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050325A_ABST
    Figure CN120050325A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for generating push information, and the method for generating the push information provided under the technical concept of the specification can generate the push information under the current service scene for the user in a personalized manner through feature extraction and fusion in combination with long-term behavior information and short-term behavior information of the user. The personalized copywriting is generated by introducing mutual complementation of long-term and short-term interests of the user, so that the copywriting seen by the user in the same service scene is more targeted, and the effectiveness of information pushing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a method and device for generating push information. Background Art

[0002] So-called information push is usually a new technology that reduces information overload by actively transmitting information required by users on the Internet through certain technical standards or protocols triggered by conditions (such as reaching a regular time interval, opening relevant pages of an application, etc.). The push technology reduces the time spent searching on the network by automatically transmitting information to users and helps users efficiently discover valuable information. Here, push information can be used to describe the candidate information to be pushed to users or the information actually pushed to users. Push information can concisely and quickly describe the target to be pushed. For example, graphic advertisements such as public service advertisements and event promotion pages are a type of push information. Graphic advertisements usually include image information and advertising slogans of the target to be pushed. With the development of artificial intelligence, in order to improve the user experience, more and more products pay more attention to personalization, that is, information is pushed to users according to user portraits.

[0003] In conventional technologies, personalized information push usually selects one or more from pre-determined candidate information for different users and displays them to users as personalized push information. However, even when pushing information to users according to user portraits, different users in the same category may have different information preferences. Therefore, how to provide more personalized and targeted push information, improve the user experience, and more effectively achieve the purpose of information push is a technical problem worthy of research. Summary of the Invention

[0004] One or more embodiments of this specification describe a method and device for generating push information to solve one or more problems mentioned in the background art.

[0005] According to a first aspect, a method for generating push information is provided. The method includes: obtaining long-term behavior information and short-term behavior information of a first user; respectively encoding the long-term behavior information and the short-term behavior information through a first network and a second network to obtain corresponding long-term features and short-term features; processing the long-term features, the short-term features, and scene features through a large language model to generate push information for pushing information to the first user; wherein, the model parameters of the first network and the second network are adjusted with the goal of maximizing the evaluation score of the push information generated for sample users.

[0006] In one embodiment, the first network and the second network are networks formed by a first structure under different model parameters. The first structure includes an embedding layer, a depth association layer, and an adaptation layer. The embedding layer is used to convert corresponding behavior information into an embedding tensor. The depth association layer is used to mine the associations in the corresponding behavior information by processing the embedding tensor to obtain corresponding behavior representations. The adaptation layer is used to adaptively adjust the behavior representations to the embedding results of the large language model.

[0007] In one embodiment, the push information includes at least one of text, image, and audio.

[0008] In one embodiment, the long-term behavior information and the short-term behavior information each include at least one of the following user behaviors: click, browse, search, forward, support, oppose, favorite, submit form, location.

[0009] In one embodiment, the scoring of the push information generated for a sample user is implemented via a pre-trained reward model. The input data of the reward model in the pre-training stage is the push information generated for the sample user, the output data is the evaluation score, and the label data is determined for the push information generated for the sample user.

[0010] In a further embodiment, the reward model, the first network, and the second network are continuously updated in the following manner: obtaining the push information generated for multiple users and the operation information of the users on the push information; converting the operation information into the score labels of the reward model for reversely adjusting the model parameters in the reward model, the first network, and the second network.

[0011] In another further embodiment, the generating of the push information for information push to the first user by processing the long-term feature, the short-term feature, and the scenario feature through the large language model includes: processing the long-term feature, the short-term feature, and the scenario feature through the large language model to output multiple candidate push information; using the reward model to score each candidate push information, and selecting the candidate push information that meets the predetermined conditions as the push information for information push to the first user according to the obtained evaluation scores.

[0012] In a further embodiment, the predetermined conditions include: the corresponding evaluation score is greater than a predetermined threshold, or the corresponding evaluation score is ranked within a predetermined number in the order from largest to smallest.

[0013] In one embodiment, the scenario feature is pre-extracted based on at least one of the following description information of the current business scenario: business domain, business object, and target to be pushed.

[0014] According to a second aspect, there is provided an apparatus for generating push information, the apparatus comprising:

[0015] An acquisition unit configured to acquire the long-term behavior information and short-term behavior information of a first user;

[0016] An encoding unit configured to encode the long-term behavior information and the short-term behavior information via a first network and a second network respectively to obtain corresponding long-term features and short-term features;

[0017] A generation unit configured to generate push information for information push to the first user by processing the long-term features, the short-term features, and the scenario features through a large language model; wherein, the model parameters of the first network and the second network are adjusted with the goal of maximizing the evaluation score of the push information generated for the sample users.

[0018] According to a third aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method of the first aspect.

[0019] According to a fourth aspect, there is provided a computing device, including a memory and a processor. An executable code is stored in the memory, and when the processor executes the executable code, the method of the first aspect is implemented.

[0020] Through the apparatus and method provided in the embodiments of the present specification, in the case of information push, the long-term behavior information and short-term behavior information of the user are respectively subjected to feature extraction through different networks to obtain corresponding long-term features and short-term features. Then, the long-term features and short-term features of a single user and the scenario features of the current business scenario are together used by the generation module to generate push information for the corresponding user. Among them, the network for extracting user features is determined by maximizing the evaluation score of the generated push information. Since the long-term features and short-term features of the user complement each other, the personalized features of the user are deeply mined, personalized copywriting is generated, so that the copywriting seen by the user for the same business scenario is more targeted, and the effectiveness of information push is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 A schematic diagram of a specific implementation architecture under the technical concept of this specification;

[0023] Figure 2 A schematic diagram of a model architecture for generating push information according to an embodiment of this specification is shown;

[0024] Figure 3 A schematic flowchart of generating push information according to an embodiment of this specification is shown;

[0025] Figure 4 A structural block diagram of a device for generating push information according to an embodiment of this specification is shown. Detailed implementation manners

[0026] The solution provided in this specification will be described below with reference to the accompanying drawings.

[0027] Figure 1 A schematic diagram of a specific implementation architecture of this specification is shown. As Figure 1 shown, in the implementation architecture provided in this specification, it includes a business terminal corresponding to the business party, a user terminal corresponding to the user, and a server. Both the business terminal and the user terminal can be various terminal devices that can perform human-computer interaction with the user, such as a smart phone, a laptop computer, a tablet computer, a smart watch, etc. The business terminal is the terminal used by the business party. In Figure 1 the shown implementation architecture, the business party can be various business entities that push relevant information of the target to be pushed to the user, such as merchants, enterprises and institutions, event organizers, etc. Correspondingly, the target to be pushed can be, for example, a commodity, a policy regulation to be publicized, an event invitation, etc. The user terminal can be the terminal device corresponding to any user who may receive the push information. The user terminal can install and run various terminal applications. For example, information push applications, news applications, shopping applications, food delivery platform applications, Q&A platform applications, etc. The server can provide business support for the terminal applications. For example, it provides business support for news applications, etc. The server can have a computing platform set locally, or be connected to a computing platform on other devices, for determining the business content provided to the user terminal. For example, the server of an information push application can generate personalized push information through the computing platform and push it to the user terminal. For example Figure 1 in, under the same trigger condition (such as the user opening the application or the business party actively broadcasting), information 1 can be pushed to user 1, and information 2 can be pushed to user 2, and so on.

[0028] It can be understood that in Figure 1 the shown architecture, the server and the business terminal can be different terminals, or can be the same terminal (that is, the business party itself is the service party that provides services for the user terminal), and the identities of the business party and the user can also be interchanged. For example, in an information push service, the business party provides the target to be pushed as the service provider, and in another information push service, the business party can receive the push information as a user. Additionally,Figure 1 The architecture shown is only a specific example. In practice, the number of business terminals and user terminals can be any possible number, such as 100,000, 100 million, etc., which are not limited here.

[0029] In this specification, a personalized push message composed of at least one of text, image, and audio is generated in an information push scenario.

[0030] In conventional technologies, examples of personalized information push solutions include:

[0031] Convert the user's personalized information into natural language and directly input it into a high-quality large model, such as GPT4, etc., to let the large model generate corresponding personalized copywriting. Since it depends on the capabilities of external large models and directly uses external application programming interfaces (APIs), there may be privacy risks;

[0032] When generating copywriting with a personalized writing style, some linguistic features of the user's historical text are statistically analyzed and used to guide the model to generate copywriting with a personalized writing style. The linguistic features in this solution pay more attention to the writing style;

[0033] Use traditional technologies such as CVAE, taking the personalized information as a condition to control the text generation process. In this solution, traditional technologies such as CVAE are difficult to extend to other data outside the training data;

[0034] And so on.

[0035] This specification provides a technical concept, which combines a large language model, uses the user's long-term features as a supplement to the user's short-term features, and together with the scenario features as the prompt information for the large model, and a large language model (which can also be other pre-trained language models with similar functions) generates personalized push messages (which can include at least one of text, pictures, videos, animations, etc.).

[0036] Among them, the large language model (LLM), that is, the large-scale language model, or the large model for short, usually has a large number of parameters, such as in the order of billions. The large language model is a natural language processing model based on deep learning. It can learn the grammar and semantics of natural language, and thus can generate human-readable text. Due to its huge corpus, the large language model can be used as a pre-trained model for various language processing scenarios, such as question-and-answer scenarios, information push scenarios, text or image generation scenarios, and so on. These scenarios cover various application fields, such as the marketing field of products or services, the news broadcast field, the event promotion field, and so on. It can be understood that the scenarios listed here are roughly divided according to their uses. In practice, these scenarios may overlap. For example, the text or image generation scenario can occur alone or be embedded in other scenarios, such as in the question-and-answer scenario or the information push scenario.

[0037] Under the technical concept of this specification, the extraction models of user long-term features and user short-term features are adjusted according to the loss determined by the generated push information. In this way, the long-term and short-term user behavior information can be used to complement each other to generate more targeted personalized push information and improve the effectiveness of information push.

[0038] The following describes the technical concept of this specification in detail with reference to the embodiments shown in the accompanying drawings.

[0039] Figure 2 It shows a schematic diagram of the model architecture for generating push information proposed in the embodiment of this specification. In Figure 2 In the shown schematic diagram of the model architecture, the content shown within the dotted connection lines or dotted boxes is optional content, that is, in some embodiments, it may not be included in the model architecture or be replaced by other structures with similar functions. Figure 2 In the schematic of, the model architecture of this specification at least includes two major parts: a feature extraction network and an information generation model. Optionally, a feature fusion network can also be performed between the two parts. Each part is not clearly divided in Figure 2 and will be described in detail below.

[0040] In the feature extraction network, it includes a user feature extraction network and a business scenario feature extraction network.

[0041] The business scenario features therein can be features for characterizing at least one piece of information among the business domain, business object, target to be pushed, etc. Specifically, the business domain can be, for example, the application domain to which specific services such as product push, news push, activity (such as public welfare activities, etc.) push belong. The business object can be, for example, the object targeted by the information push (such as office workers), the effect expected by the service (such as 40% of users can accept, etc.), etc. The target to be pushed is, for example, the specific push object (what is pushed) involved in services such as products, news, activities, etc. The business scenario feature network can adopt various extraction networks in conventional technologies, which will not be elaborated here. In an alternative embodiment, the business scenario feature network can be the (Embedding layer) of a large language model. At this time, the scenario description information can be input into the embedding layer of the large language model, and the embedding layer of the large language model processes it to obtain scenario features (such as vector representations). The scenario features of the same business scenario can be reused among different users and in different information pushes of the same user.

[0042] In this specification, user features can include long-term features and short-term features. Long-term features can be features extracted from the long-term behavior information of users, and short-term features can be features extracted from the short-term behavior information of users. User behavior information can be various operation information that users perform through terminal devices, such as the pages browsed, page browsing duration, form items clicked, products purchased, etc. In an alternative embodiment, user behavior information can also include the location information reached by the terminal device, such as path information, stay information at various locations (staying means maintaining a position for more than a predetermined time period), etc. User long-term behavior information can be information collected on the behavior of users over a relatively long time period (such as three months, denoted as the first time period), and user short-term behavior information can be information collected on the behavior of users over a relatively short time period (such as three days, denoted as the second time period).

[0043] In this specification, the process of extracting user features from user long-term behavior information and short-term behavior information is similar. In one embodiment, the long-term feature extraction network and the short-term feature extraction network can, for example, have the same network architecture and independent model parameters. Here, the model parameters are undetermined parameters during the training process. After training is completed, the model parameters generally do not change with the change of input data during the feature extraction process.

[0044] Taking the long-term feature extraction network as an example, it at least includes: an encoding module and an adaptation layer. Among them, the encoding module is used to encode the long-term behavior information to obtain the corresponding long-term behavior representation. Such as Figure 2As shown, the encoding module can be implemented based on the Embedding layer. Embedding can be a process of representing characters or words through vectors, such as through one-hot representation, word2Vec, or a pre-trained Embedding neural network. The embedding tensor obtained from the Embedding layer can be used to determine the user's long-term behavior representation. The adaptation layer can be used to adapt the encoded long-term behavior representation to the input requirements of the large language model in the generation module. For example, the large language model usually processes numerical values in the range of 0 to 1 (the numerical values of each feature dimension in the input representation are between 0 and 1). If the values of each dimension of the representation output by the encoding module are between 1 and 10, the values of each dimension can be mapped to the range of 0 to 1 to adapt to the input requirements of the large language model. The output of the adaptation layer can be used as the long-term features for extracting the user's long-term behavior information.

[0045] In an alternative implementation, a depth correlation layer can be further provided after the Embedding layer of the encoding module to deeply explore the correlation between user behaviors through the processing of each dimension of the embedding tensor. The depth correlation layer can be set as a Transformer layer or the like to mine the correlation information in the long-term behavior information.

[0046] In a possible design, preprocessing can be performed before embedding the long-term behavior information. Preprocessing is an auxiliary process for encoding, and the data after preprocessing can be better used for feature extraction or mining. In one embodiment, preprocessing can be a process of extracting numerical features from the behavior information described in natural language, such as the page ID browsed (such as the string representing the page ID obtained through URL query parameters), the browsing duration (such as 60 seconds), etc. In another embodiment, preprocessing can be to convert the long-term user behavior information into the form of natural language through a large language model or manually. For example, the key-value data in the operation log is flattened, and then the long-term user information in the obtained natural language form is refined through a large language model to obtain more concise user behavior description information.

[0047] The process of extracting features from the user's short-term behavior information is similar to that of the user's long-term behavior information, which will not be elaborated here. The features of the user's short-term behavior information can be denoted as the user's short-term features.

[0048] The generation module can be used to generate push information based on the user's long-term features, short-term features, and scenario features. The push information here can be the information pushed to the current user. The push information can include at least one of text, image, and audio. The image here can be generally understood as files in various formats containing images, such as pictures, animations, videos, etc., where animations and videos can also be combined with audio. The generation module can be implemented, for example, through various models that can achieve the generation goal, such as the Bert model for generating text, the Generative Adversarial Network (GAN) for generating images, etc.

[0049] In the embodiments of this specification, the generation module can be implemented based on a large language model. At this time, the user's long-term features, short-term features, and scenario features can be used together as the prompt information (token) of the large language model, and the large language model can generate personalized push information for the current user accordingly. It can be understood that a large language model is usually a pre-trained model that outputs information according to the prompt information. In an alternative embodiment, the large language model can be fine-tuned with sample data during the training process according to the actual business scenario.

[0050] In a possible design, the long-term features, short-term features, and scenario features can be fused in a way such as concatenation or summation through a fusion layer as the input of the generation module.

[0051] The above describes the model architecture for generating push information under the implementation architecture of this specification. It can be understood that during the model training process, a large language model is usually a pre-trained language model with a huge number of parameters. Therefore, in the case where the generation module is implemented by a large language model, the model parameters to be adjusted can only include the model parameters of the user feature extraction part.

[0052] In order to adjust the model parameters, during the model training process, the push information generated for the sample user can be evaluated and scored to obtain the evaluation score as a label. The scoring process can be implemented manually or by introducing a reward model (or called a scoring model).

[0053] Taking the introduction of the reward model as an example, its purpose is to evaluate the personalized information generated by the generation module to obtain the corresponding evaluation score to guide the adjustment of the model parameters of the user feature extraction part. It is easy to understand that the input of the reward model can be the personalized push information generated by the generation module, and the output is the evaluation score. The evaluation score here can be considered to include the model loss of generating the personalized push information. Generally, the higher the evaluation score, the smaller the model loss, and the lower the evaluation score, the greater the model loss.

[0054] Since the evaluation score reflects the relevance of the generated personalized push information to the current user, the training of the reward model can be carried out during the training process of the model for generating push information. For example, it can be connected behind the generation model to reflect the personalized characteristics of the user. During the predetermined update period (such as the first 50) at the initial stage of training, the evaluation scores of each personalized push information generated for the sample users can be manually marked and used as sample labels to supervise the generation results of the push information. Thus, the model parameters of the user feature extraction part and the reward model can be adjusted through the reverse transmission of the gradient. After the training of the reward model meets the end condition, the model parameters of the reward model can be fixed, and the evaluation scores given by the reward model can be used to continue adjusting the model parameters of the user feature extraction part. Here, the end condition can be that the gradient of the model parameters of the reward model is less than a predetermined value (such as 0.0001) in multiple consecutive training periods (such as 10 periods), or the update amplitude (the difference before and after the update) of the model parameters is less than a predetermined value (such as 0.01), and so on.

[0055] The trained model for generating push information includes at least a user feature extraction network and a generation module, and can be used to generate push information in personalized information push in specific business scenarios. Among them, in a single business scenario, each user can share a piece of scenario feature data. For the current user, the long-term feature and short-term feature are respectively extracted through the user feature extraction network (including the network for processing long-term behavior information and the network for processing short-term behavior information), and then combined with the scenario feature, and the generation module generates personalized push information to push to the current user. The following combines Figure 3 the process shown to describe the information push process.

[0056] Figure 3 shows a schematic diagram of the process for generating push information according to an embodiment. The execution subject of this process can be a computer, device, server with certain computing capabilities. More specifically, for example, it is Figure 1 the computing platform shown. When the information push condition is met, personalized push information is generated and pushed to the user through the Figure 3 process shown. Among them, the information push condition can be, for example: the arrival of the predetermined push moment (such as the fixed time point for pushing the daily step data of the user's friends circle in sports every day), the user opens the predetermined page (such as the user opens the browsing page of a certain commodity), the activity initiator initiates a business push application (such as a public welfare organization initiates a push application for a public welfare activity for citizens), and so on.

[0057] Such as Figure 3As shown in the figure, the process for generating push information provided in the embodiments of this specification may include: Step 301, obtaining the long-term behavior information and short-term behavior information of the first user; Step 302, respectively processing the above long-term behavior information and short-term behavior information via the first network and the second network to obtain corresponding long-term features and short-term features; Step 303, processing the above long-term features, short-term features, and scenario features through a large language model to generate push information for pushing information to the first user; wherein, the model parameters of the first network and the second network are adjusted with the goal of maximizing the evaluation score of the push information generated for the sample users.

[0058] First, in step 301, obtain the long-term behavior information and short-term behavior information of the first user.

[0059] It can be understood that the first user can be any network user who meets the information push conditions. The network here can be the Internet, or the telecommunications network of a telecommunications operator (such as the GSM network corresponding to the SIM card), which is not limited here. The first user can correspond to a client of the current information push service scenario and the corresponding client identifier, such as the user ID of the application, the SIM card (card number or ISMI, etc.).

[0060] The user's behavior information may include information on various behaviors performed by the user through the device running the corresponding client, such as at least one of click, browse, search, forward, support, oppose, favorite, submit form, location, etc. obtained from the device operation log. More specifically, the browse information may include information such as the browsed page and the page browsing duration, and the submit form information may include shopping information, click information, etc. When the device running the client is a portable mobile device (such as a mobile phone), the user's location information may also correspond to one or more of movement trajectory information, stay information, stay duration information, etc.

[0061] The long-term behavior information of the first user may be the behavior information within a relatively long first time period, and the short-term behavior information may be the behavior information within a relatively short second time period. It can be understood that both the first time period and the second time period can be time periods from the current time point, and their lengths are relative. When the second time period is shorter than the first time period, it can be set according to the specific service scenario. For example, when the second time period is 1 day, the first time period can be 10 days, one week, one month, etc. When the second time period is 10 days, one week, etc., the first time period can be 1 month, 3 months, etc. In an optional embodiment, the ratio of the second time period to the first time period is less than a predetermined value (such as 0.1, etc.).

[0062] The user behavior information obtained here can be in the form of natural language description, keyword description, key-value pair (Key-Value), etc., which is not limited here. The long-term behavior information of the same user can include short-term behavior information.

[0063] Next, through step 302, the above long-term behavior information and short-term behavior information are processed via the first network and the second network respectively, and the corresponding long-term features and short-term features are extracted.

[0064] Here, the feature extraction network for processing the long-term behavior information of the user to obtain the corresponding long-term features can be denoted as the first network, and the feature extraction network for processing the short-term behavior information of the user to obtain the corresponding short-term features can be denoted as the second network. The first network and the second network together constitute Figure 2 the user feature extraction network in. The first network and the second network process the long-term behavior information and the short-term behavior information respectively. Through the differences in the processed information, the long-term behavior information and the short-term behavior information complement each other, deepening the understanding of the user's interests.

[0065] As Figure 2 shown, the first network and the second network can have the same or similar network architectures (denoted as the first structure), such as each including an encoding module and an adaptation layer (Linear). The encoding module at least includes an embedding layer for obtaining its tensor representation based on the embedding of the corresponding behavior information. This tensor representation can be transformed into the input data of the large model through the adaptation layer, so as to adapt to the data characteristics acceptable by the large model, such as each value being between 0 and 1, and the number of values in a single dimension of the tensor not exceeding 512, etc. The adaptation layer can include at least one of a fully connected neural network and an activation function.

[0066] In an alternative implementation, the first network and the second network can also include a deep association layer, which performs deep association mining between features after embedding the long-term and short-term behavior information of the user respectively. The deep association layer is implemented, for example, through a Transformer layer under the attention mechanism.

[0067] In a possible design, at least one of the first network and the second network may further include a preprocessing layer. The purpose of preprocessing is to facilitate encoding. In one embodiment, the preprocessing layer may be a process of extracting numerical features from the behavioral information described in natural language, such as the page ID browsed (e.g., the string representing the page ID obtained through URL query parameters), the browsing duration (e.g., 60 seconds), etc. In another embodiment, the preprocessing may be to convert the long-term user behavioral information into the form of natural language through a large language model or manually. For example, the Key-Value data in the operation log is flattened, and then the long-term user information in the obtained natural language form is streamlined through the large language model to obtain more concise user behavioral description information.

[0068] In summary, the first network processes the long-term behavioral information of the first user, and the second network processes the short-term behavioral information of the first user, respectively obtaining the long-term features and short-term features of the first user. The long-term features and short-term features of the first user are the personalized features of the first user. It can be understood that for the same user, at different trigger times, the long-term behavioral information and short-term behavioral information may both be different, and the corresponding long-term features and short-term features are also both different.

[0069] Among them, the model parameters of the first network and the second network can refer to the description of the process of adjusting the model parameters in the previous text about Figure 2 and are determined with the goal of maximizing the evaluation score of the information pushed to the sample user. Among them, in the model training stage, the evaluation score can be determined by manual marking or generated according to a predetermined rule. In a possible design, the evaluation score can also be determined by a reward model. In an alternative embodiment, in the online prediction stage, online data can be continuously collected, such as operation information on whether the user accepts the generated push information (such as clicking, purchasing, submitting a form page, etc.), so as to determine the evaluation score and update the model parameters in the first network and the second network. In some specific examples, the evaluation score can also be determined according to the result of whether the user accepts, and then the model parameters in the reward model are updated.

[0070] Then, in step 303, the above long-term features, short-term features, and scenario features are processed by a large language model to generate push information for pushing information to the first user.

[0071] Among them, the scenario features of the current business network scenario can be extracted based on the description information of the current business scenario. The description information of the current business scenario can be information describing at least one of the business domain, business object, target to be pushed, etc. The business domain is, for example, product push, news push, event (such as public welfare activities, etc.) push, etc. The business object is, for example, the object targeted by the information push (such as office workers), the effect expected by the business (such as 40% of users can accept, etc.), etc. The target to be pushed is, for example, products, news, events, etc. The business scenario features can be extracted using various extraction methods in conventional technologies, which will not be elaborated here. In an alternative embodiment, the scenario description information can be input into the embedding layer of the large language model, and the embedding layer of the large language model processes it to obtain the scenario features.

[0072] The scenario features can be shared by multiple users, and in the same business scenario, under the triggering conditions of each user at different times, there is no need to update. Therefore, the scenario features can be determined in the current process or can be determined in advance and applied to each user until they are updated.

[0073] The generated push information can be information containing at least one of text, image, and audio, such as text information copywriting, graphic and text promotion pages, videos, animations, etc. Since the large language model has good natural language processing capabilities and can be used as a generation model for text and images, in the implementation architecture of this specification, the large language model can be used as a generation module to process the long-term features, short-term features of the first user, and the scenario features of the current business scenario, so as to generate push information for pushing information to the first user.

[0074] When using the large language model to generate push information, the long-term features, short-term features of the first user, and the scenario features of the current business scenario can be used as the prompt information (token) of the large language model and input into the large language model respectively. Alternatively, the above long-term features, short-term features, and scenario features can be fused in a way such as concatenation, summation, and attention mechanism (attention network) processing and then input into the large language model as the prompt information, and the large language model generates the corresponding push information.

[0075] It should be noted that the large language model can generate several push messages to push to the first user, or generate multiple candidate push messages, and select several push messages that meet the predetermined conditions from them to push to the first user. It can also generate candidate push messages multiple times, and select several push messages that meet the predetermined conditions from them to push to the first user. In one embodiment, the predetermined conditions can be determined based on the evaluation scores obtained by scoring each candidate push message by the reward model, such as: the corresponding evaluation score is greater than the predetermined threshold, or the corresponding evaluation scores are arranged in the top predetermined number in descending order. By combining the manual labeling of the reward model and RLHF (Reinforcement Learning from Human Feedback), problems such as poor copywriting generation effect and large model hallucinations can be effectively solved.

[0076] Reviewing the above process, the method for generating push messages provided under the technical concept of this specification can combine the long-term behavior information and short-term behavior information of the user, and through feature extraction and fusion, generate personalized push messages for the user in the current business scenario. By introducing the long-term and short-term interests of the user, personalized push messages are generated, making the copywriting seen by the user for the same business scenario more targeted. The feature extraction of the user's long-term and short-term behavior information is obtained through a deep model, rather than directly giving an embedding layer in the large model's soft prompt, which deepens the understanding of long-term and short-term interests. Moreover, the extraction networks for the user's long-term features and short-term features are the main networks for parameter adjustment during the training process, making the generated push messages more in line with the user's characteristics. In short, the technical solution provided in this specification can improve the effectiveness of information push.

[0077] According to an embodiment of another aspect, there is also provided a device for generating push messages. This device can be set in a computer, terminal, server with certain computing capabilities, more specifically, such as Figure 1 the computing platform shown. Figure 4 FIG. shows a device 400 for generating push messages according to an embodiment. As Figure 4 shown, the device 400 may include:

[0078] An obtaining unit 401, configured to obtain the long-term behavior information and short-term behavior information of the first user;

[0079] An encoding unit 402, configured to encode the above long-term behavior information and the above short-term behavior information through a first network and a second network respectively to obtain corresponding long-term features and short-term features;

[0080] A generating unit 403, configured to process the above-mentioned long-term features, the above-mentioned short-term features, and the scenario features through a large language model to generate push information for pushing information to the above-mentioned first user; wherein, the model parameters of the above-mentioned first network and the above-mentioned second network are adjusted with the goal of maximizing the evaluation score of the push information generated for the sample user.

[0081] In one embodiment, the above-mentioned first network and the above-mentioned second network are networks formed by a first structure under different model parameters. The first structure includes an embedding layer, a depth association layer, and an adaptation layer. The embedding layer is used to convert the corresponding behavior information into an embedding tensor. The depth association layer is used to mine the associations in the corresponding behavior information through the processing of the embedding tensor to obtain the corresponding behavior representation. The adaptation layer is used to adaptively adjust the behavior representation to the embedding result of the large language model.

[0082] In one embodiment, the push information includes at least one of text, image, and audio.

[0083] In one embodiment, the long-term behavior information and the short-term behavior information each include at least one of the following user behaviors: click, browse, search, forward, support, oppose, favorite, submit form, location.

[0084] In one embodiment, the scoring of the push information generated for the sample user is implemented by a pre-trained reward model. The input data of the reward model in the pre-training stage is the push information generated for the sample user, the output data is the evaluation score, and the label data is determined through the evaluation of the push information generated for the sample user.

[0085] In a further embodiment, the apparatus 400 may further include an updating unit, configured to continuously update the reward model, the first network, and the second network in the following manner: obtain the push information generated for multiple users, and the operation information of the users on the push information; convert the operation information into a score label of the reward model for reversely adjusting the model parameters in the reward model, the first network, and the second network.

[0086] In another further embodiment, the generating unit 403 may further be configured to: process the long-term features, the short-term features, and the scenario features through a large language model to output multiple candidate push information; use the reward model to score each candidate push information, and select the candidate push information that meets the predetermined conditions as the push information for pushing information to the first user according to the obtained evaluation score.

[0087] In a further embodiment, the predetermined conditions include: the corresponding evaluation score is greater than a predetermined threshold, or the corresponding evaluation score is ranked within the top predetermined number in descending order.

[0088] In one embodiment, the scenario features are pre-extracted based on at least one of the following description information of the current business scenario: business domain, business object, and target to be pushed.

[0089] It should be noted that Figure 4 The shown apparatus 400 corresponds to Figure 3 the described method, Figure 3 and the corresponding descriptions in the illustrated method embodiments are equally applicable to the apparatus 400 and will not be elaborated herein.

[0090] According to an embodiment of another aspect, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in connection with Figure 3 etc.

[0091] According to an embodiment of still another aspect, a computing device is further provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in connection with Figure 3 etc. is implemented. Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0092] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above descriptions are only specific embodiments of the technical concept of this specification and are not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions in the embodiments of this specification should be included in the protection scope of the technical concept of this specification.

Claims

1. A method for generating push information, the method comprising: Obtaining long-term behavior information and short-term behavior information of the first user; The long-term behavior information and the short-term behavior information are encoded respectively via the first network and the second network to obtain corresponding long-term features and short-term features; Processing the long-term features, the short-term features, and the scene features through a large language model to generate push information for pushing information to the first user; The model parameters of the first network and the second network are adjusted with the goal of maximizing the evaluation score of the push information generated for the sample user.

2. The method of claim 1, wherein: The first network and the second network are networks formed by the first structure under different model parameters. The first structure includes an embedding layer, a deep association layer, and an adaptation layer. The embedding layer is used to convert the corresponding behavior information into an embedding tensor. The deep association layer is used to mine the association in the corresponding behavior information by processing the embedding tensor to obtain the corresponding behavior representation. The adaptation layer is used to adaptively adjust the behavior representation to the embedding result of the large language model.

3. The method of claim 1, wherein: The push information includes at least one of text, image, and audio.

4. The method of claim 1, wherein: The long-term behavior information and the short-term behavior information each include at least one of the following user behaviors: click, browse, search, forward, support, oppose, favorite, submit form, and location.

5. The method of claim 1, wherein: The scoring of the push information generated for the sample users is achieved through a pre-trained reward model, wherein the input data of the reward model in the pre-training stage is the push information generated for the sample users, the output data is the evaluation score, and the label data is determined through the evaluation of the push information generated for the sample users.

6. The method of claim 5, wherein: The reward model and the first network and the second network are continuously updated in the following manner: Obtain push information generated for multiple users, as well as user operation information on the push information; The operation information is converted into a score label of the reward model for reversely adjusting the reward model and model parameters in the first network and the second network.

7. The method of claim 5, wherein: The step of processing the long-term features, the short-term features, and the scene features by using a large language model to generate push information for pushing information to the first user includes: Processing the long-term features, the short-term features, and the scene features through a large language model, and outputting a plurality of candidate push information; The reward model is used to score each candidate push information, and the candidate push information that meets the predetermined conditions is selected as the push information for information push to the first user according to the obtained evaluation score.

8. The method of claim 7, wherein: The predetermined conditions include: the corresponding evaluation score is greater than a predetermined threshold, or the corresponding evaluation score is arranged in descending order within a predetermined number.

9. The method of claim 1, wherein: The scenario feature is pre-extracted based on at least one of the following descriptive information of the current business scenario: business domain, business object, and target to be pushed.

10. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 9 is implemented.