Information Push Method, Device, Computer Equipment and Storage Medium

By extracting and combining coarse and fine-grained features with intermediate steps, the method addresses the accuracy issues in information push systems, enhancing feature learning and delivery precision.

CN116226501BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110898411.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-05
Publication Date
2025-07-15
Estimated Expiration
2041-08-05

AI Technical Summary

Technical Problem

In the prior art, due to the small number of positive samples of certain features, the probability estimate model has poor semantic learning effect on these features, which affects the accuracy of information push.

Method used

The information features are divided into coarse-grained features with a large number of tail value samples and fine-grained features with a small number of tail value samples. The first feature of coarse-grained features is extracted, and the second feature is obtained by combining intermediate features. Multi-level feature learning is carried out through the probability estimate model to improve the feature representation effect.

Benefits of technology

Through multi-level feature learning, the accuracy of information push is improved, especially when processing features with a small number of positive samples, the effect of information push is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116226501B_ABST
    Figure CN116226501B_ABST
Patent Text Reader

Abstract

This application relates to an information push method, device, computer device, and storage medium, and relates to the field of Internet application technologies. The method includes: extracting information features of at least two candidate information, where the information features include coarse-grained features and fine-grained features; obtaining first features of at least two candidate information respectively based on the coarse-grained features; the first features are obtained based on intermediate features; obtaining second features of at least two candidate information respectively based on the information features and the intermediate features; obtaining target information based on the first features and the second features and pushing it. This application can synchronously learn multi-level feature representations from information features, thereby improving the representation effect of the extracted features on information. When subsequent information selection and pushing are performed through the extracted first features and second features, the accuracy of information pushing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet application technologies, and particularly to an information push method, apparatus, computer device, and storage medium. Background Art

[0002] In the field of Internet information push, in order to improve the accuracy of information push, an information push platform usually uses a machine learning model to select information to be pushed.

[0003] In the related art, when information push is required, the information push platform inputs the information features of each pushable information into a trained probability prediction model to obtain the predicted probability of a specified event occurring after the information push is displayed (such as the predicted conversion rate), and then determines the information to be pushed this time according to the predicted conversion rates of each piece of information.

[0004] However, in the information push scenario, the number of positive samples of some features is small, resulting in poor semantic learning effect of such features by the probability prediction model in the related art, thus affecting the accuracy of information push. Summary of the Invention

[0005] Embodiments of this application provide an information push method, apparatus, computer device, and storage medium, which can improve the accuracy of information push. The technical solution is as follows.

[0006] On the one hand, an information push method is provided. The method includes:

[0007] Extract the information features of at least two candidate pieces of information. The information features include coarse-grained features and fine-grained features; the number of tail value samples of the coarse-grained features is greater than the number of tail value samples of the fine-grained features;

[0008] Based on the coarse-grained features of each of the at least two candidate pieces of information, obtain the first feature of each of the at least two candidate pieces of information; the first feature is obtained based on intermediate features; the intermediate features are obtained by extracting the coarse-grained features;

[0009] Based on the information features of each of the at least two candidate pieces of information and the intermediate features of each of the at least two candidate pieces of information, obtain the second feature of each of the at least two candidate pieces of information;

[0010] Based on the first feature of each of the at least two candidate pieces of information and the second feature of each of the at least two candidate pieces of information, obtain the target information among the at least two candidate pieces of information;

[0011] Push the target information.

[0012] On the other hand, an information push device is provided, and the device includes:

[0013] An information feature extraction module, configured to extract information features of at least two candidate information respectively, where the information features include coarse-grained features and fine-grained features; the number of tail value samples of the coarse-grained features is greater than the number of tail value samples of the fine-grained features;

[0014] A first feature acquisition module, configured to acquire first features of the at least two candidate information respectively based on the coarse-grained features of the at least two candidate information; the first features are acquired based on intermediate features; the intermediate features are obtained by extracting the coarse-grained features;

[0015] A second feature acquisition module, configured to acquire second features of the at least two candidate information respectively based on the information features of the at least two candidate information and the intermediate features of the at least two candidate information;

[0016] An information acquisition module, configured to acquire target information among the at least two candidate information based on the first features of the at least two candidate information and the second features of the at least two candidate information;

[0017] An information push module, configured to push the target information.

[0018] In a possible implementation manner, the first feature acquisition module is configured to,

[0019] Extract features from the coarse-grained features of the first candidate information to obtain m first intermediate features of the first candidate information; the first candidate information is any one of the at least two candidate information; m is a positive integer;

[0020] Acquire a first weight of the m first intermediate features based on the coarse-grained features of the first candidate information;

[0021] Acquire the first feature of the first candidate information based on the m first intermediate features and the first weight of the m first intermediate features.

[0022] In a possible implementation manner, the second feature acquisition module is configured to,

[0023] Extract features from the information features of the first candidate information to obtain n second intermediate features of the first candidate information; n is a positive integer;

[0024] Acquire a second weight of the n second intermediate features and a second weight of the m first intermediate features based on the information features of the first candidate information;

[0025] Obtain the second feature of the first candidate information based on the second weights of the n second intermediate features, the second weights of the m first intermediate features, the n second intermediate features, and the m first intermediate features.

[0026] In a possible implementation, the second feature obtaining module is configured to obtain the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the information feature of the first candidate information and the popularity vector of the first candidate information; the popularity vector is used to indicate the historical conversion times of the corresponding candidate information.

[0027] In a possible implementation, the second feature obtaining module is configured to,

[0028] Concatenate the information feature of the first candidate information and the popularity vector of the first candidate information to obtain the first concatenated feature of the first candidate information;

[0029] Based on the first concatenated feature of the first candidate information, obtain the second weights of the n second intermediate features and the second weights of the m first intermediate features.

[0030] In a possible implementation, the information obtaining module is configured to,

[0031] Fuse the first features of the at least two candidate information and the second features of the at least two candidate information respectively to obtain the fused features of the at least two candidate information respectively;

[0032] Based on the fused features of the at least two candidate information respectively, obtain the estimated event probabilities of the at least two candidate information respectively; the estimated event probability is the estimated probability of a specified event occurring after the corresponding information is displayed;

[0033] Based on the estimated event probabilities of the at least two candidate information respectively, obtain the target information among the at least two candidate information.

[0034] In a possible implementation, the information obtaining module is configured to,

[0035] Based on the information feature of the second candidate information, obtain the third weight of the second feature of the second candidate information; the second candidate information is any one of the at least two candidate information;

[0036] Based on the third weight of the second feature of the second candidate information, fuse the first feature of the second candidate information and the second feature of the second candidate information to obtain the fused feature of the second candidate information.

[0037] In a possible implementation, the information acquisition module is configured to

[0038] Obtain a third weight of the second feature of the second candidate information based on the information feature of the second candidate information and the popularity vector of the second candidate information.

[0039] In a possible implementation, the information acquisition module is configured to

[0040] Concatenate the information feature of the second candidate information and the popularity vector of the second candidate information to obtain a second concatenated feature of the second candidate information;

[0041] Obtain a third weight of the second feature of the second candidate information based on the second concatenated feature of the second candidate information.

[0042] In a possible implementation, the information acquisition module is configured to

[0043] Perform weighted processing on the second feature of the second candidate information based on the second feature of the second candidate information to obtain a weighted feature of the second candidate information;

[0044] Add the weighted feature of the second candidate information to the first feature of the second candidate information to obtain a fusion feature of the second candidate information.

[0045] In a possible implementation, the first feature acquisition module is configured to process the coarse-grained features of the at least two candidate information respectively through a first extraction branch in the probability prediction model to obtain the first features of the at least two candidate information respectively;

[0046] The second feature acquisition module is configured to process the information features of the two candidate information respectively and the intermediate features of the at least two candidate information respectively through a second extraction branch in the probability prediction model to obtain the second features of the at least two candidate information respectively;

[0047] The information acquisition module is configured to process the first features of the at least two candidate information respectively and the second features of the at least two candidate information respectively through a fusion branch in the probability prediction model to obtain the fusion features of the at least two candidate information respectively;

[0048] The information acquisition module is further configured to process the fusion features of the at least two candidate information respectively through a prediction branch in the probability prediction model to obtain the predicted event probabilities of the at least two candidate information respectively.

[0049] In a possible implementation, the device further includes:

[0050] The information feature extraction module is further configured to extract the information features of the sample information before extracting the information features of at least two candidate information respectively;

[0051] The first feature acquisition module is further configured to process the coarse-grained features of the sample information through the first extraction branch to obtain the first feature of the sample information;

[0052] The second feature acquisition module is further configured to process the information features of the sample information and the intermediate features of the sample information through the second extraction branch to obtain the second feature of the sample information;

[0053] The information acquisition module is further configured to process the first feature of the sample information and the second feature of the sample information through the fusion branch to obtain the fusion feature of the sample information;

[0054] The information acquisition module is further configured to process the fusion features of the at least two candidate information respectively through the estimation branch in the probability estimation model to obtain the estimated event probability of the sample information;

[0055] The device further includes:

[0056] The loss function value acquisition module is configured to obtain a loss function value based on the estimated event probability of the sample information, the event probability label of the sample information, and the training weight of the sample information; the training weight is inversely correlated with the popularity of the sample information; the event probability label is used to indicate the labeled probability of the specified event occurring after the sample information is displayed;

[0057] The parameter update module is configured to update the parameters of the probability estimation model based on the loss function value.

[0058] On the other hand, a computer device is provided, which includes a processor and a memory. At least one computer instruction is stored in the memory, and the at least one computer instruction is loaded and executed by the processor to implement the above information push method.

[0059] On another hand, a computer-readable storage medium is provided. At least one computer instruction is stored in the storage medium, and the at least one computer instruction is loaded and executed by a processor to implement the above information push method.

[0060] In another aspect, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions that are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above information pushing method.

[0061] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0062] The information features are divided into coarse-grained features with a large number of tail value samples and fine-grained features with a small number of tail value samples. The first feature is extracted from the coarse-grained features, and the second feature is extracted from the complete information features. Moreover, when extracting the second feature, the intermediate feature between the coarse-grained feature and the first feature is also combined for the extraction of the second feature. In this way, multi-level feature representations can be synchronously learned from the information features, thereby improving the representation effect of the extracted features on the information. When subsequent information selection and pushing are performed through the extracted first feature and second feature, the accuracy of information pushing can be improved.

[0063] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] The drawings here are incorporated into the specification and form a part of the specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0065] Figure 1 is a system composition diagram of an information pushing system related to various embodiments of the present application;

[0066] Figure 2 is a flowchart of an information pushing method shown according to an exemplary embodiment;

[0067] Figure 3 is Figure 2 a schematic diagram of the tail value of the feature related to the shown embodiment;

[0068] Figure 4 is a flowchart of an information pushing method shown according to an exemplary embodiment;

[0069] Figure 5 is Figure 4 a schematic diagram of the model architecture related to the shown embodiment;

[0070] Figure 6 is Figure 4 a schematic diagram of weighted summation of expert information related to the shown embodiment;

[0071] Figure 7 is Figure 4 Schematic diagram of obtaining the second weight related to the illustrated embodiment;

[0072] Figure 8 is Figure 4 Schematic diagram of the comparative experiment results related to the illustrated embodiment;

[0073] Figure 9 is Figure 4 Schematic diagram of the ablation experiment results related to the illustrated embodiment;

[0074] Figure 10 Block diagram of the structure of an information push device shown according to an exemplary embodiment;

[0075] Figure 11 Schematic diagram of the structure of a computer device shown according to an exemplary embodiment. Detailed implementation manners

[0076] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0077] Before describing the various embodiments shown in the present application, several concepts related to the present application will be introduced first.

[0078] 1) AI (Artificial Intelligence)

[0079] AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable them to have functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0080] 2) ML (Machine Learning, machine learning)

[0081] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0082] 3) Big data

[0083] Big Data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a massive, high-growth, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery ability, and process optimization ability. With the advent of the cloud era, big data has attracted more and more attention. Big data requires special technologies to effectively process a large amount of data that can tolerate the elapsed time. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.

[0084] Please refer to Figure 1 , which shows the system configuration diagram of an information push system involved in various embodiments of the present application. As Figure 1 shown, the system includes several user terminals 120 and a server 140.

[0085] The user terminal 120 can be a smart phone, a tablet computer, an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a smart wearable device, a laptop computer, a desktop computer, and so on.

[0086] The user terminal 120 is connected to the server 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.

[0087] Among them, the server 140 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0088] Optionally, the server 140 can include a server for implementing the information delivery platform 142. Optionally, the server 140 can also include a server for implementing the information push platform 144.

[0089] Optionally, the information delivery platform 142 has the functions of pushing and maintaining the information delivery interface, and receiving the information delivered by the information deliverer.

[0090] Among them, the above information is information that can be displayed in multiple different application programs at the same time, such as advertisements. In the embodiments of the present application, the advertisement can include a non-economic advertisement and an economic advertisement. The non-economic advertisement refers to an advertisement that is not for profit, also known as an effect advertisement, such as various announcements, notices, and statements of government administrative departments, social institutions, and even individuals; the economic advertisement is also known as a commercial advertisement, which refers to an advertisement for profit.

[0091] Optionally, the information push platform 144 has the functions of managing and maintaining messages, and pushing information to the user terminal.

[0092] It should be noted that the servers for implementing the information delivery platform 142 and the information push platform 144 can be independent servers from each other, or they can also be implemented in the same physical server.

[0093] Optionally, the system may further include a management device (not shown in the figure), which is connected to the server 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.

[0094] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or proprietary data communication technologies can be used to replace or supplement the above data communication technologies.

[0095] Figure 2 is a schematic flowchart of an information push method shown according to an exemplary embodiment. This method can be executed by a computer device. For example, the computer device can be a server, where the server can be the server 140 in the above Figure 1 illustrated embodiment. As Figure 2 shown, the information push method may include the following steps.

[0096] Step 201, extract the information features of at least two candidate pieces of information respectively. The information features include coarse-grained features and fine-grained features; the number of tail value samples of the coarse-grained features is greater than the number of tail value samples of the fine-grained features.

[0097] Among them, the tail value of the above features refers to one or more feature values corresponding to the categories arranged at the end position after sorting the sample information according to the feature values of a certain feature in descending order of the number of information in each category. For example, it can be the feature value corresponding to the category arranged at the end position and the number of corresponding information is less than the quantity threshold. That is to say, the number of the above tail value samples is the number of sample information in the category arranged at the end position.

[0098] For example, please refer to Figure 3 , which shows the schematic diagram of the tail value of the features involved in the embodiments of the present application. As Figure 3 shown, taking the information as an advertisement as an example, Figure 3 includes the sample quantity histogram 31 corresponding to feature 1 (such as advertisement ID (Identity)), the sample quantity histogram 32 corresponding to feature 2 (such as advertiser), and the sample quantity histogram 33 of feature 3 (such as the product type corresponding to the advertisement).

[0099] Among them, Figure 3 in the sample quantity histogram corresponding to the advertisement ID, the ordinate can represent the click / exposure / conversion times of the advertisement corresponding to the advertisement ID, and the abscissa represents each advertisement ID. Since a lot of new advertisements will be generated on the Internet, therefore, in the sample quantity histogram corresponding to the advertisement ID, the sample quantity corresponding to each advertisement ID at the tail is extremely small. For example, the maximum / minimum / average value of the sample quantity corresponding to each advertisement ID at the tail is less than 100. Therefore, the feature of advertisement ID can be listed as a fine-grained feature.

[0100] For another example, Figure 3 in the sample quantity histogram corresponding to the advertiser, the ordinate can represent the click / exposure / conversion times of the advertisement corresponding to the advertiser, and the abscissa represents the ID of each advertiser. Since there are many small advertisers on the Internet and the number of advertisements placed by these advertisers is small, therefore, in the sample quantity histogram corresponding to the advertiser ID, the sample quantity corresponding to each advertiser at the tail is extremely small. For example, the maximum / minimum / average value of the sample quantity corresponding to each advertiser at the tail is less than 100. Therefore, the feature of advertiser ID can also be listed as a fine-grained feature.

[0101] For another example, Figure 3In the histogram of the number of samples corresponding to product types, the vertical axis can represent the number of clicks / views / conversions of the advertisements corresponding to each product type, and the horizontal axis represents each product type. Since the number of product types corresponding to advertisements in the Internet is limited, and each product type usually corresponds to a large number of advertisements, even for the product types at the tail, the corresponding number of samples is large. For example, the maximum / minimum / average value of the number of samples corresponding to each product type at the tail is greater than 1000. Therefore, the feature of product type can be classified as a coarse-grained feature.

[0102] In the embodiments of the present application, only three features, namely advertisement ID, advertiser ID, and product type, are taken as examples to introduce the division of coarse-grained and fine-grained features. Among them, the above-mentioned coarse-grained and fine-grained features can be manually divided by developers according to the number of tail samples of each feature, or the above-mentioned coarse-grained and fine-grained features can also be automatically divided by computer devices according to the division rules set by developers based on the statistical results of the number of tail samples of each feature. The embodiments of the present application do not make any limitations.

[0103] In the embodiments of the present application, when an information display opportunity arrives, the computer device can obtain each piece of information that meets the information display opportunity as a group of candidate information, and extract information features from these candidate information, where these information features are divided into coarse-grained features and fine-grained features.

[0104] Step 202: Obtain the first feature of each of at least two candidate information based on the coarse-grained features of each of the at least two candidate information; the first feature is obtained based on intermediate features; the intermediate features are obtained by extracting from the coarse-grained features.

[0105] In the embodiments of the present application, for the coarse-grained features of each candidate information, the computer device can further extract features from these coarse-grained features. For example, the computer device first extracts features from the coarse-grained features to obtain intermediate features, and then processes the intermediate features corresponding to the coarse-grained features again to obtain the above-mentioned first feature.

[0106] Step 203: Obtain the second feature of each of at least two candidate information based on the information features of each of the at least two candidate information and the intermediate features of each of the at least two candidate information.

[0107] In the embodiments of the present application, in order to extract more accurate feature representations, when extracting the second feature of at least two candidate information, in addition to using the information features of each of the at least two candidate information, the intermediate features of each of the at least two candidate information are also shared, so as to be able to learn the multi-level (information overall level, coarse-grained feature level, and fine-grained feature level) feature representations in the candidate information.

[0108] Step 204: Obtain the target information among at least two pieces of candidate information based on the respective first features and the respective second features of at least two pieces of candidate information.

[0109] Step 205: Push the target information.

[0110] In summary, for the solution shown in the embodiments of the present application, the information features are divided into coarse-grained features with a large number of tail value samples and fine-grained features with a small number of tail value samples. The first features are extracted from the coarse-grained features, and the second features are extracted from the complete information features. Moreover, when extracting the second features, the intermediate features between the coarse-grained features and the first features are also combined for the extraction of the second features, so that multi-level feature representations can be synchronously learned from the information features, thereby improving the representation effect of the extracted features on the information. When selecting and pushing information through the extracted first features and second features subsequently, the accuracy of information pushing can be improved.

[0111] In the embodiments of the present application, the above Figure 2 shown solution can be implemented by a trained probability prediction model.

[0112] Figure 4 is a schematic flowchart of an information pushing method shown according to an exemplary embodiment. This method can be executed by a computer device. For example, the computer device can be a server, and the server can be the server 140 in the above Figure 1 shown embodiment. As Figure 4 shown, the information pushing method can include the following steps.

[0113] Step 401: Extract the information features of at least two pieces of candidate information. The information features include coarse-grained features and fine-grained features; the number of tail value samples of the coarse-grained features is greater than the number of tail value samples of the fine-grained features.

[0114] This step 401 can refer to the description under step 402 in the above Figure 2 shown embodiment and will not be elaborated here.

[0115] Step 402: Based on the respective coarse-grained features of at least two pieces of candidate information, obtain the respective first features of at least two pieces of candidate information; the first features are obtained based on intermediate features; the intermediate features are obtained by extracting the coarse-grained features.

[0116] In the embodiments of the present application, when the computer device extracts the first features, it can first perform feature extraction on the coarse-grained features to obtain various intermediate features, and perform weighted processing on the various intermediate features to obtain the first features.

[0117] For example, the process of obtaining the first feature of each of at least two candidate information based on the respective coarse-grained features of the at least two candidate information may include:

[0118] Performing feature extraction on the coarse-grained feature of the first candidate information to obtain m first intermediate features of the first candidate information; the first candidate information is any one of the at least two candidate information; m is a positive integer;

[0119] Based on the coarse-grained feature of the first candidate information, obtaining the first weights of the m first intermediate features;

[0120] Based on the m first intermediate features and the first weights of the m first intermediate features, obtaining the first feature of the first candidate information.

[0121] For each of the at least two candidate information, the computer device can perform the above processing respectively, that is, the first features of the at least two candidate information can be obtained respectively.

[0122] For example, in the embodiments of the present application, the above m first intermediate features may be obtained by m pre-set expert networks respectively extracting the coarse-grained features, and the computer device also obtains the first weights respectively corresponding to the m first intermediate features based on the coarse-grained features, and then performs weighted processing on the m first intermediate features based on the first weights, that is, the first features of each candidate information are obtained.

[0123] In a possible implementation manner, the process of obtaining the first feature of each of at least two candidate information based on the respective coarse-grained features of the at least two candidate information may include: processing the respective coarse-grained features of the at least two candidate information through a first extraction branch in a probability prediction model to obtain the first features of the at least two candidate information respectively.

[0124] Wherein, the first extraction branch may include three parts: a feature extraction network, a weight acquisition network, and a weighted network.

[0125] In an exemplary solution, the above feature extraction network may include m expert networks, and the m expert networks respectively process the input coarse-grained features and respectively output a piece of expert information (i.e., the above first intermediate features).

[0126] In an exemplary solution, the above weight acquisition network may be a gate network, and the gate network in the first extraction branch may process the input coarse-grained features and output the weights respectively corresponding to the m expert networks (i.e., the above first weights).

[0127] In an exemplary solution, the above-mentioned weighted network can be implemented by including a weighted layer and a tower network. The weighted layer of the weighted network in the first extraction branch can perform a weighted sum of the expert information output by m expert networks based on the weights output by the gate network in the first extraction branch. The tower network of the weighted network in the first extraction branch can perform feature extraction on the weighted sum result of the weighted layer through knowledge distillation to obtain the first feature output by the first extraction branch.

[0128] Please refer to Figure 5 , which shows the model framework diagram related to the embodiments of the present application. As Figure 5 shown, the probability prediction model includes a first extraction branch 51, and the first extraction branch 51 includes m expert networks 51a, a gate network 51b, and a tower network 51c.

[0129] In the embodiments of the present application, the first extraction branch can also be referred to as a grouping layer; among them, the purpose of the existence of the grouping layer is to learn the generalization representation of each information group, which contains the common knowledge transmitted between all information within the group. Figure 5 The part of the first extraction branch 51 in

[0130] shows the constituent elements of the grouping layer. The bottom layer is composed of some expert networks (expert network 51a), and these expert networks take the coarse-grained feature 52 as the input, and the output is specific expert information. Different expert information corresponds to different aspects of the task, and these expert information can be shared between different tasks.

[0131]

[0132] Among them, is the input feature of the grouping layer, represents that the kth expert network maps the input feature from the initial embedding space to the new space of the coefficient matrix.

[0133] In order to adaptively fuse the expert networks, in the Figure 5 shown framework, a gate network 51b is also used for selective fusion. In the embodiments of the present application, the gate network can be composed of a single-layer neural network, and the softmax is used as the activation function, and its output can be expressed as:

[0134] w g = Softmax(W2x g )

[0135] Among them, is the coefficient matrix, and m is the number of expert networks in the grouping layer.

[0136] In Figure 5 In the first extraction branch 51 shown, after the upper - layer structure performs weighted summation on the expert information, it then uses the tower network to distill the representation vector of the grouping layer, and this representation vector is as follows:

[0137] e g = h g (f g )

[0138]

[0139] Among them, h g represents the tower - shaped network of the grouping layer.

[0140] Please refer to Figure 6 , which shows a schematic diagram of weighted summation of expert information involved in the embodiments of the present application. As Figure 6 shown, m expert networks 61 ( Figure 6 4 expert networks are shown in Figure 6 ) respectively output expert information 62. After these m expert information 62 are multiplied by their respective first weights through a weighted layer (

[0141] not shown in

[0142] The embodiments of the present application adopt an asymmetric feature - sharing processing method for feature extraction. Among them, this asymmetric feature - sharing method means that when extracting the second feature, the intermediate features obtained during the process of sharing the first feature are used.

[0143] In a possible implementation manner, the process of obtaining the second features of at least two candidate information based on the information features and intermediate features of at least two candidate information respectively can be as follows:

[0144] Perform feature extraction on the information feature of the first candidate information to obtain n second intermediate features of the first candidate information; n is a positive integer;

[0145] Based on the information feature of the first candidate information, obtain the second weights of the n second intermediate features and the second weights of the m first intermediate features;

[0146] Obtain the second feature of the first candidate information based on the second weights of n second intermediate features, the second weights of m first intermediate features, the n second intermediate features, and the m first intermediate features.

[0147] For each of at least two candidate information, the computer device can perform the above processing respectively, that is, the second features of at least two candidate information can be obtained respectively.

[0148] For example, in the embodiments of the present application, the above n second intermediate features can be obtained by n pre-set expert networks extracting coarse-grained features and fine-grained features respectively. And, the computer device also obtains the second weights corresponding to the n second intermediate features respectively based on the coarse-grained features and the fine-grained features. In addition, the computer device also obtains the second weights corresponding to the m first intermediate features respectively based on the coarse-grained features and the fine-grained features. Subsequently, based on the second weights, weighted processing is performed on the m first intermediate features and the n second intermediate features, that is, the second features of each candidate information are obtained.

[0149] In a possible implementation manner, the process of obtaining the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the information features of the first candidate information may include:

[0150] Obtain the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the information features of the first candidate information and the popularity vector of the first candidate information; the popularity vector is used to indicate the historical conversion times of the corresponding candidate information.

[0151] In the embodiments of the present application, in order to more accurately learn the features of candidate information so as to improve the accuracy of subsequent information pushing, the popularity of each candidate information may also be considered when obtaining the second weights.

[0152] In a possible implementation manner, the process of obtaining the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the information features of the first candidate information and the popularity vector of the first candidate information may include:

[0153] Concatenate the information features of the first candidate information and the popularity vector of the first candidate information to obtain the first concatenated feature of the first candidate information;

[0154] Obtain the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the first concatenated feature of the first candidate information.

[0155] In an embodiment of the present application, the computer device may splice the fine-grained features, coarse-grained features, and popularity vectors of the first candidate information, and then process the spliced features to obtain the second weight mentioned above.

[0156] In a possible implementation manner, the process of obtaining the second feature of each of the at least two candidate information based on the information features and intermediate features of each of the at least two candidate information may include:

[0157] Process the information features of each of the two candidate information and the intermediate features of each of the at least two candidate information through a second extraction branch in the probability prediction model to obtain the second feature of each of the at least two candidate information.

[0158] Among them, the second extraction branch may also include three parts: a feature extraction network, a weight acquisition network, and a weighted network.

[0159] In an exemplary solution, the feature extraction network in the second extraction branch may include n expert networks, and each of the n expert networks processes the input information features (coarse-grained features + fine-grained features) and outputs an expert information (i.e., the second intermediate feature mentioned above).

[0160] In an exemplary solution, the weight acquisition network in the second extraction branch may be a gate network. The gate network in the second extraction branch may process the input information features and output the weights corresponding to the m expert networks in the second extraction branch and the n expert networks in the first extraction branch (i.e., the second weight mentioned above).

[0161] In an exemplary solution, the weighted network may be implemented by including a weighted layer and a tower network. The weighted layer in the second extraction branch may perform a weighted sum on the expert information output by the m + n expert networks based on the weights output by the gate network in the second extraction branch. The tower network of the weighted network in the second extraction branch may perform feature extraction on the weighted sum result of the weighted layer through knowledge distillation to obtain the second feature output by the second extraction branch.

[0162] As Figure 5 shown, the probability prediction model includes a second extraction branch 54, and the second extraction branch 54 includes n expert networks 54a, a gate network 54b, and a tower network 54c.

[0163] In the implementation of the present application, Figure 5 the second extraction branch in Figure 5 may also be referred to as the information layer. In Figure 5As shown in the structure of the second extraction branch 54 in , the input of the information layer not only includes the coarse-grained feature 52, but also extends to the fine-grained feature 53. Therefore, the output of the n expert networks 54a can be expressed as:

[0164]

[0165] is the input feature of the information layer, is the transformation matrix of the k-th expert network.

[0166] In Figure 5 In the information layer shown, the expert networks of the information layer and the grouping layer are not separated, but combined and sent to the tower network for distilling the represented information. This asymmetric information sharing design pattern can greatly improve the performance of the entire model.

[0167] In addition, the embodiment of the present application also differentiates the information rich in positive samples and the new information scarce in positive samples through the historical transformation times of information. In order to enable the model to learn the difference between such information popularities, a display definition and construction of its representation are performed in the gate network of the wide information layer.

[0168] For example, in the embodiment of the present application, the popularity is first binned according to the numerical range, and representation learning is performed for each bin. Considering the oligopoly effect of popularity, the numerical range of binning will expand as the popularity increases.

[0169] For example, the computer device can divide the numerical range of popularity into r numerical intervals connected end to end. Among them, for a certain candidate information, obtain the historical transformation times of the candidate information (which can be the total transformation times, or the transformation times in a recent period of time), determine the numerical interval where the historical transformation times are located (assumed to be the s-th interval), and generate a popularity vector with a dimension of r. The s-th element in the popularity vector is 1, and other dimensions are 0.

[0170] The representation of popularity is concatenated with other input features and, after transformation, serves as the output of the gate network of the information layer. Therefore, the following formula represents the output of the gate network of the information layer:

[0171]

[0172] where, e popu represents the popularity vector, is the concatenation operation, is the parameter matrix of the gate network. Based on this lightweight design, the popularity of information can more conveniently and directly affect the representation fusion.

[0173] For example, please refer to Figure 7, which shows a schematic diagram of obtaining the second weight involved in the embodiments of the present application. As Figure 7 shown, the computer device splices the fine-grained feature 71, the coarse-grained feature 72, and the popularity vector 73 to obtain a spliced feature 74, and then inputs the spliced feature 74 into the gate network 54b for processing to obtain the second weight 75 output by the gate network 54b.

[0174] The representation vector of the information layer can be obtained by the following formula:

[0175] e a = h a (f a )

[0176]

[0177] where m and n are the numbers of the expert networks in the grouping layer and the information layer, and h a represents the tower network of the information layer.

[0178] After obtaining the above first feature and second feature, the computer device can obtain the target information among at least two candidate information based on the first features of each of the at least two candidate information and the second features of each of the at least two candidate information. This process can refer to the following steps.

[0179] Step 404, fuse the first features of each of the at least two candidate information and the second features of each of the at least two candidate information to obtain the fused features of each of the at least two candidate information.

[0180] In a possible implementation manner, the process of fusing the first features of each of the at least two candidate information and the second features of each of the at least two candidate information to obtain the fused features of each of the at least two candidate information may include:

[0181] Obtain the third weight of the second feature of the second candidate information based on the information feature of the second candidate information; the second candidate information is any one of the at least two candidate information;

[0182] Fuse the first feature of the second candidate information and the second feature of the second candidate information based on the third weight of the second feature of the second candidate information to obtain the fused feature of the second candidate information.

[0183] In the embodiments of the present application, when the computer device fuses the first feature and the second feature of each candidate information, it can perform weight processing on the second feature and then fuse it with the first feature. Among them, the third weight of the second feature is obtained through the information feature (coarse-grained feature + fine-grained feature) of the candidate information.

[0184] In a possible implementation, the process of obtaining the third weight of the second feature of the second candidate information based on the information feature of the second candidate information may include:

[0185] Based on the information feature of the second candidate information and the popularity vector of the second candidate information, obtain the third weight of the second feature of the second candidate information.

[0186] In the embodiments of the present application, when calculating the third weight of the second feature of the candidate information, the influence of the popularity of the candidate information on the weight of the second feature may also be considered.

[0187] In a possible implementation, the process of obtaining the third weight of the second feature of the second candidate information based on the information feature of the second candidate information and the popularity vector of the second candidate information may include:

[0188] Concatenate the information feature of the second candidate information and the popularity vector of the second candidate information to obtain the second concatenated feature of the second candidate information;

[0189] Based on the second concatenated feature of the second candidate information, obtain the third weight of the second feature of the second candidate information.

[0190] In the embodiments of the present application, when considering the influence of the popularity of the candidate information on the weight of the second feature, the popularity vector of the candidate information may be concatenated with the information feature of the candidate information, and the third weight may be calculated based on the obtained concatenated feature.

[0191] In a possible implementation, the process of fusing the first feature of the second candidate information and the second feature of the second candidate information based on the third weight of the second feature of the second candidate information to obtain the fused feature of the second candidate information may include:

[0192] Perform weighted processing on the second feature of the second candidate information based on the second feature of the second candidate information to obtain the weighted feature of the second candidate information;

[0193] Add the weighted feature of the second candidate information to the first feature of the second candidate information to obtain the fused feature of the second candidate information.

[0194] When fusing with the first feature after performing weight processing on the second feature, the weighted result between the second feature and the third weight may be added to the first feature to obtain the fused feature.

[0195] In a possible implementation, the process of fusing the first feature of each of at least two candidate information and the second feature of each of at least two candidate information to obtain the fused feature of each of at least two candidate information may include:

[0196] Through the fusion branch in the probability prediction model, the first features and the second features of at least two candidate information are processed respectively to obtain the fusion features of at least two candidate information respectively.

[0197] In the embodiments of the present application, the process of fusing the first feature and the second feature can be referred to as dynamic representation fusion. Please refer to Figure 5 , in dynamic representation fusion, the information layer representation learns all the information between different information, while the group layer representation is particularly important for new information or the information released by information publishers with less released information. In order to combine the two, the present application can adopt a lightweight gate network (i.e., Figure 5 the gate network 55 in

[0198]

[0199]

[0200] wherein, is the final representation vector output by the model, e a , are the representation vectors of the information layer and the group layer respectively. is the coefficient matrix, is the vector element product operation, v fuse is the learned fusion weight vector (i.e., the above-mentioned third weight), is the weighted feature.

[0201] The combination of the information layer representation and the group layer representation encompasses a large amount of effective information, making the final representation of the information have stronger generalization ability. Therefore, it can reduce the impact brought by the cold start problem in the event probability prediction after information display.

[0202] In the above embodiments of the present application, the third weight is taken as an example of a weight vector for illustration. Optionally, the third weight can also be in other forms. For example, the third weight can also be a weight value.

[0203] Step 405, based on the fusion features of at least two candidate information respectively, obtain the predicted event probabilities of at least two candidate information respectively; the predicted event probability is the predicted probability of a specified event occurring after the corresponding information is displayed.

[0204] Wherein, the above-mentioned specified event can be at least one of a conversion event, a click event or an exposure event.

[0205] In an embodiment of the present application, a computer device may estimate the probability that after at least two candidate information are pushed and displayed, an effective push (i.e., an event such as conversion, click, or exposure occurs after the push) can occur based on the respective fusion features of the at least two candidate information. For example, the estimated event probability may be at least one of an estimated conversion rate, an estimated click-through rate, and an estimated exposure rate.

[0206] In a possible implementation manner, the process of obtaining the respective estimated event probabilities of at least two candidate information based on the respective fusion features of the at least two candidate information may include:

[0207] Process the respective fusion features of at least two candidate information through an estimation branch in a probability estimation model to obtain the respective estimated event probabilities of the at least two candidate information.

[0208] In an embodiment of the present application, as Figure 5 shown, the probability estimation model may further include an estimation branch 56, and the input of the estimation branch 56 includes the respective fusion features of the at least two candidate information. Optionally, the input of the estimation branch may further include other feature information, such as relevant features of the display position and relevant features of the user corresponding to the display position, etc. (i.e., the representation vector output on the user side), and the embodiments of the present application do not make any limitations in this regard.

[0209] In an embodiment of the present application, the computer device may also train the probability estimation model before obtaining at least two candidate information.

[0210] In a possible implementation manner, the training process of the probability estimation model may be as follows:

[0211] Extract the information features of the sample information;

[0212] Process the coarse-grained features of the sample information through a first extraction branch to obtain the first feature of the sample information;

[0213] Process the information features of the sample information and the intermediate features of the sample information through a second extraction branch to obtain the second feature of the sample information;

[0214] Process the first feature of the sample information and the second feature of the sample information through a fusion branch to obtain the fusion feature of the sample information;

[0215] Process the respective fusion features of at least two candidate information through an estimation branch in the probability estimation model to obtain the estimated event probability of the sample information;

[0216] Obtain the loss function value based on the estimated event probability of the sample information, the event probability label of the sample information, and the training weight of the sample information; the training weight is inversely correlated with the popularity of the sample information; the event probability label is used to indicate the labeled probability of a specified event occurring after the sample information is displayed;

[0217] Update the parameters of the probability estimation model based on the loss function value.

[0218] Among them, the computer device can regularly collect the push situation of each piece of information in the network within a certain period of time (such as within 48 hours before the current moment), such as whether it is pushed, and whether events such as clicks, exposures, and conversions occur after the push, and based on the push situation of each piece of information in the network, construct the above-mentioned sample information and the labeled probability of the sample information.

[0219] In the embodiment of the present application, the probability estimation model can focus on learning the optimal representation vector for each piece of information, and the embodiment of the present application can use a multi-layer neural network to learn the representation vector of the user. Taking the above-mentioned estimated event probability as the conversion rate prediction value as an example, the conversion rate prediction value can be expressed as:

[0220]

[0221] Among them, e u is the representation vector output by the user side.

[0222] The embodiment of the present application can use logarithmic loss as the loss function. Among them, logarithmic loss is a commonly used loss function in conversion rate estimation. In the embodiment of the present application, since the positive samples in the real dataset always gather on a few pieces of information with high popularity, in order to avoid the loss function being overly influenced by these samples, the embodiment of the present application optimizes the loss function as follows:

[0223]

[0224] Among them, y i and respectively represent the true value of user conversion and the predicted value of the conversion rate, w i is the weight value of training sample i, and N is the total number of training samples. The significance of introducing weights in the loss function is that it can appropriately reduce the sensitivity of the loss to popular advertisements and instead focus on new advertisements.

[0225] Optionally, the calculation formula for the weight of the above-mentioned training sample is:

[0226]

[0227] Among them, K i represents the popularity of training sample i. For example, Ki It may be the historical conversion times of training sample i. In the embodiments of the present application, the weight difference between advertisements with higher popularity and new advertisements with lower popularity may reach two orders of magnitude, which may lead to unsatisfactory training results. Therefore, in the embodiments of the present application, K i can be truncated. For example, the maximum value of K i is set to 20.

[0228] Step 406: Obtain the target information from at least two candidate pieces of information based on the respective estimated event probabilities of the at least two candidate pieces of information.

[0229] In the embodiments of the present application, the computer device may sort the at least two candidate pieces of information according to the estimated event probabilities, and select one or more candidate pieces of information ranked at the top as the target information.

[0230] Step 407: Push the target information.

[0231] Please refer to Figure 8 , which shows a schematic diagram of the comparative experiment results involved in the embodiments of the present application.

[0232] Among them, Figure 8 shows the results obtained by applying the solution shown in the embodiments of the present application to two different advertisement product data sets. All the experimental results consist of the mean and variance of the area under the curve (AUC) of 3 repeated experiments. The optimal results are shown in bold.

[0233] By observing Figure 8 it can be seen that:

[0234] (1) Since AutoFuse (i.e., the probability estimation model provided in the embodiments of the present application) adopts a higher-level modeling technology compared to MGQE (Multi-granular Quantized Embedding) and AutoEmb (automatic embedding model), better results are achieved on both new and old advertisements.

[0235] (2) Compared to the DeepFM (Deep Feature Embedding) model, the PNN (Product-based Neural Networks) model, and the DCN (Deep and Cross Network) model, the solution shown in the present application has obvious advantages on new advertisements and also shows equivalent competitiveness on old advertisements, which all benefit from feature grouping and asymmetric sharing.

[0236] (3) Compared with the two multi-task models, namely MMoE (Multi-gate Mixture-of-Experts) model and PLE (Progressive Layered Extraction), the underlying feature grouping structure of AutoFuse greatly reduces the training pressure of the upper-layer structure, enabling it to focus more on the representation learning of different layers, thereby improving the generalization performance.

[0237] (4) AutoFuse fully explores the pattern features between individual and group advertisements, providing an effective solution for the cold-start problem of advertising predicted conversion rate. Compared with DNN (Deep Neural Networks), it has achieved performance improvements of 0.55% and 0.46% respectively on new advertisements in two datasets. For the two datasets of old advertisements, the proposed solution in this application has also achieved improvements of 0.18% and 0.21%. On the two datasets of overall advertisements, AutoFuse has achieved improvements of 0.55% and 0.53% respectively. An AUC improvement of 0.1% in the industrial community can be considered a significant improvement. These results fully prove that the solution shown in the embodiments of this application can obtain an overall performance improvement while alleviating the cold-start problem.

[0238] Please refer to Figure 9 , which shows the schematic diagram of the ablation experiment results involved in the embodiments of this application. To further verify the AutoFuse model, more ablation experiments were conducted based on the proposed solution in this application to compare various variants of AutoFuse.

[0239] The solution shown in the embodiments of this application adopts the strategies of feature grouping and asymmetric sharing. First, the input features are grouped, and complete isolation between the information layer and the grouping layer is ensured. The expert network in the information layer only inputs the fineness features, and the gate network in the information layer only fuses the expert network in the information layer. Numerical-based fusion is adopted in the output part of the information layer and the grouping layer, so that the final output of the entire system is the weighted sum of the information layer and the grouping layer. This variant is denoted as V1.

[0240] The embodiments of this application are marked as V2 after adding asymmetric sharing on the basis of V1. A significant performance improvement is achieved from V1 to V2, proving the importance of asymmetric sharing. V2 has a 0.75% improvement compared to V1 on old advertisements, indicating that it is necessary to add coarse-grained features to the information layer. More importantly, V2 has a significant advantage compared to DNN, reflecting the rationality of the method of fusing features using asymmetric sharing.

[0241] The solution shown in the embodiments of this application also considers popularity embedding representations. The correlation between the features of the information layer and the grouping layer is complex and affected by the sample distribution. Accordingly, AutoFuse uses popularity embedding representations to adaptively guide this fusion, denoted as V3. V3 achieved a 0.26% improvement in the AUC of new advertisements compared to V2, and performed similarly to V2 for old advertisements, indicating that popularity embedding benefits the performance of new advertisements more. This phenomenon also conforms to the expectation of this application because old advertisements have a large amount of training data and can learn meaningful representations, while new advertisements require more direct guidance for knowledge acquisition and representation information fusion.

[0242] The solution shown in the embodiments of this application also adopts the strategies of dynamic fusion and adaptive loss. Dynamic fusion is to adaptively combine the representation outputs of the information layer and the grouping layer. The numerical weighted summation method can reduce the magnitude of each vector. AutoFuse adopts vector-based fusion, assigning different weights to different dimensions of the input vectors. This method is more flexible and can introduce more non-linearity. The resulting model is denoted as V4. V4 has improved compared to V3 for both new and old advertisements. AutoFuse further adds adaptive loss based on V4, and the effect on new advertisements has been further improved.

[0243] In summary, for the solution shown in the embodiments of this application, the information features are divided into coarse-grained features with a large number of tail-value samples and fine-grained features with a small number of tail-value samples. The first feature is extracted from the coarse-grained features, and the second feature is extracted from the complete information features. Moreover, when extracting the second feature, the intermediate feature between the coarse-grained feature and the first feature is also combined for the extraction of the second feature, which can synchronously learn multi-level feature representations from the information features, thereby improving the representation effect of the extracted features on the information. When subsequent information selection and pushing are performed through the extracted first feature and second feature, the accuracy of information pushing can be improved.

[0244] Among them, the solution shown in the above embodiments of this application can be implemented or executed in combination with a blockchain. For example, some or all of the steps in the above embodiments can be executed in a blockchain system; or, the data required for the execution of each step in the above embodiments or the generated data can be stored in a blockchain system. For example, the training samples used for the above model training, as well as the model input data such as candidate information during the model application process, can be obtained by a computer device from the blockchain system; for another example, the parameters of the model obtained after the above model training can be stored in the blockchain system.

[0245] Figure 10 It is a structural block diagram of an information pushing device shown according to an exemplary embodiment. This device can implement Figure 2or Figure 4 all or part of the steps in the method provided by the illustrated embodiment, the information push device includes:

[0246] An information feature extraction module 1001, configured to extract information features of at least two candidate information respectively, where the information features include coarse-grained features and fine-grained features; the number of tail value samples of the coarse-grained features is greater than the number of tail value samples of the fine-grained features;

[0247] A first feature acquisition module 1002, configured to acquire first features of the at least two candidate information respectively based on the coarse-grained features of the at least two candidate information; the first features are acquired based on intermediate features; the intermediate features are obtained by extracting the coarse-grained features;

[0248] A second feature acquisition module 1003, configured to acquire second features of the at least two candidate information respectively based on the information features of the at least two candidate information and the intermediate features of the at least two candidate information;

[0249] An information acquisition module 1004, configured to acquire target information among the at least two candidate information based on the first features of the at least two candidate information and the second features of the at least two candidate information;

[0250] An information push module 1005, configured to push the target information.

[0251] In a possible implementation manner, the first feature acquisition module 1002 is configured to,

[0252] Extract features from the coarse-grained features of the first candidate information to obtain m first intermediate features of the first candidate information; the first candidate information is any one of the at least two candidate information; m is a positive integer;

[0253] Based on the coarse-grained features of the first candidate information, obtain first weights of the m first intermediate features;

[0254] Based on the m first intermediate features and the first weights of the m first intermediate features, obtain the first features of the first candidate information.

[0255] In a possible implementation manner, the second feature acquisition module 1003 is configured to,

[0256] Extract features from the information features of the first candidate information to obtain n second intermediate features of the first candidate information; n is a positive integer;

[0257] Obtain the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the information characteristics of the first candidate information;

[0258] Based on the second weights of the n second intermediate features, the second weights of the m first intermediate features, the n second intermediate features, and the m first intermediate features, obtain the second feature of the first candidate information.

[0259] In a possible implementation manner, the second feature obtaining module 1003 is configured to obtain the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the information characteristics of the first candidate information and the popularity vector of the first candidate information; the popularity vector is used to indicate the historical conversion times of the corresponding candidate information.

[0260] In a possible implementation manner, the second feature obtaining module 1003 is used for,

[0261] Concatenate the information characteristics of the first candidate information and the popularity vector of the first candidate information to obtain the first concatenated feature of the first candidate information;

[0262] Based on the first concatenated feature of the first candidate information, obtain the second weights of the n second intermediate features and the second weights of the m first intermediate features.

[0263] In a possible implementation manner, the information obtaining module 1004 is used for,

[0264] Fuse the first features of the at least two candidate information and the second features of the at least two candidate information respectively to obtain the fused features of the at least two candidate information respectively;

[0265] Based on the fused features of the at least two candidate information respectively, obtain the estimated event probabilities of the at least two candidate information respectively; the estimated event probability is the estimated probability of a specified event occurring after the corresponding information is displayed;

[0266] Based on the estimated event probabilities of the at least two candidate information respectively, obtain the target information among the at least two candidate information.

[0267] In a possible implementation manner, the information obtaining module 1004 is used for,

[0268] Based on the information characteristics of the second candidate information, obtain the third weight of the second feature of the second candidate information; the second candidate information is any one of the at least two candidate information;

[0269] Based on the third weight of the second feature of the second candidate information, fuse the first feature and the second feature of the second candidate information to obtain the fused feature of the second candidate information.

[0270] In a possible implementation manner, the information acquisition module 1004 is configured to

[0271] Based on the information feature of the second candidate information and the popularity vector of the second candidate information, obtain the third weight of the second feature of the second candidate information.

[0272] In a possible implementation manner, the information acquisition module 1004 is configured to

[0273] Concatenate the information feature of the second candidate information and the popularity vector of the second candidate information to obtain the second concatenated feature of the second candidate information;

[0274] Based on the second concatenated feature of the second candidate information, obtain the third weight of the second feature of the second candidate information.

[0275] In a possible implementation manner, the information acquisition module 1004 is configured to

[0276] Perform weighted processing on the second feature of the second candidate information based on the second feature of the second candidate information to obtain the weighted feature of the second candidate information;

[0277] Add the weighted feature of the second candidate information to the first feature of the second candidate information to obtain the fused feature of the second candidate information.

[0278] In a possible implementation manner, the first feature acquisition module 1002 is configured to process the coarse-grained features of the at least two candidate information through the first extraction branch in the probability prediction model to obtain the first features of the at least two candidate information respectively;

[0279] The second feature acquisition module 1003 is configured to process the information features of the two candidate information and the intermediate features of the at least two candidate information through the second extraction branch in the probability prediction model to obtain the second features of the at least two candidate information respectively;

[0280] The information acquisition module 1004 is configured to process the first features and the second features of the at least two candidate information respectively through the fusion branch in the probability prediction model to obtain the fused features of the at least two candidate information respectively;

[0281] The information acquisition module 1004 is further configured to process the fusion features of the at least two candidate information respectively through the estimation branch in the probability estimation model, so as to obtain the estimated event probabilities of the at least two candidate information respectively.

[0282] In a possible implementation manner, the device further includes:

[0283] The information feature extraction module 1001 is further configured to extract the information features of the sample information before extracting the information features of the at least two candidate information respectively;

[0284] The first feature acquisition module 1002 is further configured to process the coarse-grained features of the sample information through the first extraction branch to obtain the first feature of the sample information;

[0285] The second feature acquisition module 1003 is further configured to process the information features of the sample information and the intermediate features of the sample information through the second extraction branch to obtain the second feature of the sample information;

[0286] The information acquisition module 1004 is further configured to process the first feature of the sample information and the second feature of the sample information through the fusion branch to obtain the fusion feature of the sample information;

[0287] The information acquisition module 1004 is further configured to process the fusion features of the at least two candidate information respectively through the estimation branch in the probability estimation model to obtain the estimated event probability of the sample information;

[0288] The device further includes:

[0289] The loss function value acquisition module is configured to obtain a loss function value based on the estimated event probability of the sample information, the event probability label of the sample information, and the training weight of the sample information; the training weight is inversely correlated with the popularity of the sample information; the event probability label is used to indicate the labeled probability of the specified event occurring after the sample information is displayed;

[0290] The parameter update module is configured to update the parameters of the probability estimation model based on the loss function value.

[0291] In summary, the solution shown in the embodiments of the present application divides information features into coarse-grained features with a large number of tail value samples and fine-grained features with a small number of tail value samples, extracts first features from the coarse-grained features, extracts second features from the complete information features, and, when extracting the second features, also combines intermediate features between the coarse-grained features and the first features to extract the second features, which can synchronously learn multi-level feature representations from the information features, thereby improving the representation effect of the extracted features on the information. When selecting and pushing information through the extracted first features and second features subsequently, the accuracy of information pushing can be improved.

[0292] Figure 11 FIG. 4 is a schematic structural diagram of a computer device shown according to an exemplary embodiment. The computer device can be implemented as the computer device for training the first image recognition model in the above-mentioned various method embodiments, or can be implemented as the computer device for performing midline recognition of the brain through the second image recognition model in the above-mentioned various method embodiments. The computer device 1100 includes a central processing unit (CPU, Central Processing Unit) 1101, a system memory 1104 including a random access memory (Random Access Memory, RAM) 1102 and a read-only memory (Read-Only Memory, ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The computer device 1100 further includes a basic input / output system 1106 for facilitating the transmission of information between various components within the computer, and a mass storage device 1107 for storing an operating system 1113, application programs 1114, and other program modules 1115.

[0293] The mass storage device 1107 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer-readable medium provide non-volatile storage for the computer device 1100. That is to say, the mass storage device 1107 can include computer-readable media (not shown) such as a hard disk or a compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM) drive.

[0294] Without loss of generality, the computer-readable medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes RAM, ROM, flash memory or other solid-state storage technologies, CD-ROM, or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will understand that the computer storage medium is not limited to the above several types. The above system memory 1104 and mass storage device 1107 can be collectively referred to as memory.

[0295] The computer device 1100 can be connected to the Internet or other network devices through a network interface unit 1111 connected to the system bus 1105.

[0296] The memory further includes one or more programs, and the one or more programs are stored in the memory. The central processing unit 1101 implements Figure 2 or Figure 4 all or part of the steps of any of the methods shown.

[0297] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory including a computer program (instructions). The above program (instructions) can be executed by a processor of a computer device to complete the methods shown in various embodiments of the present application. For example, the non-transitory computer-readable storage medium can be a Read-Only Memory (ROM), Random Access Memory (RAM), Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0298] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods shown in the above various embodiments.

[0299] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the claims.

[0300] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. An information push method, characterized in that, The method includes: Extracting information features of at least two candidate information respectively, where the information features include coarse-grained features and fine-grained features; the number of tail value samples of the coarse-grained features is greater than the number of tail value samples of the fine-grained features; the tail value samples refer to the sample information in one or more classifications arranged at the tail position after classifying each sample information according to the feature values and sorting them in descending order of the number of information in each classification; Based on the coarse-grained features of the at least two candidate information respectively, obtaining the first features of the at least two candidate information respectively; the first features are obtained based on intermediate features; the intermediate features are obtained by extracting the coarse-grained features; Based on the information features of the at least two candidate information respectively and the intermediate features of the at least two candidate information respectively, obtaining the second features of the at least two candidate information respectively; Based on the first features of the at least two candidate information respectively and the second features of the at least two candidate information respectively, obtaining the target information among the at least two candidate information; Pushing the target information.

2. The method according to claim 1, wherein The obtaining the first features of the at least two candidate information respectively based on the coarse-grained features of the at least two candidate information respectively includes: Performing feature extraction on the coarse-grained features of the first candidate information to obtain m first intermediate features of the first candidate information; the first candidate information is any one of the at least two candidate information; m is a positive integer; Based on the coarse-grained features of the first candidate information, obtaining the first weights of the m first intermediate features; Based on the m first intermediate features and the first weights of the m first intermediate features, obtaining the first features of the first candidate information.

3. The method according to claim 2, wherein The obtaining the second features of the at least two candidate information respectively based on the information features of the at least two candidate information respectively and the intermediate features of the at least two candidate information respectively includes: Performing feature extraction on the information features of the first candidate information to obtain n second intermediate features of the first candidate information; n is a positive integer; Based on the information features of the first candidate information, obtaining the second weights of the n second intermediate features and the second weights of the m first intermediate features; Based on the second weights of the n second intermediate features, the second weights of the m first intermediate features, the n second intermediate features, and the m first intermediate features, obtaining the second features of the first candidate information.

4. The method according to claim 3, wherein The obtaining the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the information features of the first candidate information includes: Based on the information features of the first candidate information and the popularity vector of the first candidate information, obtaining the second weights of the n second intermediate features and the second weights of the m first intermediate features; the popularity vector is used to indicate the historical conversion times of the corresponding candidate information.

5. The method according to claim 4, characterized in that, Obtaining the second weights of the n second intermediate features and the second weights of the m first intermediate features based on the information feature of the first candidate information and the popularity vector of the first candidate information includes: Concatenating the information feature of the first candidate information and the popularity vector of the first candidate information to obtain the first concatenated feature of the first candidate information; Based on the first concatenated feature of the first candidate information, obtaining the second weights of the n second intermediate features and the second weights of the m first intermediate features.

6. The method according to claim 1, wherein Obtaining the target information among the at least two candidate information based on the respective first features of the at least two candidate information and the respective second features of the at least two candidate information includes: Fusing the respective first features of the at least two candidate information and the respective second features of the at least two candidate information to obtain the respective fused features of the at least two candidate information; Based on the respective fused features of the at least two candidate information, obtaining the respective estimated event probabilities of the at least two candidate information; the estimated event probability is the estimated probability of a specified event occurring after the corresponding information is displayed; Based on the respective estimated event probabilities of the at least two candidate information, obtaining the target information among the at least two candidate information.

7. The method according to claim 6, characterized in that, The fusing the respective first features of the at least two candidate information and the respective second features of the at least two candidate information to obtain the respective fused features of the at least two candidate information includes: Based on the information feature of the second candidate information, obtaining the third weight of the second feature of the second candidate information; the second candidate information is any one of the at least two candidate information; Based on the third weight of the second feature of the second candidate information, fusing the first feature of the second candidate information and the second feature of the second candidate information to obtain the fused feature of the second candidate information.

8. The method according to claim 7, characterized in that, The obtaining the third weight of the second feature of the second candidate information based on the information feature of the second candidate information includes: Based on the information feature of the second candidate information and the popularity vector of the second candidate information, obtaining the third weight of the second feature of the second candidate information.

9. The method according to claim 8, wherein The obtaining the third weight of the second feature of the second candidate information based on the information feature of the second candidate information and the popularity vector of the second candidate information includes: Concatenating the information feature of the second candidate information and the popularity vector of the second candidate information to obtain the second concatenated feature of the second candidate information; Based on the second concatenated feature of the second candidate information, obtaining the third weight of the second feature of the second candidate information.

10. The method according to claim 7, wherein The fusing the first feature of the second candidate information and the second feature of the second candidate information based on the third weight of the second feature of the second candidate information to obtain the fused feature of the second candidate information includes: Based on the second feature of the second candidate information, performing weighted processing on the second feature of the second candidate information to obtain the weighted feature of the second candidate information; Add the weighted feature of the second candidate information to the first feature of the second candidate information to obtain the fused feature of the second candidate information.

11. The method according to any one of claims 6 to 10, characterized in that Based on the coarse-grained features of the at least two candidate information respectively, obtaining the first feature of each of the at least two candidate information, including: Processing the coarse-grained features of the at least two candidate information respectively through a first extraction branch in the probability prediction model to obtain the first feature of each of the at least two candidate information; The obtaining the second feature of each of the at least two candidate information based on the information features of the at least two candidate information respectively and the intermediate features of the at least two candidate information respectively, including: Processing the information features of the two candidate information respectively and the intermediate features of the at least two candidate information respectively through a second extraction branch in the probability prediction model to obtain the second feature of each of the at least two candidate information; The fusing the first feature of each of the at least two candidate information and the second feature of each of the at least two candidate information to obtain the fused feature of each of the at least two candidate information, including: Processing the first feature of each of the at least two candidate information and the second feature of each of the at least two candidate information through a fusion branch in the probability prediction model to obtain the fused feature of each of the at least two candidate information; The obtaining the predicted event probability of each of the at least two candidate information based on the fused feature of each of the at least two candidate information, including: Processing the fused feature of each of the at least two candidate information through a prediction branch in the probability prediction model to obtain the predicted event probability of each of the at least two candidate information.

12. The method according to claim 11, wherein Before extracting the information features of the at least two candidate information respectively, further including: Extracting the information features of the sample information; Processing the coarse-grained feature of the sample information through the first extraction branch to obtain the first feature of the sample information; Processing the information feature of the sample information and the intermediate feature of the sample information through the second extraction branch to obtain the second feature of the sample information; Processing the first feature of the sample information and the second feature of the sample information through the fusion branch to obtain the fused feature of the sample information; Processing the fused feature of each of the at least two candidate information through the prediction branch in the probability prediction model to obtain the predicted event probability of the sample information; Based on the predicted event probability of the sample information, the event probability label of the sample information, and the training weight of the sample information, obtaining a loss function value; the training weight is inversely correlated with the popularity of the sample information; the event probability label is used to indicate the labeled probability of the specified event occurring after the sample information is displayed; Updating the parameters of the probability prediction model based on the loss function value.

13. An information push device, characterized in that, The device includes: An information feature extraction module, configured to extract information features of at least two candidate pieces of information respectively, where the information features include coarse-grained features and fine-grained features; the number of tail value samples of the coarse-grained features is greater than the number of tail value samples of the fine-grained features; the tail value samples refer to the sample information in one or more categories arranged at the end position after classifying each sample information according to the feature values and sorting in descending order according to the number of information in each category; A first feature acquisition module, configured to acquire first features of the at least two candidate pieces of information respectively based on the coarse-grained features of the at least two candidate pieces of information respectively; the first features are obtained based on intermediate features; the intermediate features are obtained by extracting from the coarse-grained features; A second feature acquisition module, configured to acquire second features of the at least two candidate pieces of information respectively based on the information features of the at least two candidate pieces of information respectively and the intermediate features of the at least two candidate pieces of information respectively; An information acquisition module, configured to acquire target information among the at least two candidate pieces of information based on the first features of the at least two candidate pieces of information respectively and the second features of the at least two candidate pieces of information respectively; An information push module, configured to push the target information.

14. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one computer instruction is stored in the memory, and the at least one computer instruction is loaded and executed by the processor to implement the information push method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, At least one computer instruction is stored in the storage medium, and the at least one computer instruction is loaded and executed by a processor to implement the information push method according to any one of claims 1 to 12.

16. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are executed by a processor of a computer device to cause the computer device to execute the information push method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method and apparatus for flexible diversification of recommendation results

    CN103620592A

  • Information pushing method and device, computer equipment and storage medium

    CN112749330A