Feature extraction model processing method and device, computer equipment and storage medium

By performing multiple masking and sample combinations on the feature representation vector and expanding the training samples, the problem of insufficient training samples in the traditional feature extraction model is solved, and the robustness and accuracy of the model are improved.

CN120632404APending Publication Date: 2025-09-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410275013.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The small number of training samples in traditional feature extraction model training leads to low model accuracy, which affects the application effect of subsequent feature extraction.

Method used

By obtaining multiple sample information, extracting feature representation vectors and performing multiple masking processes, multiple different feature mask vectors are generated, combined into positive and negative sample pairs, and model training is performed to expand the training samples.

Benefits of technology

It improves the robustness and accuracy of the feature extraction model, reduces the requirement for the number of basic samples, and improves the training effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632404A_ABST
    Figure CN120632404A_ABST
Patent Text Reader

Abstract

The invention relates to an artificial intelligence feature extraction model processing method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a plurality of pieces of sample information, respectively extracting features of each piece of sample information, and acquiring a feature representation vector corresponding to each piece of sample information; performing multiple times of mask processing on each feature representation vector to obtain a plurality of different feature mask vectors corresponding to each feature representation vector; for each feature representation vector, combining a plurality of different feature mask vectors corresponding to the feature representation vector to obtain a positive example sample pair corresponding to the feature representation vector; combining the feature mask vectors in the positive example sample pairs corresponding to different feature representation vectors to obtain negative example sample pairs; and performing model training based on the positive example sample pair and the negative example sample pair to obtain a feature extraction model. By adopting the method, training samples can be expanded, and the model precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a feature extraction model processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] With the development of computer technology, information feature extraction technology has emerged. Feature extraction technology can extract key features from information to find similar information based on the key features or recommend information to users.

[0003] In traditional technology, multiple historical push information is obtained from the push platform as training samples, and the training samples are used to train the model to obtain a feature extraction model, thereby quickly extracting the features of the information through the feature extraction model.

[0004] However, some push platforms have little historical push information, so fewer training samples are used in model training, resulting in low accuracy of the feature extraction model obtained through training, which in turn affects subsequent applications of the features extracted by the feature extraction model. Summary of the Invention

[0005] Based on this, it is necessary to provide a feature extraction model processing method, device, computer equipment, computer-readable storage medium and computer program product that can amplify sample information to improve training accuracy in response to the above technical problems.

[0006] In a first aspect, the present application provides a method for processing a feature extraction model. The method comprises:

[0007] Acquire multiple sample information, extract features of each sample information respectively, and obtain a feature representation vector corresponding to each sample information;

[0008] Performing multiple masking processes on each of the feature representation vectors to obtain multiple different feature mask vectors corresponding to each of the feature representation vectors;

[0009] For each of the feature characterization vectors, combining a plurality of different feature mask vectors corresponding to the feature characterization vector to obtain a positive sample pair corresponding to the feature characterization vector;

[0010] Combine the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs;

[0011] Model training is performed based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

[0012] In a second aspect, the present application also provides a feature extraction model processing device. The device includes:

[0013] An acquisition module is used to acquire a plurality of sample information, extract the features of each sample information, and obtain a feature representation vector corresponding to each sample information;

[0014] A masking module, configured to perform multiple masking processes on each of the feature representation vectors to obtain multiple different feature mask vectors corresponding to each of the feature representation vectors;

[0015] A first combining module is configured to combine, for each of the feature representation vectors, a plurality of different feature mask vectors corresponding to the feature representation vector to obtain a positive sample pair corresponding to the feature representation vector;

[0016] A second combining module is used to combine the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs;

[0017] The training module is used to perform model training based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

[0018] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are performed:

[0019] Acquire multiple sample information, extract features of each sample information respectively, and obtain a feature representation vector corresponding to each sample information;

[0020] Performing multiple masking processes on each of the feature representation vectors to obtain multiple different feature mask vectors corresponding to each of the feature representation vectors;

[0021] For each of the feature characterization vectors, combining a plurality of different feature mask vectors corresponding to the feature characterization vector to obtain a positive sample pair corresponding to the feature characterization vector;

[0022] Combine the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs;

[0023] Model training is performed based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

[0024] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:

[0025] Acquire multiple sample information, extract features of each sample information respectively, and obtain a feature representation vector corresponding to each sample information;

[0026] Performing multiple masking processes on each of the feature representation vectors to obtain multiple different feature mask vectors corresponding to each of the feature representation vectors;

[0027] For each of the feature characterization vectors, combining a plurality of different feature mask vectors corresponding to the feature characterization vector to obtain a positive sample pair corresponding to the feature characterization vector;

[0028] Combine the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs;

[0029] Model training is performed based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

[0030] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0031] Acquire multiple sample information, extract features of each sample information respectively, and obtain a feature representation vector corresponding to each sample information;

[0032] Performing multiple masking processes on each of the feature representation vectors to obtain multiple different feature mask vectors corresponding to each of the feature representation vectors;

[0033] For each of the feature characterization vectors, combining a plurality of different feature mask vectors corresponding to the feature characterization vector to obtain a positive sample pair corresponding to the feature characterization vector;

[0034] Combine the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs;

[0035] Model training is performed based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

[0036] The above-mentioned feature extraction model processing method, apparatus, computer device, computer-readable storage medium, and computer program product obtain multiple sample information, extract the features of each sample information separately, obtain the feature representation vector corresponding to each sample information, and thus obtain the feature representation vector of each basic sample. Each feature representation vector is subjected to multiple masking processes to obtain multiple different feature mask vectors corresponding to each feature representation vector. The masking process can achieve the amplification of the representation vector. For each feature representation vector, the multiple different feature mask vectors corresponding to the feature representation vector are combined to obtain the positive sample pair corresponding to the feature representation vector. The feature mask vectors in the positive sample pairs corresponding to different feature representation vectors are combined to obtain the negative sample pairs, thereby obtaining more positive sample pairs and negative sample pairs, achieving the amplification of training samples. In addition, through sample amplification, more training samples can be obtained on the basic sample, which can reduce the number of basic samples required. Model training is performed based on the positive sample pairs and negative sample pairs obtained by amplification, so that the robustness of the trained feature extraction model is improved and the precision and accuracy of the feature extraction model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A diagram illustrating an application environment of a feature extraction model processing method in one embodiment;

[0038] Figure 2 Schematic diagram of a process flow of a feature extraction model processing method in one embodiment;

[0039] Figure 3 Schematic diagram of comparison between sample information and feature representation vectors in one embodiment;

[0040] Figure 4 is a schematic diagram of performing mask processing on a feature representation vector in another embodiment;

[0041] Figure 5 A schematic diagram of constructing positive sample pairs and negative sample pairs in one embodiment;

[0042] Figure 6 Schematic diagram of the architecture of a feature extraction model to be trained in one embodiment;

[0043] Figure 7 A schematic diagram of the architecture of a feature extraction model to be trained and an information prediction model to be trained according to an embodiment;

[0044] Figure 8 A flowchart illustrating a method for training a feature extraction model to be trained based on positive sample pairs and negative sample pairs, and training an information prediction model to be trained based on a feature representation vector of each second category sample information to obtain an information prediction model in one embodiment;

[0045] Figure 9 A schematic diagram of a process for calculating target loss in one embodiment;

[0046] Figure 10 A schematic diagram of a plurality of information generation information flows in one embodiment;

[0047] Figure 11 A schematic flow chart of a feature extraction model processing method in one embodiment;

[0048] Figure 12 Schematic diagram of the architecture of an information prediction model to be trained in one embodiment;

[0049] Figure 13 A schematic diagram illustrating the architecture of a conversion layer in one embodiment;

[0050] Figure 14 is a structural block diagram of a feature extraction model processing device in one embodiment;

[0051] Figure 15 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0053] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, data mining, etc. For example, it is applied to the field of artificial intelligence (AI) technology, where artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making. The solution provided in the embodiments of the present application relates to a feature extraction model processing method for artificial intelligence, which is specifically described through the following embodiments.

[0054] The feature extraction model processing method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other servers. Both the terminal 102 and the server 104 can independently execute the feature extraction model processing method provided in the embodiment of the present application. The terminal 102 and the server 104 can also be used in conjunction to execute the feature extraction model processing method provided in the embodiment of the present application. When the terminal 102 and the server 104 are used in conjunction to execute the feature extraction model processing method provided in the embodiment of the present application, the terminal 102 obtains multiple sample information from the server 104, extracts the features of each sample information respectively, and obtains the feature representation vector corresponding to each sample information. The terminal 102 performs multiple masking processes on each feature representation vector respectively to obtain multiple different feature mask vectors corresponding to each feature representation vector. For each feature representation vector, the terminal 102 combines the multiple different feature mask vectors corresponding to the feature representation vector to obtain a positive sample pair corresponding to the feature representation vector. The terminal 102 combines the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs. The terminal 102 performs model training based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

[0055] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 102 and the server 104 can be directly or indirectly connected through wired or wireless communication, and this application does not impose any restrictions on this.

[0056] In this embodiment, the feature extraction model processing can be deployed in the terminal 102 or in the server 104 .

[0057] In one embodiment, Figure 2 As shown, a feature extraction model processing method is provided, which is applied to Figure 1 Computer equipment in Figure 1 The following steps are used as an example to illustrate the terminal or server in the example:

[0058] Step S202 : Acquire multiple sample information, extract features of each sample information respectively, and obtain a feature representation vector corresponding to each sample information.

[0059] The sample information is the training sample used for model training. This sample information includes target information and recommended information. Recommended information includes at least one of historically pushed information and pending information. Historically pushed information refers to information that has already been pushed to a push recipient. Pending information refers to information that has not yet been pushed. A push recipient refers to the recipient of the information, for example, the user to whom the pushed information is to be sent.

[0060] The historical push information includes at least one of the following: information that has been pushed to a push recipient and converted by the push recipient, and information that has been pushed to a push recipient and not converted by the push recipient.

[0061] In this embodiment, the multiple sample information includes at least one of the following: first-category sample information or second-category sample information; the first-category sample information includes the object information of the push recipient and the converted information of the push recipient; the second-category sample information includes the pushed information of the information push platform and the object information of the push recipient of the pushed information.

[0062] In this embodiment, the first type of sample information is provided by the information owner, and the second type of sample information is provided by the information push platform. The information owner refers to the user who holds the information. For example, if the first type of sample information is an advertisement, the information owner refers to the advertiser of the advertisement.

[0063] Specifically, the computer device may obtain multiple sample information, perform feature extraction on each sample information, and obtain a feature representation vector corresponding to each sample information. The feature representation vector is a vector used to represent the features of the sample information.

[0064] The sample information includes object information and recommendation information. The feature representation vector corresponding to the sample information includes an object representation vector and an information representation vector. The object representation vector is a vector used to represent the features of the object information, and the information representation vector is a vector used to represent the features of the recommendation information. Figure 3 As shown, the sample information A includes object information A1 and recommendation information A2, and the feature representation vector a corresponding to the sample information A includes the object representation vector a1 and the information representation vector a2.

[0065] Step S204 : performing multiple masking processes on each feature representation vector to obtain multiple different feature mask vectors corresponding to each feature representation vector.

[0066] The masking process refers to shielding at least one element in the feature characterization vector, and specifically, shielding at least one element in the feature characterization vector by a preset mask. The preset mask can be represented by a preset value, and the masking process of the feature characterization vector can be to set the value of at least one element in the feature characterization vector to the preset value. Figure 4 As shown, the preset value is 0, and the element x3 at the third position and the element x5 at the fifth position in the feature representation vector x[x1, x2, x3, x4, x5] are set to 0 to obtain the feature mask vector [x1, x2, 0, x4, 0].

[0067] Specifically, for each feature representation vector, the computer device performs multiple masking processes on the feature representation vector to obtain multiple different feature mask vectors corresponding to the feature representation vector. Furthermore, for each feature representation vector, the computer device performs multiple masking processes on the feature representation vector based on a preset mask to obtain multiple different feature mask vectors corresponding to the feature representation vector. Each feature mask vector includes at least one preset mask.

[0068] Following the same process, each feature representation vector can obtain multiple different feature mask vectors.

[0069] Different feature mask vectors refer to multiple feature mask vectors that are not completely the same, that is, multiple feature mask vectors can be completely different or partially the same. Figure 4 , the feature representation vector x[x1, x2, x3, x4, x5] is masked twice to obtain the feature mask vector [x1, x2, 0, x4, 0] and the feature mask vector [x1, 0, 0, x4, x5], [x1, x2, 0, x4, 0] and [x1, 0, 0, x4, x5] are different feature mask vectors.

[0070] In this embodiment, a feature representation vector is masked at least twice to obtain at least two different feature mask vectors. Different feature mask vectors mean that each feature mask vector obtained by the masking process contains at least one different masked element. The masking process can be random masking, which means randomly masking the elements in the feature representation vector.

[0071] In this embodiment, for each feature characterization vector, the computer device performs a masking process on the feature characterization vector to obtain a feature mask vector, and performs a masking process on the feature mask vector to obtain a feature mask vector. Furthermore, these two feature mask vectors can be used as a positive sample pair.

[0072] In this embodiment, multiple masking processes are performed on each feature representation vector to obtain multiple different feature mask vectors corresponding to each feature representation vector, including:

[0073] For each feature characterization vector, multiple masking processes are performed on the feature characterization vector according to a preset mask ratio to obtain multiple different feature mask vectors corresponding to the feature characterization vector.

[0074] The mask ratio indicates the degree of masking of the feature representation vector.

[0075] Step S206 : for each feature characterization vector, combining multiple different feature mask vectors corresponding to the feature characterization vector to obtain a positive sample pair corresponding to the feature characterization vector.

[0076] Among them, the positive sample pair includes two feature mask vectors, and these two feature mask vectors correspond to the same feature representation vector.

[0077] Specifically, for each feature representation vector, when the feature representation vector has two different feature mask vectors, the different feature mask vectors are combined into a positive sample pair corresponding to the feature representation vector.

[0078] When the targeted feature characterization vector has more than two different feature mask vectors, the computer device combines the different feature mask vectors of the targeted feature characterization vector in pairs to obtain multiple positive sample pairs corresponding to the targeted feature characterization vector.

[0079] In this embodiment, multiple masking processes are performed on the feature representation vector based on a preset mask to obtain multiple different feature representation vectors corresponding to the feature representation vector. For each feature representation vector, the number of preset masks present in the multiple feature mask vectors corresponding to the feature representation vector is determined, and two feature mask vectors with the same number of masks are combined to form a positive example pair.

[0080] Step S208 : combining the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs.

[0081] Among them, the two feature mask vectors in the negative sample pair come from the positive sample pair corresponding to different feature representation vectors.

[0082] Specifically, the computer device determines positive sample pairs corresponding to different feature representation vectors, combines the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors, and obtains multiple negative sample pairs. The two feature mask vectors in the negative sample pairs correspond to different feature representation vectors.

[0083] like Figure 5As shown, feature representation vector i1 is masked three times to obtain feature mask vector j1, feature mask vector j2, and feature mask vector j3 corresponding to feature representation vector i1. Feature representation vector i2 is masked three times to obtain feature mask vector j4, feature mask vector j5, and feature mask vector j6 corresponding to feature representation vector i2. Feature mask vectors j1, j2, and j3 corresponding to feature representation vector i1 are combined in pairs to obtain positive example pairs (j1, j2), (j1, j3), and (j2, j3). Feature mask vectors j4, j5, and j6 corresponding to feature representation vector i2 are combined in pairs to obtain positive example pairs (j4, j5), (j4, j6), and (j5, j6).

[0084] Combine feature mask vector j1 with feature mask vector j4, feature mask vector j5, and feature mask vector j6 to obtain negative example pairs (j1, j4), (j1, j5), and (j1, j6). Combine feature mask vector j2 with feature mask vector j4, feature mask vector j5, and feature mask vector j6 to obtain negative example pairs (j2, j4), (j2, j5), and (j2, j6). Combine feature mask vector j3 with feature mask vector j4, feature mask vector j5, and feature mask vector j6 to obtain negative example pairs (j3, j4), (j3, j5), and (j3, j6).

[0085] Step S210: Perform model training based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

[0086] Specifically, the computer device trains the feature extraction model to be trained based on each positive sample pair and each negative sample pair, and obtains the feature extraction model when the training is completed.

[0087] Furthermore, the computer device can obtain the negative example weight corresponding to each negative example sample pair, and train the feature extraction model to be trained based on each positive example sample pair, each negative example sample pair and the corresponding negative example weight, and obtain the feature extraction model when the training is completed.

[0088] In this embodiment, the feature extraction model to be trained includes a feature representation extraction layer and a representation conversion layer. The feature representation extraction layer is used to extract features from sample information to obtain a feature representation vector. The representation conversion layer is used to mask the feature representation vector output by the feature representation extraction layer and normalize the masked feature vector to reduce computational complexity, thereby obtaining a final feature mask vector. The trained feature extraction model includes the feature representation extraction layer.

[0089] Furthermore, the feature extraction model to be trained also includes a feature embedding layer, which is used to convert sample information into corresponding sample embedding vectors. A feature representation extraction layer is used to extract features of the sample embedding vector output by the feature embedding layer to obtain a feature representation vector. The trained feature extraction model includes the feature embedding layer and the feature representation extraction layer.

[0090] In the above-mentioned feature extraction model processing method, by obtaining multiple sample information, the features of each sample information are extracted respectively, and the feature representation vector corresponding to each sample information is obtained, thereby obtaining the feature representation vector of each basic sample. Each feature representation vector is subjected to multiple masking processes respectively to obtain multiple different feature mask vectors corresponding to each feature representation vector, and the representation vector can be amplified by masking. For each feature representation vector, the multiple different feature mask vectors corresponding to the feature representation vector are combined to obtain the positive sample pair corresponding to the feature representation vector, and the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors are combined to obtain the negative sample pair, thereby obtaining more positive sample pairs and negative sample pairs, thereby achieving the amplification of the training samples. Moreover, through sample amplification, more training samples can be obtained on the basic samples, which can reduce the number requirements for the basic samples. Model training is performed based on the positive sample pairs and negative sample pairs obtained by amplification, so that the robustness of the feature extraction model obtained by training is improved, and the precision and accuracy of the feature extraction model can be improved.

[0091] like Figure 6 The figure shows the architecture of a feature extraction model to be trained in an embodiment. The feature extraction model to be trained includes a feature embedding layer, a feature representation extraction layer and a representation conversion layer. Send it to the feature embedding layer, embed it, and obtain the sample embedding vector embedding, that is, Then, through the feature representation extraction network, the feature representation vector is obtained , .

[0092] The representation conversion layer is used for local shielding and conversion of feature representation vectors. , characterizing the conversion layer according to the preset mask ratio , randomly masked feature representation vector Elements at different positions in , get a feature mask vector . In the same way, the feature representation vector Perform random masking again to obtain another feature mask vector Then, through the representation conversion layer, the feature mask vector , feature mask vector , mapped to feature mask vector and the feature mask vector .

[0093] Construction of self-supervised positive and negative sample pairs: feature mask vectors and the feature mask vector Constitute a positive example pair. The feature mask vector The feature mask vectors corresponding to other feature representation vectors are combined into negative sample pairs, and the feature mask vectors The feature mask vectors corresponding to other feature representation vectors are combined into negative sample pairs. Negative sample pairs are automatically sampled using a self-supervised paradigm.

[0094] Contrastive learning task construction: After the positive and negative sample pairs are constructed, a loss function is constructed to perform modeling. During model training, the parameters of the feature embedding layer, the parameters of the feature representation extraction network, the parameters of the representation conversion layer, and the negative example weights of the negative sample pairs are dynamically adjusted.

[0095] In this embodiment, a self-supervised contrastive learning (CL) method is used to train the feature extraction model using pairs of positive and negative samples formed by feature mask vectors. Contrastive learning is a learning method that minimizes the distance between feature mask vectors in positive sample pairs and maximizes the distance between feature mask vectors in negative sample pairs.

[0096] Self-supervised contrastive learning can extract important discriminative features in training samples. Therefore, by implementing self-supervised training of the feature extraction model through contrastive learning, important discriminative features can be extracted, thereby improving the accuracy of the features extracted by the trained feature extraction model and improving the robustness and generalization of the feature extraction model.

[0097] In one embodiment, the feature characterization vector includes multiple elements, and multiple masking processes are performed on each feature characterization vector to obtain multiple different feature mask vectors corresponding to each feature characterization vector, including:

[0098] Obtain a preset mask ratio, and determine the number of mask elements corresponding to each feature characterization vector based on the mask ratio and the number of elements in each feature characterization vector; for each feature characterization vector, according to the number of mask elements corresponding to the target feature characterization vector, randomly perform multiple masking processes on the number of mask elements in the target feature characterization vector to obtain multiple different feature mask vectors corresponding to the target feature characterization vector.

[0099] The mask ratio is used to indicate the degree of masking of the feature representation vector. The number of mask elements refers to the number of elements in the feature representation vector that need to be masked.

[0100] Specifically, the computer device obtains a preset mask ratio and determines the number of elements in each feature representation vector. For each feature representation vector, the computer device determines the number of mask elements corresponding to the feature representation vector based on the mask ratio and the number of elements in the feature representation vector. The computer device randomly performs a masking process multiple times on the elements in the feature representation vector according to the number of mask elements corresponding to the feature representation vector, thereby obtaining multiple different feature mask vectors corresponding to the feature representation vector.

[0101] In this embodiment, for each feature representation vector, the product of the mask ratio and the number of elements in the feature representation vector is used as the number of mask elements in the feature representation vector. The masking process is performed multiple times on the number of mask elements in the feature representation vector to obtain multiple different feature mask vectors. For multiple masking processes of the same feature representation vector, the masking process is performed on the number of mask elements each time, and the number of mask elements cannot be exactly the same each time to ensure that the multiple feature mask vectors obtained are different.

[0102] In this embodiment, the computer device may have multiple preset mask ratios, and different mask ratios are used to mask feature representation vectors with different numbers of elements. The more elements in the feature representation vector, the larger the corresponding mask ratio, and the smaller the number of elements in the feature representation vector, the smaller the corresponding mask ratio. For example, if the feature representation vector includes 6 elements, the mask ratio corresponding to the feature representation vector is 30%. If the feature representation vector includes 10 elements, the mask ratio corresponding to the feature representation vector is 50%.

[0103] The computer device can determine the number of elements in each feature characterization vector, and for each feature characterization vector, obtain the corresponding mask ratio based on the number of elements in the targeted feature characterization vector, and randomly perform multiple masking processes on the number of mask elements in the targeted feature characterization vector according to the number of mask elements corresponding to the targeted feature characterization vector to obtain multiple different feature mask vectors corresponding to the targeted feature characterization vector.

[0104] In this embodiment, a preset mask ratio is obtained, and based on the mask ratio and the number of elements in each feature characterization vector, the number of mask elements corresponding to each feature characterization vector is determined to determine how many elements in each feature characterization vector need to be masked. For each feature characterization vector, according to the number of mask elements corresponding to the target feature characterization vector, the number of mask elements in the target feature characterization vector is randomly masked multiple times to obtain multiple different feature mask vectors corresponding to the target feature characterization vector, so that each basic sample can be amplified through random masking to obtain more similar samples. Moreover, the elements masked each time are random, so that the obtained feature mask vector is more consistent with the acquisition of samples in model training, and the generated feature mask vector is more universal.

[0105] In one embodiment, model training is performed based on positive sample pairs and negative sample pairs to obtain a feature extraction model, including:

[0106] Determine the positive example similarity corresponding to each positive example sample pair, where the positive example similarity represents the similarity between the feature mask vectors in the positive example sample pair; determine the negative example similarity corresponding to each negative example sample pair, where the negative example similarity represents the similarity between the feature mask vectors in the negative example sample pair; perform model training based on each positive example similarity and each negative example similarity to obtain a feature extraction model.

[0107] The positive similarity corresponding to a positive sample pair refers to the similarity between the two feature mask vectors in the positive sample pair. The negative similarity corresponding to a negative sample pair refers to the similarity between the two feature mask vectors in the negative sample pair.

[0108] Both the positive example similarity and the negative example similarity can be cosine similarity, Pearson correlation coefficient, Manhattan distance, etc., but are not limited thereto.

[0109] Specifically, for each positive example sample pair, the computer device calculates the similarity between the feature mask vectors in the positive example sample pair, and uses the similarity as the positive example similarity corresponding to the positive example sample pair.

[0110] Similarly, for each negative example sample pair, the computer device calculates the similarity between the feature mask vectors in the negative example sample pair, and uses the similarity as the negative example similarity corresponding to the negative example sample pair.

[0111] The computer device performs model training based on the similarity of each positive example and the similarity of each negative example to obtain a feature extraction model.

[0112] In this embodiment, the computer device obtains the negative example weight corresponding to each negative example similarity, performs model training based on each positive example similarity, each negative example similarity and the corresponding negative example weight, and obtains a feature extraction model.

[0113] In this embodiment, the similarity between the feature mask vectors in the positive sample pairs is calculated as the positive example similarity, and the similarity between the feature mask vectors in the negative sample pairs is calculated as the negative example similarity. Model training is performed based on each positive example similarity and each negative example similarity to achieve training of the feature extraction model through self-supervised contrastive learning, so that during the training process, the difference between the feature mask vectors in the positive sample pairs is minimized and the difference between the feature mask vectors in the negative sample pairs is maximized, so that the feature extraction model obtained by training can extract key distinguishing features.

[0114] In one embodiment, model training is performed based on the similarity of each positive example and each negative example to obtain a feature extraction model, including:

[0115] Based on the negative example similarity corresponding to each negative example sample pair, the negative example weight corresponding to each negative example similarity is determined respectively; according to each positive example similarity, each negative example similarity and the corresponding negative example weight, the feature extraction loss is determined; based on the feature extraction loss, the model is trained to obtain a feature extraction model.

[0116] The negative example weight represents the model's attention to negative example pairs during training. The larger the negative example weight, the more attention the model pays to the negative example pairs, and the smaller the negative example weight, the less attention the model pays to the negative example pairs.

[0117] Specifically, for each negative example similarity, the computer device determines a negative example weight corresponding to the targeted negative example similarity based on the negative example similarity corresponding to each negative example sample pair. For example, for a negative example similarity, the computer device may determine the negative example weight corresponding to the targeted negative example similarity based on all negative example similarities. Alternatively, the computer device may also determine negative example similarities associated with the targeted negative example similarity, and determine the negative example weight corresponding to the targeted negative example similarity based on the targeted negative example similarity and these associated negative example similarities.

[0118] If the same feature mask vector exists in two negative sample pairs, it means that the two negative sample pairs are associated, and it also means that the negative similarities of the two negative sample pairs are associated.

[0119] The computer device determines a feature extraction loss based on each positive example similarity, each negative example similarity, and the corresponding negative example weight. Further, the computer device determines a contrast loss between each positive example sample pair and each negative example sample pair based on each positive example similarity, each negative example similarity, and the corresponding negative example weight, and determines a feature extraction loss based on each contrast loss.

[0120] The computer device trains the feature extraction model to be trained based on the feature extraction loss to obtain a feature extraction model.

[0121] In this embodiment, the computer device adjusts the parameters of the feature extraction model to be trained based on the feature extraction loss and continues training, stopping when a preset training condition is met to obtain the feature extraction model. Furthermore, the computer device adjusts the negative example weights of the feature extraction model to be trained based on the feature extraction loss and continues training, stopping when a preset training condition is met to obtain the feature extraction model.

[0122] The preset training condition refers to the preset condition for stopping training, which may be that the feature extraction loss is less than the preset loss, or the number of training iterations reaches the preset number of iterations, etc.

[0123] Different pairs of negative samples can provide different information. The easier it is to distinguish a negative sample pair, the less loss it will generate during training. The more difficult it is to distinguish a negative sample pair, the more loss it will generate during training. Therefore, the model should pay more attention to the difficult-to-distinguish negative sample pairs to strengthen the model during training. In this embodiment, based on the negative similarity corresponding to each negative sample pair, the negative weight corresponding to each negative similarity is determined, so that a smaller weight is assigned to the easily distinguishable negative sample pairs and a larger weight is assigned to the difficult-to-distinguish negative sample pairs. This allows the model to pay more attention to the difficult-to-distinguish negative sample pairs during training. Based on each positive similarity, each negative similarity, and the corresponding negative weight, the feature extraction loss generated jointly by the positive sample pair and each negative sample pair is determined. This allows the model to be trained based on the feature extraction loss, so that the feature extraction model obtained through training can automatically assign weights to different negative sample pairs, making it easier for the model to focus on important detail features.

[0124] In one embodiment, based on the negative example similarity corresponding to each negative example sample pair, respectively determining the negative example weight corresponding to each negative example similarity includes:

[0125] For each feature mask vector, each negative example sample pair including the targeted feature mask vector is screened out; for the negative example similarity corresponding to each screened negative example sample pair, the negative example weight corresponding to the targeted negative example similarity is determined based on the targeted negative example similarity and the negative example similarity corresponding to each screened negative example sample.

[0126] Specifically, for each feature mask vector, the computer device selects negative sample pairs including the targeted feature mask vector from all negative sample pairs, that is, each selected negative sample pair includes the same feature mask vector.

[0127] The computer device determines the negative example similarity corresponding to each negative example sample pair that is screened out, and for each determined negative example similarity, determines the negative example weight corresponding to the negative example similarity based on the negative example similarity and the negative example similarities corresponding to each negative example sample that is screened out.

[0128] In this embodiment, for each negative example sample pair that is screened out, the negative example weight corresponding to the negative example similarity is determined based on the negative example similarity and the negative example similarity corresponding to each negative example sample that is screened out, including:

[0129] The negative similarities corresponding to each negative sample pair filtered out are averaged; for the negative similarities corresponding to each negative sample pair filtered out, the ratio of the negative similarity to the average is used as the negative weight corresponding to the negative similarity.

[0130] In this embodiment, the feature extraction loss is determined based on each positive example similarity, each negative example similarity, and the corresponding negative example weight, including:

[0131] For each feature mask vector, a positive sample pair including the targeted feature mask vector and each negative sample pair including the targeted feature mask vector are screened out; based on the positive similarity corresponding to the screened positive sample pair, the negative similarity corresponding to each screened negative sample pair, and the corresponding negative weight, a contrast loss is determined, where the contrast loss represents the loss between the screened positive sample and each screened negative sample pair; a feature extraction loss is determined based on each contrast loss.

[0132] In this embodiment, for each feature mask vector, each negative example sample pair including the targeted feature mask vector is screened out. If each of the screened negative example sample pairs includes the same feature mask vector, then there is correlation between the screened negative example sample pairs, and there is also correlation between the negative example similarities of these negative example sample pairs. For the negative example similarity corresponding to each screened negative example sample pair, based on the targeted negative example similarity and the negative example similarity corresponding to each screened negative example sample, a negative example weight corresponding to the targeted negative example similarity is determined. Based on all the negative example similarities with correlation, a different negative example weight can be determined for each negative example similarity, so that during training, the model can allocate different degrees of attention to each negative example sample pair with correlation, thereby making the feature extraction model obtained through training more accurate.

[0133] In one embodiment, determining the feature extraction loss based on each positive example similarity, each negative example similarity, and the corresponding negative example weight includes:

[0134] For each feature mask vector, a positive sample pair including the targeted feature mask vector and each negative sample pair including the targeted feature mask vector are screened out; based on the positive similarity corresponding to the screened positive sample pair, the negative similarity corresponding to each screened negative sample pair, and the corresponding negative weight, a contrast loss is determined, where the contrast loss represents the loss between the screened positive sample and each screened negative sample pair; a feature extraction loss is determined based on each contrast loss.

[0135] Specifically, for each feature mask vector, the computer device selects, from all negative sample pairs, each negative sample pair that includes the targeted feature mask vector, and selects, from all positive sample pairs, each positive sample pair that includes the targeted feature mask vector, wherein each selected negative sample pair and each positive sample pair includes the same feature mask vector. For example, for feature mask vector j1, each selected negative sample pair and each positive sample pair includes the feature mask vector j1.

[0136] For each positive sample screened out, the loss between the positive sample and each negative sample pair screened out, i.e., the contrast loss, is determined based on the positive similarity corresponding to the positive sample, the negative similarity corresponding to each negative sample pair screened out, and the negative weight corresponding to the negative similarity.

[0137] Following the same processing approach, the computer device may calculate the contrast loss between each screened positive sample pair and each screened negative sample pair, and determine the feature extraction loss based on each contrast loss.

[0138] In this embodiment, determining the feature extraction loss based on each contrast loss includes: taking the average of each contrast loss as the feature extraction loss.

[0139] In this embodiment, the contrast loss is determined based on the positive similarity corresponding to the screened positive sample pairs, the negative similarity corresponding to each screened negative sample pair, and the corresponding negative example weights, including:

[0140] Calculate the sum of the product of the negative similarity and the corresponding negative weight of each negative sample pair screened out;

[0141] The contrast loss is determined based on the sum of the positive similarities and products corresponding to the filtered positive sample pairs.

[0142] In one embodiment, the negative example weight corresponding to each negative example similarity and the feature extraction loss can be calculated according to the following formula:

[0143]

[0144]

[0145] in, is a positive sample pair, is a negative sample pair. Indicates The exponential function with base , is the cosine similarity measurement function.

[0146] From the above formula, we can know that and The larger the similarity measurement score between Is for The more difficult it is to distinguish the difficult negative sample pair, so a greater weight will be assigned to this negative sample pair .

[0147] In one embodiment, the feature extraction model to be trained includes a feature embedding layer, a feature representation extraction layer, and a representation conversion layer; model training is performed based on feature extraction loss to obtain the feature extraction model, including:

[0148] Based on the feature extraction loss, the parameters of the feature embedding layer, the parameters of the feature representation extraction layer, and the parameters of the representation conversion layer are adjusted respectively to obtain an updated feature extraction model;

[0149] Based on the feature embedding layer after adjusting the parameters, each sample information is converted into a corresponding sample embedding vector. Based on the feature representation extraction layer after adjusting the parameters, feature extraction is performed on the sample embedding vector output by the feature embedding layer after adjusting the parameters to obtain a feature representation vector. Based on the representation conversion layer after adjusting the parameters, mask processing is performed on the feature representation vector output by the feature representation extraction layer after adjusting the parameters to obtain an updated feature mask vector. Based on the updated feature mask vector, the positive sample pairs and the negative sample pairs are reconstructed, and the updated feature extraction model is trained again until a trained feature extraction model is obtained.

[0150] Since the positive and negative pairs are updated, the negative similarity of the negative pairs is also updated, and therefore the negative weight of the negative similarity is also updated. Therefore, the negative weight is adjusted by adjusting the parameters of the feature embedding layer, the parameters of the feature representation extraction layer, and the parameters of the representation conversion layer.

[0151] In this embodiment, for each feature mask vector, a positive sample pair including the feature mask vector and each negative sample pair including the feature mask vector are screened out. Then, there is a correlation between each negative sample pair and the positive sample pair, and there is also a correlation between the negative similarity of these negative sample pairs and the positive similarity of the positive sample pair. Based on the positive similarity corresponding to the screened positive sample pair, the negative similarity corresponding to each negative sample pair, and the corresponding negative weight, the contrast loss between the screened positive sample and the screened negative sample pair is determined, and the feature extraction loss is determined based on each contrast loss. The model parameters are adjusted based on the feature extraction loss so that the adjusted parameters can minimize the distance between the feature mask vectors in the positive sample pair and maximize the distance between the feature mask vectors in the negative and positive sample pairs, thereby improving the precision and accuracy of the model. In addition, the model can be trained through a self-supervised contrast learning method without the need for manually annotated labels, which can reduce the requirements for training samples and make it easier to obtain samples during model training.

[0152] In one embodiment, the plurality of sample information includes first type of sample information, and the method further includes:

[0153] Acquire multiple second-category sample information, extract features of each second-category sample information, and obtain a feature representation vector for each second-category sample information;

[0154] Model training is performed based on positive and negative sample pairs to obtain a feature extraction model, including:

[0155] Based on the positive sample pairs and the negative sample pairs, the feature extraction model to be trained is trained, and based on the feature representation vector of each second category sample information, the information prediction model to be trained is trained to obtain the information prediction model, which includes the feature extraction model.

[0156] The information prediction model is used to predict the user's conversion rate to information. The conversion rate represents the probability that information will be converted by the user. User conversion to information can be, but is not limited to, clicks or browsing of information.

[0157] Specifically, the computer device obtains a plurality of first-category sample information provided by the information owner, where the first-category sample information includes object information of a push receiving object and converted information of the push receiving object.

[0158] The computer device obtains a plurality of second-category sample information from the information push platform. The second-category sample information includes information pushed by the information push platform and object information of push recipients of the pushed information.

[0159] The computer device extracts features of each first-category sample information to obtain a feature representation vector corresponding to each first-category sample information. The computer device performs multiple masking processes on the feature representation vector corresponding to each first-category sample information to obtain multiple different feature mask vectors corresponding to each feature representation vector. For each feature representation vector, the multiple different feature mask vectors corresponding to the feature representation vector are combined to obtain a positive sample pair corresponding to the feature representation vector; and the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors are combined to obtain a negative sample pair.

[0160] The computer device extracts features of each second-category sample information to obtain a feature representation vector for each second-category sample information. The computer device trains a feature extraction model to be trained based on the positive sample pairs and the negative sample, and trains an information prediction model to be trained based on the feature representation vector for each second-category sample information to obtain an information prediction model. The obtained information prediction model includes the feature extraction model.

[0161] In this embodiment, the information prediction model to be trained includes a feature extraction model to be trained. The computer device trains the feature extraction model to be trained based on positive sample pairs and negative sample pairs, and trains the information prediction model to be trained based on the feature representation vector of each second category sample information to obtain an information prediction model, which includes a feature extraction model.

[0162] like Figure 7 Figure 2 shows a schematic diagram of the architecture of a feature extraction model to be trained and an information prediction model to be trained in accordance with an embodiment. The feature extraction model to be trained includes a feature embedding layer, a feature representation extraction layer, and a representation conversion layer. The information prediction model to be trained includes a feature embedding layer and a feature representation extraction layer.

[0163] The computer device inputs multiple first-category sample information into the feature extraction model to be trained, and inputs multiple second-category sample information into the information prediction model to be trained, so as to simultaneously train the feature extraction model to be trained and the information prediction model to be trained until a trained information prediction model is obtained.

[0164] In this embodiment, the first type of sample information and the second type of sample information are obtained, and the training samples are expanded by masking the first type of sample information to obtain more training samples to train the feature extraction model. The information prediction model to be trained is trained by the second type of sample information, and the information prediction model can be locally trained and globally trained by different types of sample information, and the local training and global training are carried out simultaneously. The information gain of the first type of sample information can be effectively used to strengthen the underlying representation of the model, so that the first type of sample information and the second type of sample information share the underlying representation of the model, which can effectively alleviate the problem of sparse information of the second type of samples. In addition, although the training of the feature extraction model and the information prediction model is synchronized, their respective training is independent and does not interfere with each other, which can greatly reduce the interference caused by the difference conversion rate estimation between different source data.

[0165] In one embodiment, Figure 8 As shown, the second type of sample information includes the pushed information and the object information of the push recipient of the pushed information; based on the positive sample pairs and the negative sample, the to-be-trained feature extraction model is trained, and based on the feature representation vector of each second type of sample information, the to-be-trained information prediction model is trained to obtain the information prediction model, including:

[0166] Step S802: Determine the feature extraction loss of the feature extraction model to be trained based on the positive sample pairs and the negative sample pairs.

[0167] For determining the feature extraction loss, reference may be made to the processing in the above embodiments.

[0168] Step S804 : predicting the conversion rate of the push recipient for the pushed information based on the feature representation vector of each second-category sample information.

[0169] Specifically, the second type of sample information includes pushed information and object information of the push recipient of the pushed information. The feature representation vector of the second type of sample information includes the information representation vector of the pushed information and the object representation vector of the object information.

[0170] For each pushed message, the computer device predicts a predicted conversion rate for the pushed message for the push recipient to which the object information belongs based on the information representation vector for the pushed message and the object representation vector for the object information. Using the same processing method, the predicted conversion rate for the pushed message for the push recipient in each second-category sample message can be predicted.

[0171] Step S806: Obtain the expected conversion rate corresponding to each second-category sample information, and determine the conversion rate prediction loss based on the predicted conversion rate and the expected conversion rate.

[0172] The expected conversion rate corresponding to the second type of sample information, that is, the label corresponding to the second type of sample information, may be a manually labeled label.

[0173] The expected conversion rate refers to the probability that the expected push recipient will convert the pushed information.

[0174] Specifically, the computer device can obtain the expected conversion rate corresponding to each second type of sample information, obtain the loss function, and determine the conversion rate prediction loss based on the loss function, the predicted conversion rate and the expected conversion rate.

[0175] The loss function can be a cross entropy loss function, as shown in the following formula:

[0176]

[0177] Among them, M represents the number of the second type of sample information, is the expected conversion rate of the second type of sample information, To predict the conversion rate.

[0178] Step S808 : training the feature extraction model to be trained and the information prediction model to be trained according to the feature extraction loss and the conversion rate prediction loss to obtain an information prediction model.

[0179] Specifically, the computer device determines the target loss according to the feature extraction loss and the conversion rate prediction loss, and trains the feature extraction model to be trained and the information prediction model to be trained based on the target loss to obtain the information prediction model.

[0180] In this embodiment, the Figure 9 The target loss function shown calculates the target loss :

[0181]

[0182] in, are the preset hyperparameters. is the feature extraction loss.

[0183] In this embodiment, based on the positive sample pairs and negative sample pairs, the feature extraction loss of the feature extraction model to be trained is determined, and based on the feature representation vector of each second-category sample information, the predicted conversion rate of the push recipient for the pushed information is predicted; the expected conversion rate corresponding to each second-category sample information is obtained, and the conversion rate prediction loss is determined based on the predicted conversion rate and the expected conversion rate. Two training branches can be formed using two types of sample information of different categories, and the processing of the two training branches is independent of each other and complementary. Finally, the feature extraction model to be trained and the information prediction model to be trained are trained together based on the losses generated by the two training branches. This can effectively alleviate the problem of sparse training samples in the conversion rate estimation branch, thereby fully training the information prediction model through a large number of training samples, making the obtained information prediction model more accurate.

[0184] In one embodiment, the method further comprises:

[0185] Acquire multiple candidate information, extract the features of each candidate information through a feature extraction model, and obtain an information representation vector for each candidate information; acquire multiple candidate receiving objects, extract the features of each candidate receiving object through a feature extraction model, and obtain an object representation vector for each candidate receiving object; based on the information representation vector of each candidate information and the object representation vector of each candidate receiving object, predict the conversion rate of each candidate receiving object corresponding to each candidate information; based on the conversion rate, push at least one candidate information to each candidate receiving object.

[0186] Specifically, the computer device obtains multiple candidate information and determines multiple candidate receiving objects. The computer device obtains object information for each candidate receiving object, inputs the multiple candidate information and the multiple object receiving information into a feature extraction model, extracts features of each candidate information using the feature extraction model, and obtains an information representation vector for each candidate information. Furthermore, the computer device extracts features of each candidate receiving object using the feature extraction model to obtain an object representation vector for each candidate receiving object.

[0187] For each candidate receiving object, the computer device predicts the conversion rate of each candidate information corresponding to the candidate receiving object based on the object representation vector of the candidate receiving object and the information representation vector of each candidate information, thereby obtaining the conversion rate of each candidate information corresponding to each candidate receiving object.

[0188] For each candidate receiving object, the computer device selects at least one candidate information from the candidate information based on the conversion rate of each candidate information corresponding to the targeted candidate receiving object, and pushes the selected candidate information to the targeted object.

[0189] In this embodiment, after the computer device screens out the candidate information, the screened candidate information can be pushed to the targeted object through at least one information push platform.

[0190] In this embodiment, after training the feature extraction model, multiple candidate information can be obtained. The features of each candidate information are extracted through the feature extraction model to obtain the information representation vector of each candidate information. Multiple candidate receiving objects are obtained. The features of each candidate receiving object are extracted through the feature extraction model to obtain the object representation vector of each candidate receiving object. Therefore, based on the feature extraction model obtained through training, the features of each candidate information and the features of each candidate receiving object can be accurately extracted, so that the predicted conversion rate is more accurate. Therefore, based on the conversion rate, at least one candidate information can be pushed to each candidate receiving object, so that the pushed information is more likely to be converted by the user, effectively improving the conversion rate of information.

[0191] In this embodiment, the information prediction model obtained through training includes a feature extraction model. The computer device can obtain object information of multiple candidate information and multiple candidate receiving objects, and input the multiple candidate information and multiple object information into the information prediction model. The information prediction model can directly predict the conversion rate of each candidate receiving object corresponding to each candidate information, thereby improving the prediction efficiency and accuracy of the conversion rate.

[0192] In one embodiment, based on the conversion rate, at least one candidate information is pushed to each candidate receiving object, including:

[0193] For each candidate receiving object, based on the conversion rate of each candidate information corresponding to the targeted candidate receiving object, multiple information whose conversion rates meet the preset push conditions are screened out from multiple candidate information; an information flow corresponding to the targeted candidate receiving object is generated, and the information flow is pushed to the targeted candidate receiving object, where the information flow includes multiple information.

[0194] Specifically, for each candidate recipient, the computer device selects at least one candidate information from multiple candidate information based on the conversion rate of each candidate information corresponding to the targeted candidate recipient. When one candidate information is selected, the information is pushed to the targeted candidate recipient. When multiple candidate information is selected, an information stream is generated based on the multiple candidate information and pushed to the targeted candidate recipient.

[0195] Furthermore, for each candidate recipient, the computer device selects, based on the conversion rate of each candidate information corresponding to the candidate recipient, multiple pieces of information whose conversion rates meet a preset push condition from the multiple candidate information. The preset push condition can be selecting candidate information with a conversion rate exceeding a threshold, or selecting a preset number of pieces of information with a conversion rate exceeding the threshold.

[0196] The computer device generates an information flow based on the filtered multiple information, and the information flow includes the filtered multiple information. The generated information flow is as follows: Figure 10 shown.

[0197] In this embodiment, for each candidate receiving object, the computer device splices multiple candidate information to generate an information stream according to the conversion rate of each candidate information corresponding to the targeted candidate receiving object, and pushes the information stream to the targeted candidate receiving object.

[0198] In this embodiment, the computer device may push the information flow to the targeted candidate recipients through at least one information push platform.

[0199] In this embodiment, for each candidate receiving object, based on the conversion rate of each candidate information corresponding to the targeted candidate receiving object, multiple information whose conversion rates meet the preset push conditions are screened out from multiple candidate information, an information flow corresponding to the targeted candidate receiving object is generated, and the information flow is pushed to the targeted candidate receiving object, so that the information can be recommended according to the user's conversion rate of the information, which is conducive to improving the exposure rate and conversion rate of the information.

[0200] In one embodiment, the multiple sample information includes at least one of the first category sample information or the second category sample information; the first category sample information includes the object information of the push receiving object and the converted information of the push receiving object, and the second category sample information includes the pushed information of the information push platform and the object information of the push receiving object of the pushed information.

[0201] Specifically, the computer device may obtain at least one of the following sample information: first-category sample information or second-category sample information. The first-category sample information is different from the second-category sample information. The first-category sample information includes the object information of the push recipient and the converted information of the push recipient, while the second-category sample information includes the pushed information of the information push platform and the object information of the push recipient of the pushed information.

[0202] When the acquired sample information includes first-category sample information or second-category sample information, feature extraction, mask processing, combined feature mask vectors, etc. are performed on the first-category sample information or the second-category sample information to obtain multiple positive sample pairs and multiple negative sample pairs.

[0203] When the computer device obtains the first category sample information and the second category sample information, it also performs feature extraction, mask processing, and combined feature mask vectors on the first category sample information and the second category sample information respectively to obtain multiple positive sample pairs and multiple negative sample pairs.

[0204] The feature extraction model is obtained by model training based on multiple positive sample pairs and multiple negative sample pairs.

[0205] In this embodiment, the sources of the first type of sample information and the second type of sample information are different. The first type of sample information is provided by the information owner, and the second type of sample information is provided by the information push platform.

[0206] In this embodiment, by obtaining at least one of the first or second types of sample information, training samples are augmented based on sample information from different sources. This allows for the acquisition of more basic samples from diverse sources when training samples are sparse, thereby augmenting the training samples. Model training is performed using a large number of training samples, such as positive sample pairs and multiple negative samples, constructed from these augmented samples, thereby improving model accuracy.

[0207] In one embodiment, Figure 11 As shown, a feature extraction model processing method is provided, which is applied to a computer device, including:

[0208] Step S1102 : Acquire multiple first-category sample information, extract features of each first-category sample information, and obtain a feature representation vector corresponding to each first-category sample information.

[0209] Step S1104 : obtaining a preset mask ratio, and determining the number of mask elements corresponding to each feature representation vector based on the mask ratio and the number of elements in each feature representation vector.

[0210] Step S1106: For each feature characterization vector, according to the number of mask elements corresponding to the feature characterization vector, randomly perform multiple masking processes on the number of mask elements in the feature characterization vector to obtain multiple different feature mask vectors corresponding to the feature characterization vector.

[0211] Step S1108 , combining multiple different feature mask vectors corresponding to the feature characterization vector to obtain a positive sample pair corresponding to the feature characterization vector; combining the feature mask vectors in the positive sample pairs corresponding to different feature characterization vectors to obtain a negative sample pair.

[0212] Step S1110 , determining the positive similarity corresponding to each positive sample pair, where the positive similarity represents the similarity between the feature mask vectors in the positive sample pair; and determining the negative similarity corresponding to each negative sample pair, where the negative similarity represents the similarity between the feature mask vectors in the negative sample pair.

[0213] Step S1112 : For each feature mask vector, filter out the positive sample pairs that include the feature mask vector, and the negative sample pairs that include the feature mask vector.

[0214] Step S1114 : For each negative example sample pair that is screened out, the negative example weight corresponding to the negative example similarity is determined based on the negative example similarity and the negative example similarities corresponding to each negative example sample that is screened out.

[0215] In step S1116, based on the positive similarity corresponding to the screened positive sample pairs, the negative similarity corresponding to each screened negative sample pair, and the corresponding negative weight, a contrast loss is determined. The contrast loss represents the loss between the screened positive sample pairs and each screened negative sample pair.

[0216] Step S1118: Determine the feature extraction loss of the feature extraction model to be trained based on each contrast loss.

[0217] Step S1120 , obtaining a plurality of second-category sample information, extracting features of each second-category sample information, and obtaining a feature representation vector of each second-category sample information.

[0218] Step S1122 : predicting the conversion rate of the push recipient for the pushed information based on the feature representation vector of each second-category sample information.

[0219] Step S1124: Obtain the expected conversion rate corresponding to each second-category sample information, and determine the conversion rate prediction loss based on the predicted conversion rate and the expected conversion rate.

[0220] Step S1126 , training the feature extraction model to be trained and the information prediction model to be trained according to the feature extraction loss and the conversion rate prediction loss to obtain an information prediction model, where the information prediction model includes the feature extraction model.

[0221] It is understood that the order of executing steps S1102-S1118 and steps S1120-S1124 is not limited. For example, steps S1102-S1118 may be executed first and then steps S1120-S1124, or steps S1120-S1124 may be executed first and then steps S1102-S1118, or they may be executed simultaneously.

[0222] In one embodiment, an application scenario of feature extraction model processing is provided, which is specifically applied to the information conversion rate prediction scenario. In this embodiment, a training information prediction model is used. The architecture diagram of the training information prediction model is as follows: Figure 12As shown, the information prediction model to be trained includes a feature extraction model to be trained, which includes a feature embedding layer, a feature representation extraction layer, and a representation conversion layer. The training of the information prediction model to be trained is divided into two branches. The first training branch processes the first type of sample information, and the second training branch processes the second type of sample information. The first training branch serves as an auxiliary task, and the second training branch serves as the main task.

[0223] First, multiple first-category sample information provided by advertisers and multiple second-category sample information are obtained from the information push platform, and both the first-category sample information and the second-category sample information are input into the feature embedding layer of the feature extraction model to be trained in the information prediction model to be trained.

[0224] like Figure 12 As shown in the figure, the training information prediction model has two training branches. The first training branch inputs the first category of sample information and outputs the feature extraction loss. The second training branch inputs the second category of sample information and outputs the conversion rate prediction loss. The first training branch includes the feature embedding layer, feature representation extraction layer, and representation conversion layer of the feature extraction model to be trained, while the second training branch includes the feature embedding layer, feature representation extraction layer, and estimation task tower of the feature extraction model to be trained. That is, the second training branch does not require masking and therefore does not include the representation conversion layer.

[0225] After both the first and second category sample information are input into the feature embedding layer, the feature embedding layer converts the first and second category sample information into sample embedding vectors respectively and passes them to the feature representation extraction layer. The feature representation extraction layer extracts features from the sample embedding vectors of the first category sample information to obtain feature representation vectors. , the feature representation vector The feature representation extraction layer extracts features from the sample embedding vector of the second type of sample information, obtains the feature representation vector, and then passes the feature representation vector corresponding to the second type of sample information to the estimation task tower for conversion rate prediction.

[0226] Specifically, the first type of sample information is trained through the feature embedding layer of the feature extraction model Embedding is performed to obtain the sample embedding vector embedding of the first type of sample information, that is , and then pass through the feature representation extraction network to obtain the feature representation vector , , let the feature representation vector vector The dimension is .

[0227] Next, according to the preset mask ratio (0 ) Randomly masked feature representation vector vector The elements at different positions in , that is, the feature representation vector middle Each dimension in the dimension element has The probability is set to 0, thereby obtaining a feature mask vector in which the local element is masked , that is, the feature mask vector There are Elements are masked. Similarly, according to the preset mask ratio (0 ) Randomly masked feature representation vector Elements at different positions in can obtain another feature mask vector . Feature mask vector and the feature mask vector The acquisition process of is completely independent, there may be a dimension with the same value between them, and there may be some dimensions that are different. For example, in the feature mask vector Middle The dimension is a non-zero value that is not masked, and in the feature mask vector Middle Dimensions are masked to 0.

[0228] like Figure 13 As shown, the representation conversion layer adopts a combination structure of a multi-layer feedforward neural network + a layer normalization network, which respectively transforms the feature mask vector and the feature mask vector Input representation conversion layer to map and obtain feature mask vector and feature masks , the principle of characterizing the conversion layer mapping is shown in the following formula:

[0229]

[0230] In the above formula, For the combined network The output of the layer, It is adjustable and can take a value of 3. represents the activation function, Representation layer normalization function, 、 It is the trainable parameter matrix of the feedforward neural network layer, which is subsequently adjusted by the target loss generated. Feature mask or feature mask , the input is the feature mask hour , similarly, the input is the feature mask When obtained .

[0231] For the auxiliary task of contrastive learning for characterizing the enhanced network, the feature mask vector and the feature mask vector Two samples will be formed. These two samples are homologous and correspond to the same feature representation vector , so these two samples are constructed as a positive sample pair.

[0232] The contrastive learning task automatically constructs negative sample pairs through a self-supervised paradigm:

[0233] For the samples in a batch during model training, let After the local masking and transformation process, each original input sample can be converted into two augmented samples for contrastive learning tasks, denoted as , thus obtaining the Sample . Each sample Consider it as an anchor sample and its corresponding homologous sample As positive samples, the two constitute a positive pair , and the rest of the batch Each sample will become negative samples, then each Can be composed Negative pairs .

[0234] For example, the 2N samples in a batch are T1, T2, T3, and T4. T2 is a homologous sample of T1, so T1 and T2 constitute a positive sample pair. The remaining T3 and T4 constitute negative sample pairs with T1 respectively. Then T1 can constitute 2N-2 negative sample pairs.

[0235] Contrastive learning task construction: The goal of the contrastive learning task is to minimize the distance between positive sample pairs and maximize the distance between negative sample pairs, thereby constraining the distribution of different representation vectors in the vector space to be more reasonable. The feature extraction loss generated by the loss function will be guided by the gradient back propagation to optimize the underlying representation vector, and ultimately achieve the purpose of strengthening the underlying representation ability of the model. Since different pairs of negative samples can provide different information, the simpler negative examples that are easier to distinguish will provide less modeling loss, and the more difficult negative examples that are difficult to distinguish will provide more modeling loss. The model should pay more attention to difficult negative sample pairs, which can enhance the model modeling more fully. Therefore, the loss function in this embodiment introduces a dynamic adjustment weight for negative sample pairs. ,as follows:

[0236]

[0237]

[0238] From the above formula, we can know that and The larger the similarity measurement score between Is for The more difficult it is to distinguish the difficult sample, so a larger weight will be assigned to this negative example pair ,vice versa.

[0239] The second category of sample information is embedded in the feature embedding layer of the feature extraction model to be trained, obtaining a sample embedding vector for the second category of sample information. This is then passed through the feature representation extraction network to obtain a feature representation vector corresponding to the second category of sample information. This feature representation vector is then input into the estimation task tower of the information prediction model to be trained, which then outputs the conversion rate prediction loss.

[0240] Finally, the loss function for the conversion rate estimation task The loss function of the contrastive learning auxiliary task is added to form the target loss function of the model, where It will be multiplied by a hyperparameter Used to adjust the contribution of the two parts of the loss, the objective loss function is as follows:

[0241]

[0242] The loss function for the conversion rate estimation task generally uses the cross entropy loss function. When the model is trained, a batch contains The second type of sample information, each of which has a label of During the training process, the information prediction model has a preset conversion rate of each second-category sample information. , then the loss function of the conversion rate estimation task is As shown below:

[0243]

[0244] In this embodiment, a feature extraction model is constructed as an auxiliary task. The auxiliary task and the main task of conversion rate estimation share the bottom layer of the model for joint modeling. The bottom layer of the model is the feature embedding layer and the feature representation extraction network. Through the self-supervised paradigm, the positive and negative sample pairs and negative sample pairs of the first type of sample information are automatically constructed, and then input into the feature extraction model for comparative learning. The information gain of the first type of sample information can be effectively used to strengthen the bottom layer representation of the model. The bottom layer of the model embeds and maps the input features into vector representations. The first type of sample information and the second type of sample information share the bottom layer of the model. Therefore, this solution can use the information gain in the auxiliary task to further strengthen the bottom layer representation on the basis of the modeling of the second type of sample information, which can effectively alleviate the problem of sparse sample information in the main task of conversion rate estimation. Moreover, the upper-layer estimation module in the form of an auxiliary task and independent of the main task of conversion rate estimation does not interfere with each other, which can greatly reduce the interference caused by the differences in different source data keys to the main task of conversion rate estimation.

[0245] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0246] Based on the same inventive concept, the present application also provides a feature extraction model processing device for implementing the feature extraction model processing method involved above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more feature extraction model processing device embodiments provided below can be found in the above limitations on the feature extraction model processing method and will not be repeated here.

[0247] In one embodiment, Figure 14 As shown, a feature extraction model processing device 1400 is provided, comprising: an acquisition module 1402, a mask module 1404, a first combination module 1406, a second combination module 1408 and a training module 1410, wherein:

[0248] The acquisition module 1402 is used to acquire multiple sample information, extract features of each sample information, and obtain a feature representation vector corresponding to each sample information.

[0249] The masking module 1404 is configured to perform multiple masking processes on each feature representation vector to obtain multiple different feature mask vectors corresponding to each feature representation vector.

[0250] The first combining module 1406 is configured to combine, for each feature representation vector, a plurality of different feature mask vectors corresponding to the feature representation vector to obtain a positive sample pair corresponding to the feature representation vector.

[0251] The second combining module 1408 is configured to combine the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs.

[0252] The training module 1410 is used to perform model training based on positive sample pairs and negative sample pairs to obtain a feature extraction model.

[0253] In this embodiment, by obtaining multiple sample information, the features of each sample information are extracted respectively, and the feature representation vector corresponding to each sample information is obtained, thereby obtaining the feature representation vector of each basic sample. Each feature representation vector is subjected to multiple masking processes respectively to obtain multiple different feature mask vectors corresponding to each feature representation vector, and the representation vector can be amplified through masking. For each feature representation vector, the multiple different feature mask vectors corresponding to the feature representation vector are combined to obtain the positive sample pair corresponding to the feature representation vector. The feature mask vectors in the positive sample pairs corresponding to different feature representation vectors are combined to obtain the negative sample pairs, thereby obtaining more positive sample pairs and negative sample pairs, thereby achieving the amplification of the training samples. In addition, through sample amplification, more training samples can be obtained on the basic samples, which can reduce the number of basic samples required. Model training is performed based on the positive sample pairs and negative sample pairs obtained by amplification, so that the robustness of the feature extraction model obtained by training is improved, and the precision and accuracy of the feature extraction model can be improved.

[0254] In one embodiment, the feature characterization vector includes multiple elements, and the mask module 1404 is further used to obtain a preset mask ratio, and determine the number of mask elements corresponding to each feature characterization vector based on the mask ratio and the number of elements in each feature characterization vector; for each feature characterization vector, according to the number of mask elements corresponding to the feature characterization vector, randomly perform multiple masking processes on the number of mask elements in the feature characterization vector to obtain multiple different feature mask vectors corresponding to the feature characterization vector.

[0255] In this embodiment, a preset mask ratio is obtained, and based on the mask ratio and the number of elements in each feature characterization vector, the number of mask elements corresponding to each feature characterization vector is determined to determine how many elements in each feature characterization vector need to be masked. For each feature characterization vector, according to the number of mask elements corresponding to the target feature characterization vector, the number of mask elements in the target feature characterization vector is randomly masked multiple times to obtain multiple different feature mask vectors corresponding to the target feature characterization vector, so that each basic sample can be amplified through random masking to obtain more similar samples. Moreover, the elements masked each time are random, so that the obtained feature mask vector is more consistent with the acquisition of samples in model training, and the generated feature mask vector is more universal.

[0256] In one embodiment, the training module 1410 is further used to determine the positive example similarity corresponding to each positive example sample pair, where the positive example similarity represents the similarity between the feature mask vectors in the positive example sample pair; determine the negative example similarity corresponding to each negative example sample pair, where the negative example similarity represents the similarity between the feature mask vectors in the negative example sample pair; and perform model training based on each positive example similarity and each negative example similarity to obtain a feature extraction model.

[0257] In this embodiment, the similarity between the feature mask vectors in the positive sample pairs is calculated as the positive example similarity, and the similarity between the feature mask vectors in the negative sample pairs is calculated as the negative example similarity. Model training is performed based on each positive example similarity and each negative example similarity to achieve training of the feature extraction model through self-supervised contrastive learning, so that during the training process, the difference between the feature mask vectors in the positive sample pairs is minimized and the difference between the feature mask vectors in the negative sample pairs is maximized, so that the feature extraction model obtained by training can extract key distinguishing features.

[0258] In one embodiment, the training module 1410 is further used to determine the negative example weight corresponding to each negative example similarity based on the negative example similarity corresponding to each negative example sample pair; determine the feature extraction loss according to each positive example similarity, each negative example similarity and the corresponding negative example weight; and perform model training based on the feature extraction loss to obtain a feature extraction model.

[0259] In this embodiment, based on the negative similarity corresponding to each negative sample pair, a negative weight corresponding to each negative similarity is determined, assigning smaller weights to easily distinguishable negative sample pairs and larger weights to difficult-to-distinguish negative sample pairs. This allows the model to focus more attention on difficult-to-distinguish negative sample pairs during training. Based on each positive similarity, each negative similarity, and the corresponding negative weight, the feature extraction loss generated jointly by the positive sample pair and each negative sample pair is determined. Model training is then performed based on the feature extraction loss, allowing the trained feature extraction model to automatically assign weights to different negative sample pairs, making it easier for the model to focus on important detailed features.

[0260] In one embodiment, the training module 1410 is further used to screen out, for each feature mask vector, each pair of negative example samples that includes the targeted feature mask vector; and for each negative example sample pair that is screened out, determine the negative example weight corresponding to the targeted negative example similarity based on the targeted negative example similarity and the negative example similarity corresponding to each negative example sample that is screened out.

[0261] In this embodiment, for each feature mask vector, each negative example sample pair including the targeted feature mask vector is screened out. If each of the screened negative example sample pairs includes the same feature mask vector, then there is correlation between the screened negative example sample pairs, and there is also correlation between the negative example similarities of these negative example sample pairs. For the negative example similarity corresponding to each screened negative example sample pair, based on the targeted negative example similarity and the negative example similarity corresponding to each screened negative example sample, a negative example weight corresponding to the targeted negative example similarity is determined. Based on all the negative example similarities with correlation, a different negative example weight can be determined for each negative example similarity, so that during training, the model can allocate different degrees of attention to each negative example sample pair with correlation, thereby making the feature extraction model obtained through training more accurate.

[0262] In one embodiment, the training module 1410 is further used to screen out, for each feature mask vector, positive sample pairs including the targeted feature mask vector, and negative sample pairs including the targeted feature mask vector; determine the contrast loss based on the positive similarity corresponding to the screened positive sample pairs, the negative similarity corresponding to each screened negative sample pair, and the corresponding negative weight, the contrast loss representing the loss between the screened positive sample and each screened negative sample pair; and determine the feature extraction loss based on each contrast loss.

[0263] In this embodiment, for each feature mask vector, a positive sample pair including the feature mask vector and each negative sample pair including the feature mask vector are screened out. Then, there is a correlation between each negative sample pair and the positive sample pair, and there is also a correlation between the negative similarity of these negative sample pairs and the positive similarity of the positive sample pair. Based on the positive similarity corresponding to the screened positive sample pair, the negative similarity corresponding to each negative sample pair, and the corresponding negative weight, the contrast loss between the screened positive sample and the screened negative sample pair is determined, and the feature extraction loss is determined based on each contrast loss. The model parameters are adjusted based on the feature extraction loss so that the adjusted parameters can minimize the distance between the feature mask vectors in the positive sample pair and maximize the distance between the feature mask vectors in the negative and positive sample pairs, thereby improving the precision and accuracy of the model. In addition, the model can be trained through a self-supervised contrast learning method without the need for manually annotated labels, which can reduce the requirements for training samples and make it easier to obtain samples during model training.

[0264] In one embodiment, the plurality of sample information includes first-category sample information, and the acquisition module 1402 is further configured to acquire a plurality of second-category sample information, extract features of each second-category sample information, and obtain a feature representation vector for each second-category sample information;

[0265] The training module 1410 is also used to train the feature extraction model to be trained based on the positive sample pairs and the negative sample pairs, and to train the information prediction model to be trained based on the feature representation vector of each second category sample information to obtain the information prediction model, which includes the feature extraction model.

[0266] In this embodiment, the first type of sample information and the second type of sample information are obtained, and the training samples are expanded by masking the first type of sample information to obtain more training samples to train the feature extraction model. The information prediction model to be trained is trained by the second type of sample information, and the information prediction model can be locally trained and globally trained by different types of sample information, and the local training and global training are carried out simultaneously. The information gain of the first type of sample information can be effectively used to strengthen the underlying representation of the model, so that the first type of sample information and the second type of sample information share the underlying representation of the model, which can effectively alleviate the problem of sparse information of the second type of samples. In addition, although the training of the feature extraction model and the information prediction model is synchronized, their respective training is independent and does not interfere with each other, which can greatly reduce the interference caused by the difference conversion rate estimation between different source data.

[0267] In one embodiment, the second category of sample information includes pushed information and object information of the push recipient of the pushed information; the training module 1410 is also used to determine the feature extraction loss of the feature extraction model to be trained based on positive sample pairs and negative samples; based on the feature representation vector of each second category of sample information, predict the predicted conversion rate of the push recipient for the pushed information; obtain the expected conversion rate corresponding to each second category of sample information, and determine the conversion rate prediction loss based on the predicted conversion rate and the expected conversion rate; according to the feature extraction loss and the conversion rate prediction loss, train the feature extraction model to be trained and the information prediction model to be trained to obtain an information prediction model.

[0268] In this embodiment, based on the positive sample pairs and negative sample pairs, the feature extraction loss of the feature extraction model to be trained is determined, and based on the feature representation vector of each second-category sample information, the predicted conversion rate of the push recipient for the pushed information is predicted; the expected conversion rate corresponding to each second-category sample information is obtained, and the conversion rate prediction loss is determined based on the predicted conversion rate and the expected conversion rate. Two training branches can be formed using two types of sample information of different categories, and the processing of the two training branches is independent of each other and complementary. Finally, the feature extraction model to be trained and the information prediction model to be trained are trained together based on the losses generated by the two training branches. This can effectively alleviate the problem of sparse training samples in the conversion rate estimation branch, thereby fully training the information prediction model through a large number of training samples, making the obtained information prediction model more accurate.

[0269] In one embodiment, the device further includes a prediction module and a push module, wherein:

[0270] The acquisition module 1402 is also used to obtain multiple candidate information, extract the features of each candidate information through a feature extraction model, and obtain the information representation vector of each candidate information; obtain multiple candidate receiving objects, extract the features of each candidate receiving object through a feature extraction model, and obtain the object representation vector of each candidate receiving object.

[0271] The prediction module is used to predict the conversion rate of each candidate information corresponding to each candidate information based on the information representation vector of each candidate information and the object representation vector of each candidate receiving object.

[0272] The push module is used to push at least one candidate information to each candidate receiving object based on the conversion rate.

[0273] In this embodiment, the information prediction model obtained through training includes a feature extraction model. The computer device can obtain object information of multiple candidate information and multiple candidate receiving objects, and input the multiple candidate information and multiple object information into the information prediction model. The information prediction model can directly predict the conversion rate of each candidate receiving object corresponding to each candidate information, thereby improving the prediction efficiency and accuracy of the conversion rate.

[0274] In one embodiment, the push module is also used to, for each candidate receiving object, based on the conversion rate of each candidate information corresponding to the candidate receiving object, screen out multiple information whose conversion rates meet preset push conditions from multiple candidate information; generate an information flow corresponding to the candidate receiving object, and push the information flow to the candidate receiving object, where the information flow includes multiple information.

[0275] In this embodiment, for each candidate receiving object, based on the conversion rate of each candidate information corresponding to the targeted candidate receiving object, multiple information whose conversion rates meet the preset push conditions are screened out from multiple candidate information, an information flow corresponding to the targeted candidate receiving object is generated, and the information flow is pushed to the targeted candidate receiving object, so that the information can be recommended according to the user's conversion rate of the information, which is conducive to improving the exposure rate and conversion rate of the information.

[0276] In one embodiment, the multiple sample information includes at least one of the first category sample information or the second category sample information; the first category sample information includes the object information of the push receiving object and the converted information of the push receiving object, and the second category sample information includes the pushed information of the information push platform and the object information of the push receiving object of the pushed information.

[0277] In this embodiment, by obtaining at least one of the first or second types of sample information, training samples are augmented based on sample information from different sources. This allows for the acquisition of more basic samples from diverse sources when training samples are sparse, thereby augmenting the training samples. Model training is performed using a large number of training samples, such as positive sample pairs and multiple negative samples, constructed from these augmented samples, thereby improving model accuracy.

[0278] Each module in the feature extraction model processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0279] In one embodiment, a computer device is provided, which may be a terminal or a server. Taking the terminal as an example, its internal structure diagram may be as follows: Figure 15As shown. The computer device includes a processor, memory, input / output interface, communication interface, display unit and input device. The processor, memory and input / output interface are connected via a system bus, and the communication interface, display unit and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a feature extraction model processing method is implemented. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse, etc.

[0280] Those skilled in the art will understand that Figure 15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0281] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0282] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0283] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0284] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0285] For pushed ads, users can refuse or easily refuse ad push.

[0286] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0287] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0288] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A feature extraction model processing method, characterized in that: The method comprises: Acquire multiple sample information, extract features of each sample information respectively, and obtain a feature representation vector corresponding to each sample information; Performing multiple masking processes on each of the feature representation vectors to obtain multiple different feature mask vectors corresponding to each of the feature representation vectors; For each of the feature characterization vectors, combining a plurality of different feature mask vectors corresponding to the feature characterization vector to obtain a positive sample pair corresponding to the feature characterization vector; Combine the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs; Model training is performed based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

2. The method according to claim 1, characterized in that The feature characterization vector includes multiple elements, and the masking process is performed multiple times on each of the feature characterization vectors to obtain multiple different feature mask vectors corresponding to each of the feature characterization vectors, including: Obtaining a preset mask ratio, and determining the number of mask elements corresponding to each of the feature representation vectors based on the mask ratio and the number of elements in each of the feature representation vectors; For each of the feature characterization vectors, according to the number of mask elements corresponding to the feature characterization vector, multiple masking processes are randomly performed on the number of mask elements in the feature characterization vector to obtain multiple different feature mask vectors corresponding to the feature characterization vector.

3. The method according to claim 1, characterized in that The performing model training based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model includes: Determining a positive example similarity corresponding to each positive example pair, wherein the positive example similarity represents a similarity between feature mask vectors in the positive example pair; Determine a negative example similarity corresponding to each of the negative example sample pairs, where the negative example similarity represents a similarity between feature mask vectors in the negative example sample pairs; Model training is performed based on each of the positive example similarities and each of the negative example similarities to obtain a feature extraction model.

4. The method according to claim 3, characterized in that The performing model training based on each of the positive example similarities and each of the negative example similarities to obtain a feature extraction model includes: Based on the negative example similarity corresponding to each of the negative example sample pairs, respectively determine the negative example weight corresponding to each of the negative example similarities; Determining a feature extraction loss according to each of the positive example similarities, each of the negative example similarities, and the corresponding negative example weights; Model training is performed based on the feature extraction loss to obtain a feature extraction model.

5. The method according to claim 4, characterized in that The determining, based on the negative example similarity corresponding to each pair of negative example samples, the negative example weight corresponding to each negative example similarity includes: For each feature mask vector, filter out each negative sample pair that includes the feature mask vector; For each negative example similarity corresponding to the filtered negative example sample pair, a negative example weight corresponding to the targeted negative example similarity is determined based on the targeted negative example similarity and the negative example similarities corresponding to each of the filtered negative example samples.

6. The method according to claim 4, characterized in that The determining of the feature extraction loss according to each of the positive example similarities, each of the negative example similarities, and the corresponding negative example weights includes: For each feature mask vector, filter out the positive sample pairs that include the feature mask vector, and the negative sample pairs that include the feature mask vector; Determine a contrast loss based on the positive similarity corresponding to the screened positive sample pair, the negative similarity corresponding to each screened negative sample pair, and the corresponding negative weight, where the contrast loss represents the loss between the screened positive sample and each screened negative sample pair; A feature extraction loss is determined based on each of the contrastive losses.

7. The method according to claim 1, characterized in that The plurality of sample information includes first type of sample information, and the method further includes: Acquire a plurality of second-category sample information, extract features of each second-category sample information, and obtain a feature representation vector of each second-category sample information; The performing model training based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model includes: Based on the positive sample pairs and the negative sample pairs, the feature extraction model to be trained is trained, and based on the feature representation vector of each second category sample information, the information prediction model to be trained is trained to obtain an information prediction model, which includes a feature extraction model.

8. The method according to claim 7, characterized in that The second type of sample information includes pushed information and object information of a push recipient of the pushed information; the training of a feature extraction model to be trained based on the positive sample pairs and the negative sample, and the training of an information prediction model to be trained based on the feature representation vector of each second type of sample information to obtain an information prediction model, including: Determining a feature extraction loss of a feature extraction model to be trained based on the positive sample pairs and the negative sample pairs; Predicting a predicted conversion rate of the push recipient for the pushed information based on the feature representation vector of each second-category sample information; Obtaining an expected conversion rate corresponding to each piece of second-category sample information, and determining a conversion rate prediction loss based on the predicted conversion rate and the expected conversion rate; The feature extraction model to be trained and the information prediction model to be trained are trained according to the feature extraction loss and the conversion rate prediction loss to obtain an information prediction model.

9. The method according to claim 1, characterized in that The method further comprises: Acquire multiple candidate information, extract features of each candidate information through the feature extraction model, and obtain an information representation vector for each candidate information; Acquire multiple candidate receiving objects, extract features of each candidate receiving object through the feature extraction model, and obtain an object representation vector for each candidate receiving object; Predicting, based on the information representation vector of each candidate information and the object representation vector of each candidate receiving object, a conversion rate of each candidate receiving object corresponding to each candidate information; Based on the conversion rate, at least one candidate information is pushed to each candidate receiving object.

10. The method according to claim 9, characterized in that The pushing at least one piece of candidate information to each candidate receiving object based on the conversion rate includes: For each candidate receiving object, based on the conversion rate of each candidate information corresponding to the candidate receiving object, filter out multiple pieces of information whose conversion rate meets the preset push conditions from the multiple candidate information; An information flow corresponding to the targeted candidate receiving object is generated, and the information flow is pushed to the targeted candidate receiving object, where the information flow includes the multiple information.

11. The method according to any one of claims 1 to 10, characterized in that The multiple sample information includes at least one of the first category sample information or the second category sample information; the first category sample information includes the object information of the push receiving object and the converted information of the push receiving object, and the second category sample information includes the pushed information of the information push platform and the object information of the push receiving object of the pushed information.

12. A feature extraction model processing device, characterized in that: The device comprises: An acquisition module is used to acquire a plurality of sample information, extract the features of each sample information, and obtain a feature representation vector corresponding to each sample information; A masking module, configured to perform multiple masking processes on each of the feature representation vectors to obtain multiple different feature mask vectors corresponding to each of the feature representation vectors; A first combining module is configured to combine, for each of the feature representation vectors, a plurality of different feature mask vectors corresponding to the feature representation vector to obtain a positive sample pair corresponding to the feature representation vector; A second combining module is used to combine the feature mask vectors in the positive sample pairs corresponding to different feature representation vectors to obtain negative sample pairs; The training module is used to perform model training based on the positive sample pairs and the negative sample pairs to obtain a feature extraction model.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.