Network model training method, article recommendation method and device

A semi-normalized neural network architecture for item recommendation reduces hyperparameter adjustments, enhancing training efficiency and improving accuracy by using one normalized and one unnormalized neural network.

CN120316337APending Publication Date: 2025-07-15BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410057745.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, the double tower model requires multiple hyperparameter adjustments during the training process, resulting in large manpower and material resources consumption and low model training efficiency.

Method used

Using a semi-normalized network model, one neural network has a normalization layer and the other neural network does not have a normalization layer. User and item characterization vectors are generated through feature extraction and normalization processing, and parameter weights are updated according to similarity, and hyperparameter adjustments are omitted.

Benefits of technology

It improves model training efficiency, reduces labor costs, saves system resources, improves the accuracy and flexibility of item recommendations, and expands the structural applicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316337A_ABST
    Figure CN120316337A_ABST
Patent Text Reader

Abstract

The invention discloses a network model training method, and an article recommendation method and device, and relates to the technical field of deep learning. A specific embodiment of the training method of the network model comprises the following steps: performing feature extraction on user information and article information according to a semi-normalized network model; determining a corresponding similarity according to a feature extraction result; and updating the parameter weight of the semi-normalized network model according to the similarity. According to the embodiment, the training steps can be reduced, the sensitivity of the model to the training sample is kept, and the model training efficiency is improved. A specific embodiment of the article recommendation method comprises the steps of performing feature extraction on target user information and target article information according to a semi-normalized double-tower model; determining a corresponding similarity according to a feature extraction result; and recommending the target article to the target user under the condition that the similarity meets the condition. According to the embodiment, the feature extraction step can be reduced, the preference items of the user can be accurately recognized, and the item recommendation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly to a method for training a network model, a method and an apparatus for item recommendation. Background Art

[0002] The two-tower model is used to recommend items of interest to users from a large number of candidate items. When training the two-tower model, it is usually necessary to determine the similarity based on the user representation vector and the item representation vector, and update the parameter weights of the two-tower model based on the similarity. When using the two-tower model to recommend items, it is also necessary to determine the similarity based on the user representation vector and the item representation vector, and determine the items of interest to the user based on the similarity and recommend them to the user. Among them, both the user representation vector and the item representation vector are normalized, and multiple hyperparameter adjustments are required during the normalization process to make the sensitivity of the two-tower model to the training samples and actual data moderate.

[0003] In the process of implementing the present invention, the inventors found that the prior art has at least the following problems:

[0004] The number of hyperparameter adjustments is too large, consuming a large amount of manpower and material resources, and the model training efficiency is low. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method for training a network model, a method and an apparatus for item recommendation, which can maintain the sensitivity of the model to training samples, improve the model training efficiency, reduce the labor cost, and save system resources while reducing the training steps.

[0006] To achieve the above object, according to the first aspect of the embodiments of the present invention, a method for training a network model is provided, including:

[0007] In response to receiving user information and item information, generate a user representation vector corresponding to the user information according to a preset first neural network, and generate an item representation vector corresponding to the item information according to a preset second neural network; one of the user representation vector and the item representation vector is normalized, and the other representation vector is not normalized;

[0008] Determine the similarity between the user representation vector and the item representation vector;

[0009] According to the similarity, update the parameter weights of the first neural network and the second neural network, and use the updated first neural network and second neural network as a two-tower model.

[0010] Optionally, generating a user representation vector corresponding to the user information according to a preset first neural network, and generating an item representation vector corresponding to the item information according to a preset second neural network, including:

[0011] In response to receiving the user information, performing feature extraction on the user information according to a first neural network that does not include a normalization layer to obtain a user representation vector that has not undergone normalization processing;

[0012] In response to receiving the item information, performing feature extraction on the item information according to a second neural network that includes a normalization layer to obtain a normalized item representation vector.

[0013] Optionally, the second neural network includes: a feature extraction layer and a normalization layer; performing feature extraction on the item information according to a second neural network that includes a normalization layer to obtain a corresponding item representation vector, including:

[0014] Using the feature extraction layer to perform feature extraction on the item information to obtain an initial item representation vector;

[0015] Using the normalization layer to perform normalization processing on the initial item representation vector to obtain a normalized item representation vector, and using the normalized item representation vector as the item representation vector corresponding to the item information, where the normalization processing does not include hyperparameter adjustment.

[0016] Optionally, generating a user representation vector corresponding to the user information according to a preset first neural network, and generating an item representation vector corresponding to the item information according to a preset second neural network, including:

[0017] In response to receiving the user information, performing feature extraction on the user information according to a first neural network that includes a normalization layer to obtain a normalized user representation vector;

[0018] In response to receiving the item information, performing feature extraction on the item information according to a second neural network that does not include a normalization layer to obtain an item representation vector that has not undergone normalization processing.

[0019] Optionally, the first neural network includes: a feature extraction layer and a normalization layer; performing feature extraction on the user information according to a first neural network that includes a normalization layer to obtain a normalized user representation vector, including:

[0020] Using the feature extraction layer to perform feature extraction on the user information to obtain an initial user representation vector;

[0021] Normalize the initial user representation vector using the normalization layer to obtain a normalized user representation vector, and use the normalized user representation vector as the user representation vector corresponding to the user information. The normalization process does not include hyperparameter adjustment.

[0022] Optionally, update the parameter weights of the first neural network and the second neural network according to the similarity, including:

[0023] Divide the item information into positive item information and negative item information;

[0024] In the similarity, regard the similarity between the user information and the positive item information as the positive similarity, and regard the similarity between the user information and the negative item information as the negative similarity;

[0025] Compare the positive similarity with a pre-set first similarity threshold to obtain a first comparison result, and update the parameter weights of the first neural network and the second neural network according to the first comparison result;

[0026] Compare the negative similarity with a pre-set second similarity threshold to obtain a second comparison result, and update the parameter weights of the first neural network and the second neural network according to the second comparison result.

[0027] According to the second aspect of the embodiments of the present invention, a method for item recommendation is provided, including:

[0028] In response to receiving target user information and target item information, perform feature extraction on the target user information and the target item information according to a pre-set two-tower model to obtain corresponding target user vectors and target item vectors. The two-tower model is obtained by using any one of the methods in the first aspect of the embodiments of the present invention;

[0029] Determine the similarity between the target user vector and the target item vector;

[0030] When the similarity meets a pre-set recommendation condition, recommend the target item to the target user.

[0031] Optionally, the two-tower model includes a first neural network for obtaining the target user vector and a second neural network for obtaining the target item vector. The first neural network does not include a normalization layer. There are multiple target user vectors. Before determining the similarity between the target user vector and the target item vector, the method further includes:

[0032] Normalize multiple target user vectors to obtain multiple normalized target user vectors, which are used to determine the similarity with the target item vector, and the normalization process does not include hyperparameter adjustment.

[0033] According to a third aspect of the embodiments of the present invention, there is provided a training device for a network model, including:

[0034] A feature extraction module, configured to, in response to receiving user information and item information, generate a user representation vector corresponding to the user information according to a preset first neural network, and generate an item representation vector corresponding to the item information according to a preset second neural network; one of the user representation vector and the item representation vector has been normalized, and the other representation vector has not been normalized;

[0035] A similarity measurement module, configured to determine the similarity between the user representation vector and the item representation vector;

[0036] A parameter update module, configured to update the parameter weights of the first neural network and the second neural network according to the similarity, and use the updated first neural network and second neural network as a two-tower model.

[0037] Optionally, generating a user representation vector corresponding to the user information according to a preset first neural network, and generating an item representation vector corresponding to the item information according to a preset second neural network includes:

[0038] In response to receiving user information, perform feature extraction on the user information according to a first neural network that does not include a normalization layer to obtain an unnormalized user representation vector;

[0039] In response to receiving item information, perform feature extraction on the item information according to a second neural network that includes a normalization layer to obtain a normalized item representation vector.

[0040] Optionally, the second neural network includes: a feature extraction layer and a normalization layer; performing feature extraction on the item information according to the second neural network that includes a normalization layer to obtain a corresponding item representation vector includes:

[0041] Use the feature extraction layer to perform feature extraction on the item information to obtain an initial item representation vector;

[0042] Use the normalization layer to perform normalization processing on the initial item representation vector to obtain a normalized item representation vector, and use the normalized item representation vector as the item representation vector corresponding to the item information, and the normalization process does not include hyperparameter adjustment.

[0043] Optionally, according to a preset first neural network, a user representation vector corresponding to the user information is generated, and according to a preset second neural network, an item representation vector corresponding to the item information is generated, including:

[0044] In response to receiving the user information, feature extraction is performed on the user information according to a first neural network including a normalization layer to obtain a normalized user representation vector;

[0045] In response to receiving the item information, feature extraction is performed on the item information according to a second neural network that does not include a normalization layer to obtain an unnormalized item representation vector.

[0046] Optionally, the first neural network includes: a feature extraction layer and a normalization layer; performing feature extraction on the user information according to a first neural network including a normalization layer to obtain a normalized user representation vector, including:

[0047] Use the feature extraction layer to perform feature extraction on the user information to obtain an initial user representation vector;

[0048] Use the normalization layer to perform normalization processing on the initial user representation vector to obtain a normalized user representation vector, and use the normalized user representation vector as the user representation vector corresponding to the user information, and the normalization processing does not include hyperparameter adjustment.

[0049] Optionally, according to the similarity, update the parameter weights of the first neural network and the second neural network, including:

[0050] Divide the item information into positive item information and negative item information;

[0051] In the similarity, the similarity between the user information and the positive item information is used as the positive similarity, and the similarity between the user information and the negative item information is used as the negative similarity;

[0052] Compare the positive similarity with a preset first similarity threshold to obtain a first comparison result, and update the parameter weights of the first neural network and the second neural network according to the first comparison result;

[0053] Compare the negative similarity with a preset second similarity threshold to obtain a second comparison result, and update the parameter weights of the first neural network and the second neural network according to the second comparison result.

[0054] According to the fourth aspect of the embodiments of the present invention, a device for item recommendation is provided, including:

[0055] A feature extraction module, configured to, in response to receiving target user information and target item information, extract features from the target user information and the target item information according to a pre-set twin tower model, so as to obtain corresponding target user vectors and target item vectors, where the twin tower model is obtained by using any one of the methods in the first aspect of the embodiments of the present invention;

[0056] A similarity measurement module, configured to determine the similarity between the target user vector and the target item vector;

[0057] An item recommendation module, configured to recommend the target item to the target user when the similarity meets a pre-set recommendation condition.

[0058] Optionally, the twin tower model includes a first neural network for obtaining a target user vector and a second neural network for obtaining a target item vector, the first neural network does not include a normalization layer, there are multiple target user vectors, and the apparatus further includes:

[0059] A normalization module, configured to perform normalization processing on multiple target user vectors to obtain multiple normalized target user vectors, where the multiple normalized target user vectors are used to determine the similarity with the target item vector, and the normalization processing does not include hyperparameter adjustment.

[0060] According to a fifth aspect of the embodiments of the present invention, there is provided an electronic device, including:

[0061] One or more processors;

[0062] A storage device, configured to store one or more programs,

[0063] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the above embodiments.

[0064] According to a sixth aspect of the embodiments of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any of the above embodiments is implemented.

[0065] One embodiment of the above invention has the following advantages or beneficial effects: Based on user information and item information, two neural networks with different structures and functions are trained to obtain a dual-tower model, which can maintain the sensitivity of the dual-tower model to training samples and improve the training efficiency of the dual-tower model while reducing the training steps; the first neural network does not include a normalization layer, which can enable the first neural network to automatically learn an appropriate vector norm length, thereby maintaining an appropriate sensitivity to training samples and improving the training efficiency of the first neural network; the normalization process performed on the item representation vector does not include hyperparameter adjustment, which can omit training steps such as adjusting hyperparameters multiple times, reduce labor costs, improve the training efficiency of the second neural network, and save system resources; using the trained dual-tower model to query items of interest to the user can improve the accuracy of item recommendation; normalizing the target user vector can avoid the interference of the vector norm on similarity measurement, improve the accuracy of item recommendation, improve the flexibility of item recommendation and the structural scalability of the dual-tower model.

[0066] The further effects of the above non-conventional optional ways will be described in combination with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0068] Figure 1 is a schematic diagram of the main process of the training method of the network model according to an embodiment of the present invention;

[0069] Figure 2 is a schematic diagram of a semi-normalization network model according to a reference embodiment of the present invention;

[0070] Figure 3 is a schematic diagram of similarity calculation according to a reference embodiment of the present invention;

[0071] Figure 4 is a schematic diagram of the main process of the item recommendation method according to an embodiment of the present invention;

[0072] Figure 5 is a schematic diagram of a semi-normalization dual-tower model according to a reference embodiment of the present invention;

[0073] Figure 6 is a schematic diagram of similarity calculation according to another reference embodiment of the present invention;

[0074] Figure 7 is a schematic diagram of the main modules of the training device of the network model according to an embodiment of the present invention;

[0075] Figure 8Schematic diagram of the main modules of the device for item recommendation according to an embodiment of the present invention;

[0076] Figure 9 Exemplary system architecture diagram to which an embodiment of the present invention can be applied;

[0077] Figure 10 Schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. Detailed implementation manners

[0078] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.

[0079] It should be noted that in the technical solution of the present invention, the processing of collection, use, storage, sharing, and transfer of user personal information complies with the provisions of relevant laws and regulations, and it is necessary to inform the user and obtain the consent or authorization of the user. When applicable, technical processing such as de-identification and / or anonymization and / or encryption is performed on user personal information.

[0080] The two-tower model is used to recommend items of interest to users from a large number of candidate items. When training the two-tower model, it is usually necessary to determine the similarity based on the user representation vector and the item representation vector, and update the parameter weights of the two-tower model based on the similarity. When using the two-tower model to recommend items, it is also necessary to determine the similarity based on the user representation vector and the item representation vector, and determine the items of interest to the user based on the similarity and recommend them to the user. Among them, both the user representation vector and the item representation vector are normalized, and multiple hyperparameter adjustments are required during the normalization process to make the sensitivity of the two-tower model to training samples and actual data moderate.

[0081] Too many hyperparameter adjustments consume a large amount of human and material resources, and the model training efficiency is low. For example, the inner product of the user representation vector and the item representation vector corresponding to the positive sample (i.e., the item of interest to the user) is 1, and the inner product of the user representation vector and the item representation vector corresponding to the negative sample (i.e., the item not of interest to the user) is -1. In the case where the item information includes 1 positive sample and 10 negative samples, after passing through the normalization layer (for example, the Softmax layer, that is, the network layer that performs the normalization operation by the Softmax function), the weight ratio of the positive sample is only At this time, it is necessary to introduce the Temperature mechanism, that is, divide by a hyperparameter T before passing through the normalization layer to adjust the sensitivity of the model to samples. For example, if T = 0.1 is set, then according to the above calculation formula, the weight proportion of positive samples will be increased to At this time, the sensitivity of the model to positive samples is relatively moderate. However, in specific practice, the selection and adjustment of hyperparameters are often a repetitive process. If the hyperparameter T is set too large, the model will be less sensitive to samples; if the hyperparameter T is set too small, the model will be overly sensitive to individual samples.

[0082] In view of this, according to the first aspect of the embodiments of the present invention, a training method for a network model is provided.

[0083] Figure 1 It is a schematic diagram of the main process of the training method for the network model according to the embodiments of the present invention. As Figure 1 shown, the training method for the network model according to the embodiments of the present invention mainly includes the following steps S101 to S103.

[0084] Step S101, in response to receiving user information and item information, generate a user representation vector corresponding to the user information according to a pre-set first neural network, and generate an item representation vector corresponding to the item information according to a pre-set second neural network; one of the user representation vector and the item representation vector has undergone normalization processing, and the other representation vector has not undergone normalization processing.

[0085] The execution subject of the embodiments of the present invention trains the network model, and the trained network model is used to recommend items suitable for the user, meeting the user's needs, and of the user's interest to the user. Specifically, the execution subject of the embodiments of the present invention obtains user information and item information. Among them, user information includes: basic data such as user code, user location, etc., and behavioral data such as the user's collection records, like records, and participated marketing activities. Item information includes: data such as the price, category, style, size, model, packaging style, production date, and manufacturer of the item.

[0086] Pre-set a semi-normalization network model. The semi-normalization network model includes two neural networks for feature extraction, which are respectively used to generate a user representation vector of user information and an item representation vector of item information. The structures of these two neural networks are asymmetric. One neural network structure has a normalization layer, while the other neural network structure does not have a normalization layer. Therefore, the functions of these two neural networks are also different. One neural network has a normalization function, and the other neural network does not have a normalization function, so that among the obtained user representation vector and item representation vector, one representation vector has undergone normalization processing, and the other representation vector has not undergone normalization processing. Such a pair of neural network structures is called a semi-normalization network model.

[0087] According to a reference embodiment of the present invention, the semi-normalization network model includes: a first neural network without a normalization layer and a second neural network with a normalization layer; first, according to the first neural network, feature extraction is performed on user information to obtain a user representation vector corresponding to the user information. Among them, the first neural network does not include a normalization layer, that is, the user representation vector has not undergone normalization processing. Then, according to the second neural network, feature extraction is performed on item information to obtain an item representation vector corresponding to the item information. Among them, the second neural network includes a normalization layer, that is, the item representation vector has undergone normalization processing.

[0088] Figure 2 is a schematic diagram of a semi-normalization network model according to a reference embodiment of the present invention. Exemplarily, as Figure 2 shown, the semi-normalization network model 201 includes: a deep neural network B1, a deep neural network B2, and a normalization layer. Specifically, user information is input into the deep neural network B1, and the deep neural network B1 is used to perform feature extraction on the user information to generate a user representation vector. It should be noted that the deep neural network B1 does not include a normalization layer, and the user representation vector has not undergone normalization processing; item information is input into the deep neural network B2, and the deep neural network B2 is used to perform feature extraction on the item information, and then the feature extraction result of the item information is input into the normalization layer, and the normalization layer is used to perform normalization processing on the feature extraction result of the item information, and the normalized feature extraction result is used as the item representation vector.

[0089] It should be noted that the method described in the embodiments of the present invention can achieve the same relative magnitude relationship of the Euclidean distance, inner product, and cosine similarity between any two sets of vectors between the user representation vector and the item representation vector. Specifically, in the recall scenario, the association degree between the same user and different items is usually concerned. Therefore, for the same user representation vector, the relative magnitude relationship of the Euclidean distance, inner product, and cosine similarity between it and different item representation vectors should be consistent. For example, when the Euclidean distance between the user representation vector C1 and the item representation vector C2 is greater than the Euclidean distance between the user representation vector C1 and the item representation vector C3, the inner product (or cosine similarity) between the user representation vector C1 and the item representation vector C2 is less than the inner product (or cosine similarity) between the user representation vector C1 and the item representation vector C3.

[0090] The following gives the proof of the above "consistent relative magnitude relationship": Exemplarily, for any n-dimensional user representation vector U and any two n-dimensional item representation vectors V1 and V2, since the second neural network for feature extraction of item information retains the normalization layer (e.g., L2 Norm layer), the magnitudes of the two item representation vectors V1 and V2 are both 1. When the cosine similarity between U and V1 is less than the cosine similarity between U and V2, the cosine similarity between U and V1 is The cosine similarity between U and V2 is Obviously, there is which is equivalent to U·V1 < U·V2, that is, the inner product between U and V1 is less than the inner product between U and V2. In addition, the Euclidean distance between U and V1 is The Euclidean distance between U and V2 is U·V1 < U·V2 is equivalent to -2U·V1 > 2U·V2, which is equivalent to |U| 2 +1 - 2U·V1 > |U| 2 +1 - 2U·V2, which is equivalent to That is, the Euclidean distance between U and V1 is greater than the Euclidean distance between U and V2. In summary, the method described in the embodiments of the present invention can achieve the same relative magnitude relationship of the Euclidean distance, inner product, and cosine similarity between any two sets of vectors between the user representation vector and the item representation vector.

[0091] The neural network for feature extraction of user information is regarded as the user tower, and the neural network for feature extraction of item information is regarded as the item tower. Omitting the normalization layer in the user tower can enable the user representation vector to learn an appropriate norm length, avoiding restricting the norm length of the user representation vector within 1, making the semi-normalization network model maintain appropriate sensitivity to training samples, being applicable to determining the items that the user is more interested in from multiple items, avoiding introducing the Temperature mechanism during model training, omitting the process of hyperparameter adjustment, saving system resources, reducing R & D costs, and improving training efficiency and model accuracy.

[0092] According to another referenceable embodiment of the present invention, the second neural network includes: a feature extraction layer and a normalization layer; when performing feature extraction on item information according to the second neural network to obtain an item representation vector corresponding to the item information, first use the feature extraction layer to perform feature extraction on the item information to obtain an initial item representation vector; then use the normalization layer to perform normalization processing on the initial item representation vector to obtain a normalized item representation vector, and use the normalized item representation vector as the item representation vector corresponding to the item information. The above normalization processing does not include hyperparameter adjustment, that is, the Temperature mechanism is not introduced in the above normalization layer, and the weight ratio of positive and negative samples in the item information is not adjusted through hyperparameters.

[0093] Exemplarily, as Figure 2 shown, the normalization layer in the semi-normalization network model 201 only performs normalization processing, does not introduce hyperparameters before normalization processing, and thus does not need to use hyperparameters to adjust the weight ratio of positive and negative samples.

[0094] It should be noted that regarding the network model for feature extraction of user information as the user tower and the network model for feature extraction of item information as the item tower, in the prior art, normalization layers are used in both the user tower and the item tower. Therefore, the norm lengths of the obtained user representation vector and item representation vector are both restricted, and the inner product of the two is limited within the range of [-1, 1]. To ensure the appropriate sensitivity of the network model to each sample, it is necessary to use the Temperature mechanism to expand the range of values that the inner product can take. In the method described in the embodiment of the present invention, the normalization layer of the user tower is removed, and only the normalization layer of the item tower is retained, that is, only the norm length of the item representation vector is restricted, and the norm length of the user representation vector is not restricted. At this time, the range of values that the inner product of the two can take is no longer restricted. Therefore, it is no longer necessary to use the Temperature mechanism to adjust the norm length of the item representation vector.

[0095] The neural network for extracting features from user information is regarded as the user tower, and the neural network for extracting features from item information is regarded as the item tower. In the case of omitting the normalization layer in the user tower, the Temperature mechanism is avoided from being introduced during the process of training the item tower, the process of hyperparameter adjustment is omitted, system resources are saved, the R & D cost and maintenance cost are reduced, and the training efficiency and model accuracy are improved.

[0096] Step S102: Determine the similarity between the user representation vector and the item representation vector.

[0097] After obtaining the user representation vector and the item representation vector, calculate the similarity between the user representation vector and the item representation vector by means of inner product, Euclidean distance, cosine similarity, etc. Exemplarily, the formula for calculating the inner product between the user representation vector D1 and the item representation vector D2 is: D sim = D1·D2, that is, perform a dot product operation on the user representation vector and the item representation vector, where D sim represents the similarity calculation result. Exemplarily again, the formula for calculating the cosine similarity between the user representation vector D1 and the item representation vector D2 is: where D sim represents the similarity calculation result. Exemplarily again, the formula for calculating the Euclidean distance between the user representation vector D1 and the item representation vector D2 is: where D sim represents the similarity calculation result, D1 x represents the component of the user representation vector on the x-axis, D2 x represents the component of the item representation vector on the x-axis, D1 y represents the component of the user representation vector on the y-axis, D2 y represents the component of the item representation vector on the y-axis. According to the above various calculation formulas, determine the similarity between the user representation vector and the item representation vector.

[0098] Figure 3 is a schematic diagram of similarity calculation according to a reference embodiment of the present invention. Exemplarily, as Figure 3As shown in the figure, the three groups of vector relationships from left to right respectively represent: calculating the similarity between the user representation vector and the item representation vector using cosine similarity, calculating the similarity between the user representation vector and the item representation vector using Euclidean distance, and calculating the similarity between the user representation vector and the item representation vector using inner product. Among them, vector E1 represents the user representation vector, and vector E2 and vector E3 represent two item representation vectors. When calculating the similarity between the user representation vector and the item representation vector using cosine similarity, the included angle s2 between the user representation vector E1 and the item representation vector E2 is smaller than the included angle s1 between the user representation vector E1 and the item representation vector E3. Since the cosine similarity is proportional to the cosine value of the vector included angle, the cosine similarity between the user representation vector E1 and the item representation vector E2 is greater than the cosine similarity between the user representation vector E1 and the item representation vector E3. When calculating the similarity between the user representation vector and the item representation vector using Euclidean distance, the connection line t2 between the user representation vector E1 and the item representation vector E2 is smaller than the connection line t1 between the user representation vector E1 and the item representation vector E3. Therefore, the Euclidean distance between the user representation vector E1 and the item representation vector E2 is smaller than the Euclidean distance between the user representation vector E1 and the item representation vector E3. When calculating the similarity between the user representation vector and the item representation vector using inner product, since the inner product of two vectors is proportional to the projection distance of one vector on the other vector, and the projection distance r2 of the item representation vector E2 on the user representation vector E1 is greater than the projection distance r1 of the item representation vector E3 on the user representation vector E1, the inner product between the user representation vector E1 and the item representation vector E2 is greater than the inner product between the user representation vector E1 and the item representation vector E3. The above three calculation processes also indirectly prove that the method described in the embodiments of the present invention can achieve the same relative magnitude relationship of Euclidean distance, inner product, and cosine similarity between any two sets of vectors between the user representation vector and the item representation vector, that is, a smaller Euclidean distance between a set of vectors is equivalent to a larger inner product or cosine similarity between them.

[0099] Setting multiple calculation formulas for similarity can improve the flexibility and scalability of similarity calculation and meet different similarity calculation requirements.

[0100] Step S103, update the parameter weights of the first neural network and the second neural network according to the similarity, and use the updated first neural network and second neural network as a two-tower model.

[0101] Exemplarily, different users have different degrees of interest in the same item. The users who are interested in the item are regarded as positive users, and the corresponding user information is positive user information. The users who are not interested in the item are regarded as negative users, and the corresponding user information is negative user information. The similarities between the positive user information and the item information, and between the negative user information and the item information are respectively used as the positive similarity and the negative similarity. The positive similarity represents the similarity between the users who are interested in the item and the item. Therefore, the positive similarity should be close to 1. Adjust the parameter weights between the first neural network and the second neural network to make the positive similarity gradually approach 1. The negative similarity represents the similarity between the users who are not interested in the item and the item. Therefore, the negative similarity should be close to 0. Adjust the parameter weights between the first neural network and the second neural network to make the negative similarity gradually approach 0.

[0102] Update the parameter weights of the first neural network and the second neural network to make the similarity meet the pre-set business requirements (for example, make the positive similarity close to 1 and make the negative similarity close to 0). The updated first neural network and second neural network are used as a two-tower model. The two-tower model is used to recommend items that users are interested in and determine the users who are interested in the items.

[0103] Dividing users into those who are interested in the item and those who are not interested in the item can improve the update efficiency of the parameter weights and the training efficiency of the first neural network and the second neural network.

[0104] According to a reference embodiment of the present invention, the semi-normalized network model includes: a first neural network having a normalization layer and a second neural network without a normalization layer. According to the first neural network including the normalization layer, feature extraction is performed on the user information, and the feature extraction result of the user information is normalized, and then the user representation vector corresponding to the user information is obtained, that is, the user representation vector has been normalized. According to the second neural network without a normalization layer, feature extraction is performed on the item information to obtain the item representation vector corresponding to the item information, that is, the item representation vector has not been normalized.

[0105] Exemplarily, the semi-normalized network model includes neural network F1 and neural network F2. Among them, neural network F1 (including a normalization layer) is used to perform feature extraction and normalization processing on the user information to generate a user feature vector corresponding to the user information, and neural network F2 (without a normalization layer) is used to perform feature extraction on the item information to generate an item feature vector corresponding to the item information.

[0106] The neural network for feature extraction of user information is regarded as the user tower, and the neural network for feature extraction of item information is regarded as the item tower. Omitting the normalization layer in the item tower can enable the item representation vector to learn an appropriate norm length, avoid restricting the norm length of the item representation vector within 1, make the semi-normalized network model maintain appropriate sensitivity to training samples, and is applicable to determining users interested in the target item from multiple users.

[0107] According to another referenceable embodiment of the present invention, the first neural network includes: a feature extraction layer and a normalization layer; when performing feature extraction on user information according to the first neural network including the normalization layer to obtain the corresponding user representation vector, first use the feature extraction layer to perform feature extraction on user information to obtain the initial user representation vector; then use the normalization layer to perform normalization processing on the initial user representation vector to obtain the normalized user representation vector, and take the normalized user representation vector as the user representation vector corresponding to the user information. The above normalization processing does not include hyperparameter adjustment, that is, the Temperature mechanism is not introduced in the above normalization layer, and the weight ratio of positive and negative samples in user information is not adjusted through hyperparameters.

[0108] It should be noted that the network model for feature extraction of user information is regarded as the user tower, and the network model for feature extraction of item information is regarded as the item tower. In the method described in the embodiments of the present invention, the normalization layer of the item tower is removed, and only the normalization layer of the user tower is retained. That is, only the norm length of the user representation vector is restricted, and the norm length of the item representation vector is not restricted. At this time, the value range of the inner product of the two is no longer restricted. Therefore, it is no longer necessary to use the Temperature mechanism to adjust the norm length of the item representation vector.

[0109] The neural network for feature extraction of user information is regarded as the user tower, and the neural network for feature extraction of item information is regarded as the item tower. In the case of omitting the normalization layer in the item tower, it is possible to avoid introducing the Temperature mechanism during the training of the user tower, omit the process of hyperparameter adjustment, save system resources, reduce the R & D cost and maintenance cost, and improve the training efficiency and model accuracy.

[0110] According to a reference embodiment of the present invention, when updating the parameter weights of the first neural network and the second neural network according to the similarity, the item information is divided into positive item information and negative item information, where the positive item information is the item information that the user is interested in, and the negative item information is the item information that the user is not interested in. In the similarity, the similarity between the user information and the positive item information is used as the positive similarity, and the similarity between the user information and the negative item information is used as the negative similarity. The positive similarity is compared with a preset first similarity threshold to obtain a first comparison result. Wherein, the first similarity threshold is less than or equal to 1. According to the first comparison result, the parameter weights of the first neural network and the second neural network are updated so that the positive similarity gradually increases and finally is greater than or equal to the first similarity threshold. The negative similarity is compared with a preset second similarity threshold to obtain a second comparison result. Wherein, the second similarity threshold is greater than or equal to 0 and less than the first similarity threshold. According to the second comparison result, the parameter weights of the first neural network and the second neural network are updated so that the negative similarity gradually decreases and finally is less than or equal to the second similarity threshold.

[0111] Exemplarily, the item representation vector generated according to the item information that the user is interested in is used as the positive sample, and the item representation vector generated according to the item information that the user is not interested in is used as the negative sample. After obtaining the similarity, the parameter weights of the semi-normalized network model are updated, so as to increase the similarity between the positive sample and the user representation vector, that is, to shorten the distance between the positive sample and the user representation vector, and to decrease the similarity between the negative sample and the user representation vector, that is, to increase the distance between the negative sample and the user representation vector. For example, the first similarity threshold is 1, and the second similarity threshold is -1. By adjusting the parameter weights, the similarity between the positive sample and the user representation vector gradually approaches the first similarity threshold, and the similarity between the negative sample and the user representation vector gradually approaches the second similarity threshold until the similarity between the positive sample and the user representation vector is greater than or equal to the first similarity threshold, and the similarity between the negative sample and the user representation vector is less than or equal to the second similarity threshold.

[0112] Updating the parameter weights of the semi-normalized network model according to the similarity can improve the training efficiency and accuracy of the network model.

[0113] According to a second aspect of the embodiments of the present invention, a method for item recommendation is provided.

[0114] Figure 4 It is a schematic diagram of the main process of the method for item recommendation according to the embodiments of the present invention. As Figure 4 shown, the method for item recommendation according to the embodiments of the present invention mainly includes the following steps S401 to step S403.

[0115] Step S401, in response to receiving the target user information and the target item information, according to the pre-set dual-tower model, extract features from the target user information and the target item information to obtain the corresponding target user vector and target item vector, where the dual-tower model is obtained by using any of the methods in the first aspect of the embodiments of the present invention.

[0116] The execution subject of the embodiments of the present invention determines whether to recommend a target item to a target user according to the target user information and the target item information. After receiving the target user information and the target item information, input the target user information and the target item information into the dual-tower model, use the dual-tower model to extract features from the target user information to obtain the target user vector, and extract features from the target item information to obtain the target item vector. The dual-tower model includes two neural network structures for feature extraction, and these two neural network structures are respectively used to generate the target user vector of the target user information and the target item vector of the target item information. These two neural network structures are asymmetric, one of the neural network structures has a normalization layer, and the other neural network structure does not have a normalization layer. Therefore, the dual-tower model composed of such a pair of neural network structures is called a semi-normalized dual-tower model.

[0117] Exemplarily, the semi-normalized dual-tower model includes neural network J1 and neural network J2. Among them, neural network J1 (including a normalization layer) is used to extract features from the target user information to generate the target user vector, and neural network J2 (not including a normalization layer) is used to extract features from the target item information to generate the target item vector.

[0118] Omitting the normalization layer of the user tower or the item tower in the conventional dual-tower model can obtain a more accurate target user vector or target item vector, which is convenient for improving the accuracy of item recommendation.

[0119] According to a referenceable embodiment of the present invention, the semi-normalized dual-tower model includes: a first neural network and a second neural network; when extracting features from the target user information and the target item information according to the pre-set semi-normalized dual-tower model to obtain the corresponding target user vector and target item vector, first extract features from the target user information according to the first neural network to obtain the target user vector corresponding to the target user information, where the first neural network does not include a normalization layer, that is, the target user vector is not normalized. Then, according to the second neural network, extract features from the target item information to obtain the target item vector corresponding to the target item information, where the second neural network includes a normalization layer, that is, the target item vector is normalized.

[0120] Figure 5 It is a schematic diagram of a semi-normalized dual-tower model according to a referenceable embodiment of the present invention. Exemplarily, asFigure 5 As shown in Figure 5 , the semi-normalized two-tower model 501 includes: a deep neural network G1, a deep neural network G2, and a normalization layer. Specifically, the target user information is input into the deep neural network G1, and the deep neural network G1 is used to extract features from the target user information to generate a target user vector. It should be noted that the deep neural network B1 does not include a normalization layer, and the target user vector is not normalized; the target item information is input into the deep neural network G2, and the deep neural network G2 is used to extract features from the target item information, and then the feature extraction result of the target item information is input into the normalization layer, and the normalization layer is used to perform normalization processing on the feature extraction result of the target item information, and the normalized feature extraction result is used as the target item vector.

[0121] Regarding the neural network for extracting features from the target user information as the user tower, and the neural network for extracting features from the target item information as the item tower, omitting the normalization layer in the user tower can obtain a target user vector with an appropriate modulus length, extract more accurate user features, omit the normalization operation, save system resources, and improve the accuracy and efficiency of item recommendation.

[0122] According to another referenceable embodiment of the present invention, the second neural network includes: a feature extraction layer and a normalization layer; when, according to the second neural network, extracting features from the target item information to obtain a target item vector corresponding to the target item information, first use the feature extraction layer to extract features from the target item information to obtain an initial target item vector; use the normalization layer to perform normalization processing on the initial target item vector to obtain a normalized target item vector, and use the normalized target item vector as the target item vector corresponding to the target item information. The above normalization processing does not include hyperparameter adjustment, that is, the Temperature mechanism is not introduced in the above normalization layer, and the initial target item vector is not adjusted through hyperparameters.

[0123] Exemplarily, as Figure 5 shown in Figure 5 , the normalization layer in the semi-normalized two-tower model 501 only performs normalization processing, does not introduce hyperparameters before normalization processing, and thus does not require using hyperparameters to adjust the initial target item vector.

[0124] Regarding the neural network for extracting features from the target user information as the user tower, and the neural network for extracting features from the target item information as the item tower, in the case of omitting the normalization layer in the user tower, avoid introducing the Temperature mechanism in the item tower, omit the process of hyperparameter adjustment, save system resources, and improve the efficiency and accuracy of item recommendation.

[0125] Step S402, determine the similarity between the target user vector and the target item vector.

[0126] After obtaining the user representation vector and the item representation vector, the similarity between the user representation vector and the item representation vector is calculated by means of inner product, Euclidean distance, cosine similarity, etc. The specific implementation manner of calculating the similarity between vectors has been described in detail in the previous embodiments and will not be elaborated here.

[0127] According to another referenceable embodiment of the present invention, there are multiple target user vectors. Before determining the similarity between the target user vector and the target item vector, the method further includes: normalizing the multiple target user vectors to obtain multiple normalized target user vectors, and the multiple normalized target user vectors are used to determine the similarity with the target item vector, and the above normalization process does not include hyperparameter adjustment.

[0128] In the case where there are multiple target user vectors, the magnitudes of the multiple target user vectors of the target user may vary greatly. Without performing the normalization process, it is easy to reduce the accuracy of calculating the similarity between vectors. For example, when using the Euclidean distance as the vector similarity metric to calculate the similarity between multiple target user vectors and multiple target item vectors, the target user vector with a smaller magnitude is likely to interfere with the accuracy of the similarity calculation. Another example is that when using the inner product as the vector similarity metric to calculate the similarity between multiple target user vectors and multiple target item vectors, the user representation vector with a larger magnitude is likely to interfere with the accuracy of the similarity calculation.

[0129] Figure 6 It is a schematic diagram of similarity calculation according to another referenceable embodiment of the present invention. Exemplarily, as Figure 6 shown, vector H1 and vector H2 are two target user vectors, and vectors p1 to p10 are ten target item vectors. It should be noted that: to highlight the two target user vectors, only in Figure 6One end point of the target item vector is shown. The other end point of the target item vector is the intersection of two target user vectors, i.e., the center of the dashed circle. In the vector relationship diagram 601 where the normalization process is not performed on the target user vectors, the magnitudes of the two target user vectors are not equal and are both greater than 1. Currently, it is necessary to find 4 items with relatively strong relevance to the user among the items corresponding to the target item vectors p1 to p10. The expected target item vectors to be found are: p1, p2, p6, and p7. When using the Euclidean distance as the vector similarity metric, since the magnitude of the target user vector H2 is relatively small, the actually found target item vectors are: p1, p2, p3, and p10. It should be noted that the Euclidean distance between the target user vector H2 and the target item vector p10 is less than the Euclidean distance between the target user vector H1 and the target item vector p6, i.e., |H2 - p10| < |H1 - p6|; when using the inner product as the vector similarity metric, since the magnitude of the target user vector H1 is relatively large, the actually found target item vectors are: p5, p6, p7, and p8. It should be noted that the inner product between the target user vector H1 and the target item vector p5 is greater than the inner product between the target user vector H2 and the target item vector p2, i.e., H1·p5 > H2·p2. In summary, without performing the normalization process on the target user vectors, the target items obtained using various similarity calculation methods do not meet the expectations.

[0130] After performing the normalization process on the target user vectors, the vector relationship diagram 602 is obtained. Among them, the magnitudes of the two target user vectors are equal and both equal 1. Currently, it is necessary to find 4 items with relatively strong relevance to the user among the items corresponding to the target item vectors p1 to p10. The expected target item vectors to be found are: p1, p2, p6, and p7. At this time, the target item vectors obtained by any similarity calculation method are all p1, p2, p6, and p7, meeting the expectations.

[0131] It should be noted that the normalization process performed on the target user vectors does not contain any trainable parameters, so it can be directly used for the target user vectors without training. At the same time, the relative magnitude relationships of various similarity metrics between the target user vectors before and after performing the normalization process and different target item vectors remain unchanged. The following gives the proof of the above "relative magnitude relationships remain unchanged": Exemplarily, for any n-dimensional target user vector X and any two n-dimensional target item vectors Y1 and Y2, since the second neural network for feature extraction of the target item information retains the normalization layer (e.g., L2 Norm layer), the magnitudes of the two target item vectors Y1 and Y2 are both 1. After performing the normalization operation on the target user vector, the target user vector X becomes the user representation vector Since the magnitude of the target user vector after performing the normalization operation is 1, so, the user representation vector is still equivalent to the target user vector X; therefore, "the cosine similarity between X and Y1 is less than the cosine similarity between X and Y2" is equivalent to " the cosine similarity with Y1 is less than the cosine similarity with Y2", and "the inner product between X and Y1 is less than the inner product between X and Y2" is equivalent to " the inner product with Y1 is less than the inner product with Y2"; when the Euclidean distance between X and Y1 is less than the Euclidean distance between X and Y2, the Euclidean distance between X and Y1 is the Euclidean distance between X and Y2 is the Euclidean distance with Y1 is the Euclidean distance between X and Y2 is Therefore, is equivalent to X·Y1 > X·Y2, which is equivalent to In summary, the method described in the embodiments of the present invention can achieve that the relative magnitude relationship of various similarity metrics between the target user vector before and after performing the normalization process and different target item vectors remains unchanged.

[0132] Performing the normalization process on the target user vector can avoid the interference of the vector norm length on the similarity metric, improve the accuracy of item recommendation, and improve the flexibility of item recommendation and the structural scalability of the dual tower model.

[0133] Step S403, when the similarity meets the pre-set recommendation condition, recommend the target item to the target user.

[0134] After obtaining the similarity between the target user vector and the target item vector, determine whether the similarity meets the pre-set recommendation condition. For example, sort the similarities in descending order. When 5 items need to be recommended to the user, recommend the target items corresponding to the top 5 target item vectors with the highest similarities to the target user. For another example, compare the similarity with the pre-set similarity threshold, determine the target item vectors with similarities greater than the similarity threshold, use the determined target item vectors as the recommended item vectors, and recommend the target items corresponding to the recommended item vectors to the target user.

[0135] Setting multiple recommendation conditions can improve the flexibility and accuracy of item recommendation and meet different item recommendation requirements.

[0136] According to the third aspect of the embodiments of the present invention, a training device for a network model is provided.

[0137] Figure 7It is a schematic diagram of the main modules of a training device for a network model according to an embodiment of the present invention. As Figure 7 shown, the training device 700 of the network model mainly includes:

[0138] A feature extraction module 701, configured to, in response to receiving user information and item information, generate a user representation vector corresponding to the user information according to a preset first neural network, and generate an item representation vector corresponding to the item information according to a preset second neural network. One of the user representation vector and the item representation vector is normalized, and the other representation vector is not normalized;

[0139] A similarity measurement module 702, configured to determine the similarity between the user representation vector and the item representation vector;

[0140] A parameter update module 703, configured to update the parameter weights of the first neural network and the second neural network according to the similarity, and use the updated first neural network and second neural network as a dual tower model.

[0141] According to a referenceable embodiment of the present invention, generating a user representation vector corresponding to the user information according to a preset first neural network, and generating an item representation vector corresponding to the item information according to a preset second neural network includes:

[0142] In response to receiving user information, perform feature extraction on the user information according to a first neural network that does not include a normalization layer to obtain a user representation vector that is not normalized;

[0143] In response to receiving item information, perform feature extraction on the item information according to a second neural network that includes a normalization layer to obtain a normalized item representation vector.

[0144] According to another referenceable embodiment of the present invention, the second neural network includes a feature extraction layer and a normalization layer; performing feature extraction on the item information according to a second neural network that includes a normalization layer to obtain a corresponding item representation vector includes:

[0145] Use the feature extraction layer to perform feature extraction on the item information to obtain an initial item representation vector;

[0146] Use the normalization layer to perform normalization processing on the initial item representation vector to obtain a normalized item representation vector, and use the normalized item representation vector as the item representation vector corresponding to the item information. The normalization processing does not include hyperparameter adjustment.

[0147] According to another referenceable embodiment of the present invention, a user representation vector corresponding to the user information is generated according to a pre-set first neural network, and an item representation vector corresponding to the item information is generated according to a pre-set second neural network, including:

[0148] In response to receiving the user information, feature extraction is performed on the user information according to a first neural network including a normalization layer to obtain a normalized user representation vector;

[0149] In response to receiving the item information, feature extraction is performed on the item information according to a second neural network not including a normalization layer to obtain an unnormalized item representation vector.

[0150] According to still another referenceable embodiment of the present invention, the first neural network includes: a feature extraction layer and a normalization layer; according to the first neural network including the normalization layer, feature extraction is performed on the user information to obtain a normalized user representation vector, including:

[0151] The feature extraction layer is used to perform feature extraction on the user information to obtain an initial user representation vector;

[0152] The normalization layer is used to perform normalization processing on the initial user representation vector to obtain a normalized user representation vector, and the normalized user representation vector is used as the user representation vector corresponding to the user information, and the normalization processing does not include hyperparameter adjustment.

[0153] According to still another referenceable embodiment of the present invention, the parameter weights of the first neural network and the second neural network are updated according to the similarity, including:

[0154] The item information is divided into positive item information and negative item information;

[0155] In the similarity, the similarity between the user information and the positive item information is used as the positive similarity, and the similarity between the user information and the negative item information is used as the negative similarity;

[0156] The positive similarity is compared with a pre-set first similarity threshold to obtain a first comparison result, and the parameter weights of the first neural network and the second neural network are updated according to the first comparison result;

[0157] The negative similarity is compared with a pre-set second similarity threshold to obtain a second comparison result, and the parameter weights of the first neural network and the second neural network are updated according to the second comparison result.

[0158] It should be noted that the specific implementation content of the training device of the network model in the embodiments of the present invention has been described in detail in the above-mentioned training method of the network model, so the repeated content will not be described here again.

[0159] According to a fourth aspect of the embodiments of the present invention, there is provided a device for item recommendation.

[0160] Figure 8 It is a schematic diagram of the main modules of the device for item recommendation according to the embodiments of the present invention, as Figure 8 shown, the device 800 for item recommendation mainly includes:

[0161] A feature extraction module 801, configured to, in response to receiving target user information and target item information, extract features from the target user information and the target item information according to a pre-set two-tower model, so as to obtain corresponding target user vectors and target item vectors, where the two-tower model is obtained by using any of the methods in the first aspect of the embodiments of the present invention;

[0162] A similarity measurement module 802, configured to determine the similarity between the target user vector and the target item vector;

[0163] An item recommendation module 803, configured to recommend the target item to the target user when the similarity meets a pre-set recommendation condition.

[0164] According to another referenceable embodiment of the present invention, the two-tower model includes a first neural network for obtaining a target user vector and a second neural network for obtaining a target item vector, the first neural network does not include a normalization layer, there are multiple target user vectors, and the device 800 for item recommendation further includes:

[0165] A normalization module, configured to perform normalization processing on multiple target user vectors to obtain multiple normalized target user vectors, where the multiple normalized target user vectors are used to determine the similarity with the target item vector, and the normalization processing does not include hyperparameter adjustment.

[0166] According to the technical solution of the embodiment of the present invention, two neural networks with different structures and functions are trained based on user information and item information to obtain a dual-tower model, which can maintain the sensitivity of the dual-tower model to training samples and improve the training efficiency of the dual-tower model while reducing the training steps; the first neural network does not include a normalization layer, which can enable the first neural network to automatically learn an appropriate vector norm length, thereby maintaining an appropriate sensitivity to training samples and improving the training efficiency of the first neural network; the normalization process performed on the item representation vector does not include hyperparameter adjustment, which can omit training steps such as adjusting hyperparameters multiple times, reduce labor costs, improve the training efficiency of the second neural network, and save system resources; using the trained dual-tower model to query items of interest to the user can improve the accuracy of item recommendation; normalizing the target user vector can avoid the interference of the vector norm on similarity measurement, improve the accuracy of item recommendation, improve the flexibility of item recommendation and the structural scalability of the dual-tower model.

[0167] According to a fifth aspect of the embodiments of the present invention, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method provided in the first aspect and / or the second aspect of the embodiments of the present invention.

[0168] According to a sixth aspect of the embodiments of the present invention, there is provided a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, it implements the method provided in the first aspect and / or the second aspect of the embodiments of the present invention.

[0169] Figure 9 An exemplary system architecture 900 is shown to which the training method of the network model or the training device of the network model according to the embodiments of the present invention can be applied.

[0170] As Figure 9 shown, the system architecture 900 may include terminal devices 901, 902, 903, a network 904, and a server 905. The network 904 is used to provide a medium for communication links between the terminal devices 901, 902, 903 and the server 905. The network 904 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0171] Users can use the terminal devices 901, 902, 903 to interact with the server 905 through the network 904 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 901, 902, 903, such as model training applications, item recommendation applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0172] The terminal devices 901, 902, and 903 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and the like.

[0173] The server 905 can be a server that provides various services. For example, it can be a background management server (only for example) that supports the training requests of the network models sent by the upstream terminal devices 901, 902, and 903. The background management server can, in response to receiving user information and item information, generate a user representation vector corresponding to the user information according to a preset first neural network, and generate an item representation vector corresponding to the item information according to a preset second neural network; one of the user representation vector and the item representation vector has been normalized, and the other has not been normalized; determine the similarity between the user representation vector and the item representation vector; update the parameter weights of the first neural network and the second neural network according to the similarity, and use the updated first neural network and second neural network as a dual tower model; and feedback the training situation of the network model (only for example) to the terminal device. Alternatively, the background management server can, in response to receiving target user information and target item information, perform feature extraction on the target user information and the target item information according to a preset dual tower model to obtain corresponding target user vectors and target item vectors, where the dual tower model is obtained by using any of the methods in the first aspect of the embodiments of the present invention; determine the similarity between the target user vector and the target item vector; and recommend the target item to the target user when the similarity meets the preset recommendation conditions; and feedback the item recommendation situation (only for example) to the terminal device.

[0174] It should be noted that the training method of the network model provided by the embodiments of the present invention is generally executed by the server 905. Correspondingly, the training device of the network model is generally set in the server 905. The training method of the network model provided by the embodiments of the present invention can also be executed by the terminal devices 901, 902, and 903. Correspondingly, the training device of the network model can be set in the terminal devices 901, 902, and 903.

[0175] It should be noted that the item recommendation method provided by the embodiments of the present invention is generally executed by the server 905. Correspondingly, the item recommendation device is generally set in the server 905. The item recommendation method provided by the embodiments of the present invention can also be executed by the terminal devices 901, 902, and 903. Correspondingly, the item recommendation device can be set in the terminal devices 901, 902, and 903.

[0176] It should be understood,Figure 9 The number of terminal devices, networks, and servers in [the above] is merely illustrative. Depending on implementation requirements, there can be any number of terminal devices, networks, and servers.

[0177] Reference is now made to Figure 10 , which shows a schematic structural diagram of a computer system 1000 of a terminal device suitable for implementing an embodiment of the present invention. Figure 10 The terminal device shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present invention.

[0178] As Figure 10 shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage section 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the system 1000 are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0179] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as required. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as required so that a computer program read from it can be installed into the storage section 1008 as required.

[0180] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009 and / or installed from the removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the above functions defined in the system of the embodiments of the present invention are executed.

[0181] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the embodiments of the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0182] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer programs according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order from that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0183] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a feature extraction module, a similarity measurement module, a parameter update module, or a processor includes a feature extraction module, a similarity measurement module, and an item recommendation module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases. For example, the feature extraction module can also be described as "a module for extracting features from user information and item information".

[0184] As another aspect, the embodiments of the present invention further provide a computer-readable medium. This computer-readable medium can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device implements the following methods: in response to receiving user information and item information, generate a user representation vector corresponding to the user information according to a preset first neural network, and generate an item representation vector corresponding to the item information according to a preset second neural network; one of the user representation vector and the item representation vector has been normalized, and the other representation vector has not been normalized; determine the similarity between the user representation vector and the item representation vector; update the parameter weights of the first neural network and the second neural network according to the similarity, and use the updated first neural network and second neural network as a dual tower model. Or, the device implements the following methods: in response to receiving target user information and target item information, perform feature extraction on the target user information and the target item information according to a preset dual tower model to obtain corresponding target user vectors and target item vectors, and the dual tower model is obtained by using any of the methods in the first aspect of the embodiments of the present invention; determine the similarity between the target user vector and the target item vector; in the case where the similarity meets a preset recommendation condition, recommend the target item to the target user.

[0185] According to the technical solution of the embodiment of the present invention, two neural networks with different structures and functions are trained based on user information and item information to obtain a dual-tower model, which can maintain the sensitivity of the dual-tower model to training samples and improve the training efficiency of the dual-tower model while reducing the training steps; the first neural network does not include a normalization layer, which can enable the first neural network to automatically learn an appropriate vector norm length, thereby maintaining an appropriate sensitivity to training samples and improving the training efficiency of the first neural network; the normalization process performed on the item representation vector does not include hyperparameter adjustment, which can omit training steps such as adjusting hyperparameters multiple times, reduce labor costs, improve the training efficiency of the second neural network, and save system resources; using the trained dual-tower model to query items of interest to the user can improve the accuracy of item recommendation; normalizing the target user vector can avoid the interference of the vector norm on similarity measurement, improve the accuracy of item recommendation, improve the flexibility of item recommendation and the structural scalability of the dual-tower model.

[0186] The above specific embodiments do not limit the protection scope of the embodiments of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.

Claims

1. A training method for a network model, characterized in that, Including: In response to receiving user information and item information, generate a user representation vector corresponding to the user information according to a pre-set first neural network, and generate an item representation vector corresponding to the item information according to a pre-set second neural network; one of the user representation vector and the item representation vector has been normalized, and the other representation vector has not been normalized; Determine the similarity between the user representation vector and the item representation vector; According to the similarity, update the parameter weights of the first neural network and the second neural network, and use the updated first neural network and second neural network as a two-tower model.

2. The method according to claim 1, wherein Generating a user representation vector corresponding to the user information according to a pre-set first neural network, and generating an item representation vector corresponding to the item information according to a pre-set second neural network, including: In response to receiving user information, perform feature extraction on the user information according to a first neural network that does not include a normalization layer to obtain an unnormalized user representation vector; In response to receiving item information, perform feature extraction on the item information according to a second neural network that includes a normalization layer to obtain a normalized item representation vector.

3. The method according to claim 2, wherein The second neural network includes: a feature extraction layer and a normalization layer; performing feature extraction on the item information according to a second neural network that includes a normalization layer to obtain a normalized item representation vector, including: Use the feature extraction layer to perform feature extraction on the item information to obtain an initial item representation vector; Use the normalization layer to normalize the initial item representation vector to obtain a normalized item representation vector, and use the normalized item representation vector as the item representation vector corresponding to the item information, and the normalization process does not include hyperparameter adjustment.

4. The method according to claim 1, wherein Generating a user representation vector corresponding to the user information according to a pre-set first neural network, and generating an item representation vector corresponding to the item information according to a pre-set second neural network, including: In response to receiving user information, perform feature extraction on the user information according to a first neural network that includes a normalization layer to obtain a normalized user representation vector; In response to receiving item information, perform feature extraction on the item information according to a second neural network that does not include a normalization layer to obtain an unnormalized item representation vector.

5. The method according to claim 4, characterized in that, The first neural network includes: a feature extraction layer and a normalization layer; performing feature extraction on the user information according to a first neural network that includes a normalization layer to obtain a normalized user representation vector, including: Use the feature extraction layer to perform feature extraction on the user information to obtain an initial user representation vector; Use the normalization layer to normalize the initial user representation vector to obtain a normalized user representation vector, and use the normalized user representation vector as the user representation vector corresponding to the user information, and the normalization process does not include hyperparameter adjustment.

6. The method according to claim 1, characterized in that Updating the parameter weights of the first neural network and the second neural network according to the similarity includes: Dividing the item information into positive item information and negative item information; Among the similarities, taking the similarity between the user information and the positive item information as the positive similarity, and taking the similarity between the user information and the negative item information as the negative similarity; Comparing the positive similarity with a preset first similarity threshold to obtain a first comparison result, and updating the parameter weights of the first neural network and the second neural network according to the first comparison result; Comparing the negative similarity with a preset second similarity threshold to obtain a second comparison result, and updating the parameter weights of the first neural network and the second neural network according to the second comparison result.

7. A method for item recommendation, characterized in that, Including: In response to receiving target user information and target item information, according to a preset two-tower model, performing feature extraction on the target user information and the target item information to obtain corresponding target user vectors and target item vectors, where the two-tower model is obtained by using any one of the methods recited in claims 1 to 6; Determining the similarity between the target user vector and the target item vector; When the similarity meets a preset recommendation condition, recommending the target item to the target user.

8. The method according to claim 7, wherein The two-tower model includes a first neural network for obtaining a target user vector and a second neural network for obtaining a target item vector. The first neural network does not include a normalization layer. There are multiple target user vectors. Before determining the similarity between the target user vector and the target item vector, the method further includes: Performing normalization processing on multiple target user vectors to obtain multiple normalized target user vectors, where the multiple normalized target user vectors are used to determine the similarity with the target item vector, and the normalization processing does not include hyperparameter adjustment.

9. A training device for a network model, characterized in that, Including: A feature extraction module, configured to, in response to receiving user information and item information, generate a user representation vector corresponding to the user information according to a preset first neural network, and generate an item representation vector corresponding to the item information according to a preset second neural network; one of the user representation vector and the item representation vector has undergone normalization processing, and the other representation vector has not undergone normalization processing; A similarity measurement module, configured to determine the similarity between the user representation vector and the item representation vector; A parameter update module, configured to update the parameter weights of the first neural network and the second neural network according to the similarity, and use the updated first neural network and second neural network as a two-tower model.

10. An apparatus for item recommendation, characterized in that, Including: A feature extraction module, configured to, in response to receiving target user information and target item information, perform feature extraction on the target user information and the target item information according to a preset two-tower model to obtain corresponding target user vectors and target item vectors, where the two-tower model is obtained by using any one of the methods recited in claims 1 to 3; A similarity measurement module, configured to determine the similarity between the target user vector and the target item vector; An item recommendation module, configured to recommend the target item to the target user when the similarity meets a pre-set recommendation condition.

11. An electronic device, characterized in that, Comprising: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-7 is implemented.