An information recommendation method, system, storage medium and server

By using the teacher filtering submodule and the student filtering submodule to filter sample information in the pre-training process of the information recommendation model, the problems of large resource consumption and long training time in the existing technology are solved, and more efficient and accurate information recommendation is achieved.

CN114329175BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
CN202111261663.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-05-27
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

In order to improve the accuracy of recommendations in the existing information recommendation methods, more complex recommendation algorithms are required, resulting in large resource consumption and long training time.

Method used

The pre-trained information recommendation model is adopted. During pre-training, the model selects sample information through the teacher filtering submodule and the student filtering submodule respectively, filters out the samples that interfere with the training, reduces the number of feature calculation layers, and improves training efficiency.

Benefits of technology

Improve the accuracy of information recommendation, reduce resource consumption and training time, and achieve faster and more accurate information recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329175B_ABST
    Figure CN114329175B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose an information recommendation method, system, storage medium and server, which are applied to the technical field of information processing based on artificial intelligence. An information recommendation list corresponding to an information access request is output through a pre-trained information recommendation model. During the pre-training process of the information recommendation model, a corresponding part of sample information can be selected through a teacher filtering sub-module and a student filtering sub-module respectively to filter out training samples that interfere with the training of the teacher sub-module and the student sub-module, making the training of the information recommendation model more accurate. Moreover, since the sample information processed by the teacher sub-module and the student sub-module is less, the time spent on training the information recommendation model is less. Therefore, when the application background in the embodiments of the present invention recommends information to the application terminal, it will be more accurate and the response time will be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing based on artificial intelligence, and particularly relates to an information recommendation method, system, storage medium and server. Background Art

[0002] In the field of information recommendation, the application background will determine the user's preferences according to the user's interaction operations based on information on the application terminal (such as operations like liking or commenting on a certain article), and recommend various other types of information related to the application terminal, such as articles, videos, news, or products, etc. Such a recommendation method can be applied in various scenarios, and can take into account the actual needs of users while providing relatively rich information for users.

[0003] In the prior art, when using an information recommendation method, traditional recommendation algorithms such as collaborative filtering or logistic regression can be adopted, or a deep learning ranking algorithm can be used for recommendation. However, in the existing information recommendation methods, in order to improve the accuracy of recommendation, a relatively complex recommendation algorithm is required, and thus the resources consumed by such an information recommendation method are relatively large. Summary of the Invention

[0004] Embodiments of the present invention provide an information recommendation method, system, storage medium and server, which improve the accuracy of information recommendation.

[0005] Embodiments of the present invention on the one hand provide an information recommendation method, including:

[0006] Obtain an information access request;

[0007] Determine an information recommendation list according to the user access request and a pre-trained information recommendation model, where the information recommendation list includes multiple pieces of recommended information;

[0008] Output the information recommendation list;

[0009] Among them, when the information recommendation model is pre-trained:

[0010] Determine training samples, where the training samples include multiple pieces of sample information;

[0011] Determine an initial information recommendation model, where the initial information recommendation model includes a teacher filtering sub-module, a teacher sub-module, a student filtering sub-module, and a student sub-module;

[0012] Among them, the teacher filtering sub-module is used to select the first part of the sample information from the multiple pieces of sample information, and the teacher sub-module is used to determine whether the first part of the sample information is recommended information for a specific user; the student filtering sub-module is used to select the second part of the sample information from the multiple pieces of sample information, and the student sub-module is used to determine whether the second part of the sample information is recommended information for a specific user, and the feature calculation layer included in the student sub-module is less than the feature calculation layer included in the teacher sub-module.

[0013] Train the information recommendation model according to the training samples and the initial information recommendation model. The pre-trained information recommendation model includes a trained student sub-module.

[0014] Another aspect of the embodiments of the present invention provides an information recommendation system, including:

[0015] A request acquisition unit, configured to acquire an information access request;

[0016] An information recommendation unit, configured to determine an information recommendation list according to the user access request and the pre-trained information recommendation model, where the information recommendation list includes multiple pieces of recommended information;

[0017] An output unit, configured to output the information recommendation list;

[0018] Among them, when the information recommendation model is pre-trained: determine training samples, where the training samples include multiple pieces of sample information; determine an initial information recommendation model, where the initial information recommendation model includes a teacher filtering sub-module, a teacher sub-module, a student filtering sub-module, and a student sub-module; among them, the teacher filtering sub-module is used to select the first part of the sample information from the multiple pieces of sample information, and the teacher sub-module is used to determine whether the first part of the sample information is recommended information for a specific user; the student filtering sub-module is used to select the second part of the sample information from the multiple pieces of sample information, and the student sub-module is used to determine whether the second part of the sample information is recommended information for a specific user, and the feature calculation layer included in the student sub-module is less than the feature calculation layer included in the teacher sub-module; train the information recommendation model according to the training samples and the initial information recommendation model, and the pre-trained information recommendation model includes a trained student sub-module.

[0019] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores multiple computer programs, and the computer programs are adapted to be loaded and executed by a processor to perform the information recommendation method as described in one aspect of the embodiments of the present invention.

[0020] Another aspect of the embodiments of the present invention further provides a terminal device, including a processor and a memory;

[0021] The memory is used to store a plurality of computer programs, which are used to be loaded and executed by a processor to implement the information recommendation method as described in another aspect of the embodiment of the present invention; the processor is used to implement each of the plurality of computer programs.

[0022] It can be seen that in the information recommendation method of the embodiment of the present invention, an information recommendation list corresponding to an information access request is output through a pre-trained information recommendation model. During the pre-training process of the information recommendation model, a corresponding part of sample information can be selected through a teacher filtering sub-module and a student filtering sub-module respectively to filter out the training samples that interfere with the training of the teacher sub-module and the student sub-module, making the training of the information recommendation model more accurate. And since the sample information processed by the teacher sub-module and the student sub-module is less, the time spent on training the information recommendation model is less. Therefore, when the background in the embodiment of the present invention performs information recommendation to an application terminal, it will be more accurate and the response time will be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a schematic diagram of a system to which an information recommendation method provided by an embodiment of the present invention is applicable;

[0025] Figure 2 It is a flowchart of an information recommendation method provided by an embodiment of the present invention;

[0026] Figure 3 It is a schematic diagram of an initial information recommendation model determined in an embodiment of the present invention;

[0027] Figure 4 It is a flowchart of an information recommendation method provided by another embodiment of the present invention;

[0028] Figure 5 It is a schematic diagram of a student sub-module determined in an application embodiment of the present invention;

[0029] Figure 6 It is a schematic diagram of an information interaction interface displayed on an application terminal in an application embodiment of the present invention;

[0030] Figure 7 It is a schematic diagram of the AUC calculated in an application embodiment of the present invention;

[0031] Figure 8 It is a schematic diagram of a distributed system to which the information recommendation method in another application embodiment of the present invention is applied;

[0032] Figure 9 It is a schematic diagram of a block structure in another application embodiment of the present invention;

[0033] Figure 10 It is a schematic diagram of the logical structure of an information recommendation system provided by an embodiment of the present invention;

[0034] Figure 11 It is a schematic diagram of the logical structure of a server provided by an embodiment of the present invention. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0036] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0037] The embodiment of the present invention mainly provides an information recommendation method, which can be applied to an information recommendation system as Figure 1 shown. In this system, there are an application terminal 10 and an application background 11, where:

[0038] The application terminal 10 is an information interface provided by the information recommendation system for users. Users can access various types of information provided by the application background 11 through the application terminal 10 and can operate on various types of information, such as commenting, forwarding, or favoriting, etc.

[0039] Specifically, the application terminal 10 can display an information interaction interface, which can include information provided by the application background 11 and a user interface. The user can operate on any information through the user interface or request to obtain the latest information released by the application background 11. The application terminal 10 can be a mobile phone, a personal computer, a tablet computer, a smart voice interaction device, a smart home appliance, etc.

[0040] The application background 11 is used to provide various types of information, and according to the request of the application terminal 10, provide corresponding information to the application terminal 10, and recommend various types of information that the user is interested in to the application terminal 10.

[0041] Specifically, the application background 11 can be a news background, a video background, a commodity background, etc.

[0042] Specifically, as Figure 2 shown, the application background 11 can implement information recommendation to the application terminal 10 through the following steps:

[0043] Step 101, obtain an information access request.

[0044] It can be understood that the user can initiate a request to the application background 11 through the application terminal 10 to obtain the information released by the application background 11. For example, the user can perform a pull-down operation on the information interface displayed by the application terminal 10 or click the "Home" button on the information interface, then the application terminal 10 will initiate an information access request to the application background 11, and this information access request is used to request to obtain the information released by the application background 11. Specifically, the information access request can include information such as the user identifier of the user corresponding to the application terminal 10.

[0045] Step 102, determine an information recommendation list according to the user access request and the pre-trained information recommendation model. The information recommendation list includes multiple recommended information.

[0046] Here, the pre-trained information recommendation model is a machine learning model based on artificial intelligence. It can adopt a certain amount of training samples and be trained according to a certain training method, and the operation logic of the trained information recommendation model is stored in the application background 11 in advance. When the application background 11 initiates the process of this embodiment, it can directly call the pre-set information recommendation model to determine the information recommendation list.

[0047] Among them, Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0048] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0049] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0050] Step 103, output the information recommendation list, that is, output the information recommendation list to the application terminal 10.

[0051] Among them, when the information recommendation model is pre-trained, it can be achieved through the following steps:

[0052] Step 201, determine the training samples, and the training samples include multiple pieces of sample information.

[0053] In this embodiment, the determined training samples mainly include two parts of sample information. One part of the sample information is information with labels, which can be used to supervise the training of the information recommendation model. Specifically, this part of the sample information may include information published by the application background 11 (such as news, text, or video information), label information indicating whether a certain user likes it, and historical operation information of the corresponding user, etc. Among them, if a user performs operations such as liking, following, forwarding, or commenting on a piece of information, it can be considered that the user likes this piece of information; if after a piece of information has been published for a period of time, the user does not perform any operation on this piece of information, or the user clicks the "dislike" button on this piece of information, it is considered that the user does not like this piece of information. And the historical operation information of the user may include information on which information the user has operated on in the past, etc.

[0054] The other part of the sample information is information without labels. Specifically, this part of the information may only include various types of information published by the application background 11, and does not include label information indicating whether the user likes it. This part of the sample information is generally the latest information published by the application background 11, and no user has performed relevant operations on this information, so it cannot be determined whether a user likes this information.

[0055] Step 202, determine the initial information recommendation model, which includes a teacher filtering sub-module, a teacher sub-module, a student filtering sub-module, and a student sub-module.

[0056] It can be understood that when the application background 11 determines the initial information recommendation model, it will respectively determine the multi-layer structure included in the initial information recommendation model and the initial values of the parameters in each layer structure.

[0057] Specifically, as Figure 3 shown, the initial information recommendation model may include: a teacher filtering sub-module 110, a teacher sub-module 111, a student filtering sub-module 112, and a student sub-module 113. Among them, the teacher filtering sub-module 110 is used to select the first part of the sample information from the above-mentioned multiple pieces of sample information, the teacher sub-module 111 is used to determine whether the first part of the sample information is recommended information for a specific user; the student filtering sub-module 112 is used to select the second part of the sample information from the multiple pieces of sample information, and the student sub-module 113 is used to determine whether the second part of the sample information is recommended information for a specific user.

[0058] Among them, the functions implemented by the teacher sub-module 111 and the student sub-module 113 are the same, that is, the network structures are the same. Specifically, the features of any sample information (including a certain type of information and information related to a certain user interest, such as user historical operation information, etc.) can be extracted, and based on the extracted features, the user click-through rate corresponding to the user based on any sample information can be output, which is used to represent the probability that the corresponding user clicks on the sample information. If the user click-through rate is greater than a certain threshold, the sample information can be recommended to the corresponding user, otherwise the sample information will not be recommended to the corresponding user. However, the feature calculation layer included in the student sub-module 113 is less than that included in the teacher sub-module 111, and it can be a compressed and simplified model of the teacher sub-module 111.

[0059] Here, the first part of the sample information can be the sample information that the teacher sub-module 111 is concerned about and wants to learn, while the second part of the sample information is the sample information that the student sub-module 113 is concerned about and wants to learn. There can be an intersection or no intersection between the two parts of the sample information. In this way, the teacher filtering sub-module 110 and the student filtering sub-module 112 can respectively filter out the sample information that interferes with the training of the teacher sub-module 111 and the student sub-module 113.

[0060] In specific implementation, the teacher filtering sub-module 110 and the student filtering sub-module 112 can randomly select from multiple sample information. Or, specifically, the teacher filtering sub-module 110 determines the first scores of multiple sample information according to the user click-through rates corresponding to the multiple sample information determined by the teacher sub-module 111, and selects a part of the sample information with the highest first score from the multiple sample information as the first part of the sample information. After scoring the above multiple sample information to obtain scores, then based on the score ranking, a part of the sample information with the highest score is selected from the above multiple sample information; the student filtering sub-module 112 determines the second scores of multiple sample information according to the user click-through rates corresponding to the multiple sample information determined by the student sub-module 113, and selects a part of the sample information with the highest second score from the multiple sample information as the second part of the sample information. Generally, the higher the user click-through rate, the larger the score. In this way, the teacher filtering sub-module 110 and the student filtering sub-module 112 will respectively filter the sample information according to the actual processing of the sample information by the teacher sub-module 111 and the student sub-module 113, so that the obtained first part of the sample information and the second part of the sample information can more truly reflect the sample information that the teacher sub-module 111 and the student sub-module 113 are concerned about.

[0061] The parameters of the information recommendation initial model refer to the fixed parameters that are used in the calculation process of each layer structure in the information recommendation initial model and do not need to be assigned values at any time, such as parameters such as parameter scale, number of network layers, and user vector length.

[0062] Step 203: Train an information recommendation model based on the training samples and the initial model training information for recommendation. The pre-trained information recommendation model involved in the above Step 102 includes a trained student sub-module.

[0063] Specifically, as Figure 3 shown, after inputting the training samples into the initial information recommendation model, the error of the output results (including the results respectively output by the teacher sub-module 111 and the student sub-module 113) obtained by the initial information recommendation model through actual processing of these training samples is calculated, and the initial values of the various parameters in the initial information recommendation model are continuously adjusted, including the initial values of the various parameters in the teacher sub-module 111 and the student sub-module 113. Finally, the parameter values that can make the processing of the student sub-module 113 in the initial information recommendation model more accurate are adjusted, that is, these finally obtained parameter values are learned parameter values. Generally, the error of the output results actually obtained by the initial information recommendation model can be represented by a certain function, such as a loss function.

[0064] In this embodiment, the adjustment of the parameter values in the student sub-module 113 needs to be based on the adjustment of the parameter values in the teacher sub-module 111, so that the student sub-module 113 can learn the information in the teacher sub-module 111.

[0065] It should be noted that in specific implementation, the initial information recommendation model determined in the above Step 202 can be a Distilled reinforcement learning framework for recommendation (DRL-Rec), etc. Then, during the training process of this Step 203, when learning the parameters in the teacher sub-module 111 and the student sub-module 113, distillation learning can be performed based on the KL divergence loss of the Q value list.

[0066] Among them, distillation learning induces the training of the student sub-module 113 by introducing a soft target related to the teacher sub-module 111 as a part of the overall error of the initial information recommendation model, so as to achieve knowledge transfer from the teacher sub-module 111 to the student sub-module 113.

[0067] It can be seen that in the information recommendation process of this embodiment, an information recommendation list corresponding to an information access request is output through a pre-trained information recommendation model. During the pre-training process of the information recommendation model, a corresponding part of the sample information can be selected through the teacher filtering sub-module and the student filtering sub-module respectively to filter out the training samples that interfere with the training of the teacher sub-module and the student sub-module, making the training of the information recommendation model more accurate. Moreover, since the sample information processed by the teacher sub-module and the student sub-module is less, the time spent on training the information recommendation model is less. Therefore, when the application background in this embodiment of the present invention performs information recommendation to the application terminal, it will be more accurate and the response time will be reduced.

[0068] Another embodiment of the present invention provides an information recommendation method, which is mainly a method executed by the above-mentioned application background 11. The flowchart is as Figure 4 shown and includes:

[0069] Step 301: Determine training samples. The training samples include multiple pieces of sample information, some of which carry label information and some do not. Specifically, as described in the above embodiment, it will not be elaborated here.

[0070] Step 302: Determine an initial information recommendation model. The structure of the initial information recommendation model can be as Figure 3 shown and includes a teacher filtering sub-module, a teacher sub-module, a student filtering sub-module, and a student sub-module. Specifically, it will not be elaborated here.

[0071] Further, in this embodiment, the application background 11 can train the information recommendation model according to the above training samples and the initial information recommendation model, which can be specifically implemented through the following steps 303 to 306:

[0072] Step 303: Determine whether the first part of the sample information in the training samples belongs to the recommended information of a specific user through the initial information recommendation model, and determine whether the second part of the sample information belongs to the recommended information of a specific user.

[0073] Specifically, the teacher filtering sub-module and the student filtering sub-module in the initial information recommendation model respectively select the first part of the sample information and the second part of the sample information from multiple pieces of sample information. Then, the teacher sub-module determines whether the first part of the sample information belongs to the recommended information of a specific user, and the student sub-module determines whether the second part of the sample information belongs to the recommended information of a specific user.

[0074] Step 304: Adjust the initial information recommendation model according to the result determined by the initial information recommendation model to obtain the information recommendation model.

[0075] Specifically, the application background can adopt the following steps 3041 to 3043 to adjust the information recommendation initial model:

[0076] Step 3041: Calculate a first loss function related to the teacher sub-module, a second loss function related to the student sub-module, a third loss function for distillation from the teacher sub-module to the student sub-module, and a fourth loss function for calculating the difference between the teacher sub-module and the student sub-module according to the results determined by the information recommendation initial model.

[0077] Step 3042: Calculate the overall loss function of the information recommendation initial model according to the first loss function, the second loss function, the third loss function, and the fourth loss function. Specifically, the weighted sum value of the first loss function, the second loss function, the third loss function, and the fourth loss function can be used as the overall loss function.

[0078] Step 3043: Adjust the parameter values of the parameters in the information recommendation initial model according to the overall loss function to obtain an information recommendation model.

[0079] It can be understood that the application background will first calculate the overall loss function (also called the objective function) of the information recommendation initial model according to the results determined by the information recommendation initial model in the above step 303. The training process of the information recommendation model is to minimize the value of this overall loss function as much as possible, so that the overall loss function reaches a convergent state. This training process continuously optimizes the parameter values of the parameters in the information recommendation initial model determined in the above step 302 through a series of mathematical optimization means such as backpropagation derivative and gradient descent, and makes the calculated value of the overall loss function drop to the lowest.

[0080] Specifically, when the function value of the calculated overall loss function is relatively large, such as greater than a preset value, it is necessary to adjust the parameter values of the parameters in the current teacher sub-module and student sub-module. For example, reduce the weight value of a certain neuron connection, etc., so that the function value of the overall loss function calculated according to the adjusted parameter values decreases. In this embodiment, the adjustment of the parameter values in the student sub-module is mainly based on the adjustment of the parameter values in the teacher sub-module, so that the student sub-module can transfer the information in the teacher sub-module during the training process.

[0081] Specifically in this embodiment, the overall loss function of the information recommendation initial model can include the following loss functions:

[0082] (1) The first loss function related to the teacher sub-module, mainly for the sample information with label information among the multiple sample information of the training samples, and the sample information without label information among the multiple sample information does not participate in the calculation of the first loss function.

[0083] It can be understood that when calculating the first loss function, it is mainly for the sample information with label information in the training samples. This first loss function is used to indicate the error between whether the first part of the sample information determined by the teacher sub-module in the initial information recommendation model is the recommended information for a specific user and whether the corresponding first part of the sample information is actually the information liked by the specific user (obtained according to the label information in the first part of the sample information), such as the cross-entropy loss function, etc.

[0084] Specifically, the application background will first calculate the error related to the teacher sub-module based on whether the first part of the sample information determined by the teacher sub-module is the recommended information for a specific user and the label information of the corresponding first part of the sample information in the training samples; then calculate the first loss function related to the teacher sub-module according to this error.

[0085] (2) The second loss function related to the student sub-module is mainly for the sample information with label information among the multiple sample information of the training samples, and the sample information without label information among the multiple sample information does not participate in the calculation of the second loss function.

[0086] It can be understood that when calculating the second loss function, it is mainly for the sample information with label information in the training samples. This second loss function is used to indicate the error between whether the second part of the sample information determined by the student sub-module in the initial information recommendation model is the recommended information for a specific user and whether the corresponding second part of the sample information is actually the information liked by the specific user (obtained according to the label information in the second part of the sample information), such as the cross-entropy loss function, etc.

[0087] Specifically, the application background will first calculate the error related to the student sub-module based on whether the second part of the sample information determined by the student sub-module is the recommended information for a specific user and the label information of the corresponding second part of the sample information in the training samples; then calculate the second loss function related to the student sub-module according to this error.

[0088] (3) The third loss function for distillation from the teacher sub-module to the student sub-module is mainly for the intersection between the above-mentioned first part of the sample information and the second part of the sample information, that is, the common sample information.

[0089] It can be understood that this third loss function is used to indicate the proportion of knowledge transfer from the teacher sub-module to the student sub-module, so that the result of the student sub-module processing information can be as close as possible to the result of the teacher sub-module processing information. In practical applications, the third loss function can be the KL divergence error based on the Q-value list, etc.

[0090] (4) The fourth loss function for the difference between the teacher sub-module and the student sub-module is mainly for the intersection between the above-mentioned first part of the sample information and the second part of the sample information, that is, the common sample information.

[0091] Specifically, the application background will determine the confidence of whether the common sample information determined by the teacher sub-module is the recommended information for a specific user; calculate the difference between whether the common sample information determined by the teacher sub-module and the student sub-module is the recommended information for a specific user; and calculate the fourth loss function based on the confidence and difference of the common sample information.

[0092] Among them, the confidence level of whether the common sample information determined by the teacher submodule is the recommendation information of a specific user is used to indicate the credibility of the result obtained by the teacher submodule through processing. Since the training of the teacher submodule will guide the training of the student submodule, in this embodiment, when calculating the overall loss function, the confidence level of the result obtained by the teacher submodule through processing each common sample information is considered, which can guide the student submodule to effectively learn information from the teacher submodule. Specifically, when determining the confidence level, for the first common sample information with label information in the common sample information, the confidence level of the first common sample information can be determined based on the difference between whether the first common sample information determined by the teacher submodule is the recommendation information of a specific user and the label information in the corresponding first common sample information; and for the second common sample information without label information in the common sample information, the confidence level corresponding to the second common sample information is determined to be the average confidence level of the first common sample information.

[0093] Step 305 , determining whether the adjustment of the parameter values ​​of the information recommendation initial model in the above step 304 satisfies the preset stop condition, if so, executing step 306 ; if so, returning to executing the above step 303 .

[0094] It should be noted that the above steps 303 to 304 are the results determined by the information recommendation initial model, and are an adjustment of the parameter value in the information recommendation initial model. In practical applications, it is necessary to continuously loop through the above steps 303 to 304 until the adjustment of the parameter value meets a certain stop condition. The preset stop condition includes but is not limited to any one of the following conditions: the difference between the currently adjusted parameter value and the last adjusted parameter value is less than a threshold value, that is, the adjusted parameter value reaches convergence; and the number of parameter value adjustments is equal to the preset number of times, etc.

[0095] Step 306: The parameter values ​​adjusted in the above step 304 are used as parameter values ​​in the finally trained information recommendation model, and the student submodule in the trained information recommendation model is preset to the application background.

[0096] Step 307, when the user initiates an information access request through the application terminal, the application background, upon receiving the information access request, can call the information recommendation model preset in the current system to determine the corresponding information recommendation list, and output the information recommendation list to the application terminal.

[0097] It can be seen that in the method of this embodiment, in the process of pre-training the information recommendation model, on the one hand, a corresponding part of the sample information will be selected respectively by the teacher filtering submodule and the student filtering submodule to filter out the training samples that interfere with the training of the teacher submodule and the student submodule, so that the training of the information recommendation model is more accurate; on the other hand, the training of the student submodule will be guided according to the confidence of each sample information processed by the teacher submodule during the training process, so that the student submodule will learn certain information from the teacher submodule based on the confidence of each sample information, thereby further improving the training accuracy of the student submodule.

[0098] The following is a specific application example to illustrate the information recommendation method of the present invention. The method of this embodiment mainly includes the following two parts:

[0099] (1) Preset the information recommendation model in the application background.

[0100] In this embodiment, when the information recommendation model is pre-set in the application background, the information recommendation model is first trained by a certain method, and then the relevant information of the trained information recommendation model is stored in the application background. The specific method of pre-training the information recommendation model is as described in the above embodiment, except that in this embodiment:

[0101] (1) The structure of the initial information recommendation model determined by the application background can be as follows Figure 3 As shown, in the embodiment, the teacher filtering submodule in the information recommendation initial model scores each sample information according to the user click rate output by the teacher submodule to obtain a score, and then selects the k sample information with the highest scores as the first part of the sample information.

[0102] The student filtering submodule in the initial information recommendation model will score each sample information according to the user click rate output by the student submodule, and then select the k sample information with the highest scores as the second part of the sample information.

[0103] In other embodiments, the teacher filtering submodule and the student filtering submodule may also use other filtering methods to select corresponding partial sample information, which is not limited here.

[0104] (2) The network structures of the teacher submodule and the student submodule in the information recommendation model determined by the application background can be specifically DRL-Rec, but the feature calculation layer of the student submodule is less than that of the teacher submodule, which is a compressed and simplified model.

[0105] Specific as Figure 5As shown, taking the student sub-module as an example, in a specific implementation, the student sub-module can use a structure of reinforcement learning, such as a network structure of Double Deep Q-Learning (DDQN), etc., to determine an information recommendation list according to any sample information, and the information recommendation list includes multiple pieces of recommendation information. Specifically, in the student sub-module:

[0106] If it is necessary to determine the recommendation information at position t in the information recommendation list, first use a Transformer to perform feature interaction on the recommendation information at t - 1 positions in the already predicted information recommendation list, as shown in Formula 1 below; then a Gated Recurrent Unit (GRU) can be used to concatenate the features of t - 1 pieces of recommendation information to obtain the current state s t representation, as shown in Formula 2 below:

[0107] f i = Flatten(Transformer(F i ')) (1)

[0108] s t = GRU(seq t ) = GRU({f 1 ,..., f m}) (2)

[0109] Furthermore, then the above-obtained current state s t and the feature a of a piece of information to be processed t After calculation by a multi-layer fully connected layer (MLP), the predicted value based on the q value corresponding to the piece of information to be processed can be output, which can represent the user click-through rate of the piece of information. Specifically, it is as shown in Formula 3 below, and the q predicted values at different positions can be represented by Formula 4 below:

[0110] q t = MLP(Concat(s t , a t ) (3)

[0111]

[0112] In this way, for any piece of information to be processed, the predicted value output by the student sub-module can be represented by Formula 5 below. Similarly, the predicted value output by the teacher sub-module can be represented by Formula 6 below:

[0113] y t = r t + γQ(s t+1 , argmaxa Q(s t+1 ,a|θ S )|θ' S ) (5)

[0114] y t =r t +γQ(s t+1 ,argmax a Q(s t+1 ,a|θ T )|θ' T ) (6)

[0115] (3) During the process of the application background adjusting the parameter values in the initial information recommendation model according to the results determined by the initial model for information recommendation, when calculating the overall loss function:

[0116] The calculation of the first loss function related to the teacher sub-module can be represented by the following formula 7:

[0117]

[0118] The calculation of the second loss function related to the student sub-module can be represented by the following formula 8:

[0119]

[0120] When calculating the third loss function for distillation from the teacher sub-module to the student sub-module, it can be represented by the following formula 9, where is the q prediction value obtained by the teacher sub-module according to any common sample information i, and is the q prediction value obtained by the student sub-module according to any common sample information i, and the common sample information is the intersection between the first part of sample information selected by the teacher filtering sub-module and the second part of sample information selected by the student filtering sub-module:

[0121]

[0122] When calculating the fourth loss function for the difference between the teacher sub-module and the student sub-module, it can be represented by the following formula 10, where is the calculation result output by a certain fully connected layer in the teacher sub-module for a common sample information i, is the calculation result output by a certain fully connected layer in the student sub-module for a common sample information i, c i is the confidence of the result obtained by the teacher sub-module processing a common sample information i, and can be represented by the following formula 11:

[0123]

[0124]

[0125] It should be noted that the confidence level determined in the above formula 11 is the confidence level determined for the first common sample information with label information in the training samples, where y t is the calculation result actually output by the teacher sub-module according to the common sample information, and is the label information carried in the common sample information; for the confidence level of the second common sample information without label, it can be the average confidence level corresponding to the above first common sample information.

[0126] Therefore, the overall loss function of the information recommendation initial model can be expressed by the following formula 12:

[0127] L = λ 1 L T + λ 2 L S + λ 3 L KL + λ 4 L Hint (12)

[0128] Furthermore, after training the information recommendation model through the above method, the relevant information of the student sub-module can be pre-set in the application background.

[0129] (2) Online application of the pre-set information recommendation model.

[0130] When the information recommendation model is pre-set in the application background, the user can operate the application terminal so that the application terminal can display an information interaction interface as shown in Figure 6 . The information interaction interface can include various types of information and user interfaces, such as link interfaces and operation buttons for various information, such as the "Home" button and the "User" button, etc. When the user clicks the "Home" button or performs a pull-down operation on the information page of the home page, the application terminal can send an information access request to the application background.

[0131] When the application background receives the information access request, it can directly call the above pre-set information recommendation model. In this way, the information recommendation model will call the above pre-set information recommendation model; at the same time, the application background will obtain the user information carried in the information access request and obtain the user's historical operation information according to the obtained user information, such as the information of the user's operations in a period of time before the current moment, such as the information of the user's comments, likes or forwards, etc.

[0132] Further, in the student sub-module of the information recommendation model called by the application background, according to the user's historical operation information and the information recently released in the application background, an information recommendation list corresponding to the user is determined. The information recommendation list includes multiple pieces of recommended information. Then, the application background will send the information recommendation list to the application terminal for display.

[0133] (III) Comparative verification.

[0134] For the information recommendation model trained by using the method in the embodiment of the present invention, in terms of the model size and the calculation time-consuming for information recommendation, the student sub-module is 49.7% and 76.7% of the teacher sub-module respectively. And when the compression rate of the student sub-module is within a certain range, such as 50%, the accuracy of information recommendation has a significant improvement compared with the teacher sub-module.

[0135] Further, when using the information recommendation model trained by using the method of this embodiment for information recommendation, compared with using the information recommendation model trained by using the existing method for information recommendation, in terms of the computation cost, it can be reduced by 23.3%, and in terms of the memory cost, it can be reduced by 50.3%. However, in terms of the average click number (ACN), it can be increased by 1.38%.

[0136] Further, for information recommendation models with different network structures, such as FM, NFM, AFN, HRL-Rec and the DRL-Rec network structure in this embodiment, etc., the calculation model measurement parameters can be as shown in Table 1 below, that is, the area under the rate-distortion curve (AUC) and the related improvement of AUC, namely RelaImpr, etc. Among them, the larger the AUC and RelaImpr are, the better the information recommendation effect is. It can be seen that the effect of the information recommendation model trained in the embodiment of the present invention is the best.

[0137] Model AUC RelaImpr FM 0.7321 0% NFM 0.7384 2.71% AFN 0.7459 5.95% HRL-Rec 0.7625 13.1% The model of the embodiment of the present invention 0.7653 14.3%

[0138] Table 1

[0139] Further, when obtaining the AUC with different compression rates, and when the number of sample information output by the teacher filtering sub-module (or student filtering sub-module) is different, as Figure 7 shown, when the compression rate is above 50%, the AUC can maintain a significant improvement effect, and when the number of filtered sample information is 3, the AUC effect is the best. This shows the importance of using the teacher filtering sub-module and the student filtering sub-module for noise control when training the information recommendation model.

[0140] It can be seen that in the method of this embodiment, a corresponding part of the sample information will be selected through the teacher filtering sub-module and the student filtering sub-module respectively to filter out the training samples that interfere with the training of the teacher sub-module and the student sub-module, making the training of the information recommendation model more accurate. On the other hand, the training of the student sub-module will be guided according to the confidence of the teacher sub-module in processing each sample information during the training process, so that the student sub-module will learn certain information from the teacher sub-module based on the confidence of each sample information, that is, the training accuracy of the student sub-module is further improved based on confidence-guided distillation.

[0141] The following uses another specific application example to illustrate the information recommendation method in the present invention. The information recommendation system in the embodiment of the present invention is mainly a distributed system 100, and the distributed system may include a client 300 and multiple nodes 200 (any form of computing device connected to the network, such as a server, a user terminal), and the client 300 and the nodes 200 are connected in the form of network communication.

[0142] Taking the distributed system as a blockchain system as an example, see Figure 8 It is an optional structural schematic diagram of the distributed system 100 provided by the embodiment of the present invention applied to the blockchain system, formed by multiple nodes 200 (any form of computing device connected to the network, such as a server, a user terminal) and a client 300. A peer-to-peer (P2P) network is formed between the nodes. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In the distributed system, any machine such as a server or a terminal can join and become a node. The node includes a hardware layer, an intermediate layer, an operating system layer, and an application layer.

[0143] See Figure 8 The functions of each node in the blockchain system shown involve the following functions:

[0144] 1) Routing, which is the basic function of a node and is used to support communication between nodes.

[0145] In addition to the routing function, a node may also have the following functions:

[0146] 2) Application, which is used to be deployed in the blockchain, implement specific services according to actual business needs, record the data related to the implemented functions to form record data, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes verify the source and integrity of the record data successfully, the record data is added to the temporary block.

[0147] For example, the services implemented by the application include the code for implementing an information recommendation function, which mainly includes:

[0148] Obtain an information access request; determine an information recommendation list according to the user access request and a pre-trained information recommendation model, where the information recommendation list includes multiple pieces of recommended information; output the information recommendation list; where, when the information recommendation model is pre-trained: determine training samples, where the training samples include multiple pieces of sample information; determine an initial information recommendation model, where the initial information recommendation model includes a teacher filtering sub-module, a teacher sub-module, a student filtering sub-module, and a student sub-module; where, the teacher filtering sub-module is used to select a first part of the sample information from the multiple pieces of sample information, the teacher sub-module is used to determine whether the first part of the sample information is recommended information for a specific user; the student filtering sub-module is used to select a second part of the sample information from the multiple pieces of sample information, the student sub-module is used to determine whether the second part of the sample information is recommended information for a specific user, and the feature calculation layer included in the student sub-module is less than the feature calculation layer included in the teacher sub-module; train the information recommendation model according to the training samples and the initial information recommendation model, and the pre-trained information recommendation model includes a trained student sub-module.

[0149] 3) A blockchain, including a series of blocks (Blocks) that are sequentially connected in the order of generation. Once a new block is added to the blockchain, it will not be removed again. The block records the record data submitted by the nodes in the blockchain system.

[0150] See Figure 9 An optional schematic diagram of the block structure provided by the embodiment of the present invention. Each block includes the hash value of the transaction record stored in this block (the hash value of this block), and the hash value of the previous block. The blocks are connected through the hash values to form a blockchain. In addition, the block may also include information such as the timestamp when the block is generated. A blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains relevant information for verifying the validity of its information (anti-counterfeiting) and generating the next block.

[0151] The embodiment of the present invention also provides an information recommendation system, such as the above-mentioned application background, and its structural schematic diagram is as Figure 10 shown, and specifically may include:

[0152] A request acquisition unit 20, configured to obtain an information access request.

[0153] An information recommendation unit 21, configured to determine an information recommendation list according to the user access request obtained by the request obtaining unit 20 and a pre-trained information recommendation model, where the information recommendation list includes multiple pieces of recommendation information.

[0154] An output unit 22, configured to output the information recommendation list determined by the information recommendation unit 21.

[0155] Furthermore, the information recommendation system of this embodiment may further include:

[0156] A training unit 23, configured to pre-train the information recommendation model. Specifically, training samples are determined, where the training samples include multiple pieces of sample information; an initial information recommendation model is determined, and the initial information recommendation model includes a teacher filtering sub-module, a teacher sub-module, a student filtering sub-module, and a student sub-module. Among them, the teacher filtering sub-module is configured to select a first part of the sample information from the multiple pieces of sample information, and the teacher sub-module is configured to determine whether the first part of the sample information is recommendation information for a specific user; the student filtering sub-module is configured to select a second part of the sample information from the multiple pieces of sample information, and the student sub-module is configured to determine whether the second part of the sample information is recommendation information for a specific user, and the feature calculation layer included in the student sub-module is less than the feature calculation layer included in the teacher sub-module; the information recommendation model is trained according to the training samples and the initial information recommendation model, and the pre-trained information recommendation model includes the trained student sub-module. In this way, the information recommendation unit 21 will determine the information recommendation list according to the information recommendation model pre-trained by the training unit 23.

[0157] Among them, the teacher sub-module is specifically configured to determine the user click-through rate of each piece of sample information. Then, the teacher filtering sub-module is specifically configured to respectively determine the first scores of the multiple pieces of sample information according to the user click-through rates of the multiple pieces of sample information determined by the teacher sub-module, and select a part of the sample information with the highest first score from the multiple pieces of sample information as the first part of the sample information; the student sub-module is specifically configured to determine the user click-through rate of each piece of sample information. Then, the student filtering sub-module is specifically configured to respectively determine the second scores of the multiple pieces of sample information according to the user click-through rates of the multiple pieces of sample information determined by the student sub-module, and select a part of the sample information with the highest second score from the multiple pieces of sample information as the second part of the sample information. In this way, the above-mentioned information recommendation unit 21 will determine the information recommendation list according to the information recommendation model trained by the training unit 23.

[0158] Optionally, when training the information recommendation model based on the training samples and the information, the training unit 23 is specifically configured to determine, by using the initial information recommendation model, whether the first part of sample information belongs to the recommended information of a specific user, and whether the second part of sample information belongs to the recommended information of a specific user; and adjust the initial information recommendation model according to the determination result of the initial information recommendation model, so as to obtain the information recommendation model.

[0159] Wherein, when adjusting the initial information recommendation model according to the determination result of the initial information recommendation model, the training unit 23 is specifically configured to calculate a first loss function related to the teacher sub-module, calculate a second loss function related to the student sub-module, calculate a third loss function for distillation from the teacher sub-module to the student sub-module, and calculate a fourth loss function for the difference between the teacher sub-module and the student sub-module according to the determination result of the initial information recommendation model; calculate an overall loss function of the initial information recommendation model according to the first loss function, the second loss function, the third loss function, and the fourth loss function; and adjust the parameter values of the parameters in the initial information recommendation model according to the overall loss function, so as to obtain the information recommendation model.

[0160] Wherein, if the intersection between the first part of sample information and the second part of information is common sample information, when calculating the fourth loss function for the difference between the teacher sub-module and the student sub-module, the training unit 23 is specifically configured to determine the confidence level of whether the common sample information determined by the teacher sub-module is the recommended information of a specific user; calculate the difference between whether the common sample information determined by the teacher sub-module and the student sub-module is the recommended information of a specific user, respectively; and calculate the fourth loss function according to the confidence level and the difference of the common sample information.

[0161] Wherein, when determining the confidence level of whether the common sample information determined by the teacher sub-module is the recommended information of a specific user, the training unit 23 is specifically configured to, for the first common sample information with label information in the common sample information, determine the confidence level of the first common sample information according to the difference between whether the first common sample information determined by the teacher sub-module is the recommended information of a specific user and the label information in the corresponding first common sample information; and for the second common sample information without label information in the common sample information, determine the confidence level corresponding to the second common sample information as the average confidence level of the first common sample information.

[0162] Further, the training unit 23 is further configured to stop adjusting the parameter values when the number of adjustments of the parameter values is equal to a preset number, or if the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold.

[0163] In the system of this embodiment, the information recommendation list corresponding to the information access request is output by the information recommendation model pre-trained by the training unit 23. During the pre-training process of the information recommendation model, the training unit 23 can respectively select a corresponding part of the sample information through the teacher filtering sub-module and the student filtering sub-module to filter out the training samples that interfere with the training of the teacher sub-module and the student sub-module, making the training of the information recommendation model more accurate. And because the sample information processed by the teacher sub-module and the student sub-module is less, the time spent on training the information recommendation model is less. Therefore, when the application background in the embodiment of the present invention performs information recommendation to the application terminal, it will be more accurate and the response time will be reduced.

[0164] The embodiment of the present invention also provides a server, and its structural schematic diagram is as Figure 11 shown. The server may vary greatly due to configuration or performance, and may include one or more central processing units (CPUs) 30 (for example, one or more processors) and a memory 31, and one or more storage media 32 for storing application programs 321 or data 322 (for example, one or more mass storage devices). Among them, the memory 31 and the storage media 32 can be transient storage or persistent storage. The program stored in the storage media 32 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 30 can be set to communicate with the storage media 32 and execute a series of instruction operations in the storage media 32 on the server.

[0165] Specifically, the application program 321 stored in the storage media 32 includes an application program for information recommendation, and this program may include the request acquisition unit 20, the information recommendation unit 21, the output unit 22, and the training unit 23 in the above information recommendation system, which will not be elaborated here. Further, the central processing unit 30 can be set to communicate with the storage media 32 and execute a series of operations corresponding to the information recommendation application program stored in the storage media 32 on the server.

[0166] The server may also include one or more power supplies 33, one or more wired or wireless network interfaces 34, one or more input / output interfaces 35, and / or one or more operating systems 323, such as WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0167] The steps performed by the information recommendation system (such as the application background) described in the above method embodiment can be based on this Figure 11The structure of the server shown.

[0168] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium storing a plurality of computer programs, which are adapted to be loaded and executed by a processor to perform the information recommendation method executed by the application background as described above.

[0169] Another aspect of the embodiments of the present invention further provides a server, including a processor and a memory;

[0170] The memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by the processor to perform the information recommendation method executed by the application background as described above; the processor is used to implement each computer program in the plurality of computer programs.

[0171] In addition, according to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the information recommendation method provided in the above various optional implementation manners.

[0172] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0173] The above has introduced in detail an information recommendation method, system, storage medium and server provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An information recommendation method, characterized in that, comprising: obtaining an information access request; determining an information recommendation list according to the information access request and a pre-trained information recommendation model, where the information recommendation list includes multiple pieces of recommended information; outputting the information recommendation list; wherein, when the information recommendation model is pre-trained: determining training samples, where the training samples include multiple pieces of sample information; determining an initial information recommendation model, where the initial information recommendation model includes a teacher filtering sub-module, a teacher sub-module, a student filtering sub-module, and a student sub-module; wherein, the teacher filtering sub-module is used to select a first part of the sample information from the multiple pieces of sample information, and the teacher sub-module is used to determine whether the first part of the sample information is recommended information for a specific user; the student filtering sub-module is used to select a second part of the sample information from the multiple pieces of sample information, and the student sub-module is used to determine whether the second part of the sample information is recommended information for a specific user, and the feature calculation layer included in the student sub-module is less than the feature calculation layer included in the teacher sub-module; training the information recommendation model according to the training samples and the initial information recommendation model, where the pre-trained information recommendation model includes a trained student sub-module; wherein, during the training process, the teacher sub-module guides the training of the student sub-module so that the student sub-module learns the information in the teacher sub-module.

2. The method according to claim 1, characterized in that, the teacher sub-module is specifically used to determine the user click-through rate of each piece of sample information, and the teacher filtering sub-module is specifically used to respectively determine the first scores of the multiple pieces of sample information according to the user click-through rates corresponding to the multiple pieces of sample information determined by the teacher sub-module, and select a part of the sample information with the highest first score from the multiple pieces of sample information as the first part of the sample information; the student sub-module is specifically used to determine the user click-through rate of each piece of sample information, and the student filtering sub-module is specifically used to respectively determine the second scores of the multiple pieces of sample information according to the user click-through rates corresponding to the multiple pieces of sample information determined by the student sub-module, and select a part of the sample information with the highest second score from the multiple pieces of sample information as the second part of the sample information.

3. The method according to claim 1, characterized in that, the training the information recommendation model according to the training samples and the initial information recommendation model specifically includes: determining whether the first part of the sample information belongs to the recommended information for a specific user through the initial information recommendation model, and determining whether the second part of the sample information belongs to the recommended information for a specific user; adjusting the initial information recommendation model according to the result determined by the initial information recommendation model to obtain the information recommendation model.

4. The method according to claim 3, characterized in that, the adjusting the initial information recommendation model according to the result determined by the initial information recommendation model specifically includes: Calculate a first loss function related to the teacher sub-module, a second loss function related to the student sub-module, a third loss function for distillation from the teacher sub-module to the student sub-module, and a fourth loss function for calculating the difference between the teacher sub-module and the student sub-module according to the result determined by the initial model for information recommendation based on the said information; Calculate the overall loss function of the initial model for information recommendation according to the first loss function, the second loss function, the third loss function, and the fourth loss function; Adjust the parameter values of the parameters in the initial model for information recommendation according to the overall loss function to obtain the information recommendation model.

5. The method according to claim 4, wherein, if the intersection between the first part of sample information and the second part of information is the common sample information, then the fourth loss function for calculating the difference between the teacher sub-module and the student sub-module specifically includes: Determine the confidence that the common sample information determined by the teacher sub-module is the recommended information for a specific user; Calculate the difference between whether the common sample information determined by the teacher sub-module and the student sub-module respectively is the recommended information for a specific user; Calculate the fourth loss function according to the confidence of the common sample information and the difference.

6. The method according to claim 5, wherein, the determination of the confidence that the common sample information determined by the teacher sub-module is the recommended information for a specific user specifically includes: For the first common sample information with label information in the common sample information, determine the confidence of the first common sample information according to the difference between whether the first common sample information determined by the teacher sub-module is the recommended information for a specific user and the label information in the corresponding first common sample information; For the second common sample information without label information in the common sample information, determine the confidence corresponding to the second common sample information as the average confidence of the first common sample information.

7. The method according to any one of claims 4 to 6, wherein, When the number of times of adjusting the parameter value is equal to the preset number of times, or if the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold, then stop adjusting the parameter value.

8. An information recommendation system, wherein, comprises: A request acquisition unit, configured to acquire an information access request; An information recommendation unit, configured to determine an information recommendation list according to the information access request and a pre-trained information recommendation model, where the information recommendation list includes multiple pieces of recommended information; An output unit, configured to output the information recommendation list; Among them, when the information recommendation model is pre-trained: training samples are determined, and the training samples include multiple pieces of sample information; an initial information recommendation model is determined, and the initial information recommendation model includes a teacher filtering sub-module, a teacher sub-module, a student filtering sub-module, and a student sub-module; wherein, the teacher filtering sub-module is used to select a first part of the sample information from the multiple pieces of sample information, and the teacher sub-module is used to determine whether the first part of the sample information is recommended information for a specific user; the student filtering sub-module is used to select a second part of the sample information from the multiple pieces of sample information, and the student sub-module is used to determine whether the second part of the sample information is recommended information for a specific user, and the feature calculation layer included in the student sub-module is less than the feature calculation layer included in the teacher sub-module; the information recommendation model is trained according to the training samples and the initial information recommendation model, and the pre-trained information recommendation model includes the trained student sub-module; wherein, during the training process, the teacher sub-module guides the training of the student sub-module, so that the student sub-module learns the information in the teacher sub-module.

9. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores multiple computer programs, and the computer programs are adapted to be loaded and executed by a processor to perform the information recommendation method according to any one of claims 1 to 7.

10. A server, characterized in that, comprises a processor and a memory; the memory is used to store multiple computer programs, and the computer programs are used to be loaded and executed by a processor to perform the information recommendation method according to any one of claims 1 to 7; the processor is used to implement each of the computer programs in the multiple computer programs.

11. A computer program product, characterized in that, the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the information recommendation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for training model for predicting click rate on line and recommendation system

    CN112182362A

  • Neural network model compression method and apparatus, and storage medium and chip

    WO2021042828A1

Cited By

  • Method for dynamically generating and pushing personalized aided teaching content based on AI (artificial intelligence) large model

    CN121682136A

  • Method for dynamically generating and pushing personalized auxiliary teaching content based on AI large model

    CN121682136B