Methods, apparatus, media, and program products for generating an online prediction model
By using online learning algorithms and attention mechanisms to generate model fusion strategies, the problems of dynamic adaptation and real-time optimization in recommendation systems are solved, enabling multi-model fusion and personalized learning, thereby improving the overall efficiency of recommendation systems.
Patent Information
- Application Number
- CN202210204981.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-02
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2042-03-02
AI Technical Summary
In existing technologies, recommendation systems lack dynamic adaptation and real-time optimization at each stage, resulting in error propagation and low efficiency, making it difficult to achieve personalized multi-task and multi-objective learning.
The feature data is learned online by multiple online learning algorithms, and the real-time online feedback data is learned by combining the attention mechanism to generate a model fusion strategy. Multiple models are dynamically fused in real time to generate an online prediction model.
It achieves dynamic adaptive multi-model fusion, reduces error propagation in intermediate links, improves the efficiency of each link, promotes personalized learning, and optimizes the end-to-end performance of the recommendation system.
Smart Images

Figure CN114676753B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication, and in particular to a technology for generating an online prediction model. BACKGROUND
[0002] In the prior art, artificial intelligence (AI) is a comprehensive technology in computer science, which studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline and involves a wide range of fields, such as natural language processing technology and machine learning / deep learning. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. SUMMARY
[0003] An object of the present application is to provide a method, device, medium and program product for generating an online prediction model.
[0004] According to an aspect of the present application, a method for generating an online prediction model is provided, which comprises:
[0005] performing online learning on feature data by multiple online learning algorithms to obtain a plurality of corresponding models;
[0006] learning online real-time backflow data based on an attention mechanism to obtain a model fusion strategy corresponding to the plurality of models;
[0007] performing real-time dynamic fusion on the plurality of models according to the model fusion strategy to generate an online prediction model.
[0008] According to an aspect of the present application, a first device for generating an online prediction model is provided, which comprises:
[0009] a first module for performing online learning on feature data by multiple online learning algorithms to obtain a plurality of corresponding models;
[0010] a second module for learning online real-time backflow data based on an attention mechanism to obtain a model fusion strategy corresponding to the plurality of models;
[0011] a third module for performing real-time dynamic fusion on the plurality of models according to the model fusion strategy to generate an online prediction model.
[0012] According to an aspect of the present application, a computer device for generating an online prediction model is provided, which comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the operations of any of the above methods.
[0013] According to an aspect of the present application, a computer readable storage medium is provided, having stored thereon a computer program, characterized in that the computer program is executed by a processor to implement the operations of any of the above methods.
[0014] According to an aspect of the present application, a computer program product is provided, comprising a computer program, which is executed by a processor to implement the steps of any of the above methods.
[0015] Compared with the prior art, the present application learns feature data respectively by multiple online learning algorithms to obtain corresponding multiple models, learns online real-time backflow data based on an attention mechanism to obtain a model fusion strategy corresponding to the multiple models, and dynamically fuses the multiple models in real time according to the model fusion strategy to generate an online prediction model. The present application realizes a multi-task online learning system based on dynamic adaptive multi-model fusion, can learn from the strengths of each model, better promotes the fusion of multiple models, and at the same time, fuses the unique learning mode of each model based on time length, click rate, likes, collections and other multi-task multi-target, which can play the role of personalized learning. Based on the online learning system of the present application, the recall, rough sorting, fine sorting, cold start and other links of the recommendation system can be connected and punched through, achieving unified optimization of the whole link, reducing error transmission of intermediate links, and improving the efficiency of each link. BRIEF DESCRIPTION OF DRAWINGS
[0016] Other features, objects, and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings:
[0017] Figure 1 A flow chart of a method for generating an online prediction model is shown according to an embodiment of the present application;
[0018] Figure 2 An architecture diagram of a multi-task online learning system based on dynamic adaptive multi-model fusion is shown according to an embodiment of the present application;
[0019] Figure 3 A first device structure diagram for generating an online prediction model is shown according to an embodiment of the present application;
[0020] Figure 4 An exemplary system that can be used to implement various embodiments described herein is shown.
[0021] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0022] The application is described in further detail below with reference to the drawings.
[0023] In one typical configuration of the present application, the terminal, the device of the service network and the trusted party each include one or more processors (e.g., a Central Processing Unit (CPU)), input / output interfaces, network interfaces, and memories.
[0024] The memory can include non-persistent memory and / or other volatile memory, representing examples of the computer readable media. The memory is an example of computer readable media.
[0025] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device.
[0026] The device referred to in the present application includes, but is not limited to, a terminal, a network device, or a device formed by integrating a terminal and a network device through a network. The terminal includes, but is not limited to, any kind of mobile electronic product capable of human-computer interaction (for example, human-computer interaction through a touch panel), such as a smart phone, a tablet computer, etc. The mobile electronic product can adopt any operating system, such as an Android operating system, an iOS operating system, etc. The network device includes an electronic device capable of automatically performing numerical calculation and information processing according to a pre-set or stored instruction. The hardware of the network device includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc. The network device includes, but is not limited to, a computer, a network host, a single network server, a plurality of network servers, or a cloud formed by a plurality of servers. The cloud is formed by a large number of computers or network servers based on cloud computing. The cloud computing is a kind of distributed computing, which is a virtual supercomputer formed by a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, a wireless Ad Hoc network, etc. Preferably, the device can also be a program running on the terminal, the network device, or a device formed by integrating a terminal and a network device, a network device, a touch terminal, or a device formed by integrating a touch terminal and a network device through a network.
[0027] Of course, those skilled in the art should understand that the above device is only an example, and other existing or future devices, such as devices that can be applicable to the present application, should also be included in the protection scope of the present application, and are hereby included by reference.
[0028] In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.
[0029] Figure 1A method flowchart for generating an online prediction model is shown according to one embodiment of the present application, which includes steps S11, S12 and S13. In step S11, the first device performs online learning on feature data by multiple online learning algorithms respectively to obtain corresponding multiple models; in step S12, the first device learns online real-time backflow data based on an attention mechanism to obtain a model fusion strategy corresponding to the multiple models; and in step S13, the first device dynamically fuses the multiple models in real time according to the model fusion strategy to generate an online prediction model.
[0030] In step S11, the first device performs online learning on feature data by multiple online learning algorithms respectively to obtain corresponding multiple models. In some embodiments, the first device can be a user device, or also a network device. In some embodiments, the online learning algorithm is a kind of algorithm that uses historical data up to the current time to make decisions, and is also known as a stream algorithm. Each machine applies incremental training to learn one instance at a time. When new data becomes available, the algorithm does not need to be retrained on all data, because they continue to improve the existing model incrementally. In some embodiments, online learning is a training method of a model. Online learning can quickly adjust the model in real time according to online feedback data, so that the model reflects the changes on the line in time and improves the accuracy of online prediction. The process of online learning includes presenting the prediction results of the model to the user, then collecting the feedback data of the user, and then training the model to form a closed-loop system. In some embodiments, the traditional learning algorithm has a relatively long update cycle after the model goes online (usually one day, and the efficiency is high when it is one hour). Such a model is generally static after going online (it will not change within a period of time) and will not interact with the online situation. If the prediction is wrong, it can only be corrected at the next update time. Online learning algorithm is different. It dynamically adjusts the model according to the prediction results online. If the model prediction is wrong, it will be corrected in time. Therefore, online learning algorithm can reflect the changes on the line more timely. In some embodiments, the feature data is obtained by feature extraction on the data of each user (e.g., user browsing data, data of items browsed by the user). For example, the feature data includes but is not limited to user browsing behavior features, item features browsed by the user, and some context features, etc. In some embodiments, online learning is performed on the feature data by each online learning algorithm respectively to obtain a model corresponding to each online learning algorithm respectively.
[0031] In step S12, the first device learns the online real-time backflow data based on an attention mechanism to obtain a model fusion strategy corresponding to the plurality of models. In some embodiments, the essence of the attention mechanism comes from the human visual attention mechanism. When people perceive things, they generally do not look at everything from beginning to end in one scene each time, but often observe and pay attention to a specific part according to the needs. When people find that a scene often appears in a part where they want to observe, people will learn to put their attention on that part in the future when similar scenes appear. The attention mechanism is similar to the human visual attention mechanism, that is, it focuses attention on important points in a large amount of information, selects key information, and ignores other unimportant information. The principle of attention is to calculate the matching degree of the current input sequence and the output vector. The higher the matching degree, the higher the relative score of the attention focus point. The matching degree weight calculated by the attention is limited to the current sequence pair, not the overall weight like the network model weight. In some embodiments, the attention mechanism is a data processing method in machine learning, which is widely used in various types of machine learning tasks such as natural language processing, image recognition, and speech recognition. The attention mechanism simulates the internal process of biological observation behavior, that is, a mechanism that aligns internal experience and external feeling to increase the observation accuracy of a certain area, to automatically learn and calculate the contribution size of input data to output data. In some embodiments, the online real-time backflow data includes the prediction value of the online prediction model of the application client online real-time backflow to the application server and the related business indicators of the application after using the prediction value in the application (for example, presenting the information item with the highest click rate (for example, a certain book) in a certain recommendation position). The business indicators can be user behavior feedback data in the application for the recommendation position, which includes but is not limited to whether the user clicks the recommendation position to enter the reading page of a certain recommended book, whether the user collects the recommended book corresponding to the recommendation position, the number of chapters the user has watched in the recommended book corresponding to the recommendation position, whether the user unlocks the paid chapter in the recommended book corresponding to the recommendation position by watching the incentive video, and the like. In some embodiments, the model fusion strategy can be a weighting coefficient corresponding to each model, or it can also be a model fusion method corresponding to the plurality of models and related fusion parameters, etc., wherein the model fusion method includes but is not limited to self-aggregation method, boosting method, stacking method, and combination of at least two of the above three methods.
[0032] In step S13, the first device dynamically fuses the plurality of models in real time according to the model fusion strategy to generate an online prediction model. In some embodiments, the plurality of models can be dynamically fused in real time in a model weighting manner according to the weighting weight coefficients corresponding to each model in the model fusion strategy. In some embodiments, the plurality of models can also be dynamically fused in real time at the model level in the model fusion manner specified in the model fusion strategy. In some embodiments, the model fusion strategy corresponding to the plurality of models is obtained by learning the online real-time backflow data (for example, online past business indicator effects), and the plurality of models are dynamically and adaptively fused in real time according to the model fusion strategy, which can better promote the fusion of the plurality of models. In some embodiments, the online prediction model is generated by dynamically fusing the plurality of models in real time, and the input of the online prediction model includes but is not limited to user information (for example, user portrait information) of a certain user, feature information obtained by performing feature extraction on the user information, related information (for example, author information, classification information, title information, and introduction information) of a certain information item (for example, a recommended book), and feature information obtained by performing feature extraction on the related information. In some embodiments, the output of the online prediction model includes but is not limited to predicting the user click rate of a certain information item (for example, a recommended book), predicting the user reading average duration of the information item, predicting whether a certain user will click the information item, predicting the duration of reading the information item by the user, and the like. In some embodiments, the online prediction model can balance the indicators and performance by combining the plurality of models. In some embodiments, the fusion can be performed at the model training and learning end, or the fusion can also be performed at the parameter server end, which is used to store the model parameters learned by the model, or the fusion can also be performed at the online prediction module end. In some embodiments, the dynamic adaptive module obtains the model fusion strategy corresponding to the plurality of models by learning the online real-time backflow data, dynamically updates the expert system module according to the model fusion strategy, and then dynamically fuses the plurality of models in real time through the expert system module to generate an online prediction model.The application realizes a multi-task online learning system based on dynamic self-adaptive multi-model fusion, which can learn from the strengths of each model, better promote the fusion of multi-model, and at the same time, fuse the unique multi-task multi-object learning mode of each model based on time length, click rate, likes, collections, etc., which can achieve the effect of personalized learning. Based on the online learning system of the application, the recall, rough sorting, fine sorting, cold start and other links of the recommendation system can be connected and optimized, the error transmission of the intermediate links can be reduced, and the efficiency of each link can be improved. Each link in the recommendation system involves some related models, and the application can use the online learning system in the application to optimize all links, which can save a lot of resources and costs.
[0033] In some embodiments, the online learning algorithm includes a streaming logistic regression algorithm; a factorization machine algorithm. In some embodiments, the model based on ftrl (Follow The Regularized Leader) algorithm online learning fuses many expert knowledge cross features, has good online prediction efficiency under the premise of ensuring model accuracy. In some embodiments, the model based on fm (Factorization Machine) algorithm online learning learns hidden cross features, and through theoretical transformation, the model prediction efficiency can be guaranteed under the premise of ensuring model accuracy. In some embodiments, the application is not only used for the fusion of the model based on ftrl algorithm online learning and the model based on fm algorithm online learning, but also suitable for the fusion of any model, not only suitable for the fusion of two models, but also suitable for the fusion of three models or more models.
[0034] In some embodiments, the step S11 includes: the first device, for each online learning algorithm in the plurality of online learning algorithms, performing online learning on the feature data in multiple dimensions based on the online learning algorithm through multiple views, to obtain a model corresponding to the online learning algorithm. In some embodiments, for each online learning algorithm in the plurality of online learning algorithms, based on the multi-head attention (Multi Head Attention) mechanism in the field of deep learning, a plurality of views are opened through the online learning algorithm to learn feature data in multiple different dimensions, to obtain a model corresponding to the online learning algorithm. The model can capture more rich features from different dimensions and different views, to achieve the purpose of improving model generalization and accelerating model convergence speed.
[0035] In some embodiments, the online learning of the feature data in multiple dimensions based on multiple perspectives by the online learning algorithm comprises: learning of multiple tasks in parallel by the online learning algorithm; and online learning of the feature data in multiple dimensions based on multiple perspectives for each task. In some embodiments, multi-task learning is a machine learning method opposite to single-task learning. In the field of machine learning, the standard algorithm theory is to learn one task at a time, that is, the output of the system is a real number. A complex learning problem is first decomposed into theoretically independent sub-problems, then each sub-problem is learned, and finally a mathematical model of the complex problem is established by combining the learning results of the sub-problems. Multi-task learning is a kind of joint learning, multiple tasks are learned in parallel, and the results influence each other. Multi-task learning is to solve multiple problems at the same time. In some embodiments, multiple related tasks are placed together for parallel learning. For example, the model simultaneously learns multiple tasks such as predicting user click rate and predicting user average reading time. The model can output multiple prediction results corresponding to multiple tasks for one input. For example, the model can output the user click rate and the user average reading time corresponding to an information item for the input of the information item, so that the model can take into account the estimation of the user click rate and the estimation of the user average reading time at the same time, achieving the effect of multi-task learning.
[0036] In some embodiments, the learning of the plurality of tasks in parallel by the online learning algorithm further comprises adjusting the learning weights of the plurality of tasks according to at least one model learning objective. In some embodiments, the learning weight of each task in model learning needs to be adjusted according to at least one model learning objective, i.e., each task corresponds to a different learning weight, and each task needs to be weighted in model learning according to the learning weight according to at least one model learning objective, so that in multi-task learning, some tasks are optimized as much as possible, but do not affect the optimization of other tasks, i.e., the optimization degree or optimization priority of a task in multi-task learning is proportional to the learning weight of the task, the higher the learning weight of the task, the higher the optimization degree or optimization priority of the task in multi-task learning. In some embodiments, it is necessary to first determine whether each task has an impact on the completion of the model learning objective and the degree of the impact, if a task has an impact on the completion of the model learning objective or the degree of the impact of the task is greater than or equal to a predetermined degree threshold or the degree of the impact of the task is greater than the degree of the impact of other tasks on the completion of the model learning objective, the learning weight of the task can be appropriately increased, if a task has no impact on the completion of the model learning objective or the degree of the impact of the task is less than a predetermined degree threshold or the degree of the impact of the task is less than the degree of the impact of other tasks on the completion of the model learning objective, the learning weight of the task can be appropriately increased, the learning weight of the task can be appropriately reduced. In some embodiments, a plurality of model learning objectives are added to weight the learning weights of the plurality of tasks such as predicting user click rate and predicting user average reading time, achieving the effect of multi-task multi-objective learning.
[0037] In some embodiments, the plurality of tasks comprises user click rate estimation and user average reading time estimation. In some embodiments, the model simultaneously learns to predict user click rate and predict user average reading time, for example, the relevant information of an information item (for example, a recommended book) (for example, the author information, classification information, title information, and introduction information of the recommended book) or the feature information obtained by feature extraction on the relevant information is input into the model, and the estimated user click rate and the estimated user average reading time corresponding to the information item can be output simultaneously, i.e., the probability of each user clicking the information item and the average time of each user reading the information item are estimated.
[0038] In some embodiments, the at least one model learning target comprises at least one of the following: improving user click rate; improving user average reading time; improving user like number; improving user collection number. In some embodiments, the model learning target comprises, but is not limited to, improving user click rate and / or a specific improvement value, improving user average reading time and / or a specific improvement value, improving user like number and / or a specific improvement value, improving user collection number and / or a specific improvement value. For example, the model learning target can be to improve user collection number and improve user average reading time at the same time. For example, the model learning target can be that if the information item (for example, a certain book) with the highest user click rate output by the learned model is presented on a certain recommendation position, the user collection number of the recommendation position or the information item and the user average reading time can be improved at the same time, and the specific improvement value is greater than or equal to a predetermined value threshold.
[0039] In some embodiments, the learning of the online real-time backflow data based on the attention mechanism comprises: learning the online real-time backflow data based on the attention mechanism using a deep neural network. In some embodiments, in the dynamic adaptive module, based on the principle of self-attention mechanism, a deep neural network (DNN) can be used to perform real-time learning calculation on online real-time backflow data (for example, online past business indicator effect) to calculate a strategy for model fusion in the next step. In some embodiments, the dynamic adaptive module dynamically updates the expert system module by performing real-time learning calculation on online real-time backflow data, and the expert system module dynamically fuses the plurality of modules in real time to learn the advantages of each model.
[0040] In some embodiments, the model fusion strategy comprises a weighting weight coefficient corresponding to each model. In some embodiments, by performing real-time learning calculation on online real-time backflow data (for example, online past business indicator effect), a weighting weight coefficient corresponding to each model can be calculated to dynamically adjust the adaptive expert system module, and then the expert system module uses the weighting weight coefficient to perform real-time dynamic fusion of the plurality of models in a model weighting manner, thereby better promoting the fusion of multiple models.
[0041] In some embodiments, the learning of the online real-time backflow data based on the attention mechanism includes learning based on a Bayesian network to construct a probabilistic graphical model based on the online real-time backflow data. In some embodiments, the probabilistic graphical model is a theory that uses graphs to represent the probability dependence of variables, combines the knowledge of probability theory and graph theory, and uses graphs to represent the joint probability distribution of variables related to the model. The probabilistic graphical model is a general term for a model that uses graphical patterns to express probability-based relationships. In some embodiments, the expert system module based on the Bayesian network learns and calculates the online real-time backflow data (e.g., the effect of online past business indicators) in real time based on the probabilistic graphical model constructed based on the online real-time backflow data, and calculates the strategy for model fusion in the next step. In some embodiments, the fusion strategy is updated in real time and dynamically based on the online performance of the online prediction model.
[0042] In some embodiments, the model fusion strategy includes a model fusion method. In some embodiments, the model fusion strategy includes a model fusion method and related fusion parameters, where the model fusion method includes but is not limited to a bootstrap aggregation method, a boosting method, a stacking method, and a combination of at least two of the above three methods. In some embodiments, the subsequent model fusion method and related fusion parameters can be calculated by real-time learning and calculation of online real-time backflow data (e.g., the effect of online past business indicators), and the adaptive expert system module is dynamically adjusted. Then, the subsequent expert system module uses the model fusion method to dynamically fuse the multiple models in real time according to the model fusion method and related fusion parameters, thereby better promoting the fusion of multiple models.
[0043] In some embodiments, the real-time dynamic fusion of the multiple models according to the model fusion strategy includes using a model weighting method to dynamically fuse the multiple models in real time according to the weighted weight coefficient corresponding to each model in the model fusion strategy. In some embodiments, the multiple models are dynamically fused in real time using a model weighting method according to the weighted weight coefficient corresponding to each model. The specific method can be to dynamically fuse the multiple models in real time at the feature level, i.e., the input level of the model, or it can also be to dynamically fuse the multiple models in real time at the model result level, i.e., the output level of the model.
[0044] In some embodiments, the model weighting manner is used to dynamically fuse the plurality of models in real time, including any one of the following: the model weighting manner is used to dynamically fuse the plurality of models in real time at a feature level; the model weighting manner is used to dynamically fuse the plurality of models in real time at a model result level. In some embodiments, the model weighting manner can be used to dynamically fuse the plurality of models in real time at a feature level, i.e., an input level of the model, for example, the feature data is divided and distributed to each model as the input of the model according to the corresponding weighting weight coefficient of the model. In some embodiments, the model weighting manner can also be used to dynamically fuse the plurality of models in real time at a model result level, i.e., an output level of the model, for example, the output result of each model is weighted according to the corresponding weighting weight coefficient of the model, and then the output result of the online prediction model is obtained according to the weighted output result corresponding to each model, for example, the sum of the weighted output results corresponding to each model is taken as the output result of the online prediction model.
[0045] In some embodiments, the plurality of models are dynamically fused in real time according to the model fusion strategy, including: the model fusion manner in the model fusion strategy is used to dynamically fuse the plurality of models in real time at a model level; wherein the model fusion manner includes any one of the following: self-aggregation manner; boosting manner; stacking manner; combination of at least two of the above three manners. In some embodiments, the model fusion manner specified in the model fusion strategy is used to dynamically fuse the plurality of models in real time at a model level, wherein the fusion at a model level includes but is not limited to stacking and designing of the plurality of models, for example, a multi-layer model is constructed based on the plurality of models, and the output result of one model is taken as the feature input of another model. In some embodiments, the model fusion manner includes but is not limited to Bagging, Boosting, Stacking, combination of at least two of the above three manners, etc., wherein Bagging usually considers homogeneous weak learners, learns these weak learners independently in parallel, and combines them according to a certain deterministic averaging process, Boosting also usually considers homogeneous weak learners. It learns these weak learners sequentially in a highly adaptive manner (each base model depends on the previous model), and combines them according to a certain deterministic strategy, Stacking usually considers heterogeneous weak learners, learns them in parallel, and combines them by training a meta-model, and outputs a final prediction result according to the prediction result of different weak models.
[0046] Figure 2 An architecture diagram of a multi-task online learning system based on dynamic adaptive multi-model fusion according to an embodiment of the present application is shown.
[0047] As shown in Figure 2 The feature service in the online learning system is used to store feature data required for online learning of respective models. For each of a plurality of online learning algorithms, online learning of the feature data is performed in multiple dimensions by the online learning algorithm based on multiple views, to obtain a model (multi-view real-time training model 1, multi-view real-time training model 2) corresponding to the online learning algorithm. The parameter server 1 and the parameter server 2 are used to respectively store model parameters learned by each model. Then, the dynamic adaptive model updates the online prediction module (i.e., the expert system module) in real time / timing by learning online real-time backflow data, and generates an online prediction model by real-time dynamic fusion of the plurality of models through the online prediction module.
[0048] Figure 3 A first device structure diagram for generating an online prediction model according to an embodiment of the present application is shown. The device includes a first module 11, a second module 12, and a third module 13. The first module 11 is used to perform online learning of feature data by a plurality of online learning algorithms, to obtain a plurality of corresponding models. The second module 12 is used to learn online real-time backflow data based on an attention mechanism, to obtain a model fusion strategy corresponding to the plurality of models. The third module 13 is used to perform real-time dynamic fusion of the plurality of models according to the model fusion strategy, to generate an online prediction model.
[0049] The first module 11 is configured to perform online learning on the feature data by using a plurality of online learning algorithms respectively to obtain a plurality of corresponding models. In some embodiments, the first device can be a user device or a network device. In some embodiments, the online learning algorithm is an algorithm that uses historical data up to the current time to make decisions. The online learning algorithm is also known as a stream algorithm. Each machine applies incremental training to learn one instance at a time. When new data becomes available, the algorithm does not need to be retrained on all data because it continues to improve the existing model incrementally. In some embodiments, online learning is a training method of a model. Online learning can quickly adjust the model in real time according to online feedback data, so that the model reflects the changes on the line in time and improves the accuracy of online prediction. The process of online learning includes displaying the prediction result of the model to the user, then collecting the feedback data of the user, and then using the feedback data to train the model to form a closed-loop system. In some embodiments, the traditional learning algorithm has a relatively long update cycle (usually one day, and the efficiency is one hour). After the model is put online, the model is generally static (it will not change within a period of time) and will not interact with the online situation. If the prediction is wrong, it can only be corrected at the next update time. The online learning algorithm is different. It dynamically adjusts the model according to the prediction result on the line. If the model prediction is wrong, it will be corrected in time. Therefore, the online learning algorithm can reflect the changes on the line more timely. In some embodiments, the feature data is obtained by performing feature extraction on the data of each user (for example, the user's browsing data, the data of the item browsed by the user). For example, the feature data includes but is not limited to user browsing behavior features, user browsing item features, and some context features, etc. In some embodiments, the feature data is learned by each online learning algorithm respectively to obtain a model corresponding to each online learning algorithm respectively.
[0050] A second module 12 is configured to learn from the online real-time backflow data based on an attention mechanism to obtain a model fusion strategy corresponding to the plurality of models. In some embodiments, the essence of the attention mechanism comes from the human visual attention mechanism. When people perceive things, they generally do not look at everything from beginning to end in one scene every time, but often observe and pay attention to a specific part according to the needs. When people find that a scene often appears in a part where they want to observe, people will learn to put their attention on that part in the future when similar scenes appear. The attention mechanism is similar to the human visual attention mechanism, that is, it focuses on important points in a large amount of information, selects key information, and ignores other unimportant information. The principle of attention is to calculate the matching degree of the current input sequence and the output vector. The higher the matching degree, the higher the relative score of the attention focus point. The matching degree weight calculated by the attention is limited to the current sequence pair, not the overall weight like the network model weight. In some embodiments, the attention mechanism is a data processing method in machine learning, which is widely used in various types of machine learning tasks such as natural language processing, image recognition, and speech recognition. The attention mechanism simulates the internal process of biological observation behavior, that is, a mechanism that aligns internal experience and external feeling to increase the observation accuracy of a certain area, to automatically learn and calculate the contribution size of input data to output data. In some embodiments, the online real-time backflow data includes the prediction value of the online prediction model provided by the application client to the application server in real time, and the related business indicators of the application after using the prediction value in the application (for example, presenting the information item with the highest click rate (for example, a certain book) in a certain recommendation position). The business indicators can be the user behavior feedback data in the application for the recommendation position, which includes but is not limited to whether the user clicks the recommendation position to enter the reading page of a certain recommended book, whether the user collects the recommended book corresponding to the recommendation position, the number of chapters the user has watched in the recommended book corresponding to the recommendation position, whether the user unlocks the paid chapter in the recommended book corresponding to the recommendation position by watching the incentive video, etc. In some embodiments, the model fusion strategy can be a weighted weight coefficient corresponding to each model, or it can also be a model fusion method corresponding to the plurality of models and related fusion parameters, etc. The model fusion method includes but is not limited to self-aggregation method, boosting method, stacking method, and combination of at least two of the above three methods.
[0051] a third module 13 configured to dynamically fuse the plurality of models in real time according to the model fusion strategy to generate an online prediction model. In some embodiments, the plurality of models can be dynamically fused in real time in a model weighting manner according to the weighting weight coefficients corresponding to each model in the model fusion strategy. In some embodiments, the plurality of models can also be dynamically fused in real time at the model level in a model fusion manner specified in the model fusion strategy. In some embodiments, the model fusion strategy corresponding to the plurality of models can be obtained by learning the online real-time backflow data (e.g., online past business indicator effect), and the plurality of models can be dynamically and adaptively fused in real time according to the model fusion strategy, which can better promote the fusion of the plurality of models. In some embodiments, the online prediction model generated by dynamically fusing the plurality of models in real time can have an input including but not limited to user information (e.g., user portrait information) of a certain user, feature information obtained by performing feature extraction on the user information, related information (e.g., author information, classification information, title information, and introduction information) of a certain information item (e.g., a recommended book), and feature information obtained by performing feature extraction on the related information. In some embodiments, the output of the online prediction model can include but is not limited to predicting the user click rate of a certain information item (e.g., a recommended book), predicting the user reading average duration of the information item, predicting whether a certain user will click the information item, predicting the duration of reading the information item by the user, and the like. In some embodiments, the online prediction model can combine the plurality of models to complement each other and balance the indicators and performance. In some embodiments, the fusion can be performed at a model training and learning end, or can also be performed at a parameter server end configured to store the model parameters learned by the model, or can also be performed at an online prediction module end. In some embodiments, the dynamic adaptive module can obtain the model fusion strategy corresponding to the plurality of models by learning the online real-time backflow data, dynamically update the expert system module according to the model fusion strategy, and then dynamically fuse the plurality of models in real time through the expert system module to generate an online prediction model.The application realizes a multi-task online learning system based on dynamic self-adaptive multi-model fusion, which can learn from the advantages of each model, better promote the fusion of multi-model, and at the same time, fuse the unique multi-task multi-object learning mode of each model based on time length, click rate, likes, collection, etc., which can play the role of personalized learning. Based on the online learning system of the application, the recall, rough sorting, fine sorting, cold start and other links of the recommendation system can be connected and optimized, the error transmission of the intermediate links is reduced, and the efficiency of each link is improved. Each link in the recommendation system involves some related models, and the application can use the online learning system in the application to optimize all links, which can save a lot of resources and costs.
[0052] In some embodiments, the online learning algorithm includes: a streaming logistic regression algorithm; a factorization machine algorithm. Here, the related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference. Figure 1 The related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference.
[0053] In some embodiments, the one-to-one module 11 is configured to: for each online learning algorithm in the plurality of online learning algorithms, perform online learning on the feature data in multiple dimensions based on multiple perspectives by the online learning algorithm, to obtain a model corresponding to the online learning algorithm. Here, the related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference. Figure 1 The related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference.
[0054] In some embodiments, the online learning on the feature data in multiple dimensions based on multiple perspectives by the online learning algorithm includes: learning on multiple tasks in parallel by the online learning algorithm; for each task, performing online learning on the feature data in multiple dimensions based on multiple perspectives. Here, the related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference. Figure 1 The related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference.
[0055] In some embodiments, the learning on multiple tasks in parallel by the online learning algorithm further includes: adjusting the learning weight of the multiple tasks according to at least one model learning target. Here, the related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference. Figure 1 The related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference.
[0056] In some embodiments, the multiple tasks include: user click rate estimation; user average reading time length estimation. Here, the related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference. Figure 1 The related operations are the same as or similar to those in the embodiments shown in the description, and are not described again here by way of reference.
[0057] In some embodiments, the at least one model learning objective comprises at least one of: improving user click rate; improving user average reading duration; improving user like number; improving user collection number. Here, the related operations are the same or similar to those in the embodiments shown in the above Figure 1 and will not be described here again, which are hereby included herein by reference.
[0058] In some embodiments, the learning of the online real-time backflow data based on the attention mechanism comprises: learning of the online real-time backflow data based on the attention mechanism by using a deep neural network. Here, the related operations are the same or similar to those in the embodiments shown in the above Figure 1 and will not be described here again, which are hereby included herein by reference.
[0059] In some embodiments, the model fusion strategy comprises a weighting weight coefficient corresponding to each model. Here, the related operations are the same or similar to those in the embodiments shown in the above Figure 1 and will not be described here again, which are hereby included herein by reference.
[0060] In some embodiments, the learning of the online real-time backflow data based on the attention mechanism comprises: learning of the online real-time backflow data based on the attention mechanism by constructing a probabilistic graphical model based on the online real-time backflow data through a Bayesian network. Here, the related operations are the same or similar to those in the embodiments shown in the above Figure 1 and will not be described here again, which are hereby included herein by reference.
[0061] In some embodiments, the model fusion strategy comprises a model fusion manner. Here, the related operations are the same or similar to those in the embodiments shown in the above Figure 1 and will not be described here again, which are hereby included herein by reference.
[0062] In some embodiments, the real-time dynamic fusion of the plurality of models according to the model fusion strategy comprises: real-time dynamic fusion of the plurality of models in a model weighting manner according to the weighting weight coefficient corresponding to each model in the model fusion strategy. Here, the related operations are the same or similar to those in the embodiments shown in the above Figure 1 and will not be described here again, which are hereby included herein by reference.
[0063] In some embodiments, the real-time dynamic fusion of the plurality of models in the model weighting manner comprises any one of: real-time dynamic fusion of the plurality of models in the model weighting manner at a feature level; real-time dynamic fusion of the plurality of models in the model weighting manner at a model result level. Here, the related operations are the same or similar to those in the embodiments shown in the above Figure 1 and will not be described here again, which are hereby included herein by reference.
[0064] In some embodiments, the real-time dynamic fusion of the plurality of models according to the model fusion strategy comprises: adopting a model fusion manner in the model fusion strategy to perform real-time dynamic fusion of the plurality of models at a model layer; and the model fusion manner comprises any one of the following: a self-help aggregation manner, a boosting manner, a stacking manner, or a combination of at least two of the above three manners. Here, the related operations are the same as or similar to those of the embodiments shown in the foregoing embodiments, and thus are not described again, and are hereby included by reference. Figure 4 The embodiments shown in the foregoing embodiments are the same or similar to those of the embodiments shown in the foregoing embodiments, and thus are not described again, and are hereby included by reference.
[0065] In addition to the methods and devices described in the foregoing embodiments, the present application also provides a computer-readable storage medium storing computer code, which, when executed, performs the method of any one of the foregoing.
[0066] The present application also provides a computer program product, which, when executed by a computer device, performs the method of any one of the foregoing.
[0067] The present application also provides a computer device, which comprises:
[0068] one or more processors;
[0069] a memory for storing one or more computer programs;
[0070] when the one or more computer programs are executed by the one or more processors, the one or more processors implement the method of any one of the foregoing.
[0071] Figure 4 An exemplary system that can be used to implement various embodiments described in the present application is shown;
[0072] As shown in some embodiments, system 300 can function as any of the devices in the various embodiments. In some embodiments, system 300 can include one or more computer-readable media (e.g., system memory or NVM / storage 320) having instructions and one or more processors (e.g., processor(s) 305) coupled with the one or more computer-readable media and configured to execute the instructions to implement modules to perform the actions described in the present application.
[0073] For one embodiment, system control module 310 can include any suitable interface controllers to provide any suitable interface between each of the one or more processors 305 and / or any suitable device or component in communication with system control module 310.
[0074] The system control module 310 can include a memory controller module 330 to provide an interface to system memory 315. The memory controller module 330 can be a hardware module, a software module, and / or a firmware module.
[0075] The system memory 315 can be used, for example, to load and store data and / or instructions for the system 300. For one embodiment, the system memory 315 can include any suitable volatile memory, such as suitable DRAM. In some embodiments, the system memory 315 can include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0076] For one embodiment, the system control module 310 can include one or more input / output (I / O) controllers to provide an interface to the NVM / storage device 320 and the communication interface(s) 325.
[0077] The NVM / storage device 320 can be used, for example, to store data and / or instructions. The NVM / storage device 320 can include any suitable non-volatile memory (e.g., flash memory) and / or can include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).
[0078] The NVM / storage device 320 can include a storage resource that is physically part of the device on which the system 300 is installed, or that is accessed via the device but not necessarily physically part of the device. For example, the NVM / storage device 320 can be accessed over a network via the communication interface(s) 325.
[0079] The communication interface(s) 325 can provide an interface to the system 300 to communicate over one or more networks and / or with any other suitable device. The system 300 can communicate wirelessly with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols.
[0080] For one embodiment, at least one of the processor(s) 305 can be packaged together with logic for one or more controllers of the system control module 310 (e.g., a memory controller module 330). For one embodiment, at least one of the processor(s) 305 can be packaged together with logic for one or more controllers of the system control module 310 to form a system in a package (SiP). For one embodiment, at least one of the processor(s) 305 can be integrated on the same die with logic for one or more controllers of the system control module 310. For one embodiment, at least one of the processor(s) 305 can be integrated on the same die with logic for one or more controllers of the system control module 310 to form a system on a chip (SoC).
[0081] In various embodiments, the system 300 can be, but is not limited to, a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet, a netbook, etc.). In various embodiments, the system 300 can have more or less components, and / or different architectures. For example, in some embodiments, the system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including touch screen displays), non- volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and speakers.
[0082] It is noted that the present application can be implemented in software and / or in a combination of software and hardware, e.g., application specific integrated circuit (ASIC), general purpose computer or any other similar hardware devices. In one embodiment, the software program of the present application is implemented by the processor to perform predetermined functions or tasks. Also, the software program of the present application (including related data structures) can be stored in a computer readable recording medium, e.g., RAM memory, magnetic or optical drive or diskette, and so on. Further, some of the steps or functions noted in the application can be implemented as circuitry, which work in conjunction with the processor to do these steps or functions, etc.
[0083] In addition, part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present application can be called or provided. Those skilled in the art should understand that the form of computer program instructions in computer readable medium includes but is not limited to source file, executable file, installation package file and the like, and accordingly, the way of computer program instructions executed by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.
[0084] Communication media includes any medium through which computer readable instructions, data structures, program modules, or other data is communicated from one system to another, e.g., according to the present application. Communication media can include wired media such as a wired network or direct-wired connection carrying one or more data signals, and wireless media such as acoustic, electromagnetic, RF, microwave, and infrared, capable of propagating energy patterns into space. Computer readable instructions, data structures, program modules, or other data can be embodied as modulated data signals, e.g., carrier waves, such as those implemented as part of an extended spectrum technique. The term "modulated data signal" refers to a signal that has one or more of its characteristics changed or set in a manner to encode information in the signal. The modulation can be analog, digital, or a combination of the two.
[0085] By way of example, and not limitation, computer readable storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. For example, computer readable storage media includes, but is not limited to, random access memory (RAM), such as dynamic RAM (DRAM), static RAM (SRAM), and / or non-volatile memory, such as read only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic and / or optical disks, other magnetic media, and / or optical media. The computer readable storage media can also be other means for storing information, such as cache, network accessible storage, or the like. However, the computer readable storage media does not include communication media.
[0086] Here, according to one embodiment of the present application includes a device, the device includes a memory for storing computer program instructions and a processor for executing program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to run the method and / or technical solutions based on the foregoing according to the plurality of embodiments of the present application.
[0087] It will be obvious to a person skilled in the art that the application is not limited to the details of the above-described exemplary embodiments, but that the application can be implemented in other embodiments without departing from the scope of the application. The embodiments are therefore to be seen as illustrative and not restrictive, the scope of the application being defined by the appended claims rather than by the description above, which is therefore intended merely as explanatory and not restrictive. No reference signs in the claims should be considered as limiting the scope of the claims in any way. Furthermore, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. Multiple units or devices also can be presented by a single unit or device, for example by software or hardware. The words "first", "second", etc. do not imply any special ordering, but are used for naming purposes only.
Claims
1. A method for generating an online prediction model, wherein, The method comprises: For each online learning algorithm in the plurality of online learning algorithms, performing online learning of a plurality of dimensions on feature data based on a plurality of perspectives by the online learning algorithm to obtain a plurality of corresponding models; learning online real-time backflow data based on an attention mechanism to obtain a model fusion strategy corresponding to the plurality of models, wherein the online real-time backflow data comprises a prediction value of an online prediction model fed back by an application client to an application server in real time and a related business indicator of the application after the prediction value is used in the application, and the model fusion strategy comprises a model fusion manner; performing real-time dynamic fusion of the plurality of models according to the model fusion strategy to generate an online prediction model, wherein the input of the online prediction model comprises at least one of user information of a user, feature information obtained by performing feature extraction on the user information, related information of an information item, and feature information obtained by performing feature extraction on the related information, and the output of the online prediction model comprises at least one of a user click rate of the information item, a predicted average reading time of the user for the information item, a prediction of whether the user will click the information item, and a prediction of the time length for which the user reads the information item, and the information item comprises a recommended book; wherein the real-time dynamic fusion of the plurality of models according to the model fusion strategy comprises: performing real-time dynamic fusion of the plurality of models at a model level using a model fusion manner in the model fusion strategy, wherein the fusion at the model level comprises stacking and designing of the plurality of models; wherein the model fusion manner comprises any one of the following: a bootstrap aggregation manner; a boosting manner; a stacking manner; a combination of at least two of the above three manners; wherein the learning of online real-time backflow data based on an attention mechanism comprises any one of the following: learning online real-time backflow data based on a deep neural network using an attention mechanism; learning online real-time backflow data based on a Bayesian network to construct a probabilistic graphical model.
2. The method of claim 1, wherein, The online learning algorithm comprises: an ftrl algorithm; an fm algorithm.
3. The method of claim 1, wherein, The online learning of a plurality of dimensions on the feature data based on a plurality of perspectives by the online learning algorithm comprises: performing learning on a plurality of tasks in parallel by the online learning algorithm; for each task, performing online learning of a plurality of dimensions on the feature data based on a plurality of perspectives.
4. The method of claim 3, wherein, The performing learning on a plurality of tasks in parallel by the online learning algorithm further comprises: adjusting learning weights of the plurality of tasks according to at least one model learning objective.
5. The method of claim 4, wherein, The plurality of tasks comprises: user click rate estimation; average reading time estimation.
6. The method of claim 5, wherein, The at least one model learning objective comprises at least one of the following: improving user click rate; improving average reading time of the user; improving user likes; improving user collections.
7. The method of claim 1, wherein, The model fusion strategy comprises a weighting weight coefficient corresponding to each model.
8. The method of claim 1, wherein, The model fusion strategy comprises a model fusion manner.
9. The method of claim 1, wherein, The real-time dynamic fusion of the plurality of models according to the model fusion strategy comprises: According to a weighting weight coefficient corresponding to each model in the model fusion strategy, the multiple models are dynamically fused in real time in a model weighting manner.
10. The method of claim 9, wherein, The dynamically fusing the multiple models in real time in a model weighting manner includes any one of the following: The multiple models are dynamically fused in real time in a model weighting manner at a feature level. The multiple models are dynamically fused in real time in a model weighting manner at a model result level. 11.A computer device for generating an online prediction model, comprising a memory, a processor and a computer program stored on the memory, wherein, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.
12. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the steps of the method according to any one of claims 1 to 10.
13. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Fusion sorting model training method and device, search sorting method and device and equipment
CN112507196A
Fusion method of machine learning model based on attention mechanism
CN112633396A