Predictive model management
By segmenting the training process and updating the parameter step size of the prediction model based on the offset, the problem that the existing prediction model cannot adapt to data changes in time is solved, and the adaptability and accuracy of the model are improved.
Patent Information
- Application Number
- CN202480005457.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-07
- Filing Date
- 2024-06-18
- Publication Date
- 2025-07-22
AI Technical Summary
Existing prediction models are difficult to learn knowledge from recent data effectively during training, resulting in the model being unable to accurately reflect the current data distribution, especially when objects change rapidly.
By dividing the training process into multiple time periods, the offset is determined within each time period, and the parameter step size of the prediction model is periodically updated based on the offset and gradient information, emphasizing the importance of recent sample data and ignoring the influence of historical data.
The predictive model can learn knowledge from recent sample data faster and more efficiently, improving the accuracy and adaptability of the model to current trends.
Smart Images

Figure CN120359524A_ABST
Abstract
Description
Cross - reference
[0001] This application claims the benefit of U.S. Patent Application No. 18 / 219,158, titled "PREDICTION MODEL MANAGEMENT", filed on July 7, 2023, which is hereby incorporated by reference in its entirety. Technical Field
[0002] The present disclosure generally relates to prediction model management, and more particularly, to methods, devices, and computer program products for prediction model management based on periodic reset during a training process. Background Art
[0003] Nowadays, machine learning techniques have been widely applied to data processing. For example, in a recommendation environment, objects such as articles, advertisements, messages, audio, video, games, etc. can be provided to users. Then, users can subscribe to channels that provide articles, purchase products recommended in advertisements, etc. At this time, events (such as subscription events, purchase events, etc.) between users and corresponding objects can be detected. Solutions for training a prediction model using sample data associated with users, objects, and events have been proposed, and then this prediction model can be used to output the future event trend between users and objects. However, the prediction model is gradually trained by historical data covering a relatively long duration and cannot accurately reflect the most recent data distribution in the training data. At this time, how to enable the prediction model to learn knowledge from the most recent data has become a hot topic of concern. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for managing a prediction model is provided. In this method, gradient information associated with the prediction model is obtained based on sample data of time slots within a predetermined time period. The offset of the time slot within the predetermined time period is obtained. Based on the gradient information, the offset, and historical gradient information determined based on historical sample data of a historical time slot group before the time slot, a step size for updating the parameters of the prediction model is determined.
[0005] In a second aspect of the present disclosure, an electronic device is provided. The electronic device includes: a computer processor coupled to a computer - readable memory unit, the memory unit including instructions that, when executed by the computer processor, implement the method according to the first aspect of the present disclosure.
[0006] In a third aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer - readable storage medium having program instructions embodied thereon, the program instructions being executable by an electronic device to cause the electronic device to execute the method according to the first aspect of the present disclosure.
[0007] The present invention content is provided to introduce a selection of concepts in a simplified form, which are further described in the following detailed implementation. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other objects, features, and advantages of the present disclosure will become more apparent from the more detailed description of some embodiments of the present disclosure in the accompanying drawings, where the same reference numerals generally refer to the same components in the embodiments of the present disclosure.
[0009] Figure 1 An example environment for predictive model management according to machine learning techniques is shown;
[0010] Figure 2 An example graph of sample data for various time slots within a predetermined time period according to an embodiment of the present disclosure is shown;
[0011] Figure 3 An example graph for managing a predictive model according to an embodiment of the present disclosure is shown;
[0012] Figure 4 An example graph of sample data according to an embodiment of the present disclosure is shown;
[0013] Figure 5 An example graph for determining a step size of a parameter for updating a predictive model according to an embodiment of the present disclosure is shown;
[0014] Figure 6 An example flowchart of a method for updating a predictive model according to an embodiment of the present disclosure is shown;
[0015] Figure 7 An example flowchart of a method for managing a predictive model according to an embodiment of the present disclosure is shown; and
[0016] Figure 8 A block diagram of a computing device in which various embodiments of the present disclosure can be implemented is shown. DETAILED DESCRIPTION
[0017] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and helps those skilled in the art to understand and implement the present disclosure, without implying any limitation to the scope of the present disclosure. The disclosure described herein can be implemented in various ways other than the ways described below.
[0018] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0019] References in this disclosure to "one embodiment", "an embodiment", "example embodiment", etc., indicate that the described embodiment may include a particular feature, structure, or characteristic, but not necessarily every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is considered within the knowledge of one of ordinary skill in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0020] It should be understood that although the terms "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be termed a second element, and similarly, a second element may be termed a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0021] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms as well. It will also be understood that when used herein, the terms "comprises", "comprising", "has", "having", "includes", and / or "including" specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0022] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and helps one of ordinary skill in the art to understand and implement this disclosure, without implying any limitation on the scope of this disclosure. The disclosure described herein can be implemented in various ways other than those described below. In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0023] It can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations, and related rules.
[0024] It can be understood that before using the technical solutions disclosed in various embodiments of the present disclosure, the types, usage scopes, and usage scenarios of personal information involved in the present disclosure should be notified to users in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0025] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly notify the user that the requested operation will require obtaining and using the user's personal information. Therefore, the user can independently choose whether to provide personal information to the software or hardware (such as an electronic device, an application, a server, or a storage medium) that performs the operations of the technical solutions of the present disclosure based on the prompt message.
[0026] As an optional but non-limiting embodiment, the manner of sending a prompt message to the user in response to receiving an active request from the user may include, for example, a pop-up window, and the prompt message may be presented in text form in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It can be understood that the above process of notifying and obtaining user authorization is only an example and does not limit the embodiments of the present disclosure. Other methods that comply with relevant laws and regulations are also applicable to the embodiments of the present disclosure.
[0028] For the purpose of description, the following paragraphs will provide more details using a recommendation system as an example environment. In a recommendation system, various objects (such as articles, advertisements, messages, audio, video, games, etc.) can be sent to users. Sometimes, a user is interested in an object and then performs a subscription or places an order. If the user is not interested in the object, he / she can skip the object without doing anything. So far, solutions for generating a prediction model for predicting future events related to users and objects have been provided. In the following, reference will be made to Figure 1 to obtain more details about the prediction model, where Figure 1 an example environment for prediction model management according to machine learning techniques is shown.
[0029] In Figure 1 it, a model 130 can be provided for event prediction. Here, the environment 100 includes a training system 150 and an application system 152. Figure 1The upper part shows the training process, while the lower part shows the application process. Before the training process, the model 130 can be configured with untrained or partially trained parameters (such as initial parameters or pre-trained parameters). During the training process, the model 130 can be trained in the training system 150 based on a training data set 110 that includes a plurality of training samples 112. Here, each training sample 112 can have a binary tuple format and can include data 120 (e.g., data associated with a user and an object) and a label 122 of an event. Specifically, a large number of training samples 112 can be used to iteratively implement the training process. After the training process, the parameters of the model 130 can be updated and optimized, and a model 130' with trained parameters can be obtained. At this time, the model 130' can be used to implement a prediction task during the application process. For example, the data 140 to be processed can be input into the application system 152, and then the corresponding prediction 144 can be output.
[0030] In Figure 1 , the model training system 150 and the model application system 152 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device can refer to any type of mobile device, fixed terminal, or portable device, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablets, media computers, multimedia tablets, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server can include, but is not limited to, mainframes, edge computing nodes, computing devices in a cloud environment, etc. It should be understood that Figure 1 the components and arrangements in the environment 100 are only examples, and the computing systems suitable for implementing the example embodiments described in the present disclosure can include one or more different components and other components. For example, the training system 150 and the application system 152 can be integrated in the same system or device.
[0031] As Figure 1 shown, the training data set 110 can include historical training samples 112 collected from the data log of the recommendation system according to the requirements of corresponding laws, regulations, and related rules. However, the performance of the model 130' is closely related to whether the features (such as embeddings) of the training samples 112 can correctly reflect all aspects of the training samples 112.
[0032] A variety of solutions have been proposed for training a model 130 using a historical training data set 110. For example, a loss can be determined between a label 122 and a prediction of the label 122 determined from the model 130 and the data 120. Additionally, gradient information can be determined based on the loss, and then the parameters of the model 130 can be updated based on the gradient information. Over time, the model gradually acquires knowledge about events over a long duration. However, in a recommendation system, typically the objects change rapidly. For example, the objects can be advertisements related to various camera products, and the product development cycle is short and new camera products are developed quickly. At this time, knowledge about old and outdated camera products is no longer important compared to knowledge about recent camera products. At this time, it is desirable to learn more knowledge about recent camera products in a faster and more efficient manner.
[0033] In view of the above, the present disclosure proposes a prediction model management solution based on periodic resets during the training process. Specifically, the entire duration of the training process can be divided into multiple time periods (e.g., the time period can have a length of one day, one week, etc.). Reference will be made Figure 2 to obtain the sample data collected during the training process, where Figure 2 FIG. 200 shows an example graph of sample data for various time slots in a predetermined time period according to an embodiment of the present disclosure. As Figure 2 shown, the entire training duration can be divided into time period 210, time period 212, etc., and multiple sample data 220, 222,..., 224,..., and 226 can be collected within time period 210.
[0034] Time period 210 can have a predetermined duration. For example, time period 210 can include one day, and time period 210 can be divided into multiple time slots. For example, sample data 220 can be collected from time slot T1, sample data 222 can be collected from time slot T2,..., sample data 224 can be collected from time slot T i collected,..., sample data 226 can be collected from time slot T n collected. Here, time period 210 can be divided into, for example, 24 time slots, so each time slot can cover a time length of one hour. Alternatively and / or additionally, the time period and / or the time slot can have a time length different from the above example.
[0035] In the context of the present disclosure, the sample data 220, 222,..., 224,..., and 226 can be used to train a prediction model 250. A corresponding offset can be determined for each time slot and then considered during the training process. For example, regarding time slot T i, the offset 212 can be determined and then used when determining the step size for updating the parameter(s) of the prediction model 250. In an embodiment of the present disclosure, more or less sample data can be collected within each time slot. It should be understood that Figure 2 only an example of the sample data within the time period 210 is shown, and the time period 212 can have a similar structure, and more or less sample data can be collected within the time period 212.
[0036] Hereinafter, reference is made to Figure 3 for more details on prediction model management, where Figure 3 FIG. 300 shows an example for managing the prediction model 250 according to an embodiment of the present disclosure. As Figure 3 shown, each duration can be equally divided into time slots 312,..., 314, 316,..., and 318, and the gradient information associated with the prediction model 250 can be obtained based on the sample data of each time slot in a predetermined time period. For example, the gradient information 322 can be obtained for the time slot 316. Then, the offset 330 of the time slot 316 in the predetermined time period can be obtained. In addition, based on the gradient information 322, the offset 330, and the historical gradient information 320 determined based on the historical sample data of the historical time slot group before this time slot, the step size 340 for updating the parameter(s) of the prediction model 250 can be determined. As Figure 3 shown, the historical gradient information 320 for the time slot 316 can be determined based on the historical sample data for the time slots 312,..., and 314. In addition, the step size 340 can be used to update the parameter(s) of the prediction model 250.
[0037] Through the embodiments of the present disclosure, the offset for each time slot can be considered when determining the step size, so the step size can be periodically reset in each predetermined time period. In this way, the offset can be used to control the importance of the historical gradient information 320 and the gradient information 322. For example, at the beginning of each time period (i.e., when the current time slot is the first in the time period), the historical gradient information can be ignored, and only the gradient information for the current time slot is considered when determining the step size 340. In this way, more knowledge can be learned from the most recent sample data in a faster and more efficient manner.
[0038] After providing a general description of the solution, the following paragraphs will provide more details on prediction model management. In an embodiment of the present disclosure, the corresponding gradient information can be determined from each sample data. Reference is made to Figure 4 for details on the sample data, where Figure 4 FIG. 400 shows an example of the sample data according to an embodiment of the present disclosure. As Figure 4As shown, the sample data 220 can include two parts: the data part 410 can represent features related to the user 412 and the object 414; and the label part 420 can represent the event 422 between the user 412 and the object 414.
[0039] It should be understood that all information about users, data, and events does not include any sensitive information. For example, all information can be collected according to the requirements of the corresponding laws, regulations, and relevant rules, and then can be converted into an invisible format (such as embedding) for protection purposes. In the context of the present disclosure, the event 422 can include various types, for example, click events and / or conversion events.
[0040] In the context of the present disclosure, the object 414 (such as an advertisement, message, audio, video, game, etc.) can be provided to the user 412. Then the user 412 can click and open the object 414, at which time, a click event is detected. In another example, the conversion event here can indicate that the user behavior is converted towards a deeper interaction with the recommendation system. Generally, the conversion rate (CVR) is a key factor in measuring whether the object 414 attracts the user's attention. Here, the conversion event can include a subscription event, an order placement event, a download event, an add-to-cart event, a follow event, or a comment event, and the conversion event occurs after the click event. In an online shopping environment, the conversion event can include an order placement event, an add-to-cart event, etc. In a multimedia service environment, the conversion event can include a download event, etc. Through these implementations, the sample data can include rich information about the events between the user and the object. Therefore, the prediction model 250 can learn rich knowledge from the sample data and then provide accurate predictions for future events.
[0041] In an embodiment of the present disclosure, in order to obtain the gradient information, based on the prediction model 250, a prediction for the label part 420 in the sample data 220 can be obtained based on the data part 410 in the sample data 220. Specifically, the data part 410 can be input into the prediction model 250, and then the prediction model 250 can output a prediction for the label part 420 based on the (multiple) current parameters of the prediction model 250. The loss between the prediction for the label part 420 and the label part 420 can be determined based on a predetermined loss function. At this time, the gradient information can be obtained based on the gradient of the loss and the parameters of the prediction model 250. Specifically, the gradient information g for time slot t t can be determined based on the following formula 1:
[0042] In the above formula 1, g t represents the gradient information for time slot t, and W tdenote the (multiple) parameters of the prediction model at time slot t, J() denotes the loss function, denote the gradient operation associated with the (multiple) parameters, and α denotes the weight decay coefficient for L2 regularization. It should be understood that Equation 1 above is only an example for determining the gradient information for time slot t. Additionally and / or alternatively, Equation 1 can be modified by considering more or fewer variables in the equation. For example, the weight α can be omitted to simplify the calculation. At this time, the gradient information can be determined in an easy and efficient manner according to the mathematical operations.
[0043] In an embodiment of the present disclosure, the predetermined time period can have a fixed length of one day or more. For example, the time period can be set to one day. In another example, the time period can have a different time length such as two days, or other values. Additionally, the time period can be determined based on the change frequency of the object. In a recommendation environment, if the object indicates an advertisement for a camera product, the length of the time period can be decreased as the change frequency of the camera product increases. In other words, if a new camera product is developed over a relatively long duration (i.e., a lower change frequency), the length of the time period can be set to a relatively long length; and if a new camera product is developed over a relatively short duration (i.e., a higher change frequency), the length of the time period can be set to a relatively short length. Through these embodiments, the time period can be determined in a flexible and dynamic manner, which can help to learn more knowledge about recent sample data.
[0044] In an embodiment of the present disclosure, an offset can be obtained for each time slot. Here, the offset can represent the difference between the time slot and the first time slot in the time period. For example, the offset can be measured by the difference between the sequence numbers associated with the two time slots. Alternatively and / or additionally, the offset can be measured by the time difference between the time points of the two time slots. Then, the step size can be determined based on the gradient information, the offset, and the historical gradient information associated with the group of historical time slots before the time slot.
[0045] Here, the offset can reflect whether the current sample data is recent sample data. If the offset is equal to zero, it indicates that the current sample data was collected at the beginning of the time period. At this time, the influence of the sample data can be enhanced when updating the prediction model 250 (e.g., only considering the current sample data and excluding the historical sample data when determining the step size). If the offset is not equal to zero, it indicates that the sample was not collected at the beginning of the time period, so both the current sample data and the historical sample data are considered when determining the step size.
[0046] Figure 5 FIG. 500 shows an example diagram for determining the step size for updating the parameters of a prediction model according to an embodiment of the present disclosure. As Figure 5As shown, to determine the step size 340, the weight 510 for the historical gradient information 320 can be determined based on the offset 330. For example, the weight 510 can be determined within a predefined region and increase with the offset 330, and then the step size can be generated based on the gradient information, the historical gradient information, and the weight for the historical gradient information. Specifically, the step size can be determined based on the following formulas 2.1 and 2.2: Δ t = f(β·v t-1 , g t ) Formula 2.1 β = f1(o t , area1) Formula 2.2
[0047] In the above Formula 2.1, Δ t represents the step size for updating the parameters of the prediction model during time slot t, v t-1 represents the historical gradient information for time slot t, which is associated with a group of time slots before time slot t, g t represents the gradient information for time slot t, β represents the weight for the historical gradient information, and f() represents a function associated with the gradient information g t , the historical gradient information v t-1 and the weight β for the historical gradient information.
[0048] In addition, the weight β can be determined within a region (denoted as area1 such as [0, 1]) based on the offset (denoted as o t ). f1 can represent a function associated with the offset o t and the predefined region area1. For example, if the offset indicates that time slot t is the first in a predefined time period, then β can be set to 0 (i.e., the lower limit of the region, or a relatively small value in the region, such as 0.01 or other values). If the offset indicates that time slot t is not the first in a predefined time period, then β can be set to 1 (i.e., the upper limit of the region, or a relatively large value in the region, such as 0.99 or other values). In another example, β can increase with the offset within the region, for example, in proportion to the offset.
[0049] It should be understood that the above paragraphs only provide example values for the weight. Alternatively and / or additionally, if the time slot is not the first in the time period, then β can be set to 0.5 or other values. Through these embodiments, the weight can be determined in an easy and effective manner, which can help learn more knowledge about the recent sample data at the beginning of the time period.
[0050] In an embodiment of the present disclosure, the historical gradient information can be related to the cumulative gradient information for historical time slots before the time slot. For example, v t-1It can be determined based on the historical gradient information for the corresponding historical time slots [1, t - 1] before the time slot t, and the corresponding gradient information for the corresponding historical time slots [1, t - 1] can be determined based on the above formula 1. To determine the historical gradient information, the corresponding gradient information can be determined based on the corresponding historical sample data for the historical time slot group before the time slot, and then the historical gradient information can be obtained based on the corresponding historical gradient information.
[0051] Here, the historical sample data for the historical time slots before the time slot is used to determine the historical gradient information. Specifically, the historical sample data for the time slot t is determined from the historical sample data for the corresponding historical time slots [1, t - 1]. If the time slot t is the first time slot in the time period, there is no historical time slot before the time slot t; and if the time slot t is the second time slot or a subsequent time slot in the time period, the historical sample data for the corresponding historical time slots [1, t - 1] can be used to determine the historical gradient information for the time slot t.
[0052] In an embodiment of the present disclosure, the historical gradient information can be determined in an iterative manner. Specifically, to obtain the historical gradient information, the corresponding squares of the corresponding gradient information associated with the corresponding historical time slots in the historical time slot group can be determined, and then the historical gradient information can be determined based on the sum of the corresponding squares. For example, the historical gradient information for the time slot t + 1 can be determined based on the following formula 3.1:
[0053] In this formula, represents the square of the gradient information g t (determined as in formula 1), and v t-1 represents the historical gradient information for the time slot t. Here, β has the same meaning as in formula 2, where β is determined based on whether the time slot t is the first in the time period. In one example, when the time slot t is the first, β = 0 (or a relatively small value); and when the time slot t is not the first, β = 1 (or a relatively large value). Similarly, the historical gradient information for the time slot t can be determined based on the following formula 3.2:
[0054] At this time, the historical gradient information can be determined in an iterative manner, and thus the step size determined as in formula 2.1 is also determined in an iterative manner. Through these embodiments, the determination of the step size is converted into a mathematical operation, and thus the step size can be determined in a simple and effective manner.
[0055] It should be understood that the (multiple) historical time slot groups can be within the current time period. In other words, during the current time period, the sample data related to the historical time slots in the previous time period before the current time period is excluded from the determination of the step size. Therefore, the importance of the most recent data within the current time period is emphasized when updating the prediction model, so that the prediction model can accurately reflect the trend of the changing situation in the recommendation system.
[0056] In an embodiment of the present disclosure, the function f() in Formula 2.1 can be defined in various ways. For example, Formula 2.1 can be refined into the following Formula 4:
[0057] In this formula, the intermediate parameter v can be determined based on the gradient information and the weighted historical gradient information determined based on the historical gradient information and the weights. t (which is associated with time slot t). Specifically, v t can be determined iteratively according to Formula 3.1, and then the step size can be determined based on the intermediate parameter v t and the gradient information g t In other words, the intermediate parameter v for the current time slot t t can be determined based on the historical gradient information v for the current time slot t t-1 Then, the intermediate parameter v for the current time slot t t can be used as the historical gradient information for the next time slot t + 1. η represents the learning rate of the prediction model. Through these embodiments, the intermediate parameter of the previous time slot t - 1 can be reused in time slot t, so the computational complexity for determining the step size can be reduced.
[0058] In an embodiment of the present disclosure, the intermediate parameter v can be determined according to Formula 3.1 t , so Formula 4 can be transformed into Formula 5:
[0059] In this formula, Δ t represents the step size for updating the parameters of the prediction model during time slot t. v t-1 represents the historical gradient information associated with the (multiple) time slot groups before time slot t, and here v t-1 can be determined iteratively based on the (multiple) corresponding sample data in the (multiple) corresponding time slots in this time period. g trepresents the gradient information for the time slot t, and it can be determined based on Equation 1. ∈ represents a constant epsilon value used to ensure that the denominator part in Equation 5 will not be zero. α represents the weight for the historical gradient information, and it can be determined based on the offset of the time slot t in the time period (i.e., whether the time slot t is the first in the time period). Through these embodiments, the step size Δ for updating the prediction model t can be controlled by the offset of the time slot. Therefore, the prediction model can be updated in a direction that provides more accurate predictions in newly occurring situations.
[0060] In an embodiment of the present disclosure, Equation 5 can be modified in various ways. For example, the symbol ∈ can have a fixed value selected from a region such as [10 -6 , 1] (or other regions such as [10 -5 , 1] or [10 -5 , 0.5], etc.). In another example, the symbol ∈ can be a dynamic decay factor based on the offset as an intermediate parameter. Specifically, ∈ can be determined based on the following Equation 6: ∈ = f2(o t , area2) Equation 6
[0061] In this equation, ∈ represents a decay factor selected from a predefined region (denoted as area2) based on the offset o of the time slot t t , and f2() represents a function associated with o t and area2. At this time, the step size can be determined based on the gradient information and the decay intermediate parameter determined based on the intermediate parameter and the decay factor. Equation 5 can be converted to Equation 7 based on Equation 6, and all symbols have the same meanings as in the previous equations:
[0062] Through these embodiments, the step size for updating the prediction model depends on the offset o t . At this time, the importance of historical sample data can be reduced during the training process within the time slot t, so the importance of recent sample data can be increased during the training process within the time slot t. Therefore, the step size can be determined in a direction that better matches the recent sample data.
[0063] In an embodiment of the present disclosure, once the step size is determined, the step size can be used to update the parameter(s) of the prediction model. Specifically, the prediction model can be updated according to Equation 8: W t,1 = W t + Δ t Equation 8
[0064] In this equation, W tdenote(s) the parameter(s) of the prediction model for time slot t, and Δ t denotes the step size for updating the parameter(s) of the prediction model during time slot t, W t,1 denote(s) the updated parameter(s) of the prediction model (such as the weight(s) of a machine learning model, and these weights can be used as the parameter(s) for time slot t+1). Through these embodiments, the parameter(s) of the prediction model can be updated by the most recent sample data in a direction to better match the newly occurring situation, so that the updated prediction model can work well in the newly occurring situation.
[0065] The previous paragraphs have provided details of the individual steps in prediction model management. In the following, reference Figure 6 is made to obtain a training process by considering the offset in updating the prediction model. Figure 6 FIG. shows an example flowchart of a method 600 for updating a prediction model according to an embodiment of the present disclosure. As Figure 6 shown, at block 610, the prediction model can be started. For example, the parameter(s) of the prediction model can be set to initial values. Alternatively and / or additionally, the parameter(s) can be set to partially optimized values. At block 620, new sample data can be retrieved from a data source 622 for storing newly collected sample data. If new sample data arrives, method 600 can proceed to block 630 for training the prediction model; otherwise, if no new sample data arrives, method 600 can wait until new sample data arrives.
[0066] Multiple steps can be implemented at block 630. First, an offset related to the sample data can be determined at block 632, and then a corresponding branch can be selected based on the determined offset. If the determined offset is equal to zero (i.e., the sample data is related to the first time slot in the time period), method 600 can proceed to block 636, and thus the step size can be determined based on an intermediate parameter v t which depends on a weight β, historical gradient information v t-1 and gradient information g t . Here, the weight β is determined by the offset, and the weight β can be set to zero or a relatively small value close to zero. If the offset is not equal to zero (i.e., the sample data is not related to the first time slot in the time period), method 600 can proceed to block 634. At this time, the step size can be determined based on an intermediate parameter v t which depends on both historical gradient information v t-1 and gradient information g t . In this way, in each time period, the influence of historical sample data is reduced by the weight β at the beginning of the time period, so that the step size can take into account more influence of the most recent sample data within the time period.
[0067] In addition, based on backpropagation, the (multiple) parameters of the prediction model can be updated based on the determined step size. At block 640, if the training is completed (e.g., the training process reaches a predetermined convergence condition), the training process can end. If the training is not completed, method 600 can repeat the steps in block 630 until the predetermined convergence condition is met. Although the previous paragraphs have provided details on determining the step size through individual sample data associated with each time slot, each time slot can involve a group of sample data, which can then be used as a batch for training the prediction model. Here, the prediction model can be effectively updated in multiple batches in an iterative manner in an optimized direction.
[0068] Once the training process ends, new data (e.g., only including the data portion as Figure 4 shown) can be input into the prediction model, so that the prediction model can output a prediction for the new data. For example, if the prediction model is trained to learn whether a user will purchase a camera product based on historical samples, the prediction model will output the probability that a specific user will purchase a specific camera.
[0069] Through the embodiments of the present disclosure, the second-order momentum of the prediction model can be reset in a periodic pattern, which shapes the effective learning rate of the optimizer into a cosine annealing restart type. In addition, past local optima can be discarded with the reset. Thereafter, the convergence and generalization of the prediction model can better fit the recent data distribution, especially the data snapshot dumped yesterday. At the same time, even if the daily data snapshots are not fully shuffled, since the effective learning rate decays rapidly to the end of the day, the optimizer will not significantly overfit the last batch of training data (e.g., the training data within 23:00 - 24:00).
[0070] The above paragraphs have described the details of feature management. According to the embodiments of the present disclosure, a method for feature management is provided. More details about the method will be referred to Figure 7 for more details about the method, where Figure 7 FIG. 7 shows an example flowchart of a method 700 for managing a prediction model according to an embodiment of the present disclosure. At block 710, gradient information associated with the prediction model is obtained based on sample data of time slots within a predetermined time period. At block 720, the offset of the time slot within the predetermined time period is obtained. At block 730, a step size for updating the parameters of the prediction model is determined based on the gradient information, the offset, and historical gradient information determined based on historical sample data of a historical time slot group before the time slot.
[0071] In an embodiment of the present disclosure, determining the step size includes: determining a weight for the historical gradient information based on the offset; and generating a step size based on the gradient information, the historical gradient information, and the weight for the historical gradient information.
[0072] In an embodiment of the present disclosure, the weight is within a predefined region and increases with the offset.
[0073] In an embodiment of the present disclosure, generating a step size includes: determining an intermediate parameter associated with a time slot based on gradient information and weighted historical gradient information determined based on historical gradient information and weights; and creating a step size based on the intermediate parameter and the gradient information.
[0074] In an embodiment of the present disclosure, creating a step size includes: obtaining an attenuation factor for the intermediate parameter based on the offset; and determining the step size based on the gradient information and the attenuated intermediate parameter determined based on the intermediate parameter and the attenuation factor.
[0075] In an embodiment of the present disclosure, obtaining gradient information includes: obtaining a prediction for a label part in sample data based on a data part in the sample data and a prediction model; determining a loss between the prediction for the label part and the label part; and obtaining gradient information based on the gradient of the loss and the parameters of the prediction model.
[0076] In an embodiment of the present disclosure, the data part represents features associated with a user and an object, the label part represents an event between the user and the object, and the predefined time period has a length of one or more days.
[0077] In an embodiment of the present disclosure, the method further includes determining historical gradient information by: obtaining corresponding gradient information based on corresponding historical sample data for a historical time slot group before the time slot; and obtaining historical gradient information based on the obtained corresponding gradient information.
[0078] In an embodiment of the present disclosure, obtaining historical gradient information includes: determining corresponding squares of the corresponding gradient information associated with the corresponding historical time slots in the historical time slot group, the historical time slot group being within a predefined time period; and determining historical gradient information based on the sum of the corresponding squares.
[0079] In an embodiment of the present disclosure, the method further includes: updating the parameters of the prediction model using the step size.
[0080] According to an embodiment of the present disclosure, an apparatus for predicting model management is provided. The apparatus includes: an acquisition unit configured to obtain gradient information associated with a prediction model based on sample data of time slots within a predetermined time period; an acquisition unit configured to obtain an offset of the time slot within the predetermined time period; and a determination unit configured to determine a step size for updating parameters of the prediction model based on the gradient information, the offset, and historical gradient information determined based on historical sample data of a historical time slot group before the time slot. In addition, the apparatus may include other units for implementing other steps in method 700.
[0081] According to an embodiment of the present disclosure, an electronic device for implementing method 700 is provided. The electronic device includes: a computer processor coupled to a computer-readable memory unit, the memory unit including instructions that, when executed by the computer processor, implement a method for managing a prediction model. The method includes: obtaining gradient information associated with a prediction model based on sample data of time slots within a predetermined time period; obtaining an offset of the time slot within the predetermined time period; and determining a step size for updating parameters of the prediction model based on the gradient information, the offset, and historical gradient information determined based on historical sample data of a historical time slot group before the time slot.
[0082] In an embodiment of the present disclosure, determining the step size includes: determining a weight for the historical gradient information based on the offset; and generating a step size based on the gradient information, the historical gradient information, and the weight for the historical gradient information.
[0083] In an embodiment of the present disclosure, the weight is within a predefined region and increases with the offset.
[0084] In an embodiment of the present disclosure, generating the step size includes: determining an intermediate parameter associated with the time slot based on the gradient information and a weighted historical gradient information determined based on the historical gradient information and the weight; and creating a step size based on the intermediate parameter and the gradient information.
[0085] In an embodiment of the present disclosure, creating the step size includes: obtaining an attenuation factor for the intermediate parameter based on the offset; and determining the step size based on the gradient information and a decayed intermediate parameter determined based on the intermediate parameter and the attenuation factor.
[0086] In an embodiment of the present disclosure, obtaining the gradient information includes: obtaining a prediction for a label part in the sample data based on a data part in the sample data and the prediction model; determining a loss between the prediction for the label part and the label part; and obtaining the gradient information based on the gradient of the loss and the parameters of the prediction model.
[0087] In an embodiment of the present disclosure, the data portion represents features associated with a user and an object, the label portion represents an event between the user and the object, and the predetermined time period has a length of one or more days.
[0088] In an embodiment of the present disclosure, the method further includes determining historical gradient information by: obtaining corresponding gradient information(s) based on corresponding historical sample data for a historical time slot group before the time slot; and obtaining historical gradient information based on the obtained corresponding gradient information(s).
[0089] In an embodiment of the present disclosure, obtaining historical gradient information includes: determining corresponding squares of the corresponding gradient information(s) associated with the corresponding historical time slots in the historical time slot group within a predefined time period; and determining historical gradient information based on the sum of the corresponding squares.
[0090] In an embodiment of the present disclosure, the method further includes: updating parameters of the prediction model using a step size.
[0091] Figure 8 A block diagram of a computing device 800 in which various embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 8 the illustrated computing device 800 is for illustrative purposes only and does not imply any limitation to the functions and scope of the present disclosure in any way. The computing device 800 may be used to implement the above method 700 in an embodiment of the present disclosure. As Figure 8 shown, the computing device 800 may be a general-purpose computing device. The computing device 800 may include at least one or more processors or processing units 810, a memory 820, a storage unit 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.
[0092] The processing unit 810 may be a physical or virtual processor and may implement various processes based on programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 800. The processing unit 810 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0093] Computing device 800 generally includes various computer storage media. Such media can be any media accessible to computing device 800, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 830 can be any removable or non-removable media and can include machine-readable media, such as memory, flash drive, disk, or other media, which can be used to store information and / or data in computing device 800 and can be accessed.
[0094] Computing device 800 may also include additional removable / non-removable, volatile / non-volatile memory media. Although not shown in Figure 8 it, a disk drive for reading from and / or writing to a removable and non-volatile disk and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In such a case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0095] Communication unit 840 communicates with another computing device via a communication medium. Additionally, the functions of the components in computing device 800 may be implemented by a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, computing device 800 may operate in a networked environment using a logical connection to one or more other servers, networked personal computers (PCs), or other general network nodes.
[0096] Input device 850 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 860 can be one or more of various output devices, such as a display, speaker, printer, etc. With the help of communication unit 840, computing device 800 can also communicate with one or more external devices (not shown) (such as storage devices and display devices), one or more devices that enable a user to interact with computing device 800, or any device that enables computing device 800 to communicate with one or more other computing devices (such as network cards, modems, etc.), if needed. Such communication may be performed via an input / output (I / O) interface (not shown).
[0097] In some embodiments, instead of being integrated in a single device, some or all components of computing device 800 may also be arranged in a cloud computing architecture. In a cloud computing architecture, the components may be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which will not require the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses appropriate protocols to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. The software or components and corresponding data of the cloud computing architecture may be stored on a server at a remote location. The computing resources in a cloud computing environment may be consolidated or distributed at the location of a remote data center. The cloud computing infrastructure may provide services through a shared data center, although they appear as a single access point for the user. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.
[0098] The functions described herein may be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0099] The program code for performing the methods of the subject matter described herein may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely or partially on the machine, partially as a stand-alone software package on the machine, partially on a remote machine, or entirely on a remote machine or server.
[0100] In the context of the present disclosure, a machine-readable medium can be any tangible medium that can contain or store a program used by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0101] Moreover, although the operations are shown in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Also, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular implementations. Certain features described in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features described in a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.
[0102] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the above specific features and acts are disclosed as example forms of implementing the claims.
[0103] From the foregoing, it can be understood that the present disclosure has described specific implementations of the presently disclosed technology for purposes of illustration, but various modifications can be made without departing from the scope of the present disclosure. Accordingly, the presently disclosed technology is not limited except as by the appended claims.
[0104] Embodiments of the subject matter and the functional operations described in this disclosure can be implemented in various systems, in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and structural equivalents thereof) or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances implementing a machine-readable propagated signal, or one or more of them. The term “data processing unit” or “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also include code that creates an execution environment for the computer program being discussed, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them, in addition to the hardware.
[0105] A computer program (also called a program, software, software application, script, or code) can be written in any form of programming language, including a compiled or interpreted language, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers distributed at one site or across multiple sites and interconnected by a communication network.
[0106] Processors suitable for executing computer programs include, by way of example, any one or more processors of general and special purpose microprocessors as well as any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include or be operatively coupled to receive data from or transfer data to one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0107] The specification and drawings are intended to be considered only as exemplary, where exemplary means an example. As used herein, the use of "or" is intended to include "and / or" unless the context clearly dictates otherwise.
[0108] Although the present disclosure contains many details, these details should not be construed as limitations on the scope of any disclosure or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular disclosure. Certain features that are described in the context of separate embodiments in the present disclosure may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be excised from the combination, and the claimed combination may be directed to a sub-combination or a variation of the sub-combination.
[0109] Similarly, although operations are shown in the figures in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve a desired result. Additionally, the separation of various system components in the embodiments described in the present disclosure should not be understood as requiring such separation in all embodiments. Only a few embodiments and examples have been described, and other embodiments, enhancements, and variations may be made based on what is described and illustrated in the present disclosure.
Claims
1. A method for managing a prediction model, comprising: obtaining gradient information associated with the prediction model based on sample data of time slots within a predetermined time period; acquiring an offset of the time slots within the predetermined time period; and determining a step size for updating parameters of the prediction model based on the gradient information, the offset, and historical gradient information, the historical gradient information being determined based on historical sample data of a historical time slot group prior to the time slots.
2. The method according to claim 1, wherein determining the step size comprises: determining a weight for the historical gradient information based on the offset; and generating the step size based on the gradient information, the historical gradient information, and the weight for the historical gradient information.
3. The method according to claim 2, wherein the weight is within a predefined region and increases with the offset.
4. The method according to claim 2, wherein generating the step size comprises: determining an intermediate parameter associated with the time slot based on the gradient information and weighted historical gradient information, the weighted historical gradient information being determined based on the historical gradient information and the weight; and creating the step size based on the intermediate parameter and the gradient information.
5. The method according to claim 4, wherein creating the step size comprises: obtaining an attenuation factor for the intermediate parameter based on the offset; and determining the step size based on the gradient information and the attenuated intermediate parameter, the attenuated intermediate parameter being determined based on the intermediate parameter and the attenuation factor.
6. The method according to claim 1, wherein obtaining the gradient information comprises: obtaining a prediction for a label part in the sample data based on a data part in the sample data and the prediction model; determining a loss between the prediction for the label part and the label part; and obtaining the gradient information based on a gradient of the loss and the parameters of the prediction model.
7. The method according to claim 6, wherein the data part represents features associated with a user and an object, the label part represents an event between the user and the object, and the predetermined time period has a length of one or more days.
8. The method according to claim 1, further comprising determining the historical gradient information by: obtaining corresponding gradient information based on corresponding historical sample data of the historical time slot group prior to the time slots; and obtaining the historical gradient information based on the obtained corresponding gradient information.
9. The method according to claim 8, wherein obtaining the historical gradient information comprises: determining a corresponding square of the corresponding gradient information associated with the corresponding historical time slots in the historical time slot group within the predefined time period; and determining the historical gradient information based on a sum of the corresponding squares.
10. The method according to claim 1 further comprises: Updating the parameters of the prediction model using the step size.
11. An electronic device includes a computer processor coupled to a computer-readable memory unit, the memory unit including instructions that, when executed by the computer processor, implement a method for managing a prediction model, the method including: Obtaining gradient information associated with the prediction model based on sample data of time slots within a predetermined time period; Obtaining an offset of the time slots within the predetermined time period; And Determining a step size for updating parameters of the prediction model based on the gradient information, the offset, and historical gradient information, the historical gradient information being determined based on historical sample data of a historical time slot group prior to the time slots.
12. The device according to claim 11, wherein determining the step size includes: Determining a weight for the historical gradient information based on the offset; And Generating the step size based on the gradient information, the historical gradient information, and the weight for the historical gradient information.
13. The device according to claim 12, wherein the weight is within a predefined region and increases with the offset.
14. The device according to claim 12, wherein generating the step size includes: Determining an intermediate parameter associated with the time slot based on the gradient information and weighted historical gradient information, the weighted historical gradient information being determined based on the historical gradient information and the weight; and Creating the step size based on the intermediate parameter and the gradient information.
15. The device according to claim 14, wherein creating the step size includes: Obtaining an attenuation factor for the intermediate parameter based on the offset; And Determining the step size based on the gradient information and the attenuated intermediate parameter, the attenuated intermediate parameter being determined based on the intermediate parameter and the attenuation factor.
16. The device according to claim 11, wherein obtaining the gradient information includes: Obtaining a prediction for a label portion of the sample data based on a data portion of the sample data and the prediction model; Determining a loss between the prediction for the label portion and the label portion; And Obtaining the gradient information based on a gradient of the loss and the parameters of the prediction model.
17. The apparatus according to claim 16, wherein the data portion represents a feature associated with a user and an object, the tag portion represents an event between the user and the object, and the predetermined time period has a length of one or more days, and the method further comprises: Updating the parameters of the prediction model using the step size.
18. The device according to claim 11, wherein the method further includes determining the historical gradient information by: Obtaining corresponding gradient information based on corresponding historical sample data of the historical time slot group prior to the time slots; and Obtaining the historical gradient information based on the obtained corresponding gradient information.
19. The device according to claim 18, wherein obtaining the historical gradient information includes: Determining a corresponding square of the corresponding gradient information associated with the corresponding historical time slots in the historical time slot group within the predefined time period; And Determining the historical gradient information based on a sum of the corresponding squares.
20. A non-transitory computer program product, the non-transitory computer program product comprising a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by an electronic device to cause the electronic device to perform a method for managing a prediction model, the method comprising: Obtaining gradient information associated with the prediction model based on sample data of time slots within a predetermined time period; Obtaining an offset of the time slots within the predetermined time period; And Determining a step size for updating parameters of the prediction model based on the gradient information, the offset, and historical gradient information, the historical gradient information being determined based on historical sample data of a historical time slot group prior to the time slots.