Object prediction model training method and apparatus, and computing device

By using state lists to manage user behavior data in the training method of object prediction model and building tag data at the preset time, the confusion of real-time user behavior data is solved, and the accuracy and user experience of the model are improved.

CN120013612APending Publication Date: 2025-05-16XINGIN INFORMATION TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510118709.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the prior art builds an object prediction model in real time, the confusion and inconsistency caused by user behavior data from different sources leads to insufficient accuracy of label data, affecting the accuracy of the model and user experience.

Method used

By obtaining user behavior data, the data is recorded in the status list based on the user identification. When the preset trigger time is reached, the user behavior data is taken out from the status list, the corresponding tag data is constructed, and the object prediction model is trained.

Benefits of technology

This method effectively manages real-time streaming user behavior data, reduces data confusion and inconsistency, improves the accuracy of label data, and thus improves the accuracy and user experience of object prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013612A_ABST
    Figure CN120013612A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a training method and device of an object prediction model and computing equipment, and the training method of the object prediction model comprises the steps that at least one piece of user behavior data is acquired, and any piece of user behavior data is used for representing an interaction behavior event of a user for an object; based on the user identifier of the at least one piece of user behavior data, recording the at least one piece of user behavior data in a corresponding state list; under the condition that the preset triggering time of a first state list is reached, the recorded first user behavior data are taken out from the first state list, first label data corresponding to the first user behavior data are constructed, and the first state list is any state list; and training an object prediction model based on the first tag data. The accuracy of the label data is improved, the object prediction model obtained through training has high accuracy, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of data processing, and more particularly to a training method, apparatus, and computing device for an object prediction model. Background Art

[0002] With the development of data processing technology, the user experience and platform interaction rate have been greatly improved by analyzing user behavior data and providing users with objects that meet their needs. For example, e-commerce platforms predict products of interest to users by analyzing users' purchase history and browsing behavior of products. Another example is that note-sharing platforms predict related notes for users by analyzing users' browsing, liking, commenting and sharing behaviors of notes.

[0003] At present, it has become a standard practice for many platforms to construct label data for real-time user behavior data and train object prediction models in real time. However, since real-time user behavior data comes from different sources, it is difficult to accurately and effectively construct label data, and it cannot truly meet user perception, resulting in insufficient accuracy of the label data constructed in real time, and insufficient accuracy of the object prediction model trained based on it, which reduces the user experience. Summary of the invention

[0004] In view of this, an embodiment of this specification provides a method for training an object prediction model. One or more embodiments of this specification also relate to a training device for an object prediction model, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a method for training an object prediction model is provided, comprising:

[0006] Acquire at least one user behavior data, wherein any user behavior data is used to characterize an interactive behavior event of a user with respect to an object;

[0007] Based on a user identifier of the at least one user behavior data, recording the at least one user behavior data in a corresponding status list;

[0008] When a preset trigger time of the first state list is reached, taking out the recorded first user behavior data from the first state list, and constructing first label data corresponding to the first user behavior data, wherein the first state list is any state list;

[0009] Based on the first label data, an object prediction model is trained.

[0010] According to a second aspect of an embodiment of this specification, a training device for an object prediction model is provided, comprising:

[0011] An acquisition module is configured to acquire at least one user behavior data, wherein any user behavior data is used to characterize an interactive behavior event of a user with respect to an object;

[0012] a recording module configured to record at least one user behavior data in a corresponding state list based on a user identifier of the at least one user behavior data;

[0013] a construction module configured to, when a preset trigger time of a first state list is reached, take out the recorded first user behavior data from the first state list and construct first tag data corresponding to the first user behavior data, wherein the first state list is any state list;

[0014] The training module is configured to train the object prediction model based on the first label data.

[0015] According to a third aspect of an embodiment of this specification, a computing device is provided, including:

[0016] Memory and processor;

[0017] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the training method of the above-mentioned object prediction model are implemented.

[0018] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the training method of the above-mentioned object prediction model are implemented.

[0019] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the steps of the training method of the above-mentioned object prediction model.

[0020] In one embodiment of the present specification, at least one user behavior data is obtained, wherein any user behavior data is used to characterize the user's interactive behavior event with respect to an object; based on the user identifier of at least one user behavior data, at least one user behavior data is recorded in a corresponding state list; when the trigger time of the first state list is reached, the recorded first user behavior data is taken out from the first state list, and the first label data corresponding to the first user behavior data is constructed, wherein the first state list is any state list; based on the first label data, the object prediction model is trained. By adopting the state list constructed according to the user identifier, the orderly and efficient management of the real-time streaming user behavior data is completed, the confusion and inconsistency caused by the user behavior data from different sources is reduced, the label data is constructed accurately and effectively, and it is ensured that the label data of the constructed user behavior data truly conforms to the user perception, and the accuracy of the label data is improved. The object prediction model obtained by training has high accuracy, and the user experience is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart of a method for training an object prediction model provided by an embodiment of this specification;

[0022] Figure 2 It is a data flow diagram of a training method of an object prediction model provided by an embodiment of this specification;

[0023] Figure 3 It is a schematic diagram of an interactive behavior event in a training method of an object prediction model provided by an embodiment of this specification;

[0024] Figure 4 It is a schematic diagram of triggering time in a training method of an object prediction model provided by an embodiment of this specification;

[0025] Figure 5 It is a flowchart of order preservation processing in a training method of an object prediction model provided by an embodiment of this specification;

[0026] Figure 6 It is a flowchart of deduplication processing in a training method of an object prediction model provided by an embodiment of this specification;

[0027] Figure 7 It is a flowchart of summary processing in a training method of an object prediction model provided by an embodiment of this specification;

[0028] Figure 8 It is a flowchart of distribution processing in a training method of an object prediction model provided by an embodiment of this specification;

[0029] Fig. 9is a processing flow chart of a training method for an object prediction model applied to a note sharing platform provided by an embodiment of the present specification;

[0030] Fig.10 It is a structural schematic diagram of a training device for an object prediction model provided by an embodiment of this specification;

[0031] Fig.11 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0032] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0033] The terms used in one or more embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present invention. The singular forms of "a", "said" and "the" used in one or more embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of the present invention refers to and includes any or all possible combinations of one or more associated listed items.

[0034] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present invention, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0035] In addition, it should be noted that the data involved in one or more embodiments of the present invention are information and data authorized by the user or fully authorized by all parties, and the statistics, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0036] First, the terms involved in one or more embodiments of this specification are explained.

[0037] Session Labeler: A framework for building labeled data that is used to extract and build labeled data from user behavior data to facilitate subsequent object prediction model training. Session Labeler generates labeled data that reflects user interests and preferences by processing user behavior events within a specific time period.

[0038] Interaction behavior: various operations performed by users on objects when using applications or platforms, including but not limited to browsing, clicking, liking, sharing, following, commenting, and blocking. These behavior data reflect the interests and preferences of users and are an important basis for building user portraits and recommendation systems.

[0039] Priority queue: A special data structure in which each element has a priority and elements with higher priority are processed first.

[0040] Ordinary queue: A first-in-first-out (FIFO) data structure used to store and manage elements.

[0041] Sliding window: A data processing mechanism that uses a sliding time window according to certain rules to collect and process user behavior data within a specific time period. Sliding windows can help the system process and analyze user behavior in real time, ensuring the timeliness and accuracy of data.

[0042] Timer: A timed trigger mechanism that triggers a piece of logic or task at a specific time point. It is used to monitor the arrival time and processing time of user behavior data to ensure that data processing logic is triggered at the appropriate time point, such as the construction of tag data and the update of status information.

[0043] Flink: A distributed real-time processing engine for processing large-scale stream and batch data. Flink provides low-latency, high-performance data processing capabilities, suitable for real-time data analysis and stream processing scenarios. In this embodiment, Flink is used to process user behavior data in real time to ensure the real-time and accuracy of the data.

[0044] Kafka: A distributed message queue system used to deliver messages between different applications. Kafka has the characteristics of high throughput, scalability, and persistence, and is suitable for big data processing and real-time data stream transmission. In this embodiment, Kafka is used to transmit user behavior data and tag data between different components to ensure reliable data transmission and processing.

[0045] RocksDB: A stand-alone key-value (KV) database supported by Flink.

[0046] In this specification, a training method for an object prediction model is provided. This specification also relates to a training device for an object prediction model, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0047] See also Figure 1 , Figure 1 A flowchart of a method for training an object prediction model provided by an embodiment of the present specification is shown, and includes the following specific steps:

[0048] Step 102: Acquire at least one user behavior data, wherein any user behavior data is used to characterize an interactive behavior event of a user with respect to an object.

[0049] This specification applies to application platforms where users can interact with objects, such as e-commerce platforms, social media platforms, note sharing platforms, etc., and is not limited here.

[0050] User behavior data is data generated by users interacting with objects, which characterizes the interactive behavior events of the users with respect to the objects. User behavior data is generated in real time. User behavior data is the basis for building user portraits and recommendation systems. By analyzing these data, we can understand users' interests, preferences, and behavior patterns, thereby providing more relevant object predictions. For example, for e-commerce platforms, user behavior data includes but is not limited to: users browsing products, searching for keywords, adding to shopping carts, purchasing products, etc. For another example, for social media platforms, user behavior data includes but is not limited to: users posting updates, liking, commenting, sharing, following other users, etc. For another example, for note sharing platforms, user behavior data includes but is not limited to: users browsing notes, liking, commenting, collecting, sharing notes, etc.

[0051] Objects are objects of user interaction in the application platform, including but not limited to: products, articles, videos, pictures, and notes.

[0052] Interaction events are the interactions between users and objects on the application platform. These operations generate user behavior data. For example, browsing: users view a page or object. Clicking: users click on a link, button or object. Like: users like or agree with an object. Sharing: users share an object with other users or social networks. Following: users follow a user, brand or topic. Commenting: users comment on or give feedback on an object. Blocking: users set an object or user as uninterested or undesirable.

[0053] Obtain at least one user behavior data, wherein any user behavior data is used to characterize the user's interactive behavior event with respect to an object. An optional method is to extract at least one user behavior data from the log data information of the application platform. Another optional method is to extract at least one user behavior data from the real-time streaming data. Another optional method is to receive at least one user behavior data transmitted by a third-party platform through a data interface. This is not limited here.

[0054] For example, three user behavior data are extracted from the log data information of the note sharing platform:

[0055] {"userid":"u1","itemid":"a1","timestamp":400000","action":"like","eventid":0001};

[0056] {"userid":"u2","itemid":"a1","timestamp":415000","action":"comment","eventid":0004};

[0057] {"userid":"u1","itemid":"a2","timestamp":410000","action":"share","eventid":0007}.

[0058] Acquire at least one user behavior data, wherein any user behavior data is used to characterize the user's interactive behavior events toward the object, providing a data basis for the subsequent orderly and efficient management of the real-time streaming user behavior data, and providing original sample data support for the subsequent construction of label data to train the object prediction model.

[0059] Step 104: Based on the user identification of the at least one user behavior data, record the at least one user behavior data in a corresponding status list.

[0060] The user ID of user behavior data is an identifier that uniquely identifies the user and is used to distinguish different users on the application platform. In user behavior data, the user ID is a necessary field of user behavior data and is used to associate the user with the user behavior data. For example, the user ID (userid) in user behavior data ({"userid":"u1","itemid":"a1","timestamp":400000","action":"like","eventid":0001}) is "u1".

[0061] The state list corresponding to the user ID is a data structure used to store and manage the user behavior data of the user. Each user ID corresponds to a state list, which records all the user behavior data of the user in a specific time period. The state list is an important tool for managing and processing user behavior data. Through the state list, the user's behavior data can be managed in an orderly and efficient manner, which is convenient for subsequent data processing and label construction.

[0062] Based on the user identification of at least one user behavior data, at least one user behavior data is recorded in a corresponding status list. One optional way is: based on the user identification and object identification of at least one user behavior data, at least one user behavior data is recorded in a corresponding status list. Another optional way is: based on the user identification and event type of at least one user behavior data, at least one user behavior data is recorded in a corresponding status list. Another optional way is: based on the user identification and event identification of at least one user behavior data, at least one user behavior data is recorded in a corresponding status list. There is no limitation here.

[0063] Exemplarily, based on the user identifier and the object identifier of the at least one user behavior data, the at least one user behavior data is recorded in a corresponding state list:

[0064] Status list of user u1:

[0065] {"userid":"u1","itemid":"a1","timestamp":400000","action":"like","eventid":0001};

[0066] {"userid":"u1","itemid":"a2","timestamp":410000","action":"share","eventid":0007}.

[0067] Among them, the status list of user u1 records the like behavior data of user u1 on note a1 and the sharing behavior data of user u1 on note a2.

[0068] Status list of user u2:

[0069] {"userid":"u2","itemid":"a1","timestamp":415000,"action":"comment","eventid":0004}.

[0070] Among them, the status list of user u2 records the comment behavior data of user u2 on note a1.

[0071] Based on the user identification of at least one user behavior data, at least one user behavior data is recorded in a corresponding status list. User behavior data from different sources can be managed uniformly through the user identification. The user behavior data of each user is in an independent status list, which is convenient for subsequent data processing and analysis, while reducing data confusion and inconsistency problems and ensuring data consistency and accuracy.

[0072] Step 106: When a preset trigger time of the first state list is reached, the recorded first user behavior data is taken out from the first state list, and first label data corresponding to the first user behavior data is constructed, wherein the first state list is any state list.

[0073] The trigger time of the first status list is the trigger time for retrieving data pre-set for the first status list. When the trigger time point or trigger time period is reached, the recorded first user behavior data is retrieved from the first status list to ensure the timeliness and accuracy of the data, and to process and analyze the user behavior data in a timely manner to avoid data backlog and expiration, thereby losing the effect of real-time collection. The trigger time can be flexibly set according to business needs. For example, a timer registered for the first status list.

[0074] The first user behavior data is the user behavior data recorded in the first status list.

[0075] The first label data corresponding to the first user behavior data is the label data constructed for the first user behavior data. The first label data is used to describe the user's interests and preferences. It is the reference data for constructing user portraits and training object prediction models. The first label data can truly reflect the user's perception. For example, user u1 collected note a1, and the label data indicates that user u1 is more willing to personally collect a certain type of note represented by note a1. For another example, user u1 shared note a2, and the label data indicates that user u1 is more willing to share a certain type of note represented by note a1 with others.

[0076] An optional way to construct the first label data corresponding to the first user behavior data is to construct the first label data corresponding to the first user behavior data based on a queue mechanism.

[0077] Exemplarily, when the preset trigger time (30s) of the first state list is reached, the recorded first user behavior data is retrieved from the first state list:

[0078] {"userid":"u1","itemid":"a1","timestamp":400000","action":"like","eventid":0001};

[0079] {"userid":"u1","itemid":"a2","timestamp":410000","action":"share","eventid":0007},

[0080] Based on the queue mechanism, the first label data corresponding to the first user behavior data is constructed:

[0081] Interest tags: Like note a1: User u1 likes note a1, indicating that the user is interested in note a1. Share note a2: User u1 shares note a2, indicating that the user is very interested in note a2 and is willing to share it with others.

[0082] Behavior tags: Like: User u1 has performed 1 like behavior. Share: User u1 has performed 1 share behavior.

[0083] When the preset trigger time of the first state list is reached, the recorded first user behavior data is taken out from the first state list, and the first label data corresponding to the first user behavior data is constructed to ensure the orderliness and efficient management of the data, accurately and effectively construct the label data, ensure that the constructed label data of the user behavior data truly conforms to the user perception, and improve the accuracy of the label data.

[0084] Step 108: Based on the first label data, train an object prediction model.

[0085] The object prediction model is a machine learning model that predicts objects that are relevant to users. The object prediction model analyzes user behavior data and predicts objects that are highly relevant to users to improve user satisfaction and platform interaction rate.

[0086] Based on the first label data, the object prediction model is trained. An optional method is: based on the first user behavior data and the first label data, the object prediction model is supervised and trained.

[0087] Exemplarily, based on the first user behavior data and the first label data:

[0088] {"userid":"u1","itemid":"a1","timestamp":400000","action":"like","eventid":0001};

[0089] Like note a1: User u1 clicked the like button on note a1, indicating that the user is interested in note a1. Like: User u1 clicked the like button once.

[0090] {“userid”: “u1”, “itemid”: “a2”, “timestamp”: 410000, “action”: “share”, “eventid”: 0007}, supervised training of the note prediction model;

[0091] Sharing note a2: User u1 shared note a2, indicating that the user is very interested in note a2 and is willing to share it with others. Sharing: User u1 performed 1 sharing behavior.

[0092] In the embodiments of the present specification, a status list constructed based on user identification is used to achieve orderly and efficient management of real-time streaming user behavior data, reduce the confusion and inconsistency caused by user behavior data from different sources, accurately and effectively construct label data, ensure that the constructed label data of user behavior data truly conforms to user perception, improve the accuracy of label data, and thereby train the object prediction model with high accuracy, thereby improving user experience.

[0093] Figure 2 A data flow diagram of a training method for an object prediction model provided by an embodiment of this specification is shown, such as Figure 2 As shown:

[0094] User interaction events on objects on the application platform generate user behavior data. Use a label data construction framework, such as Session Labeler, to construct label data corresponding to user behavior data. Based on the label data, train the object prediction model.

[0095] Figure 3 A schematic diagram of an interactive behavior event in a training method for an object prediction model provided in one embodiment of the present specification is shown. Figure 3 As shown:

[0096] On the target page of the application platform, multiple note covers are displayed: including picture and text note A, video note B, picture and text note C, video note D, picture and text note E, and picture and text note F. The user clicks on a note to enter the note details page. The note details page displays the note content details, note comments: user 1, user 1's avatar and user comment 1, and a comment input box.

[0097] Among them, the target page can be the note recommendation page, nearby recommendation page, note search page and author details page.

[0098] In an optional embodiment of the present specification, after step 104, the following specific steps are also included:

[0099] When the first user behavior data recorded in the first status list is empty, setting the triggering time of the first status list;

[0100] When the first user behavior data recorded in the first status list is not empty, preprocessing is performed on the first user behavior data recorded in the first status list.

[0101] Preprocessing the first user behavior data recorded in the first status list, one optional way is: performing order preservation processing on the first user behavior data recorded in the first status list, another optional way is: performing deduplication processing on the first user behavior data recorded in the first status list, another optional way is: performing data cleaning processing on the first user behavior data recorded in the first status list, another optional way is: performing data conversion processing on the first user behavior data recorded in the first status list, which is not limited here.

[0102] Exemplarily, the first user behavior data recorded in the first status list is empty, and the trigger time of the first status list is set to 30s.

[0103] The first user behavior data recorded in the first status list is empty, and the trigger time of the first status list is set to 30s.

[0104] {"userid":"u1","trigger_time":400030}

[0105] The first user behavior data recorded in the first status list is not empty, and the first user behavior data recorded in the first status list is sequence-preserving and de-duplicating.

[0106] {"userid": "u1", "itemid": "a1", "timestamp": 400000, "action": "like", "eventid": 0001},

[0107] {"userid": "u1", "itemid": "a2", "timestamp": 410000, "action": "share", "eventid": 0007};

[0108] {"userid": "u1", "itemid": "a1", "timestamp": 405000, "action": "like", "eventid": 0002}.

[0109] In the embodiments of the present specification, when the first user behavior data recorded in the first status list is empty, the trigger time of the first status list is set; when the first user behavior data recorded in the first status list is not empty, the first user behavior data recorded in the first status list is preprocessed, and the trigger time is set to ensure timely system response and avoid idle resources; the order preservation processing ensures that the data is processed in chronological order to avoid analysis errors caused by disorder, and the data quality of the user behavior data is improved through preprocessing.

[0110] In an optional embodiment of the present specification, preprocessing the first user behavior data recorded in the first status list includes the following specific steps: sorting the first user behavior data based on the timestamp of the first user behavior data recorded in the first status list.

[0111] The timestamp of the first user behavior data is recorded in the user behavior data in the first state list, and represents the numerical value of the specific time when the behavior occurs. The timestamp is usually an integer, representing a fixed time point. The timestamp is used to accurately record the time when the user behavior occurs.

[0112] The first user behavior data is sorted based on the timestamp of the first user behavior data recorded in the first status list. The sorting may be from front to back or from back to front, which is not limited here.

[0113] Exemplarily, a state list including three states is created for sorting the user behavior data to ensure that the user behavior data conforms to a real interaction behavior sequence.

[0114] State 1: The device time of a single user's behavior data is used to count a single user in the clear state. The timestamp is used for order-keeping state monitoring.

[0115] State 2: state information of user behavior data, the key (Key) is used to record the state information of the previous timer, and the value (Value) is the number of user behavior data in the state list corresponding to the user identifier.

[0116] State three: used to save the complete information in the user behavior data in the state list corresponding to the user ID.

[0117] The timer is triggered to distribute the user behavior data that meets the rules and is arranged to the downstream label construction operator, and the user behavior data that does not meet the rules is saved again in the status list.

[0118] Specifically, if the state list records user behavior data, the information of state 2 of the user is retrieved, and state 2 is updated at the same time, with the key updated to the latest timestamp and the value updated to the amount of user behavior data.

[0119] After the update of state 2 is completed, the complete information in the user behavior data is saved in the state list of state 3.

[0120] If there is no user behavior data recorded in the state list, register a timer to set the trigger time (current time plus 30s) to initialize state 1, which is used to count the start time and interval time of the user behavior data in the trigger time when the timer is triggered.

[0121] Initialize state 2, the key is updated to the latest timestamp, and the value is updated to the number of user behavior data.

[0122] After the update of state 2 is completed, the complete information in the user behavior data is saved in the state list of state 3.

[0123] In the embodiments of this specification, timestamp sorting ensures that data is processed in time order, avoiding analysis errors caused by disorder, and status monitoring and timers ensure real-time processing and updating of data, thereby improving the data quality of user behavior data.

[0124] In an optional embodiment of the present specification, after taking out the recorded first user behavior data from the first status list in step 106, the following specific steps are also included: adding the taken out first user behavior data to a priority queue, wherein the first user behavior data recorded in the priority queue is arranged according to timestamps;

[0125] Correspondingly, constructing the first label data corresponding to the first user behavior data in step 106 includes the following specific steps: when the timestamp of the first user behavior data is less than the trigger time of the preset priority queue, taking the first user behavior data from the priority queue and constructing the first label data corresponding to the first user behavior data.

[0126] The priority queue is a queue of first user behavior data sorted by timestamp, where each first user behavior data has a certain priority. The order in which the first user behavior data is dequeued depends not only on its order of entry, but also on the priority of the first user behavior data, ensuring that those behavior data that are earlier in time can be processed first. For example, user u1 liked note a1 at time 400000, and then shared note a2 at 410000. These two behavior data will be arranged in the priority queue in timestamp order, even if the sharing operation data was entered later than the like operation.

[0127] The trigger time of the priority queue is the predetermined time point or time interval for taking out and processing the head element from the queue. The trigger time of the priority queue can be used as a time window to determine the frequency of checking the queue and performing subsequent data processing, which helps to control the speed of data processing and prevent data backlog or processing delays. For example, the trigger time of the priority queue is every 15 seconds. This means that every 15 seconds, the oldest user behavior data in the queue (that is, the one with the earliest timestamp) will be checked and taken out for the next step of processing, such as building label data. If at a certain trigger, the timestamp of the oldest user behavior data in the queue is 400000, and the current time is 400015, then this data will be taken out and processed.

[0128] For example, the trigger time for sorted user behavior data to exit the priority queue is set to 15s, which is the upper limit of the queue. The user behavior data of state 3 is taken out from the state list and added to the priority queue sorted by timestamp. The priority queue is traversed and the first user behavior data in the priority queue sorted by timestamp is continuously taken out:

[0129] If the timestamp of the user behavior data is less than the upper bound of the sorted queue, the user behavior data is set as the sorted user behavior data, and label data corresponding to the user behavior data is constructed.

[0130] Figure 4 A schematic diagram of triggering time in a training method for an object prediction model provided in one embodiment of this specification is shown, such as Figure 4 As shown:

[0131] Step 1: Register the timer: The trigger time of the status list is set to 30s.

[0132] Step 2: Use the time window to sort the user behavior data: the trigger time of the priority queue is set to 15s.

[0133] Step 3: When the time window is reached, the downstream label construction operator distributes the user behavior data value: when the timestamp of the user behavior data is less than the trigger time of the preset priority queue, the label data corresponding to the user behavior data is constructed.

[0134] The trigger time of the status list overlaps with the trigger time of the priority queue, ensuring that user behavior data in each time period can be processed in a timely manner without data omissions due to delays in a certain link. At the same time, the trigger time of the status list sets a larger time window to ensure regular data collection, while the trigger time of the priority queue sets a smaller time window to ensure timely data processing. This dual mechanism ensures the continuity and real-time nature of data processing.

[0135] In the embodiments of this specification, the dual mechanism of the trigger time of the status list and the trigger time of the priority queue is used to ensure timely, continuous and efficient processing of user behavior data, avoid data backlog and omission, and improve the efficiency and accuracy of subsequent construction of label data.

[0136] In an optional embodiment of the present specification, after adding the retrieved first user behavior data to the priority queue, the following specific steps are also included: when the timestamp of the first user behavior data is greater than a preset trigger time of the priority queue, the first user behavior data is transferred from the priority queue to the ordinary queue; when the preset trigger time of the ordinary queue is reached, the first user behavior data recorded in the ordinary queue is returned to the first status list, and the ordinary queue is cleared.

[0137] The common queue is a queue of the first user behavior data arranged according to the time of entry, in which each first user behavior data has no priority, and the order of exiting the queue of the first user behavior data depends only on its order of entry, and it is exited in a first-in-first-out manner. For example, user u1 liked note a1 at time 400000, and then shared note a2 at 410000. The entry time of the sharing operation data is later than that of the like operation, and the user behavior data of the like operation is before the user behavior data of the share operation.

[0138] The trigger time of the ordinary queue is the predetermined time point or time interval for taking out and processing the head element from the ordinary queue. The trigger time of the ordinary queue can be used as a time window to control when to take out the user behavior data from the ordinary queue and put it back into the first state list for processing, which helps to control the speed of data processing and prevent data backlog or processing delays. For example, the trigger time of the ordinary queue is 15 seconds, which means that every 15 seconds, the oldest user behavior data in the ordinary queue (that is, the one with the earliest timestamp) is checked, and these data are taken out and put back into the first state list for processing. If at a certain trigger, the timestamp of the oldest user behavior data in the ordinary queue is 410000, and the current time is 410015, then these data will be taken out and put back into the first state list, waiting for the next processing.

[0139] Exemplarily, if the timestamp of the user behavior data is greater than the upper bound of the sorted queue, the user behavior data is added to the general queue. In the general queue, if the general queue is not empty, the general queue is set to the state list result of state three, and state two is updated to save the number of user behavior data in state three, and the trigger time of the general queue is increased by 15s as the time point of the next window trigger. In the general queue, if the general queue is empty, state one, state two, and state three are cleared, and the difference between the initialization time and the current time of the data in state one is recorded for statistical data age in the sorting operator.

[0140] In the embodiments of the present specification, by transferring user behavior data that exceeds the trigger time of the priority queue to the ordinary queue and putting it back into the first state list when the trigger time of the ordinary queue is reached, comprehensive coverage and orderly processing of the data are ensured, data omissions and backlogs are avoided, and the efficiency and real-time performance of data processing are improved.

[0141] Figure 5 A schematic diagram of the process of preserving order in a training method for an object prediction model provided in one embodiment of the specification is shown. Figure 5 As shown:

[0142] Start. Get user behavior data. Based on the user ID, record the user behavior data in the corresponding state list. Trigger the execution of the order-preserving operator. Query the database to determine whether the state list is empty; if so, initialize: set the trigger time of the state list, initialize the state one, state two and state three of the recorded user behavior data; if not, take out state two and state three, and update state two and state three. Determine whether the trigger time of the state list has been reached; if so, take out state one, state two and state three; set the priority queue, the trigger time of the priority queue and the ordinary queue; add the user behavior data recorded in state three to the priority queue and traverse. Determine whether the trigger time of the priority queue has been reached; if so, take out the user behavior data from the priority queue and construct the label data corresponding to the user behavior data; if not, transfer the user behavior data from the priority queue to the ordinary queue. Determine whether the ordinary queue is empty; if so, clear the state list and record the data age; if not, return the user behavior data recorded in the ordinary queue to the state list and clear the ordinary queue.

[0143] like Figure 5 As shown, by sorting the real behavior sequence of the same user behavior in time series, it is ensured that the label data of the constructed user behavior data truly conforms to the user perception, is real and conforms to the time series, and avoids the disorder of data time series when user behavior data from different sources are transmitted, thereby improving the accuracy of the label data. The model can consume accurate real user interaction data of various business scenarios. The object prediction model trained in this way has high accuracy, which improves the user experience.

[0144] In an optional embodiment of the present specification, the first user behavior data recorded in the first status list is preprocessed, including the following specific steps: based on the event identifier of the first user behavior data recorded in the first status list, duplicate first user behavior data is deduplicated.

[0145] The event identifier of the first user behavior data is recorded in the user behavior data in the first state list and is used to uniquely identify the field of the interactive behavior event. The event identifier is usually a unique string or number, which is used to distinguish different user behavior data and ensure that each user behavior data is unique. For example, on a note sharing platform, user u1 has performed multiple like operations on note a1, and the generated user behavior data is as follows:

[0146] User u1 liked note a1 at time 400000:

[0147] {"userid": "u1", "itemid": "a1", "timestamp": 400000, "action": "like", "eventid": 0001}

[0148] User u1 liked note a1 again at time 405000:

[0149] {"userid": "u1", "itemid": "a1", "timestamp": 405000, "action": "like", "eventid": 0002}

[0150] In these two user behavior data, the "eventid" field records 0001 and 0002 respectively. These two event identifiers ensure that even if user u1 likes the same note a1 multiple times, each operation is unique in the system and can be accurately deduplicated and tracked.

[0151] Exemplarily, a state list including two states is created to deduplicate user behavior data, thereby ensuring that the user behavior data conforms to a real interaction behavior sequence.

[0152] State 1: The device time of a single user's behavior data is used to count a single user in the clear state. The timestamp is used for order-keeping state monitoring.

[0153] State 2: used to save the complete information in the user behavior data in the state list corresponding to the user ID.

[0154] Trigger the timer and continue to save the user behavior data that meets the rules in the status list.

[0155] Specifically, if there is no user behavior data for the event identifier in the state list, the data is directly handed over to the downstream label construction operator for processing, the event identifier and timestamp are saved in state two, and it is determined whether the user has set state one. If not, a timer is registered to set the trigger time and initialize state one.

[0156] If the status list contains user behavior data identified by the event, duplicate user behavior data is removed and not handed over to the downstream label construction operator for processing.

[0157] In the embodiments of the present specification, by deduplicating the first user behavior data recorded in the first state list, the uniqueness and accuracy of each user behavior data is ensured, analysis errors caused by duplicate data are avoided, the efficiency and accuracy of data processing are improved, the authenticity and reliability of user behavior data are ensured, and high-quality data support is provided for the subsequent construction of label data and training object prediction models.

[0158] In an optional embodiment of the present specification, after taking out the recorded first user behavior data from the first status list in step 106, the following specific steps are also included: when the timestamp of the first user behavior data reaches the trigger time of the preset status list, the first user behavior data is added to the general queue; if the general queue is not empty, based on the event identifier of the first user behavior data recorded in the general queue, the corresponding first user behavior data recorded in the first status list is deleted, and the trigger time of the status list is updated; if the general queue is empty, the first status list is cleared.

[0159] Exemplarily, a common queue is initialized to store the user behavior data that needs to be eliminated. A common timestamp is initialized to store the minimum timestamp that has not been eliminated in state 2.

[0160] Take out the data in state 2 and traverse it: if the time when the current timer is triggered is greater than the timestamp of the user behavior data plus 30 minutes, it meets the elimination rule and the event identifier is added to the general queue. If the time when the current timer is triggered is less than the timestamp of the user behavior data plus 30 minutes, update to determine whether the user behavior data is the smallest timestamp of a single user and assign a value.

[0161] Traverse the normal queue: If there are elements in the normal queue, remove the user behavior data with the same event ID in state 2. At the same time, set a new timer to update the trigger time of the state list: the normal timestamp plus 30 minutes is the latest window trigger time.

[0162] If there are no elements in the normal queue, state 2 and state 3 are cleared.

[0163] In the embodiments of the present specification, by moving expired user behavior data to the general queue and removing the corresponding data in the first state list based on the event identifier, the timeliness and accuracy of the data are ensured, data redundancy is reduced, data processing efficiency and accuracy are improved, data backlogs and repeated processing are avoided, and efficient and accurate data deduplication is ensured.

[0164] Figure 6 FIG. 1 shows a flow chart of deduplication processing in a training method for an object prediction model provided by an embodiment of the present specification, such as Figure 6 As shown:

[0165] Start. Get user behavior data. Based on the user ID, record the user behavior data in the corresponding state list. Trigger the execution of the deduplication operator. Query the database to determine whether the state list is empty; if so, initialize: set the trigger time of the state list, initialize the state one of the recorded user behavior data; if not, take out state two. Determine whether the state list has repeated event identifiers; if so, perform deduplication; if not, take out the user behavior data, hand it over to the downstream tag to build the calculation group, and save the event identifier and timestamp. Determine whether the trigger time of the state list has been reached; if so, traverse state two; when the timestamp of the user behavior data reaches the preset trigger time of the state list, add the user behavior data to the ordinary queue; traverse state two and remove the elements that need to be removed. Determine whether the ordinary queue is empty; if so, based on the event identifier of the user behavior data recorded in the ordinary queue, remove the corresponding user behavior data recorded in the state list, and update the trigger time of the state list; if not, clear the state list.

[0166] like Figure 6 As shown, by deduplicating the real behavior sequence of the same user behavior, it is ensured that the label data of the constructed user behavior data truly conforms to the user perception, is real and conforms to the time series, and avoids data duplication when different user behavior data are transmitted, thereby improving the accuracy of the label data. The model can consume accurate real user interaction data of various business scenarios, and the object prediction model trained in this way has high accuracy, thereby improving the user experience.

[0167] In an optional embodiment of the present specification, before step 104, the following specific steps are also included: determining the business scenario to which each user behavior data belongs based on the interactive behavior attribute of at least one user behavior data;

[0168] Correspondingly, step 108 includes the following specific steps: based on the first label data, training an object prediction model for the first business scenario.

[0169] Interaction behavior attributes are attributes that describe the specific characteristics and details of the user's interaction behavior on the application platform. These attributes may include but are not limited to event type, timestamp, object identifier, user identifier, etc. The interaction behavior is recorded in the user behavior data in the first state list, which is used to record and describe each interaction behavior of the user in detail. For example, the first user behavior data includes the interaction behavior attribute "source": recommendation_page", which indicates that the source is the recommendation page.

[0170] A business scenario is a specific function or service module on an application platform. Each business scenario has its own unique user behavior pattern and data characteristics. The division of business scenarios helps to analyze and process user behavior data in a targeted manner and improve the efficiency and accuracy of data processing. Figure 3 As shown in the figure, users can enter the note details page from any of the note recommendation page, nearby recommendation page, note search page, and author details page, corresponding to the recommendation scenario, search scenario, and details page scenario. For example, a user clicks on a note from the recommendation page to enter the note details page. According to the interactive behavior attribute in the user behavior data, "source": recommendation_page", the business scenario is the recommendation scenario.

[0171] The object prediction model used for the first business scenario is a machine learning model specially trained for a specific business scenario. The model predicts objects that users may be interested in and makes highly relevant predictions by analyzing user behavior data in the business scenario. The user behavior patterns and data characteristics of each business scenario may be different, so training a prediction model for each business scenario separately can improve the accuracy and relevance of the prediction. For example, on a note sharing platform, an object prediction model was trained for the business scenario of "note recommendation page". The model predicts notes that users may be interested in and makes highly relevant note recommendations by analyzing user behavior data on the note recommendation page (such as likes, comments, shares, etc.).

[0172] In the embodiments of this specification, by determining the business scenario to which each user behavior data belongs based on the interactive behavior attributes of the user behavior data, and training a dedicated object prediction model for each business scenario, the technical effects of improving prediction accuracy, enhancing high-correlation predictions, improving data processing efficiency and reducing data processing complexity are achieved, thereby improving user satisfaction and the platform's interaction rate.

[0173] In an optional embodiment of the present specification, after step 104, the following specific steps are further included: based on the business scenario to which the first user behavior data belongs, aggregating and obtaining a batch of first user behavior data corresponding to each business scenario;

[0174] Correspondingly, after taking out the recorded first user behavior data from the first status list in step 106, the following specific steps are also included: distributing each batch of the first user behavior data to a message queue;

[0175] Correspondingly, constructing the first label data corresponding to the first user behavior data in step 106 includes the following specific steps: taking out the first user behavior data of a batch corresponding to the first business scenario recorded from the message queue, and constructing the first label data corresponding to the first user behavior data.

[0176] The business scenario corresponding batch is to aggregate the user behavior data belonging to the same business scenario into one batch to facilitate subsequent data processing and model training. Each batch contains user behavior data within a certain time range, and these data have the same business scenario attributes. By dividing user behavior data into batches according to business scenarios, data can be managed and processed more efficiently, ensuring that the data of each business scenario maintains consistency and integrity during processing and analysis. For example, on a note sharing platform, user behavior data is divided into batches according to business scenarios:

[0177] Recommended scenarios corresponding batches:

[0178] {"userid": "u1", "itemid": "a1", "timestamp": 400000, "action": "click", "eventid": 0001, "source": "recommendation_page"};

[0179] {"userid": "u2", "itemid": "a2", "timestamp": 405000, "action": "like", "eventid": 0002, "source": "recommendation_page"}

[0180] Search scenario corresponding batch:

[0181] {"userid":"u3","itemid":"a3","timestamp":400000,"action":"click","eventid":0003,"source":"search_page","query":"Programming tips"};

[0182] {"userid":"u4","itemid":"a4","timestamp":405000,"action":"like","eventid":0004,"source":"search_page","query":"Programming Tips"}

[0183] Message queue is an asynchronous communication mechanism in distributed systems, used to transmit messages in distributed systems. Message queue can achieve decoupling between producers and consumers, ensuring reliable transmission and processing of messages. Common message queue systems include Kafka, RabbitMQ, etc. In the user behavior data processing process, message queue is used to distribute the first batch of user behavior data corresponding to each business scenario to different processing nodes. Through message queue, asynchronous data processing can be achieved, the throughput and stability of the system can be improved, and the orderliness and consistency of data can be ensured.

[0184] Exemplarily, the batch of first user behavior data corresponding to each business scenario is distributed to the Kafka message queue. The batch of first user behavior data corresponding to the first business scenario recorded is retrieved from the Kafka message queue, and first label data corresponding to the first user behavior data is constructed.

[0185] In the embodiments of this specification, by distributing the first user behavior data of the batch corresponding to each business scenario to the message queue, and taking out the first user behavior data of the batch corresponding to the first business scenario recorded from the message queue, and constructing the first label data corresponding to the first user behavior data, it is possible to improve data processing efficiency and ensure data consistency and integrity, and use asynchronous message queues to improve throughput and stability, thereby optimizing platform performance and user experience.

[0186] In an optional embodiment of the present specification, based on the business scenario to which the first user behavior data belongs, the first user behavior data of batches corresponding to each business scenario is summarized, including the following specific steps: if the event type of the first user behavior data recorded in the first status list is a basic event type, the basic attribute information under the first business scenario is updated, and the first business scenario is identified as a business scenario with basic event information; if the event type of the first user behavior data recorded in the first status list is a click event type, check whether the trigger time of the first status list is updated, if not, update the trigger time of the first status list, and update the first user behavior data recorded in the first status list.

[0187] Basic event types are basic interactive behavior events performed by users on the application platform. These events are usually used to initialize or update the basic behavior data of users, providing basic support for subsequent advanced behavior analysis and prediction model training. Basic event types usually include basic operations such as user registration, login, and browsing. These events are the most basic part of user behavior data, reflecting the user's activity and basic behavior patterns. By recording and analyzing these basic events, you can get a preliminary understanding of the user's activities and basic preferences.

[0188] Basic attribute information is the basic attributes and behavior characteristics of users on the application platform. This information is usually initialized and updated through data of basic event types. Basic attribute information includes the user's registration time, login frequency, browsing history, etc. This information reflects the user's activity and basic behavior patterns, and provides an important reference for subsequent advanced behavior analysis and prediction model training. For example, registration time: the time when the user first registered. Login frequency: the number of times a user logs in within a certain period of time. Browsing history: the list of objects browsed by the user.

[0189] Click event types are click operations performed by users on the application platform. These operations usually involve user interactions with specific objects or functions. Click event type data is a very important part of user behavior data, reflecting user interests and preferences. Click event types include operations in which users click on a link, button, or object. These event data can help the platform understand users' interests and behavior patterns, and provide important basis for high-correlation predictions and object prediction optimization.

[0190] Exemplarily, if it is a basic event: such as a note exposure event, the basic attribute information is updated, and the business scenario is identified as a business scenario with basic event information; if it is not a basic event: the behavior type is determined, and the behavior information of the first user behavior data is updated:

[0191] If it is a click event: Check if the 20-minute timer has been extended. If it has been extended, skip it. Check if the 20-minute timer has been extended. If it has not been extended, re-register a 20-minute timer based on the timestamp of the click event and delete the old 20-minute timer. Take out the old status result information, update the behavior information based on the behavior type, and save the updated result back to the status.

[0192] Figure 7 FIG. 1 shows a flow chart of a summary process in a training method for an object prediction model provided by an embodiment of the present specification, such as Figure 7 As shown:

[0193] Start. Get user behavior data. Based on the user ID, record the user behavior data in the corresponding state list. Trigger the execution of the deduplication operator. Determine the business scenario to which the user behavior data belongs. Based on the business scenario to which the user behavior data belongs, summarize the user behavior data of the corresponding batches of each business scenario. Query the database to determine whether the state list is empty; if so, initialize: set the trigger time of the state list and initialize the state list; if not, retrieve the state information. Determine whether the user behavior data is a click event; if so, determine whether the trigger time of the check state list is updated; if so, skip; if not, update the trigger time of the state list and update the user behavior data recorded in the state list; if not, update the state information and mark the user behavior data to determine that the trigger time of the state list has been reached; if so, retrieve the summarized user behavior data of the corresponding batches of each business scenario; determine whether there is a basic event; if not, clear the state list; if so, distribute each batch of user behavior data to the message queue, perform data compression, and distribute the timestamp to the message queue.

[0194] like Figure 7 As shown, by aggregating the real behavior sequence of the same user behavior, it is ensured that the label data of the constructed user behavior data truly conforms to the user perception, is real and consistent with the time series, and conforms to the model training label specifications, thereby improving the accuracy of the model, improving the model quality, and thus improving the user experience.

[0195] In an optional embodiment of the present specification, after step 104, the following specific steps are also included:

[0196] Distribute the first user behavior data to the downstream label construction operator of each business scenario according to the distribution rules corresponding to each business scenario;

[0197] Correspondingly, constructing the first label data corresponding to the first user behavior data in step 106 includes the following specific steps:

[0198] Execute the downstream label construction operator of each business scenario to construct the first label data corresponding to the first user behavior data.

[0199] The distribution rules corresponding to the business scenarios are rules for distributing user behavior data to the processing nodes or components of the corresponding business scenarios based on the business scenario attributes of the user behavior data. These rules ensure that the user behavior data can be correctly routed to the label construction operator or other processing logic responsible for processing the business scenario. Distribution rules are usually formulated based on key attributes in user behavior data (such as event type, source, timestamp, etc.). Through these rules, data can be ensured to be transmitted efficiently and orderly in the distributed system, avoiding confusion and duplication in data processing.

[0200] The downstream label building operator of a business scenario refers to the component or algorithm responsible for processing user behavior data in a specific business scenario and building the corresponding label data. These operators analyze and process user behavior data according to the characteristics and needs of the business scenario to generate label data that describes user interests and preferences. Downstream label building operators usually include steps such as data cleaning, feature extraction, and label generation. Each business scenario may have different label building logic, so it is necessary to design special operators to process data for specific scenarios. For example, on a note sharing platform, business scenarios are divided into prediction scenarios, search scenarios, and detail page scenarios. The downstream label building operators can be as follows:

[0201] 1. Label construction operator for prediction scenarios:

[0202] Input: User behavior data (such as clicks, likes, shares, etc.).

[0203] Processing logic: Clean the data and remove invalid or abnormal data. Extract features such as user ID, object ID, behavior type, timestamp, etc. Generate corresponding label data according to the behavior type, such as interest tags, behavior tags, etc.

[0204] Output: Interest tags: the user's interest in a specific note. Behavior tags: the user's specific behavior (such as likes, shares, etc.).

[0205] 2. Label construction operator for search scenarios:

[0206] Input: User behavior data (such as searches, clicks on search results, etc.).

[0207] Processing logic: Clean the data and remove invalid or abnormal data. Extract features, such as user ID, query keywords, clicked object ID, etc. Generate corresponding label data based on query keywords and click behaviors, such as search interest tags, search behavior tags, etc.

[0208] Output: Search interest tags: the user's interest in a specific keyword. Search behavior tags: the user's specific search behavior (such as clicking on search results, etc.).

[0209] 3. Label construction operator for detail page scenario:

[0210] Input: User behavior data (such as comments, favorites, etc.).

[0211] Processing logic: Clean the data and remove invalid or abnormal data. Extract features, such as user ID, object ID, behavior type, comment object, etc. Generate corresponding label data according to the behavior type, such as comment label, collection label, etc.

[0212] Output: Comment label: the user's comment object and its sentiment. Collection label: the user's collection behavior of a specific note.

[0213] Exemplarily, according to the distribution rules corresponding to each business scenario, the first user behavior data is distributed to the downstream label construction operator of each business scenario:

[0214] Recommended scenarios:

[0215] defrecommendation_label_builder(user_behavior_data):

[0216] ifuser_behavior_data['action']=='click':

[0217] return{'interest':f"User {user_behavior_data['userid']} is interested in note {user_behavior_data['itemid']}"}

[0218] elifuser_behavior_data['action']=='like':

[0219] return{'behavior':f"User {user_behavior_data['userid']} liked the note {user_behavior_data['itemid']}"}

[0220] elifuser_behavior_data['action']=='share':

[0221] return{'behavior':f"User {user_behavior_data['userid']} shared the note {user_behavior_data['itemid']}"}

[0222] Search scenarios:

[0223] defsearch_label_builder(user_behavior_data):

[0224] ifuser_behavior_data['action']=='click':

[0225] return{'search_interest':f"User {user_behavior_data['userid']} is interested in the note {user_behavior_data['itemid']} in the search results for keyword '{user_behavior_data['query']}'"}

[0226] elifuser_behavior_data['action']=='search':

[0227] return{'search_behavior':f"User {user_behavior_data['userid']} searched for keyword '{user_behavior_data['query']}'"}

[0228] Details page scenario:

[0229] defdetail_label_builder(user_behavior_data):

[0230] ifuser_behavior_data['action']=='comment':

[0231] return{'comment':f"User {user_behavior_data['userid']} commented on the note {user_behavior_data['itemid']}, the comment object is '{user_behavior_data['content']}'"}

[0232] elifuser_behavior_data['action']=='collect':

[0233] return{'behavior':f"User {user_behavior_data['userid']} collected the note {user_behavior_data['itemid']}"}

[0234] In the embodiments of this specification, by distributing user behavior data to the downstream label construction operator of each business scenario according to the distribution rules corresponding to each business scenario, it is ensured that the data is transmitted efficiently and orderly in the distributed system, avoiding confusion and duplication of data processing. The label construction operator of each business scenario generates label data describing user interests and preferences according to specific processing logic, which improves the accuracy and comprehensiveness of the label data and provides high-quality data support for the training of the object prediction model, thereby improving the accuracy and relevance of the prediction system and optimizing the user experience and the interaction rate of the platform.

[0235] In an optional embodiment of the present specification, according to the distribution rules corresponding to each business scenario, distributing the first user behavior data to the downstream label construction operator of each business scenario includes the following specific steps:

[0236] Traverse the distribution rules corresponding to each business scenario;

[0237] If the distribution rule of the second business scenario is hit, the first user behavior data belonging to the second business scenario is distributed to the downstream label construction operator of the second business scenario, where the second business scenario is any business scenario.

[0238] Exemplarily, it is determined whether the user behavior data arrives at the distribution operator for the first time. If it is the first time, a 20-minute timer is registered and user-dimensional status information is initialized for the user.

[0239] According to the data logic, determine whether the data is missing some information. If it is missing, fill it in from the user dimension status and update the user's status information: save the user behavior data of browsing the last 300 notes for 10 minutes. Save the user behavior data of clicking the last 300 notes for 10 minutes.

[0240] Traverse all distribution rules:

[0241] Further determine whether the basic event logic is hit.

[0242] Assemble the data, extract the user ID and note ID, and determine whether it is basic data information and hits a business scenario rule information. Judge the data, and determine the distribution rule corresponding to the business scenario according to the business scenario to which the data belongs. If they are the same, the data is determined to be the business scenario data, otherwise, the data is discarded.

[0243] In the embodiments of this specification, by traversing the distribution rules of each business scenario, the user behavior data is accurately distributed to the corresponding downstream label construction operator to ensure the efficiency and accuracy of data processing. The first arriving data registers the timer and initializes the user status, and ensures data integrity by filling in missing information and updating the status. Save the user's recent behavior data to support real-time analysis. After hitting the basic event logic, distribute the data according to the business scenario rules to ensure the accuracy and consistency of the label data, and improve the high relevance of object prediction and user experience.

[0244] In an optional embodiment of the present specification, after traversing the distribution rules corresponding to each business scenario, the following specific steps are also included:

[0245] If the distribution rule of any business scenario is not matched, the signal control is traversed to determine whether the time window of the downstream label construction operator needs to be adjusted;

[0246] If user behavior data is generated within the time window reaching the trigger time of the first state list, the time window is updated;

[0247] If no user behavior data is generated within the time window reaching the triggering time of the first status list, the first status list is cleared.

[0248] The time window of the downstream label building operator is the time range or period used to process and build user behavior data labels. It is a time period set according to business needs to determine when to start processing data, building labels, and outputting the results to the next processing stage or storage. The setting of the time window is crucial for real-time data processing. It not only affects the real-time performance of data processing, but also affects the rational allocation and utilization of resources. By setting the time window reasonably, it can be ensured that data processing is neither too frequent to waste resources nor too sparse to cause data delays. For example, the note sharing platform sets a time window for processing user behavior data every 10 minutes. Whenever this 10-minute time window ends, the downstream label building operator will start, process all user behavior data collected during this period, build corresponding labels, and use these label data for subsequent prediction model training or high-correlation predictions.

[0249] The time window of the trigger time of the first status list is the time interval from the creation of the status list or the last processing to the next triggering of data processing. This time interval is pre-set to ensure that the data in the status list can be processed within the specified time to avoid data backlog and expiration. The time window of the trigger time is an important parameter in the status list management mechanism, which determines the frequency and timeliness of data processing. Reasonable trigger time settings can ensure the timeliness and accuracy of data while avoiding excessive consumption of resources. For example, the note sharing platform sets a time window that triggers the status list processing every 30 seconds. Whenever this 30-second time window ends, the system checks the data in the status list, extracts and processes the data, builds labels, and uses the label data for subsequent prediction model training or high-correlation predictions. If no new user behavior data is generated within these 30 seconds, the status list is cleared and ready to receive new data.

[0250] For example, all signal configuration logics are traversed:

[0251] Determine whether the signal event logic is hit, and send a signal to the downstream summary window based on the hit signal information to drive whether the downstream window is changed.

[0252] Determine whether the user has any interactive behavior in the last 5 minutes. If so, continue to extend the user's time window. If there is no user interaction record within 5 minutes when the time window is triggered, clear the first state list.

[0253] The user ID, note ID, and business scenario are used as the unique key for a single user for a single note, and Flink's group by is used for aggregation.

[0254] In the embodiments of this specification, the time window of the downstream tag construction operator and the time window of the trigger time of the first state list are reasonably set and managed to ensure the efficiency and real-time performance of data processing. The specific technical effects include: dynamically adjusting the time window, updating in real time according to user behavior, avoiding resource waste and data backlog, clearing useless data in time, reducing storage overhead, and improving system performance; ensuring the accuracy and timeliness of data processing through signal control and user behavior monitoring, and improving the high relevance of object prediction and user experience.

[0255] In an optional embodiment of the present specification, before step 104, the following specific steps are also included: determining whether the advance distribution rule is met; if so, distributing the first user behavior data to the downstream label construction operator of each business scenario;

[0256] Correspondingly, step 104 includes the following specific steps: if not, taking out the recorded first user behavior data from the first status list.

[0257] The advance distribution rule is a rule that distributes the data to the downstream tag building operators of each business scenario in advance according to preset conditions or logic after the user behavior data arrives at the system. This rule allows the system to process and distribute data immediately in some cases without waiting for the trigger time of the status list, so as to achieve faster response and more timely data processing. The advance distribution rule is designed to deal with some urgent or important user behavior data to ensure that the data can be processed and used quickly. By distributing data in advance, the real-time performance and response speed of the system can be improved, especially when the user behavior data has a great impact on the business, the prediction strategy can be adjusted in time to improve the user experience. For example, suppose that on a note sharing platform, there is an advance distribution rule that stipulates that if a user has multiple interactions (such as likes, comments, shares, etc.) on a note within a short period of time (such as within 5 minutes), the behavior data will be immediately distributed to the downstream tag building operator of the corresponding business scenario instead of waiting for the trigger time of the status list.

[0258] Exemplarily, traverse all advance distribution rules:

[0259] Determine whether the advance distribution rule is hit. If so, further determine whether the basic event logic is hit:

[0260] Assemble the data, extract the user ID and note ID, and determine whether it is basic data information and hits a business scenario rule information. Judge the data, and determine the distribution rule corresponding to the business scenario according to the business scenario to which the data belongs. If they are the same, the data is determined to be the business scenario data, otherwise, the data is discarded.

[0261] In the implementation of this specification, by introducing advance distribution rules, the system can immediately judge and process important data after the user behavior data arrives, without waiting for the trigger time of the status list. This not only improves the real-time performance and response speed of the system, ensuring that key user behavior data can be quickly utilized, but also can adjust the prediction strategy in time to improve user experience and platform interaction rate. Through dynamic judgment and rapid distribution, the system effectively reduces data processing delays and enhances the timeliness of object prediction.

[0262] Figure 8 FIG. 1 shows a flow chart of distribution processing in a training method for an object prediction model provided by an embodiment of the present specification, such as Figure 8 As shown:

[0263] Start. Get user behavior data. Based on the user ID, record the user behavior data in the corresponding state list. Trigger the execution of the deduplication operator. Determine the business scenario to which the user behavior data belongs. Trigger the execution of the distribution operator. Query the database to determine whether the advance distribution rule is met: if so, initialize: load and initialize the business data filtering configuration, advance data filtering configuration, regular data filtering configuration filtering configuration; if not, retrieve the state information to trigger the regular distribution logic. According to the user ID, complete the user behavior data. Update the state information. Traverse the distribution rules corresponding to each business scenario. Determine whether the distribution rule is hit; if so, assemble the user behavior data to determine whether it is the user behavior data belonging to the business scenario; if so, distribute the user behavior data belonging to the business scenario to the downstream label construction operator of the business scenario; if not, traverse the signal control to determine whether the time window of the downstream label construction operator needs to be adjusted; if not, determine whether the time window of the downstream label construction operator needs to be adjusted; if yes, determine whether the user behavior data is generated within the time window of the trigger time of the state list; if yes, update the time window; if not, clear the state list.

[0264] like Figure 8 As shown, by distributing the real behavior sequence of the same user behavior according to the scenario information, it is ensured that when the label system processes the user behavior data, it can distribute one piece of user behavior data to multiple business scenarios, thereby improving the accuracy of the label data of each scenario, thereby improving the accuracy of the model, improving the model quality, and thus improving the user experience.

[0265] The following combination Fig. 9 Taking the application of the object prediction model training method provided in this specification in the note sharing platform as an example, the object prediction model training method is further described. Fig. 9 A flowchart of a processing process of a training method for an object prediction model applied to a note sharing platform provided by an embodiment of the present specification is shown, including the following specific steps:

[0266] Step 902: Obtain at least one user behavior data on the note sharing platform, wherein any user behavior data is used to characterize the user's interactive behavior event with respect to the note.

[0267] Step 904: Based on the interactive behavior attribute of at least one user behavior data, determine the business scenario to which each user behavior data belongs.

[0268] Step 906: Based on the user identifier of the at least one user behavior data, record the at least one user behavior data in a corresponding status list.

[0269] Step 908: When the user behavior data recorded in the status list is empty, set the trigger time of the status list.

[0270] Step 910: Sort the user behavior data based on the timestamps of the user behavior data recorded in the status list.

[0271] Step 912: based on the event identifiers of the user behavior data recorded in the status list, duplicate user behavior data are deduplicated.

[0272] Step 914: Based on the business scenarios to which the user behavior data belongs, the batches of user behavior data corresponding to each business scenario are summarized.

[0273] Step 916: When the preset trigger time of the status list is reached, the recorded user behavior data is taken out from the status list, and each batch of user behavior data is distributed to the message queue.

[0274] Step 918: Distribute the user behavior data to the downstream label construction operator of each business scenario according to the distribution rules corresponding to each business scenario.

[0275] Step 920: Take out the batch of user behavior data corresponding to the recorded business scenario from the message queue, and construct label data corresponding to the user behavior data.

[0276] Step 922: Based on the label data, train a note recommendation model for each business scenario.

[0277] In the embodiments of this specification, by preprocessing the user behavior data, the accuracy of the model training labels is significantly improved, ensuring that the user behavior sequences consumed downstream are all sequenced and deduplicated, avoiding the disorder and duplication problems caused by the transmission of data between different systems, which not only reduces the confusion of different business scenarios in understanding and using data, ensuring that the user's real behavior sequence is consistent with the user's perception and expectations, but also improves the accuracy of the note recommendation model and enhances the user experience. In addition, through data aggregation, the secondary processing of data in different downstream scenarios is avoided, reducing the possibility of ambiguity in the understanding and use of data in different business scenarios. Through the consistent distribution of business scenarios, the distribution efficiency and accuracy of the model training labels are significantly improved, ensuring that the user behavior sequences consumed downstream are all data-completed, and a piece of data is distributed to multiple business scenarios, ensuring the data consistency of multiple business scenarios, thereby improving the accuracy of the model and user satisfaction.

[0278] Corresponding to the above method embodiment, this specification also provides an embodiment of a training device for an object prediction model. Fig.10 FIG. 1 is a schematic diagram showing a structure of a training device for an object prediction model provided by an embodiment of the present specification. Fig.10 As shown, the device comprises:

[0279] The acquisition module 1002 is configured to acquire at least one user behavior data, wherein any user behavior data is used to characterize an interactive behavior event of a user with respect to an object;

[0280] The recording module 1004 is configured to record at least one user behavior data in a corresponding state list based on a user identifier of the at least one user behavior data;

[0281] The construction module 1006 is configured to retrieve the recorded first user behavior data from the first state list and construct first tag data corresponding to the first user behavior data when a preset trigger time of the first state list is reached, wherein the first state list is any state list;

[0282] The training module 1008 is configured to train the object prediction model based on the first label data.

[0283] Optionally, the device also includes: a preprocessing module, configured to set the trigger time of the first status list when the first user behavior data recorded in the first status list is empty; and preprocess the first user behavior data recorded in the first status list when the first user behavior data recorded in the first status list is not empty.

[0284] Optionally, the preprocessing module is further configured to: sort the first user behavior data based on the timestamp of the first user behavior data recorded in the first status list.

[0285] Optionally, the device further comprises: a priority queue module configured to add the retrieved first user behavior data to a priority queue, wherein the first user behavior data recorded in the priority queue is arranged according to timestamps;

[0286] Correspondingly, the construction module 1006 is further configured to: when the timestamp of the first user behavior data is less than the preset trigger time of the priority queue, take out the first user behavior data from the priority queue and construct the first label data corresponding to the first user behavior data.

[0287] Optionally, the device also includes: a first ordinary queue module, configured to transfer the first user behavior data from the priority queue to the ordinary queue when the timestamp of the first user behavior data is greater than the preset trigger time of the priority queue; when the preset trigger time of the ordinary queue is reached, return the first user behavior data recorded in the ordinary queue to the first status list and clear the ordinary queue.

[0288] Optionally, the preprocessing module is further configured to: deduplicate the duplicate first user behavior data based on event identifiers of the first user behavior data recorded in the first status list.

[0289] Optionally, the device also includes: a second general queue module, configured to add the first user behavior data to the general queue when the timestamp of the first user behavior data reaches a preset trigger time of the state list; if the general queue is not empty, based on the event identifier of the first user behavior data recorded in the general queue, remove the corresponding first user behavior data recorded in the first state list, and update the trigger time of the state list; if the general queue is empty, clear the first state list.

[0290] Optionally, the device further comprises: a scenario determination module configured to determine the business scenario to which each user behavior data belongs based on the interactive behavior attribute of at least one user behavior data;

[0291] Correspondingly, the training module 1008 is further configured to: train an object prediction model for a first business scenario based on the first label data, wherein the first business scenario is any business scenario.

[0292] Optionally, the device further includes: a summarizing module configured to summarize, based on the business scenarios to which the first user behavior data belongs, first user behavior data of batches corresponding to each business scenario;

[0293] Correspondingly, the device further includes: a first distribution module configured to distribute each batch of first user behavior data to a message queue;

[0294] Correspondingly, the construction module 1006 is further configured to: retrieve the first batch of user behavior data corresponding to the first business scenario recorded from the message queue, and construct first label data corresponding to the first user behavior data.

[0295] Optionally, the summary module is further configured to: if the event type of the first user behavior data recorded in the first status list is a basic event type, update the basic attribute information under the first business scenario, and identify the first business scenario as a business scenario with basic event information; if the event type of the first user behavior data recorded in the first status list is a click event type, check whether the trigger time of the first status list is updated, if not, update the trigger time of the first status list, and update the first user behavior data recorded in the first status list.

[0296] Optionally, the device further includes: a second distribution module configured to distribute the first user behavior data to a downstream label construction operator of each business scenario according to a distribution rule corresponding to each business scenario;

[0297] Correspondingly, the construction module 1006 is further configured to: execute the downstream label construction operator of each business scenario to construct the first label data corresponding to the first user behavior data.

[0298] Optionally, the second distribution module is further configured to: traverse the distribution rules corresponding to each business scenario; if the distribution rule of the second business scenario is hit, distribute the first user behavior data belonging to the second business scenario to the downstream label construction operator of the second business scenario, where the second business scenario is any business scenario.

[0299] Optionally, the device also includes: a signal control module, which is configured to traverse the signal control and determine whether it is necessary to adjust the time window of the downstream label construction operator if the distribution rule of any business scenario is not hit; if user behavior data is generated within the time window of the trigger time of the first state list, update the time window; if no user behavior data is generated within the time window of the trigger time of the first state list, clear the first state list.

[0300] Optionally, the device further includes: an advance distribution module configured to determine whether an advance distribution rule is met; if so, distribute the first user behavior data to downstream label construction operators of each business scenario;

[0301] Correspondingly, the construction module 1006 is further configured to: if not, retrieve the recorded first user behavior data from the first status list.

[0302] In the embodiments of the present specification, a status list constructed based on user identification is used to achieve orderly and efficient management of real-time streaming user behavior data, reduce the confusion and inconsistency caused by user behavior data from different sources, accurately and effectively construct label data, ensure that the constructed label data of user behavior data truly conforms to user perception, improve the accuracy of label data, and thereby train the object prediction model with high accuracy, thereby improving user experience.

[0303] The above is a schematic scheme of a training device for an object prediction model of this embodiment. It should be noted that the technical scheme of the training device for the object prediction model and the technical scheme of the training method for the object prediction model belong to the same concept, and the details not described in detail in the technical scheme of the training device for the object prediction model can be found in the description of the technical scheme of the training method for the object prediction model.

[0304] Fig.11 The structure block diagram of a computing device provided by an embodiment of the present specification is shown. The components of the computing device 1100 include but are not limited to a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and the database 1150 is used to store data.

[0305] The computing device 1100 also includes an access device 1140, which enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, and a Near Field Communication (NFC).

[0306] In one embodiment of the present specification, the above components of the computing device 1100 and Fig.11 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig.11 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0307] The computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1100 may also be a mobile or stationary server.

[0308] The processor 1120 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the training method of the above-mentioned object prediction model.

[0309] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the training method of the object prediction model described above belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the training method of the object prediction model described above.

[0310] An embodiment of the present specification also provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the training method of the above-mentioned object prediction model when executed by a processor.

[0311] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the training method of the object prediction model described above belong to the same concept, and the detailed objects not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the training method of the object prediction model described above.

[0312] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned object prediction model training method when executed by a processor.

[0313] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the training method of the object prediction model described above belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be found in the description of the technical scheme of the training method of the object prediction model described above.

[0314] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0315] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the contents contained in the computer readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer readable media do not include electric carrier signals and telecommunication signals.

[0316] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0317] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0318] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A training method for an object prediction model, characterized in that: include: Acquire at least one user behavior data, wherein any user behavior data is used to characterize an interactive behavior event of a user with respect to an object; Based on the user identifier of the at least one user behavior data, recording the at least one user behavior data in a corresponding status list; When a preset trigger time of a first state list is reached, taking out the recorded first user behavior data from the first state list, and constructing first tag data corresponding to the first user behavior data, wherein the first state list is any state list; Based on the first label data, an object prediction model is trained.

2. The method according to claim 1, characterized in that After recording the at least one user behavior data in a corresponding state list based on the user identification of the at least one user behavior data, the method further includes: When the first user behavior data recorded in the first status list is empty, setting a trigger time of the first status list; When the first user behavior data recorded in the first status list is not empty, preprocessing is performed on the first user behavior data recorded in the first status list.

3. The method according to claim 2, characterized in that The preprocessing of the first user behavior data recorded in the first status list includes: The first user behavior data is sorted based on the timestamp of the first user behavior data recorded in the first status list.

4. The method according to claim 3, characterized in that After taking out the recorded first user behavior data from the first status list, the method further includes: Adding the retrieved first user behavior data to a priority queue, wherein the first user behavior data recorded in the priority queue is arranged according to timestamps; The constructing the first label data corresponding to the first user behavior data includes: When the timestamp of the first user behavior data is less than the preset trigger time of the priority queue, the first user behavior data is taken out from the priority queue, and first label data corresponding to the first user behavior data is constructed.

5. The method according to claim 4, characterized in that After adding the retrieved first user behavior data to the priority queue, the method further includes: When the timestamp of the first user behavior data is greater than the preset trigger time of the priority queue, transferring the first user behavior data from the priority queue to the common queue; When a preset trigger time of a common queue is reached, the first user behavior data recorded in the common queue is returned to the first status list, and the common queue is cleared.

6. The method according to claim 2, characterized in that The preprocessing of the first user behavior data recorded in the first status list includes: Based on the event identifier of the first user behavior data recorded in the first status list, duplicate first user behavior data is deduplicated.

7. The method according to claim 6, characterized in that After taking out the recorded first user behavior data from the first status list, the method further includes: When the timestamp of the first user behavior data reaches a preset trigger time of the status list, adding the first user behavior data to a common queue; If the common queue is not empty, based on the event identifier of the first user behavior data recorded in the common queue, the corresponding first user behavior data recorded in the first status list is deleted, and the triggering time of the status list is updated; If the common queue is empty, clear the first status list.

8. The method according to any one of claims 1 to 7, characterized in that: Before recording the at least one user behavior data in a corresponding state list based on the user identification of the at least one user behavior data, the method further includes: Determining the business scenario to which each user behavior data belongs based on the interactive behavior attribute of the at least one user behavior data; The training of the object prediction model based on the first label data includes: Based on the first label data, an object prediction model for the first business scenario is trained, wherein the first business scenario is any business scenario.

9. The method according to claim 8, characterized in that After recording the at least one user behavior data in a corresponding state list based on the user identification of the at least one user behavior data, the method further includes: Based on the business scenarios to which the first user behavior data belongs, aggregating and obtaining batches of first user behavior data corresponding to each business scenario; After taking out the recorded first user behavior data from the first status list, the method further includes: Distribute the first user behavior data of each batch to the message queue; The constructing the first label data corresponding to the first user behavior data includes: The first batch of user behavior data corresponding to the first business scenario is retrieved from the message queue, and first label data corresponding to the first user behavior data is constructed.

10. The method according to claim 9, characterized in that The first user behavior data of a batch corresponding to each business scenario is obtained by aggregating the business scenarios to which the first user behavior data belongs, including: If the event type of the first user behavior data recorded in the first status list is a basic event type, updating the basic attribute information under the first business scenario, and identifying the first business scenario as a business scenario with basic event information; If the event type of the first user behavior data recorded in the first status list is a click event type, check whether the trigger time of the first status list is updated; if not, update the trigger time of the first status list and update the first user behavior data recorded in the first status list.

11. The method according to claim 8, characterized in that After taking out the recorded first user behavior data from the first status list, the method further includes: Distributing the first user behavior data to downstream label construction operators of each business scenario according to the distribution rules corresponding to each business scenario; The constructing the first label data corresponding to the first user behavior data includes: Execute the downstream label construction operator of each business scenario to construct the first label data corresponding to the first user behavior data.

12. The method according to claim 11, characterized in that The distributing the first user behavior data to the downstream label construction operator of each business scenario according to the distribution rules corresponding to each business scenario includes: Traversing the distribution rules corresponding to each business scenario; If the distribution rule of the second business scenario is hit, the first user behavior data belonging to the second business scenario is distributed to the downstream label construction operator of the second business scenario, wherein the second business scenario is any business scenario.

13. The method according to claim 12, characterized in that After traversing the distribution rules corresponding to each business scenario, the method further includes: If the distribution rule of any business scenario is not matched, the signal control is traversed to determine whether the time window of the downstream label construction operator needs to be adjusted; If user behavior data is generated within the time window reaching the trigger time of the first state list, updating the time window; If no user behavior data is generated within the time window when the trigger time of the first status list is reached, the first status list is cleared.

14. The method according to claim 12, characterized in that Before taking out the recorded first user behavior data from the first status list, the method further includes: Determine whether the advance distribution rule is met; If so, distributing the first user behavior data to the downstream label construction operators of each business scenario; The step of retrieving the recorded first user behavior data from the first status list includes: If not, the recorded first user behavior data is retrieved from the first status list.

15. A training device for an object prediction model, characterized in that: include: An acquisition module is configured to acquire at least one user behavior data, wherein any user behavior data is used to characterize an interactive behavior event of a user with respect to an object; a recording module configured to record the at least one user behavior data in a corresponding state list based on a user identifier of the at least one user behavior data; a construction module, configured to, when a preset trigger time of a first state list is reached, take out the recorded first user behavior data from the first state list and construct first tag data corresponding to the first user behavior data, wherein the first state list is any state list; The training module is configured to train an object prediction model based on the first label data.

16. A computing device, characterized in that: include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 14 are implemented.

17. A computer-readable storage medium, characterized in that: It stores a computer program / instruction, which implements the steps of the method described in any one of claims 1 to 14 when executed by a processor.

18. A computer program product, characterized in that The method comprises a computer program / instruction which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 14.