Model training method and device, behavior coding method and device, equipment and medium
Through the execution time and categories of coding behaviors, combined with time representation and category representation, the data representation model is trained, which solves the problem of unevenness of behavior sequence applications in the game field, and improves the prediction effect.
Patent Information
- Application Number
- CN202410040879.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-11
AI Technical Summary
Existing natural language models cannot be effectively applied to behavior sequences in the gaming domain because behavior execution time intervals are uneven and the contextual relationships of behavior execution time cannot be fully utilized.
By encoding the execution time and category of behavior, combining time representation and category representation, the data representation model is trained, considering the context relationship of behavior and the context relationship of execution time, and using a weighted mask strategy and a Transformer architecture for model training.
Improves the accuracy of the model's behavior prediction effect in the gaming field, especially in social recommendations, product recommendations, and game matching tasks.
Smart Images

Figure CN120296405A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and particularly relates to a model training method, a behavior encoding method, a device, a device and a medium. Background Art
[0002] In the game field, an object will generate rich behavior sequences during the daily game process. For example, a behavior sequence includes logging in to the game, logging out of the game, purchasing items, playing games, etc.
[0003] However, the natural language model provided by the related technology assumes that the positions between the behavior tokens (basic input units of the model) of the behavior sequence are evenly distributed, and uses an incrementing method for position encoding. However, in the behavior sequence in the game field, the intervals of the execution times of the behaviors are usually uneven, and the natural language model of the related technology cannot be well applied to the behavior sequence in the game field. Summary of the Invention
[0004] The present application provides a model training method, a behavior encoding method, a device, a device and a medium. The present application considers the context relationship of the behavior execution time, can make full use of the behavior execution time information, and the trained model will have a better prediction effect. The technical solutions include the following contents.
[0005] According to one aspect of the present application, a model training method is provided, and the method includes the following steps.
[0006] Obtain a plurality of object behavior sequences, and any one of the plurality of object behavior sequences includes at least two behaviors executed by an object in an application program in chronological order;
[0007] In a data representation model, for each object behavior sequence in the plurality of object behavior sequences, obtain the execution time of each of the plurality of behaviors in each object behavior sequence;
[0008] Encode the execution time of each of the plurality of behaviors to obtain a plurality of time representations; and, encode the categories of each of the plurality of behaviors to obtain a plurality of category representations; combine the plurality of time representations and the plurality of category representations according to the belonging behaviors to obtain the behavior representations of each of the plurality of behaviors;
[0009] Based on the behavior representations of each of the plurality of behaviors in each object behavior sequence, train a data representation model, and the data representation model is used to output the behavior representations of behaviors.
[0010] According to one aspect of the present application, a behavior encoding method is provided, and the method includes the following steps.
[0011] Obtain a target object behavior sequence, where the target object behavior sequence includes at least two behaviors executed by the object in the application in chronological order;
[0012] Obtain the execution time of each of the multiple behaviors in the target object behavior sequence;
[0013] Encode the execution time of each of the multiple behaviors to obtain multiple time representations; and, encode the category of each of the multiple behaviors to obtain multiple category representations; combine the multiple time representations and the multiple category representations according to the behavior to which they belong to obtain the behavior representation of each of the multiple behaviors.
[0014] According to another aspect of the present application, there is provided a model training device, and the device includes the following modules.
[0015] An acquisition module, configured to acquire multiple object behavior sequences, where any one of the multiple object behavior sequences includes at least two behaviors executed by the object in the application in chronological order;
[0016] The acquisition module is further configured to, in the data representation model, for each object behavior sequence among the multiple object behavior sequences, acquire the execution time of each of the multiple behaviors in each object behavior sequence;
[0017] An encoding module, configured to encode the execution time of each of the multiple behaviors to obtain multiple time representations; and, encode the category of each of the multiple behaviors to obtain multiple category representations;
[0018] A processing module, configured to combine the multiple time representations and the multiple category representations according to the behavior to which they belong to obtain the behavior representation of each of the multiple behaviors;
[0019] A training module, configured to train the data representation model based on the behavior representations of each of the multiple behaviors in each object behavior sequence, and the data representation model is used to output the behavior representation of the behavior.
[0020] According to another aspect of the present application, there is provided a behavior encoding device, and the device includes the following modules.
[0021] An acquisition module, configured to acquire a target object behavior sequence, where the target object behavior sequence includes at least two behaviors executed by the object in the application in chronological order;
[0022] The acquisition module is further configured to acquire the execution time of each of the multiple behaviors in the target object behavior sequence;
[0023] An encoding module, configured to encode the execution time of each of the multiple behaviors to obtain multiple time representations; and, encode the category of each of the multiple behaviors to obtain multiple category representations;
[0024] A processing module, configured to combine multiple time representations and multiple category representations according to the belonging behaviors, so as to obtain behavior representations of respective behaviors.
[0025] According to one aspect of the present application, there is provided a computer device, which includes: a processor and a memory. The memory stores a computer program, and the computer program is loaded and executed by the processor to implement the above model training method.
[0026] According to another aspect of the present application, there is provided a computer-readable storage medium, which stores a computer program, and the computer program is loaded and executed by the processor to implement the above model training method or behavior encoding method.
[0027] According to another aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above model training method or behavior encoding method.
[0028] The beneficial effects brought by the technical solutions provided by the embodiments of the present application at least include the following content.
[0029] By encoding the execution time of the behavior, the time representation of the behavior is obtained, and then the time representation of the behavior is combined with the category representation (such as ID representation) to obtain the behavior representation of the behavior.
[0030] In the present application, not only the context relationship of the behavior is considered by encoding the categories of multiple behaviors in the object behavior sequence, but also the context relationship of the execution time of the behavior is considered by encoding the execution time of multiple behaviors in the object behavior sequence. Furthermore, better behavior representations are obtained. When the behavior representations are applied to downstream tasks (such as social recommendation, commodity recommendation, game matching), better prediction effects will be achieved. For example, when the behavior representations are applied to the task of commodity recommendation in a game, virtual commodities more suitable for the object will be recommended. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1 It is a schematic diagram of the principle of the model training method provided by an embodiment of the present application.
[0033] Figure 2 It is a flowchart of a model training method provided by an embodiment of the present application.
[0034] Figure 3 It is a flowchart of a model training method provided by an embodiment of the present application.
[0035] Figure 4 It is a flowchart of a model training method provided by an embodiment of the present application.
[0036] Figure 5 It is a schematic diagram of a model training method provided by an embodiment of the present application.
[0037] Figure 6 It is a schematic diagram of a model training method provided by an embodiment of the present application.
[0038] Figure 7 It is a schematic diagram of positional encoding performed in the related art and the present application.
[0039] Figure 8 It is a schematic diagram of some stages in the model training process provided by an embodiment of the present application.
[0040] Figure 9 It is a flowchart of a behavior encoding method provided by an embodiment of the present application.
[0041] Figure 10 It is a structural block diagram of a model training device provided by an embodiment of the present application.
[0042] Figure 11 It is a structural block diagram of a behavior encoding device provided by an embodiment of the present application.
[0043] Figure 12 It is a structural block diagram of a computer device provided by an embodiment of the present application.
[0044] Figure 13 It is a structural block diagram of a computer device provided by another embodiment of the present application. Detailed implementation manners
[0045] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0046] First, a brief introduction to the nouns involved in the embodiments of the present application will be given.
[0047] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0048] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0049] Serialized log data is an important data format for characterizing players in games. Game players will generate rich serialized log data during their daily gaming processes, such as logins, logouts, purchases, matches, etc. Based on the serialized log data, game developers can build a series of data science applications, such as player matching, social recommendations, product recommendations, etc., to improve the player gaming experience.
[0050] Self-supervised learning is a machine learning method whose main idea is to automatically generate labels from unlabeled data and then use these generated labels to train the model. Compared with traditional supervised learning, self-supervised learning does not require a large amount of manual annotation of data, but learns by leveraging the inherent structure or features of the data itself. For example, in the field of natural language processing, the text can be randomly masked in the form of masked learning to learn the masked text and thus mine the context features of the text sequence.
[0051] Fine-tuning is a transfer learning technique used to further train a specific task based on a pre-trained model. In fine-tuning, we use a model pre-trained on a large-scale dataset, usually trained on similar but not exactly the same tasks. The benefit of fine-tuning is that it can utilize the general features learned by the pre-trained model on large-scale data, thereby training on a relatively small specific task dataset, reducing the need for a large amount of labeled data. This can improve the generalization ability and performance of the model, and also achieve good results with limited resources.
[0052] The models provided by related technologies are mainly pre-trained for text sequences of natural language and cannot be directly applied to the scenario of pre-training serialized log data. The models of related technologies have the following disadvantages.
[0053] (1) Temporal insensitivity. Traditional language models assume that the positions between tokens are uniformly distributed and use an incremental method for position encoding. However, in serialized log data, the time intervals between behaviors are usually non-uniform, and time information plays an important role in serialized log data, implying a lot of player information. For example, if a long time has passed from the start to the end of a game, it implies that the player has experienced an exciting game. Logging in to the game very late often implies that the user may work overtime severely and only has time to relax late at night.
[0054] (2) Uniform masking strategy. Existing large language models such as BERT (Bidirectional Encoder Representation from Transformers, pre-trained language representation model) usually adopt a uniform masking strategy. The behavior distribution of users in the game is usually unbalanced. For example, behaviors such as game matches and match settlements are usually more frequent, while behaviors such as recharges are usually less frequent. At the same time, the behaviors of users in the game are also divided into active and passive. Among them, passive behaviors mainly include task settlements, match settlements, etc., and active behaviors mainly include logins, matches, purchases, etc. Since the active behaviors of users represent the intentions of users, active behaviors are more important in the game. However, the masking method of the large model is random and does not consider the frequency and importance of the occurrence of behaviors, which results in the model being able to effectively learn high-frequency behaviors, but it is difficult to learn low-frequency behaviors.
[0055] (3) Coarse-grained representation method. The large language models provided by related technologies only encode behavior IDs. However, in the game scenario, the descriptions behind the behavior IDs are also very important, revealing more fine-grained features of the behaviors.
[0056] Figure 1The figure shows a schematic diagram of the principle of the model training method provided by an exemplary embodiment of the present application. Figure 1 The left half shows the model training framework. Figure 1 The left half shows the pre-training stage and the fine-tuning stage of the data representation model 101. The data representation model 101 is pre-trained through multiple object behavior sequences 102. After pre-training, the data representation model 101 is fine-tuned on the datasets of each target prediction task 103 to obtain a data representation model applicable to each target prediction task.
[0057] The object behavior sequence 102 includes at least two behaviors executed by an object in the application in chronological order. For example, if the application is a game application, the object behavior sequence includes at least two game behaviors executed by the player in the game in chronological order. The object behavior sequence is the behavior sequence recorded in the game log, and the object behavior sequence includes "log in to the game, play a game, add game friends, recharge, log out of the game". The data representation model 101 is used to encode behaviors to obtain the behavior representations of the behaviors. The target prediction task 103 can be tasks such as player matching, social recommendation, and product recommendation in a game application. According to the pre-trained general model, model fine-tuning is performed in each downstream task, so as to achieve representation migration on different tasks and obtain a model applicable to the downstream tasks.
[0058] Figure 1 The right half shows the pre-training stage of the data representation model 101. In the pre-training stage, for an object behavior sequence that includes at least two behaviors, each behavior in the object behavior sequence is weighted and masked to obtain a masked behavior sequence.
[0059] Combined with reference Figure 1 to the right half of Figure 1 The right half shows an object behavior sequence 104. The object behavior sequence 104 includes behaviors A to E. According to the occurrence frequency of each behavior, the masking probability of each behavior is calculated. For example, if behavior A is a behavior with a higher occurrence frequency, the masking probability of behavior A is smaller; if behavior B is a behavior with a lower occurrence frequency, the masking probability of behavior B is larger. Figure 1 The right half shows the masked behavior sequence 105. At this time, behaviors B and D are masked.
[0060] For the unmasked behaviors (i.e., behaviors A, C, and E), time encoding is performed respectively to obtain the respective time representations 106 of the masked behavior sequence. Time encoding refers to encoding the execution time of behaviors.
[0061] Perform categorical encoding on the unmasked behaviors respectively to obtain the respective categorical representations 107 of the behavior sequence after masking. The category can be, optionally, the behavior ID. For example, the ID of the behavior "log in to the game" is 0001, and the ID of the behavior "log out of the game" is 1111.
[0062] Perform text encoding on the description texts of the unmasked behaviors respectively to obtain the respective categorical representations 108 of the behavior sequence after masking. The description text is the specific introduction content behind the behavior. For example, if the behavior is "obtain a vehicle", the description text is "obtain a car with high moving speed and more remaining fuel".
[0063] In an optional embodiment, perform time encoding on the behavior sequence 105 after masking, input the encoded sequence into the Transformer architecture for categorical encoding, and fuse the result of time encoding and the result of categorical encoding in the Transformer architecture to obtain an intermediate representation. Through a large language model (such as LLAMA) provided by related technologies, perform text encoding on the description text of the unmasked behavior to obtain a text representation. Fuse the text representation and the intermediate representation to obtain the final behavior representation.
[0064] For a behavior, combine the time representation, categorical representation, and text representation of the behavior to obtain the behavior representation of the behavior. Figure 1 The right half shows the behavior representation A (the behavior representation of behavior A), the behavior representation C (the behavior representation of behavior C), and the behavior representation E (the behavior representation of behavior E).
[0065] Predict each masked behavior according to the behavior representations of multiple unmasked behaviors in the behavior sequence after masking. Figure 1 shows predicting behavior B and behavior D according to the behavior representation of behavior A, the behavior representation of behavior C, and the behavior representation of behavior E. Specifically, predict the categories (behavior IDs) of behavior B and behavior D.
[0066] Train the data representation model based on the loss constructed based on the prediction accuracy.
[0067] In summary, the large model technologies provided by related art (such as ChatGPT, LLaMa, ChatGLM, etc.) are architectures designed for text sequences. Due to the lack of consideration for the characteristics of serialized log data, they cannot be directly applied to scenarios such as the gaming field. This application introduces a training framework for large models (pre-training + fine-tuning). By pre-training a general data representation model and then fine-tuning it on the datasets of various target prediction tasks, the dataset for pre-training is a large-scale dataset, and the dataset for fine-tuning is a small-scale dataset for specific tasks. Overall, it reduces the need for a large amount of data, and can improve the generalization ability and effect of the data representation model, and can also achieve good results under limited resources.
[0068] Moreover, in the above pre-training stage, time encoding is performed on the execution time of behaviors. The timestamp information is added to the position encoding method, so that the position encoding can well reflect the time information of behaviors.
[0069] Moreover, in the above pre-training stage, weighted masking is performed on the occurrence frequency of behaviors. Weighted masking is performed on the occurrence frequency of each behavior, so that the model adopts different masking strategies for different behaviors. Behaviors with more occurrences have a smaller probability of being masked, and behaviors with fewer occurrences have a larger probability of being masked.
[0070] Moreover, in the above pre-training stage, text encoding is performed on the description text of behaviors, revealing finer-grained features of behaviors, and finally obtaining better behavior representations.
[0071] In some embodiments, Figure 1 the shown model training process is executed by a computer device. The computer device can be a terminal device or a server. The device types of the terminal device include at least one of smart phones, smart watches, in-vehicle terminals, wearable devices, smart TVs, tablets, e-book readers, MP3 players, MP4 players, laptop computers, and desktop computers. The terminal device includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, in-vehicle terminals, aircraft, etc.
[0072] In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0073] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the execution time, category and description text of the behaviors involved in this application are obtained under full authorization.
[0074] Moreover, regarding the relevant information, the relevant information processors will follow the principles of legality, legitimacy and necessity, clarify the purpose, method and scope of relevant information processing, obtain the consent of the relevant information subjects, and take necessary technical and organizational measures to ensure the security of relevant information.
[0075] Figure 2 The flowchart of the model training method provided by an exemplary embodiment of this application is shown. Taking the method as being executed by a computer device as an example, the method includes:
[0076] Step 210, obtain multiple object behavior sequences, where any one of the multiple object behavior sequences includes at least two behaviors executed by an object in an application program in chronological order;
[0077] The object behavior sequence includes at least two behaviors executed by an object in an application program in chronological order. Schematically, the application program is a game application program, and the object behavior sequence is serialized log data. For example, the object behavior sequence includes "log in to the game, play a game session, add game friends, recharge, log out of the game".
[0078] Schematically, the application program is a payment application program, and the object behavior sequence is serialized log data. For example, the object behavior sequence includes "take the subway, pay for breakfast at xx breakfast shop, pay for dinner at xx restaurant, pay for coffee at xx coffee shop, take the subway".
[0079] Schematically, the application program is a music playback application program, and the object behavior sequence is serialized log data. For example, the object behavior sequence includes "listen to traffic radio station FM, listen to personal favorite songs, listen to recommended songs, listen with friends".
[0080] Step 220, in the data representation model, for each object behavior sequence among the multiple object behavior sequences, obtain the execution time of each of the multiple behaviors in each object behavior sequence;
[0081] The data representation model is used to encode behaviors to obtain behavior representations. In the embodiments of this application, the execution time of behaviors will be encoded to obtain time representations.
[0082] For example, for a sequence of object behaviors, the sequence of object behaviors includes 5 behaviors: "log in to the game, play a game, add a game friend, recharge, log out of the game". Player A logged in to the game at 19:10, started a game at 19:12, added player B as a game friend at 19:25, player A recharged game vouchers at 19:30, and player A logged out of the game at 19:35.
[0083] In one embodiment, multiple behaviors in each sequence of object behaviors are obtained. For the first behavior, the relative execution time of the first behavior and the initial behavior in the object behavior sequence where it is located is obtained, and the first behavior is any one of the multiple behaviors.
[0084] For example, continuing with the above example, the relative execution time of the behavior "log in to the game" is 0, the relative execution time of "play a game" is 2, the relative execution time of "add a game friend" is 15, the relative execution time of "recharge" is 20, and the relative execution time of "log out of the game" is 25.
[0085] Step 230, encode the execution time of each of the multiple behaviors to obtain multiple time representations;
[0086] Encode the execution time of each of the multiple behaviors in the object behavior sequence to obtain the time representation of each behavior. In one embodiment, encode the relative execution time of each of the multiple behaviors in the object behavior sequence to obtain the time representation of each behavior.
[0087] In one embodiment, the time representation is a high-dimensional vector. Based on the relative execution time of the first behavior and in combination with the dimension label of the time representation, the time representation of the first behavior is obtained, and the first behavior is any one of the multiple behaviors. The time representation being a high-dimensional vector enables the time representation to cover more fine-grained information, which is not only beneficial for improving the prediction ability of the model but also for enhancing the prediction effect of the model.
[0088] Schematically, for the 2i-th dimension of the time representation of the first behavior, based on the relative execution time of the first behavior and the dimension label (2i) of the 2i-th dimension, the value of the 2i-th dimension is obtained through a sine function, where i is a non-negative integer. It is represented by the formula as follows:
[0089]
[0090]
[0091] where, t j represents the execution time of the j-th behavior, Represents the time elapsed for the nth behavior relative to the initial behavior. Scale represents the log function for scaling the time; sin represents the sine function. D is a scalar (optionally 10,000), and d represents the number of dimensions of the time representation (optionally 256). TE(t n ,2i) represents the value of the 2i-th dimension of the time representation of the nth behavior.
[0092] Schematically, for the (2i + 1)-th dimension of the time representation of the first behavior, based on the relative execution time of the first behavior and the dimension label of the 2i-th dimension, the value of the (2i + 1)-th dimension is obtained through the cosine function. It is represented by the formula as follows:
[0093]
[0094] Where cos represents the cosine function, D is a scalar (optionally 10,000), and d is the number of dimensions of the time representation (optionally 256). TE(t n ,2i + 1) represents the value of the (2i + 1)-th dimension of the time representation of the nth behavior.
[0095] Concatenate each dimension in the dimension order to obtain the time representation of the first behavior.
[0096] It can be understood that by encoding the even dimensions according to the above sine formula and the odd dimensions according to the above cosine formula, the product-to-sum and sum-to-product properties of the sine and cosine functions can be utilized. When encoding the time of a certain dimension, the results of the time encoding of other dimensions can be used.
[0097] Step 240: Encode the categories of multiple behaviors respectively to obtain multiple category representations;
[0098] The category, optionally, is represented by the behavior ID. For example, the ID of the behavior "log in to the game" is 0001, and the ID of the behavior "log out of the game" is 1111.
[0099] In one embodiment, the categories of multiple behaviors are encoded through the Transformer architecture to obtain multiple category representations.
[0100] Step 250: Combine multiple time representations and multiple category representations according to the belonging behaviors to obtain the behavior representations of multiple behaviors respectively;
[0101] For one behavior, combine the time representation and the category representation of this behavior (optionally, by numerical addition) to obtain the behavior representation of this behavior.
[0102] It can be understood that the present application does not adopt the uniform position encoding in the related art, but additionally designs a time-aware position encoding. In the Transformer architecture, the result of the additional time encoding is combined with the result of the category encoding (optionally, by numerical addition) to obtain the behavior representation.
[0103] Step 260: Based on the behavior representations of multiple behaviors in each object behavior sequence, train a data representation model, where the data representation model is used to output the behavior representation of a behavior.
[0104] In one embodiment, the data representation model is trained in a supervised learning manner. For example, based on the behavior representations obtained above, continue to perform downstream prediction tasks, and train the data representation model according to the error between the result of the prediction task and the label. For example, perform a commodity recommendation task according to the behavior representation, and train the data representation model according to the error between the predicted commodity and the labeled commodity.
[0105] In one embodiment, the data representation model is trained in a self-supervised manner. For example, a masking strategy is used for model training. Optionally, the masking strategy is a uniform masking strategy or a weighted masking strategy. The uniform masking strategy will perform a masking operation or not perform a masking operation on each behavior with the same masking probability. For the weighted masking strategy, weights will be assigned to each behavior, and each behavior may perform a masking operation or not perform a masking operation with different masking probabilities (which will be introduced in detail below).
[0106] In summary, by encoding the execution time of a behavior, the time representation of the behavior is obtained, and then the time representation of the behavior is combined with the category representation (such as the ID representation) to obtain the behavior representation of the behavior.
[0107] That is, the present application performs time encoding for the execution time of a behavior. While considering the context relationship of the behavior, it also considers the context relationship of the execution time of the behavior, and thus a better behavior representation can be obtained. When the behavior representation is applied to downstream tasks (such as social recommendation, commodity recommendation, game matching), it will have a better prediction effect. For example, when the behavior representation is applied to the task of commodity recommendation in a game, more suitable virtual commodities for the object will be recommended.
[0108] Next, the content of weighted masking & masking prediction will be introduced.
[0109] Based on Figure 2 In the optional embodiment shown, multiple behaviors are replaced by multiple unmasked behaviors. Figure 3 FIG. shows a flowchart of a model training method according to an exemplary embodiment of the present application. Before step 220, steps 310 to 330 are further included; step 260 can be replaced by step 340 and step 350.
[0110] Step 310, in the data representation model, mask each object behavior sequence to obtain a masked behavior sequence, which includes multiple unmasked behaviors and multiple masked behaviors;
[0111] Illustratively, the object behavior sequence includes five behaviors: "logging in to the game, playing a game, adding a game friend, recharging, and logging out of the game". After masking, the behaviors "playing a game" and "recharging" are masked, and during the subsequent behavior encoding process, the behaviors "playing a game" and "recharging" will not be encoded.
[0112] Step 320, during the masking process, calculate the masking probability of the second behavior based on the occurrence frequency of the second behavior, where the second behavior is any behavior in each object behavior sequence;
[0113] For an object behavior sequence, calculate the masking probability of each behavior in the object behavior sequence, and determine whether to mask the behavior according to its respective masking probability.
[0114] In one embodiment, the occurrence frequency is at least one of the inverse term frequency (ITF) and the inverse document frequency (IDF). Optionally, the occurrence frequency is the product of the inverse term frequency and the inverse document frequency.
[0115] The inverse term frequency represents the reciprocal of the frequency of the second behavior among the multiple behaviors included in the multiple object behavior sequences. For example, if the multiple object behavior sequences contain a total of 10 behaviors and the second behavior appears 3 times, the inverse term frequency is 10 / 3. The inverse document frequency represents the reciprocal of the occurrence frequency of the second behavior in the multiple object behavior sequences. For example, if the second behavior appears in 2 out of 10 object behavior sequences, the inverse document frequency is 5.
[0116] Step 330, perform a masking operation or not on the second behavior according to the masking probability of the second behavior;
[0117] Illustratively, the masking probability of the behavior "logging in to the game" is 0.3. Obtain a random number within the range of 0 to 1. If the random number is less than 0.3, perform a masking operation on the behavior "logging in to the game".
[0118] Step 340, predict the categories of the multiple masked behaviors based on the multiple behavior representations of the multiple unmasked behaviors; construct a prediction loss based on the prediction accuracy;
[0119] For each masked behavior in the masked behavior sequence, the class probability that the class of each masked behavior is the true class is predicted through the behavior representations of multiple unmasked behaviors. If the prediction is correct, the class probability is involved in the accumulation; if the prediction is incorrect, the class probability is discarded; the accumulation result of the multiple class probabilities corresponding to the multiple masked behaviors in the masked behavior sequence is calculated, and the multiple class probabilities correspond to the multiple masked behaviors one by one; based on the accumulation result, a prediction loss is constructed.
[0120] When making a masked prediction, the masked behavior will be predicted through the unmasked behaviors.
[0121] Exemplarily, if the probability that the class of a certain masked behavior is class A is 20%, the probability of class B is 30%, and the probability of class C is 50%, then the prediction result is that the class of this masked behavior is class C (with the highest probability). If the true class (label) of this masked behavior is class C, then 50% is involved in the probability accumulation.
[0122] For another example, if the probability that the class of another masked behavior is class A is 10%, the probability of class B is 30%, and the probability of class C is 80%, then the prediction result is that the class of this masked behavior is class C (with the highest probability). If the true class (label) of this masked behavior is class A, then 80% is discarded.
[0123] The multiple class probabilities corresponding to the multiple masked behaviors are accumulated to obtain an accumulation result; the accumulation result is negated to construct a prediction loss.
[0124] Schematically, the negative log-likelihood function is used to construct the loss, which is expressed by the formula:
[0125]
[0126] where L MLM represents the loss value, V represents the set of masked behaviors included in an object behavior sequence, p i represents the probability that the class of the masked behavior is the true class. If the prediction is correct, y i takes the value of 0. If the prediction is incorrect, y i takes the value of 1.
[0127] Step 350, based on the prediction loss, train the data representation model.
[0128] Based on the prediction loss, adjust the network parameters in the data representation model until the loss value of the prediction loss is less than the loss threshold, or until the number of training times reaches the number threshold.
[0129] In summary, in the above embodiments, weighted masking is performed on the occurrence frequency of behaviors, that is, the model adopts different masking strategies for behaviors with different frequencies, which is beneficial for the model to effectively learn both high-frequency behaviors and low-frequency behaviors, and finally can improve the generalization ability of the model.
[0130] Based on Figure 3 In the optional embodiment shown, for step 320, the masking probability of the second behavior can be obtained through the following calculation process.
[0131] I. Calculate the occurrence frequency of each of the multiple behaviors in the behavior set. Optionally, the occurrence frequency is the product of the inverse term frequency and the inverse document frequency. This makes behaviors with more occurrences have a smaller masking probability, and behaviors with fewer occurrences have a larger masking probability.
[0132] The behavior set is a set of multiple behaviors included in multiple object behavior sequences.
[0133] Schematically, the masking probability of the second behavior is calculated using the following formula.
[0134]
[0135]
[0136] Among them, x i (i.e., itf i *idf i ) represents the occurrence frequency of the second behavior, and I is the behavior set.
[0137] II. Based on the ratio of the occurrence frequency of each of the multiple behaviors in the behavior set to the temperature coefficient, obtain their respective updated occurrence frequencies;
[0138] With reference to the above formula, T represents the temperature coefficient, which is used to balance the numerical value of the final masking probability and avoid the value being too large or too small.
[0139] III. Accumulate the updated occurrence frequencies of each of the multiple behaviors to obtain the total occurrence frequency.
[0140] With reference to the above formula, represents the total occurrence frequency, represents the updated occurrence frequency of the second behavior.
[0141] IV. Based on the ratio of the occurrence frequency of the second behavior to the total occurrence frequency, calculate the masking probability of the second behavior. The second behavior is one of the behaviors in the behavior set.
[0142] With reference to the above formula, Softmax(x i ,T) represents the masking probability of the second behavior.
[0143] It is understandable that when considering the masking probability of a single behavior, the ratio to the total occurrence frequency will be calculated. That is, the final masking probability of a single behavior obtained is a relative value, and the expression meaning of the masking probability is more accurate, rather than simply calculating for a single behavior alone.
[0144] 5. Based on the masking probability of the second behavior, combined with the length of the object behavior sequence where the second behavior is located and the preset masking ratio, obtain the updated masking probability of the second behavior.
[0145] Combined with the above formula, the length of the object behavior sequence where the second behavior is located is l, and the preset masking ratio is β. The preset masking ratio represents the proportion of behaviors masked in multiple object behavior sequences, such as 30%. In the formula, the sequence length l and the preset masking ratio are also considered to ensure that the proportion of behaviors masked in each object behavior sequence is the same.
[0146] Based on Figure 2 In the optional embodiment shown, Figure 4 The flowchart of the model training method shown in an exemplary embodiment of the present application is shown. Before step 260, steps 410 and 420 are also included.
[0147] Step 410, encode the description texts of multiple behaviors respectively to obtain multiple text representations;
[0148] The description text is the specific descriptive text behind the behavior. For example, if the behavior is "obtain a vehicle", the description text of the behavior "obtain a vehicle" is "obtain a car with high moving speed and more remaining fuel". For example, if the behavior is "end of the game", the description text of the behavior "end of the game" is "the game fails but breaks the record".
[0149] In one embodiment, the description text is encoded by a large language model provided by related technologies to obtain a text representation. For example, the description text is encoded by a pre-trained LLAMA model.
[0150] Step 420, combine the behavior representations of multiple behaviors with their respective text representations respectively to obtain the updated behavior representations of multiple behaviors.
[0151] Optionally, add and fuse the behavior representation and the text representation to obtain the updated behavior representation.
[0152] In summary, the above embodiments also perform text encoding on the descriptive text of behaviors, revealing more fine-grained information of behaviors, and thus obtaining better behavior representations. When the behavior representations are applied to downstream tasks (such as social recommendation, product recommendation, game matching), better prediction effects will be achieved. For example, when the behavior representation is applied to the task of product recommendation in a game, virtual products more suitable for the object will be recommended.
[0153] Figure 5 FIG. shows a schematic diagram of a model training framework provided by an exemplary embodiment of the present application. In the model training framework of the present application, pre-training will be performed based on the serialized log data inside the game, and the pre-trained model will be fine-tuned. The fine-tuned model can be applied to multiple game science downstream tasks, such as social recommendation, game matching, product recommendation, etc.
[0154] Schematically, according to the serialized log data of a large number of games, a general data representation model is pre-trained. Then, according to the small dataset of the "social recommendation" task, the general data representation model is fine-tuned to obtain a data representation model applied to the "social recommendation" task.
[0155] Schematically, according to the serialized log data of a large number of games, a general data representation model is pre-trained. Then, according to the small dataset of the "game matching" task, the general data representation model is fine-tuned to obtain a data representation model applied to the "game matching" task.
[0156] Schematically, according to the serialized log data of a large number of games, a general data representation model is pre-trained. Then, according to the small dataset of the "product recommendation" task, the general data representation model is fine-tuned to obtain a data representation model applied to the "product recommendation" task.
[0157] Figure 6 FIG. shows a schematic diagram of a model training framework provided by an exemplary embodiment of the present application.
[0158] In the data collection stage, serialized log data 610 and text sequences 620 are collected. The present application will collect the logs of players inside the game, obtain the serialized log data 610 of multiple players, and sort the multiple serialized log data 610 according to the time sequence and player ID. For one player, the serialized log data 610 of the player and the corresponding text sequence 620 will be processed separately.
[0159] In the behavior ID representation learning stage 630, a weighted masking strategy 631, time-aware position encoding 632, and Transformer encoding 633 will be sequentially executed to obtain a behavior ID representation 634.
[0160] The weighted masking strategy 631 will mask different behaviors with different probabilities according to the frequency of occurrence of the behaviors. To take into account the impact of the uneven distribution of behaviors on model learning, this application adopts a weighted masking strategy. For behaviors with fewer occurrences, the masking probability will be greater; for behaviors with more occurrences, the masking probability will be smaller. The way of mask calculation adopts the following formula:
[0161]
[0162]
[0163] Among them, itf represents the inverse term frequency, idf represents the inverse document frequency, T represents the temperature coefficient (used to balance the size of the masking probability), β represents the proportion of tokens that are masked in the overall, L represents the length of the sequence, I represents the set of behavior tokens, i represents the i-th behavior token, represents the masking probability; the Softmax(x i ,T) function is used for probability calculation. The denominator represents the sum of all behavior tokens, and the numerator is the itf*idf of the i-th behavior token.
[0164] Time-aware positional encoding 632, after weighted masking, the masked serialized log data is obtained. The time information of the behavior is input into the sequence for positional encoding to obtain the sequence after time encoding. The positional information encoded by the traditional Transformer architecture is as shown in part (A) of Figure 7 , Figure 7 part (A) of which shows uniform positional information, losing the representation of the time when the behavior occurs. After modification, this application uses relative positional information to embed the time representation, as shown in part (B) of Figure 7 .
[0165] The representation form of time encoding is shown in the following formula:
[0166]
[0167]
[0168]
[0169] Among them, t j represents the execution time of the j-th behavior, represents the time elapsed from the n-th behavior to the initial behavior, Scale represents the log function, which scales the time; sin represents the sine function, D is a scalar (optionally, 10000), and d represents the number of dimensions of the time representation (optionally, 256). TE(tn , where \(t_{n,2i}\) represents the value of the \(2i\)-th dimension of the temporal representation of the \(n\)-th behavior; \(\cos\) represents the cosine function, and \(TE(t n , 2i + 1)\) represents the value of the \((2i + 1)\)-th dimension of the temporal representation of the \(n\)-th behavior.
[0170] Concatenate each dimension in the dimension order to obtain the temporal representation of the \(n\)-th behavior.
[0171] It can be understood that by encoding the even dimensions according to the above sine formula and the odd dimensions according to the above cosine formula, the product-to-sum and sum-to-product properties of the sine and cosine functions can be utilized. When encoding a certain dimension in terms of time, the results of the temporal encodings of other dimensions can be utilized.
[0172] Transformer Encoding 633, input the sequence after temporal encoding into the Transformer architecture for encoding. In the Transformer architecture, encode the ID of the serialized log data to obtain an ID representation, and then combine the ID representation with the temporal representation obtained from the temporal encoding to obtain the behavior representation 634.
[0173] In the text representation learning stage 640, this application directly uses the pre-trained large language model LLAMA 641 to obtain the representation of the natural language text, that is, the text representation 642.
[0174] Fuse the behavior representation 634 and the text representation 642 to obtain the final behavior representation 643, and then use the unmasked behaviors to predict the masked behaviors, thereby realizing the mining of the context.
[0175] In order to be able to simultaneously realize the mining of fine-grained and coarse-grained behavior representations, this application performs additive fusion on the text representation and the behavior representation to obtain the final behavior representation. Referring to Figure 8 , Figure 8 shows the descriptive texts of two behaviors. The first descriptive text: Obtain a car with a high moving speed and a relatively large amount of remaining fuel. The second descriptive text: Achieved victory in the game by defeating all opponents. Encode these two descriptive texts through the large language model LLAMA respectively to obtain their respective text representations. Figure 8 It also shows that encoding the behavior ID 174 (corresponding to the first descriptive text) and the behavior ID 803 (corresponding to the second behavior) through the transformer architecture to obtain the behavior representation. Additively fuse the behavior representation and the text representation to obtain the final behavior representation.
[0176] Mask Learning: The loss function for model training is as follows, using the negative log-likelihood function. The formula is as follows:
[0177]
[0178] Among them, L MLM represents the loss value, V represents the set of behaviors included in an object behavior sequence, and p i represents the probability that the category of the masked behavior is the true category. If the prediction is correct, y i takes the value of 0. If the prediction is incorrect, y i takes the value of 1.
[0179] The model trained by the model training method provided in this application can reach 73.3% in the downstream behavior prediction task. Compared with the large language model architecture of the related technology, the accuracy rate is increased by 12.2%.
[0180] Figure 9 The flowchart of the behavior encoding method provided by an exemplary embodiment of this application is shown. This method is applied to the usage stage of the data representation model. Taking this method being executed by a computer device as an example, this method includes:
[0181] Step 910, obtain a target object behavior sequence, where the target object behavior sequence includes at least two behaviors executed by an object in an application program in chronological order;
[0182] The object behavior sequence includes at least two behaviors executed by an object in an application program in chronological order. Schematically, the application program is a game application program, and the object behavior sequence can be serialized log data. For example, the object behavior sequence includes "log in to the game, play a game, add game friends, recharge, log out of the game".
[0183] During the model inference process, the target object behavior sequence will be encoded to obtain the behavior representations of each behavior in the target object behavior sequence. The target object behavior sequence is an object behavior sequence to be encoded.
[0184] Step 920, obtain the execution time of each of the multiple behaviors in the target object behavior sequence;
[0185] In the embodiment of this application, time encoding will be performed according to the execution time of the behavior to obtain the time representation.
[0186] In one embodiment, obtain multiple behaviors in the target object behavior sequence. For any one of the behaviors, obtain the relative execution time of the behavior and the initial behavior in the object behavior sequence where it is located.
[0187] Step 930, encode the execution time of each of the multiple behaviors to obtain multiple time representations;
[0188] In one embodiment, the present application encodes each of the multiple behaviors in the target object behavior sequence according to their respective relative execution times to obtain the time representation of each behavior.
[0189] For more details on time encoding, please refer to the detailed introduction under step 230 above.
[0190] Step 940: Encode the categories of each of the multiple behaviors to obtain multiple category representations;
[0191] The category, optionally, is represented by a behavior ID. For example, the ID of the behavior "log in to the game" is 0001, and the ID of the behavior "log out of the game" is 1111.
[0192] In one embodiment, the categories of each of the multiple behaviors are encoded through a Transformer architecture to obtain multiple category representations.
[0193] Step 950: Encode the description text of each of the multiple behaviors to obtain multiple text representations;
[0194] The description text is the specific descriptive text behind the behavior. In one embodiment, the description text is encoded through a large language model provided by related technologies to obtain the text representation. For example, the description text is encoded through a pre-trained LLAMA model.
[0195] Step 960: Combine the multiple time representations, multiple category representations, and multiple text representations according to the behaviors to which they belong to obtain the behavior representations of each of the multiple behaviors.
[0196] For one behavior, the time representation, category representation, and text representation obtained for this behavior are numerically added and fused to obtain the behavior representation of this behavior.
[0197] In summary, the time representation of the behavior is obtained by encoding the execution time of the behavior; the text representation of the behavior is obtained by encoding the description text of the behavior. Then, the time representation, category representation (such as ID representation), and text representation of the behavior are combined to obtain the behavior representation of the behavior.
[0198] That is, the present application performs time encoding on the execution time of the behavior, taking into account both the context relationship of the behavior and the context relationship of the execution time of the behavior, and thus better behavior representations can be obtained. When the behavior representations are applied to downstream tasks, better prediction effects will be achieved.
[0199] Moreover, the present application also performs text encoding on the description text of behaviors, revealing more fine-grained information of behaviors, and thus better behavior representations can be obtained. When the behavior representations are applied to downstream tasks (such as social recommendation, product recommendation, game matching), better prediction effects will be achieved. For example, when the behavior representations are applied to the task of product recommendation in a game, virtual products more suitable for the object will be recommended.
[0200] Figure 10 The structural block diagram of a model training device provided by an exemplary embodiment of the present application is shown. The device includes:
[0201] An acquisition module 1001, configured to acquire a plurality of object behavior sequences, where any one of the plurality of object behavior sequences includes at least two behaviors executed by an object in an application program in chronological order;
[0202] The acquisition module 1001 is further configured to, in a data representation model, for each object behavior sequence among the plurality of object behavior sequences, acquire the execution time of each of the plurality of behaviors in each object behavior sequence;
[0203] An encoding module 1002, configured to encode the execution time of each of the plurality of behaviors to obtain a plurality of time representations; and encode the category of each of the plurality of behaviors to obtain a plurality of category representations;
[0204] A processing module 1003, configured to combine the plurality of time representations and the plurality of category representations according to the behaviors to which they belong to obtain the behavior representation of each of the plurality of behaviors;
[0205] A training module 1004, configured to train a data representation model based on the behavior representation of each of the plurality of behaviors in each object behavior sequence, and the data representation model is used to output the behavior representation of a behavior.
[0206] In an optional embodiment, the time representation is a high-dimensional vector. The acquisition module 1001 is configured to acquire a plurality of behaviors in each object behavior sequence; for the first behavior, acquire the relative execution time of the first behavior and the initial behavior in the object behavior sequence to which the first behavior belongs, and the first behavior is any one of the plurality of behaviors.
[0207] The encoding module 1002 is configured to, based on the relative execution time of the first behavior, combine with the dimension label of the time representation to obtain the time representation of the first behavior.
[0208] In an optional embodiment, the encoding module 1002 is configured to, for the 2i-th dimension of the time representation of the first behavior, based on the relative execution time of the first behavior and the dimension label of the 2i-th dimension, obtain the value of the 2i-th dimension through a sine function, where i is a non-negative integer;
[0209] For the (2i + 1)-th dimension of the time representation of the first behavior, based on the relative execution time of the first behavior and the dimension label of the 2i-th dimension, the value of the (2i + 1)-th dimension is obtained through a cosine function.
[0210] Concatenate each dimension in the dimension order to obtain the time representation of the first behavior.
[0211] In an optional embodiment, the multiple behaviors include multiple unmasked behaviors. The apparatus further includes a masking module 1005. The masking module 1005 is configured to mask each object behavior sequence in the data representation model to obtain a masked behavior sequence, where the masked behavior sequence includes multiple unmasked behaviors and multiple masked behaviors; during the masking process, based on the occurrence frequency of the second behavior, calculate the masking probability of the second behavior, where the second behavior is any behavior in each object behavior sequence; perform a masking operation or not perform a masking operation on the second behavior according to the masking probability of the second behavior.
[0212] The training module 1004 is further configured to predict the categories of the multiple masked behaviors respectively based on the multiple behavior representations of the multiple unmasked behaviors; construct a prediction loss based on the prediction accuracy; and train the data representation model based on the prediction loss.
[0213] In an optional embodiment, the masking module 1005 is further configured to calculate the occurrence frequency of each of the multiple behaviors in the behavior set, where the second behavior is a behavior in the behavior set.
[0214] Based on the ratio of the occurrence frequency of the second behavior to the total occurrence frequency, calculate the masking probability of the second behavior, where the total occurrence frequency is used to indicate the sum of the occurrence frequencies of each of the multiple behaviors in the behavior set.
[0215] In an optional embodiment, the masking module 1005 is further configured to obtain the updated occurrence frequency of each based on the ratio of the occurrence frequency of each of the multiple behaviors in the behavior set to the temperature coefficient.
[0216] Accumulate the updated occurrence frequencies of each of the multiple behaviors to obtain the total occurrence frequency.
[0217] In an optional embodiment, the masking module 1005 is further configured to obtain the updated masking probability of the second behavior based on the masking probability of the second behavior, in combination with the length of the object behavior sequence where the second behavior is located and a preset masking ratio.
[0218] In an optional embodiment, the training module 1004 is further configured to, for each masked behavior in the masked behavior sequence, predict the category probability that the category of each masked behavior is the true category through the multiple behavior representations of the multiple unmasked behaviors.
[0219] If the prediction is correct, the class probabilities are accumulated; if the prediction is incorrect, the class probabilities are discarded;
[0220] Calculate the cumulative results of the multiple class probabilities corresponding to the multiple masked behaviors in the masked behavior sequence. The multiple class probabilities correspond to the multiple masked behaviors one by one;
[0221] Based on the cumulative results, a prediction loss is constructed.
[0222] In an optional embodiment, the encoding module 1002 is further configured to encode the description text of each of the multiple behaviors to obtain multiple text representations;
[0223] The processing module 1003 is further configured to combine the behavior representations of each of the multiple behaviors with their respective text representations to obtain updated behavior representations for each of the multiple behaviors.
[0224] In an optional embodiment, the application is a game application, and the object behavior sequence is the behavior sequence recorded in the game log. The training module 1004 is further configured to fine-tune the data representation model based on the dataset of the target prediction task existing in the game application to obtain a data representation model suitable for the target prediction task.
[0225] In summary, by encoding the execution time of the behavior, the time representation of the behavior is obtained, and then the time representation of the behavior is combined with the class representation (such as the ID representation) to obtain the behavior representation of the behavior.
[0226] That is, this application performs time encoding on the execution time of the behavior. While considering the context relationship of the behavior, it also considers the context relationship of the execution time of the behavior, and thus a better behavior representation can be obtained. When the behavior representation is applied to downstream tasks (such as social recommendation, product recommendation, game matching), it will have a better prediction effect. For example, when the behavior representation is applied to the task of product recommendation in a game, virtual products more suitable for the object will be recommended.
[0227] Figure 11 The structural block diagram of a behavior encoding device provided by an exemplary embodiment of this application is shown. The device includes:
[0228] An acquisition module 1101, configured to acquire a target object behavior sequence, where the target object behavior sequence includes at least two behaviors executed by an object in an application in chronological order;
[0229] The acquisition module 1101 is further configured to acquire the execution time of each of the multiple behaviors in the target object behavior sequence;
[0230] An encoding module 1102, configured to encode the execution time of each of the multiple behaviors to obtain multiple time representations; and encode the class of each of the multiple behaviors to obtain multiple class representations;
[0231] A processing module 1103 is configured to combine multiple time representations and multiple category representations according to the belonging behaviors to obtain behavior representations of respective behaviors.
[0232] In an optional embodiment, the time representation is a high-dimensional vector; the acquisition module 1101 is further configured to acquire multiple behaviors in the behavior sequence of the target object; for any behavior in the behavior sequence of the target object, acquire the relative execution time of the behavior and the initial behavior in the object behavior sequence where it is located.
[0233] The encoding module 1102 is further configured to obtain the time representation of the behavior based on the relative execution time of the behavior and in combination with the dimension label of the time representation.
[0234] In an optional embodiment, the encoding module 1102 is further configured to, for the 2i-th dimension of the time representation of the behavior, obtain the value of the 2i-th dimension through a sine function based on the relative execution time of the behavior and the dimension label of the 2i-th dimension, where i is a non-negative integer;
[0235] For the (2i + 1)-th dimension of the time representation of the behavior, obtain the value of the (2i + 1)-th dimension through a cosine function based on the relative execution time of the first behavior and the dimension label of the 2i-th dimension;
[0236] Concatenate each dimension in the dimension order to obtain the time representation of the behavior.
[0237] In an optional embodiment, the encoding module 1102 is further configured to encode the description texts of respective behaviors to obtain multiple text representations;
[0238] The processing module 1103 is further configured to combine the behavior representations of respective behaviors with their respective text representations to obtain updated behavior representations of respective behaviors.
[0239] In an optional embodiment, the application is a game application, and the object behavior sequence is the object behavior sequence recorded in the game log.
[0240] In summary, by encoding the execution time of the behavior, the time representation of the behavior is obtained, and then the time representation of the behavior is combined with the category representation (such as the ID representation) to obtain the behavior representation of the behavior.
[0241] That is, the present application performs time encoding on the execution time of the behavior, taking into account both the context relationship of the behavior and the context relationship of the execution time of the behavior, and thus a better behavior representation can be obtained. When the behavior representation is applied to downstream tasks, it will have a better prediction effect.
[0242] Figure 12It is a schematic structural diagram of a computer device shown according to an exemplary embodiment. The computer device 1200 includes a Central Processing Unit (CPU) 1201, a system memory 1204 including a Random Access Memory (RAM) 1202 and a Read-Only Memory (ROM) 1203, and a system bus 1205 connecting the system memory 1204 and the central processing unit 1201. The computer device 1200 also includes a basic Input / Output (I / O) system 1206 for facilitating information transmission between various components within the computer device, and a mass storage device 1207 for storing an operating system 1213, application programs 1214, and other program modules 1215.
[0243] The basic Input / Output system 1206 includes a display 1208 for displaying information and input devices 1209 such as a mouse, keyboard, etc. for user input of information. Both the display 1208 and the input devices 1209 are connected to the central processing unit 1201 through an input / output controller 1210 connected to the system bus 1205. The basic Input / Output system 1206 may also include an input / output controller 1210 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1210 also provides outputs to a display screen, printer, or other types of output devices.
[0244] The mass storage device 1207 is connected to the central processing unit 1201 through a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1207 and its associated computer-readable medium provide non-volatile storage for the computer device 1200. That is to say, the mass storage device 1207 may include computer-readable media (not shown) such as a hard disk or a Compact Disc Read-Only Memory (CD-ROM) drive.
[0245] Without loss of generality, the computer device-readable medium may include a computer device storage medium and a communication medium. The computer device storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer device-readable instructions, data structures, program modules, or other data. The computer device storage medium includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, digital video disc (DVD), or other optical storage, magnetic tape cartridges, magnetic tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will understand that the computer device storage medium is not limited to the above several types. The above-mentioned system memory 1204 and mass storage device 1207 may be collectively referred to as memory.
[0246] According to various embodiments of the present disclosure, the computer device 1200 may also operate by connecting to a remote computer device on a network through a network such as the Internet. That is, the computer device 1200 may be connected to the network 1211 through the network interface unit 1212 connected to the system bus 1205, or rather, the network interface unit 1212 may also be used to connect to other types of networks or remote computer device systems (not shown).
[0247] The memory further includes one or more programs, and the one or more programs are stored in the memory. The central processing unit 1201 implements all or part of the steps of the above-mentioned model training method or behavior coding method by executing the one or more programs.
[0248] Figure 13 The structural block diagram of a computer device 1300 provided by an exemplary embodiment of the present application is shown. The computer device 1300 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, or a desktop computer. The computer device 1300 may also be referred to by other names such as a user device, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0249] Generally, the computer device 1300 includes a processor 1301 and a memory 1302.
[0250] The processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1301 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0251] The memory 1302 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1302 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1302 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1301 to implement the model training method or behavior encoding method provided in the method embodiments of the present application.
[0252] In some embodiments, the computer device 1300 may further optionally include a peripheral device interface 1303 and at least one peripheral device. The processor 1301, the memory 1302, and the peripheral device interface 1303 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1303 through a bus, signal lines, or a circuit board. Exemplarily, the peripheral device may include at least one of a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, and a power supply 1308.
[0253] The peripheral device interface 1303 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1301 and the memory 1302. In some embodiments, the processor 1301, the memory 1302, and the peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1301, the memory 1302, and the peripheral device interface 1303 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.
[0254] The radio frequency circuit 1304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1304 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1304 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1304 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1304 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0255] The display screen 1305 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1305 is a touch display screen, the display screen 1305 also has the ability to collect touch signals on or above the surface of the display screen 1305. The touch signals can be input to the processor 1301 as control signals for processing. At this time, the display screen 1305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1305, which is provided on the front panel of the computer device 1300; in other embodiments, there may be at least two display screens 1305, which are respectively provided on different surfaces of the computer device 1300 or are in a folding design; in other embodiments, the display screen 1305 may be a flexible display screen, which is provided on a curved surface or a folding surface of the computer device 1300. Even, the display screen 1305 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1305 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0256] The camera module 1306 is used to collect images or videos. Optionally, the camera module 1306 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera respectively, to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 1306 may also include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to a combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0257] The audio circuit 1307 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1301 for processing, or input to the radio frequency circuit 1304 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the computer device 1300. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1301 or the radio frequency circuit 1304 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1307 may further include a headphone jack.
[0258] The power supply 1308 is used to supply power to each component in the computer device 1300. The power supply 1308 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1308 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.
[0259] In some embodiments, the computer device 1300 further includes one or more sensors 1309. The one or more sensors 1309 include but are not limited to: an acceleration sensor 1310, a gyroscope sensor 1311, a pressure sensor 1312, an optical sensor 1313, and a proximity sensor 1314.
[0260] The acceleration sensor 1310 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the computer device 1300. For example, the acceleration sensor 1310 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1301 can control the display screen 1305 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1310. The acceleration sensor 1310 can also be used for collecting game or user's motion data.
[0261] The gyroscope sensor 1311 can detect the body direction and rotation angle of the computer device 1300. The gyroscope sensor 1311 can cooperate with the acceleration sensor 1310 to collect the 3D actions of the user on the computer device 1300. Based on the data collected by the gyroscope sensor 1311, the processor 1301 can achieve the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0262] The pressure sensor 1312 can be disposed on the side frame of the computer device 1300 and / or the lower layer of the display screen 1305. When the pressure sensor 1312 is disposed on the side frame of the computer device 1300, it can detect the holding signal of the user on the computer device 1300, and the processor 1301 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1312. When the pressure sensor 1312 is disposed on the lower layer of the display screen 1305, the processor 1301 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1305. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0263] The optical sensor 1313 is used to collect the ambient light intensity. In one embodiment, the processor 1301 can control the display brightness of the display screen 1305 according to the ambient light intensity collected by the optical sensor 1313. Exemplarily, when the ambient light intensity is high, the display brightness of the display screen 1305 is increased; when the ambient light intensity is low, the display brightness of the display screen 1305 is decreased. In another embodiment, the processor 1301 can also dynamically adjust the shooting parameters of the camera assembly 1306 according to the ambient light intensity collected by the optical sensor 1313.
[0264] The proximity sensor 1314, also known as a distance sensor, is usually disposed on the front panel of the computer device 1300. The proximity sensor 1314 is used to collect the distance between the user and the front of the computer device 1300. In one embodiment, when the proximity sensor 1314 detects that the distance between the user and the front of the computer device 1300 is gradually decreasing, the processor 1301 controls the display screen 1305 to switch from the lit state to the off state; when the proximity sensor 1314 detects that the distance between the user and the front of the computer device 1300 is gradually increasing, the processor 1301 controls the display screen 1305 to switch from the off state to the lit state.
[0265] Those skilled in the art can understand that Figure 13 the structure shown in does not limit the computer device 1300, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0266] The present application also provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the model training method or the behavior encoding method provided by the above method embodiment.
[0267] The present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the model training method or the behavior encoding method provided in the above method embodiments.
[0268] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0269] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0270] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A model training method, characterized in that, The method includes: Obtaining a plurality of object behavior sequences, where any one of the plurality of object behavior sequences includes at least two behaviors executed by an object in an application in chronological order; In a data representation model, for each object behavior sequence in the plurality of object behavior sequences, obtaining the execution time of each of the multiple behaviors in the each object behavior sequence; Encoding the execution time of each of the multiple behaviors to obtain a plurality of time representations; and, encoding the categories of each of the multiple behaviors to obtain a plurality of category representations; combining the plurality of time representations and the plurality of category representations according to the belonging behaviors to obtain the behavior representations of each of the multiple behaviors; Training the data representation model based on the behavior representations of each of the multiple behaviors in the each object behavior sequence, where the data representation model is used to output the behavior representation of a behavior.
2. The method according to claim 1, characterized in that, The time representation is a high-dimensional vector; The obtaining the execution time of each of the multiple behaviors in the each object behavior sequence includes: Obtaining the multiple behaviors in the each object behavior sequence; for a first behavior, obtaining the relative execution time of the first behavior and the initial behavior in the object behavior sequence where the first behavior is located, where the first behavior is any one of the multiple behaviors; The encoding the execution time of each of the multiple behaviors to obtain a plurality of time representations includes: Based on the relative execution time of the first behavior, combining with the dimension label of the time representation, obtaining the time representation of the first behavior.
3. The method according to claim 2, wherein The obtaining the time representation of the first behavior based on the relative execution time of the first behavior and combining with the dimension label of the time representation includes: For the 2i-th dimension of the time representation of the first behavior, based on the relative execution time of the first behavior and the dimension label of the 2i-th dimension, obtaining the value of the 2i-th dimension through a sine function, where i is a non-negative integer; For the (2i + 1)-th dimension of the time representation of the first behavior, based on the relative execution time of the first behavior and the dimension label of the 2i-th dimension, obtaining the value of the (2i + 1)-th dimension through a cosine function; Concatenating each dimension in dimension order to obtain the time representation of the first behavior.
4. The method according to any one of claims 1 to 3, characterized in that The multiple behaviors include multiple unmasked behaviors; The method further includes: In the data representation model, masking each object behavior sequence to obtain a masked behavior sequence, where the masked behavior sequence includes multiple unmasked behaviors and multiple masked behaviors; During the masking process, based on the occurrence frequency of a second behavior, calculating the masking probability of the second behavior, where the second behavior is any one of the behaviors in each object behavior sequence; performing a masking operation or not performing a masking operation on the second behavior according to the masking probability of the second behavior; The training the data representation model based on the behavior representations of each of the multiple behaviors in the each object behavior sequence includes: Predicting the categories of each of the multiple masked behaviors based on the behavior representations of the multiple unmasked behaviors; constructing a prediction loss based on the prediction accuracy; Training the data representation model based on the prediction loss.
5. The method according to claim 4, wherein Calculating the masking probability of the second behavior based on the occurrence frequency of the second behavior includes: Calculating the occurrence frequency of each of multiple behaviors in the behavior set, where the second behavior is a behavior in the behavior set; Calculating the masking probability of the second behavior based on the ratio of the occurrence frequency of the second behavior to the total occurrence frequency, where the total occurrence frequency is used to indicate the sum of the occurrence frequencies of each of the multiple behaviors in the behavior set.
6. The method according to claim 5, wherein The method further includes: Obtaining the updated occurrence frequency of each based on the ratio of the occurrence frequency of each of the multiple behaviors in the behavior set to the temperature coefficient; Accumulating the updated occurrence frequencies of each of the multiple behaviors to obtain the total occurrence frequency.
7. The method according to claim 5, wherein The method further includes: Based on the masking probability of the second behavior, combining the length of the object behavior sequence to which the second behavior belongs and a preset masking ratio to obtain the updated masking probability of the second behavior.
8. The method according to claim 4, wherein Predicting the categories of the multiple masked behaviors based on the multiple behavior representations of the multiple unmasked behaviors includes: For each masked behavior in the masked behavior sequence, predicting, through the multiple behavior representations of the multiple unmasked behaviors, the category probability that the category of each masked behavior is the true category; Constructing a prediction loss based on the prediction accuracy includes: If the prediction is correct, accumulating the category probability; if the prediction is incorrect, discarding the category probability; Calculating the cumulative result of the multiple category probabilities corresponding to the multiple masked behaviors in the masked behavior sequence, where the multiple category probabilities correspond one-to-one to the multiple masked behaviors; Constructing the prediction loss based on the cumulative result.
9. The method according to any one of claims 1 to 3, characterized in that The method further includes: Encoding the description text of each of the multiple behaviors to obtain multiple text representations; Combining the behavior representation of each of the multiple behaviors with its respective text representation to obtain the updated behavior representation of each of the multiple behaviors.
10. The method according to any one of claims 1 to 3, characterized in that The application program is a game application program, and the object behavior sequence is the behavior sequence recorded in the game log; The method further includes: Fine-tuning the data representation model based on the data set of the target prediction task existing in the game application program to obtain a data representation model applicable to the target prediction task.
11. A behavior encoding method, characterized in that, The method includes: Obtaining a target object behavior sequence, where the target object behavior sequence includes at least two behaviors executed by the object in the application program in chronological order; Obtaining the execution time of each of the multiple behaviors in the target object behavior sequence; Encoding the execution time of each of the multiple behaviors to obtain multiple time representations; and encoding the category of each of the multiple behaviors to obtain multiple category representations; Combining the multiple time representations and the multiple category representations according to the belonging behaviors to obtain the behavior representation of each of the multiple behaviors.
12. A model training device, characterized in that, The device includes: An acquisition module, configured to acquire multiple object behavior sequences, where any one of the multiple object behavior sequences includes at least two behaviors executed by the object in the application program in chronological order; The obtaining module is further configured to obtain, in the data representation model, for each of the multiple object behavior sequences, the execution time of each of the multiple behaviors in each object behavior sequence; The encoding module is configured to encode the execution time of each of the multiple behaviors to obtain multiple time representations; and encode the category of each of the multiple behaviors to obtain multiple category representations; The processing module is configured to combine the multiple time representations and the multiple category representations according to the behaviors to which they belong to obtain the behavior representations of each of the multiple behaviors; The training module is configured to train the data representation model based on the behavior representations of each of the multiple behaviors in each object behavior sequence, and the data representation model is used to output the behavior representations of behaviors.
13. A computer device, characterized in that, The computer device includes: a processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the model training method according to any one of claims 1 to 10, or the behavior encoding method according to claim 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by the processor to implement the model training method according to any one of claims 1 to 10, or the behavior encoding method according to claim 11.
15. A computer program product, characterized in that, The computer program product stores a computer program, and the computer program is loaded and executed by the processor to implement the model training method according to any one of claims 1 to 10, or the behavior encoding method according to claim 11.