Method and device for representing user characteristics based on representation model
By introducing feature representation model and expert network into the characterization model, combining gated subnet and expert subnet, the problem of insufficient adaptability of the existing technology in multitasking and complex scenarios is solved, and better user feature representation data generation is achieved.
Patent Information
- Application Number
- CN202510243966.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art lacks adaptability in multitasking and complex scenarios, making it difficult to provide better user feature characterization data.
Using a method based on the characterization model, user characteristics are obtained through the feature characterization model and expert network, and feature characterization data is determined. Combined with the gated subnet and expert subnet, weight and feature characterization data are combined to generate more suitable feature characterization data.
Through this method, more adaptive feature representation data can be generated in multitasking and complex scenarios, improving the performance of the model in these environments.
Smart Images

Figure CN120180357A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of data processing, and in particular, to a method and apparatus for characterizing user features based on a characterization model. Background Art
[0002] With the rapid development of fields such as natural language processing (NLP) technology and computer vision technology, deep learning-based models have become the key technologies for solving complex problems in these fields. In the solution for processing sequential data, the self-attention mechanism performs excellently and has been widely applied in natural language processing, computer vision, and other fields, but it has some limitations, such as insufficient adaptability in multi-task and complex scenarios.
[0003] Then, how to provide an improved method for characterizing user features to better adapt to multi-task and complex scenarios has become an urgent problem to be solved. Summary of the Invention
[0004] One or more embodiments of this specification provide a method and apparatus for characterizing user features based on a characterization model to obtain better feature characterization data corresponding to user features.
[0005] According to a first aspect, there is provided a method for characterizing user features based on a characterization model. The characterization model includes a feature characterization model, and the feature characterization model includes a first characterization network and a first expert network. The first expert network includes a gating sub-network and a plurality of expert sub-networks arranged in parallel. The method includes:
[0006] Obtain the first user feature of the first user;
[0007] Based on the first user feature, determine first feature characterization data through the first characterization network;
[0008] Based on the first feature characterization data, obtain the first weight corresponding to each expert sub-network through the gating sub-network;
[0009] Based on the first feature characterization data, obtain the second feature characterization data corresponding to each expert sub-network through each expert sub-network;
[0010] Combine the first weight and the second feature characterization data corresponding to each expert sub-network to determine third feature characterization data.
[0011] According to a second aspect, there is provided an apparatus for characterizing user features based on a characterization model. The characterization model includes a feature characterization model, and the feature characterization model includes a first characterization network and a first expert network. The first expert network includes a gating sub-network and a plurality of expert sub-networks arranged in parallel. The apparatus includes:
[0012] A first acquisition module, configured to acquire first user characteristics of a first user;
[0013] A first determination module, configured to determine first feature representation data based on the first user characteristics through the first representation network;
[0014] A first obtaining module, configured to obtain first weights corresponding to each expert sub-network based on the first feature representation data through the gating sub-network;
[0015] A second obtaining module, configured to obtain second feature representation data corresponding to each expert sub-network based on the first feature representation data through each expert sub-network;
[0016] A second determination module, configured to determine third feature representation data by combining the first weights corresponding to each expert sub-network and the second feature representation data.
[0017] According to a third aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in the first aspect.
[0018] According to a fourth aspect, there is provided a computing device, including a memory and a processor. Among them, an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0019] According to the method and device for representing user characteristics based on a representation model provided in the embodiments of this specification, the representation model includes a feature representation model, the feature representation model includes a first representation network and a first expert network, the first expert network includes a gating sub-network and a plurality of expert sub-networks arranged in parallel. The method includes: acquiring first user characteristics of a first user; determining first feature representation data based on the first user characteristics through the first representation network; obtaining first weights corresponding to each expert sub-network based on the first feature representation data through the gating sub-network; obtaining second feature representation data corresponding to each expert sub-network based on the first feature representation data through each expert sub-network; determining third feature representation data by combining the first weights corresponding to each expert sub-network and the second feature representation data.
[0020] In the above process, after determining the first feature representation data of the first user feature through the first representation network, the first weights corresponding to each expert sub-network obtained by the joint gating sub-network and each expert sub-network are used to obtain the second feature representation data corresponding to each expert sub-network, and the third feature representation data is jointly determined, so as to allocate the most suitable expert sub-network for different tasks or scenarios through the gating sub-network, process the data through its most suitable expert sub-network, and obtain better feature representation data more adapted to the corresponding task or scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0022] Figure 1 Schematic diagram of the implementation framework of an embodiment disclosed in this specification;
[0023] Figure 2 Schematic flow chart of a method for representing user features based on a representation model provided by an embodiment;
[0024] Figure 3A Schematic structural diagram of a representation model provided by an embodiment;
[0025] Figure 3B Another schematic flow chart of a method for representing user features based on a representation model provided by an embodiment;
[0026] Figure 4 Schematic flow chart of a process for determining the first feature representation data provided by an embodiment;
[0027] Figure 5 Schematic flow chart of a process for determining the first feature representation data provided by an embodiment;
[0028] Figure 6 Schematic flow chart of a process for determining the first feature representation data provided by an embodiment
[0029] Figure 7 Schematic flow chart of a process for determining the first feature representation data provided by an embodiment;
[0030] Figure 8 Schematic flow chart of a discrimination process based on a representation model provided by an embodiment;
[0031] Figure 9A Schematic diagram of a training process of a representation model provided by an embodiment;
[0032] Figure 9B A flowchart of the training process of the characterization model provided for the embodiment;
[0033] Figure 10 A schematic block diagram of the device for characterizing user features based on the characterization model provided for the embodiment. Detailed implementation manners
[0034] Next, the technical solutions of the embodiments of the present specification will be described in detail with reference to the accompanying drawings.
[0035] The embodiments of the present specification disclose a method and device for characterizing user features based on a characterization model. First, the application scenario and technical concept of the method will be introduced as follows:
[0036] As described above, in the solutions for processing sequence data, the self-attention mechanism performs excellently and has been widely applied in fields such as natural language processing and computer vision. However, it has some limitations, such as insufficient adaptability in multi-task and complex scenarios. Then, how to provide an improved method for characterizing user features to better adapt to multi-task and complex scenarios has become an urgent problem to be solved.
[0037] In view of this, the inventors propose a method for characterizing user features based on a characterization model, Figure 1 Fig. shows a schematic diagram of an implementation scenario according to an embodiment disclosed in the present specification. In this implementation scenario, the characterization model includes a feature characterization model, and the feature characterization model includes a first characterization network and a first expert network. The first expert network includes a gating sub-network and a plurality of expert sub-networks arranged in parallel (such as Figure 1 shown including expert sub-network 1, expert sub-network 2... expert sub-network z). In the process of characterizing user features based on the characterization model, obtain the first user feature of the first user; based on the first user feature, determine the first feature characterization data through the first characterization network; based on the first feature characterization data, obtain the first weight corresponding to each expert sub-network through the gating sub-network; based on the first feature characterization data, obtain the second feature characterization data corresponding to each expert sub-network through each expert sub-network (such as Figure 1 the second feature characterization data 1 corresponding to expert sub-network 1, the second feature characterization data 2 corresponding to expert sub-network 2... the second feature characterization data z corresponding to expert sub-network z shown); combine the first weight and the second feature characterization data corresponding to each expert sub-network to determine the third feature characterization data.
[0038] In the above process, after determining the first feature representation data of the first user feature through the first representation network, the first weights corresponding to each expert sub-network obtained by the joint gating sub-network and each expert sub-network are used to obtain the second feature representation data corresponding to each expert sub-network, and the third feature representation data is jointly determined, so as to allocate the most suitable expert sub-network for different tasks or scenarios through the gating sub-network, and process the data through its most suitable expert sub-network to obtain better feature representation data that is more adapted to the corresponding task or scenario.
[0039] The following describes in detail the method for representing user features based on the representation model provided in this specification in combination with specific embodiments.
[0040] Figure 2 The flowchart of the method for representing user features based on the representation model in an embodiment of this specification is shown. This method is executed by an electronic device, and the electronic device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities.
[0041] The representation model includes a feature representation model, and the feature representation model includes a first representation network and a first expert network. Among them, the first expert network includes a gating sub-network and multiple parallel expert sub-networks. Exemplarily, the feature representation model can be understood as an encoder model for encoding inputs, and its output can be used to perform subsequent specified tasks, such as the discrimination process in the subsequent mentioned specified discrimination task. In some possible examples, a single expert sub-network can be implemented as a feedforward neural network, a DNN (Deep Neural Network), a multi-layer perceptron, etc.
[0042] In some implementation manners, the feature representation model can be implemented as a model based on the Transformer structure, such as the Bert model, or can be implemented as other models that can deploy the first representation network and the first expert network. In some possible examples, the first representation network can be a network based on the attention mechanism.
[0043] The following introduces the process of representing user features based on the representation model. In the process of representing user features based on the representation model, as Figure 2 shown, the method includes the following steps S210 - S250:
[0044] In step S210, obtain the first user feature of the first user.
[0045] It can be understood that, in one implementation, the process of characterizing user features based on the characterization model can be a subprocess in the training process of the characterization model. That is to say, the feature characterization data corresponding to the user features obtained through the process of characterizing user features based on the characterization model can continue to be combined with the labels corresponding to the user features to train the characterization model. Correspondingly, the first user can be any sample user in the training set used to train the characterization model.
[0046] In another implementation, the process of characterizing user features based on the characterization model can be a subprocess in the discrimination process based on the characterization model. That is to say, the feature characterization data corresponding to the user features obtained through the process of characterizing user features based on the characterization model can continue to be used to determine the discrimination result of the user features under a specified discrimination task, that is, the discrimination result of the user indicated by the user features under the specified discrimination task. Correspondingly, the first user can be any user who needs to be discriminated under a specified discrimination task.
[0047] Among them, the feature characterization data here can refer to a vector or a matrix. The feature characterization data can refer to the third feature characterization data corresponding to the user features mentioned later or the aggregated feature characterization data mentioned later.
[0048] Exemplarily, the first user features can include, but are not limited to, features corresponding to user information such as the first user portrait, transaction records, scene activity, historical risk information, the operation behavior of the first user for a specified application, and the information of the group to which the first user belongs. The specified application can include, but is not limited to, electronic payment platforms, electronic trading platforms, and financial platforms, etc.
[0049] In some implementations, the feature characterization model may further include an embedding network, which is arranged before the first characterization network. In step S210, the electronic device can obtain the first user information of the first user. In this embedding network, word embedding processing (i.e., vectorization processing) can be performed on the first user information, and then the corresponding first user features can be obtained.
[0050] In some possible embodiments, the embedding network can be implemented as an embedding network based on any word embedding algorithm in the related art. In some cases, in the discrimination process based on the representation model, the parameters of the embedding network have been trained based on a large amount of text corpus (including the characteristics of sample users and their corresponding labels). Correspondingly, word vectors of each word involved in the user's characteristics are obtained through training, that is, a trained word vector table is formed; alternatively, in the training process of the representation model, the embedding network has been trained based on the previous iteration process of the current iteration process (taking the k-th training iteration process as an example), that is, the k - 1-th iteration process, to obtain the corresponding word vectors of each word, that is, a trained word vector table corresponding to the current iteration process is formed. Thus, the first user feature can be determined by referring to the corresponding word vector table, where the word vectors corresponding to each word in the first user information can be included.
[0051] Next, in step S220, based on the first user feature, the first feature representation data is determined through the first representation network.
[0052] In this step, the first user feature can be input into the first representation network to process the first user feature through the first representation network and determine the first feature representation data.
[0053] In some possible examples, the first representation network can be a network based on the attention mechanism, and its specific structure will be introduced later.
[0054] In some possible embodiments, there may be a large amount of data of the first user feature (for example, the sequence length of the first user feature is long). When the first representation network is a network based on the attention mechanism, the subsequent data calculation load is large. To reduce the consumption of computing resources in the process of representing user features to a certain extent, the feature representation model can further include a sequence segmentation network; the sequence segmentation network is arranged before the first representation network. Correspondingly, before step S220, step 11 can also be included:
[0055] In step 11, the first user feature is segmented through the sequence segmentation network to obtain a plurality of first user feature blocks. Exemplarily, a maximum segmentation length is set in the sequence segmentation network. After the electronic device obtains the first user feature, the first user feature can be input into the sequence segmentation network, and the sequence segmentation network evenly segments or divides the first feature sequence into multiple blocks according to the maximum segmentation length to form the first user feature blocks. Among them, the sizes of the plurality of first user feature blocks are as consistent as possible to minimize the context information that may be lost at the boundaries of the plurality of first user feature blocks. The maximum segmentation length is used to limit the maximum length of the segmented feature blocks.
[0056] In some possible examples, in the sequence segmentation network, the number of chunks can be determined first based on the length of the first user feature and the maximum segmentation length; then the first user feature is segmented based on the number of chunks, so that the lengths of the obtained multiple first user feature chunks are as consistent as possible.
[0057] Next, in step S220, it specifically includes: respectively based on each first user feature chunk, through the first representation network, determining the feature representation data corresponding to each first user feature chunk, and then jointly determining the first feature representation data.
[0058] In this step, for each first user feature chunk (taking the first user feature chunk O as an example for illustration), the first user feature chunk O can be input into the first representation network, and the first representation network processes the first user feature chunk O to obtain the feature representation data corresponding to the first user feature chunk O, and so on, to obtain the feature representation data corresponding to each first user feature chunk. Then, the feature representation data corresponding to each first user feature chunk are spliced to obtain the first feature representation data.
[0059] In the above implementation, when the first representation network calculates the attention weights of each feature within each first user feature chunk respectively, compared with calculating the attention weights of each feature within the complete first user feature, the amount of calculation becomes less, reducing the consumption of computing resources, significantly reducing the computational complexity and improving the processing speed of the representation model; to a certain extent, it can also enable the first representation network to avoid processing the features of long sequences and deploying more parameters.
[0060] After obtaining the first feature representation data, in step S230, based on the first feature representation data, through the gating sub-network, the first weights corresponding to each expert sub-network are obtained.
[0061] In one implementation, the gating sub-network can be a network layer based on a deep learning algorithm, and its parameters can be adjusted. The gating sub-network learns how to make effective allocations during the training process, that is, determining the weights corresponding to each expert sub-network for the input, so that the expert network model can manage the complex decision-making process more precisely, thereby enabling the entire representation model to achieve relatively better performance in multiple tasks or scenarios.
[0062] In another implementation, the gating sub-network can include a series of preset probability determination strategies, so as to adaptively determine the weights corresponding to each expert sub-network based on the input through the preset probability determination strategies.
[0063] In step S240, based on the first feature representation data, second feature representation data corresponding to each expert sub-network is obtained through each expert sub-network. In this step, the first feature representation data is input into each expert sub-network respectively, so as to process the first feature representation data through each expert sub-network respectively, and obtain the second feature representation data corresponding to each expert sub-network.
[0064] It should be noted that Figure 2 In the shown process, the order of first executing step S230 and then executing step S240 is shown. In practice, step S240 can also be executed first, then step S230, or step S230 and step S240 can be executed simultaneously. This specification does not make any limitation on this.
[0065] Next, in step S250, the third feature representation data is determined by combining the first weights corresponding to each expert sub-network and the second feature representation data. In this step, for each expert sub-network (taking expert sub-network h as an example for illustration), the product of the first weight corresponding to expert sub-network h and the second feature representation data is calculated to obtain the product result corresponding to expert sub-network h. Then, the product results corresponding to each expert sub-network are aggregated to obtain the third feature representation data, which is used to determine the discrimination result under the specified discrimination task mentioned later. In some possible examples, this aggregation may refer to superposition, or may refer to aggregation based on the attention mechanism.
[0066] In the above process, after the first feature representation data of the first user feature is determined by the first representation network, the first weights corresponding to each expert sub-network obtained by combining the gating sub-network and the second feature representation data obtained by each expert sub-network are combined to determine the third feature representation data, so as to allocate the most suitable expert sub-network for different tasks or scenarios through the gating sub-network, and process the data through its most suitable expert sub-network to obtain better feature representation data more adapted to the corresponding task or scenario.
[0067] In the above process, by integrating the module of the expert network in the feature representation model, the ability of the model to handle multi-task environments is better enhanced. This module enables the model to manage complex decision-making processes more precisely and effectively allocate tasks among numerous expert sub-networks through the gating mechanism, improving the representation ability of the model.
[0068] In some possible examples, considering that in the actual inference process or training process, there may be a situation where some of the user information (i.e., user features) collected for the user is missing. To better facilitate the management of user information, the user information can be grouped. Among them, the user information can be grouped from multiple perspectives. Exemplarily, it can be classified according to the type of user information. For example, it can include but is not limited to user basic attribute information (such as height, weight, and occupation, which can also be called user portraits), user transaction flow information, scene activity, historical risk information, user credit information, user social information, user operation behaviors for a specified application, and user gang information, etc. The embodiments of this specification do not limit the grouping method of user information, and any method that can group user information can be applied to the embodiments of this specification.
[0069] In the case of grouping user information, correspondingly, the structure of the representation model can be set according to the grouping situation of the user information, so as to meet the requirements of the inference process or training process and obtain feature representation data that is more conducive to the accuracy of the discrimination result of the subsequent specified discrimination task. As Figure 3A shown, the representation model can include a feature representation model. The feature representation model includes a first representation network and a first expert network corresponding to the first feature group, a second representation network and a second expert network corresponding to several second feature groups (as Figure 3A shown, including a second representation network 1 and a second expert network 1 corresponding to the second feature group 1... a second representation network n and a second expert network n corresponding to the second feature group n), and an aggregation network. Among them, the first expert network includes a gating sub-network and multiple parallel expert sub-networks (not shown in the figure), and each second expert network includes a gating sub-network and multiple parallel expert sub-networks (not shown in the figure). Correspondingly, as Figure 3B shown, the method for representing user features based on the representation model can include the following steps S310 - 380:
[0070] In step S310, obtain the first user feature of the first user.
[0071] In step S320, based on the first user feature, determine the first feature representation data through the first representation network.
[0072] In step S330, based on the first feature representation data, obtain the first weights corresponding to each expert sub-network through the gating sub-network;
[0073] In step S340, based on the first feature representation data, obtain the second feature representation data corresponding to each expert sub-network through each expert sub-network;
[0074] In step S350, the first weights corresponding to each expert sub-network and the second feature representation data are combined to determine the third feature representation data.
[0075] Among them, the implementation principles of steps S310 - S350 are similar to those of the foregoing steps S210 - S250, and the implementation process can refer to the implementation process of the foregoing steps S210 - S250, which will not be elaborated here.
[0076] In step S360, a number of second user features corresponding to a number of second feature groups of the first user are obtained. Among them, the implementation principle of step S360 is similar to that of the foregoing step S210, and the implementation process can refer to the implementation process of the foregoing step S210, which will not be elaborated here.
[0077] In step S370, the second user features corresponding to each second feature group are processed respectively through the second representation network and the second expert network corresponding to each second feature group to obtain a number of fourth feature representation data.
[0078] In this step, each second representation network can be a network based on the attention mechanism. In some cases, the structure of each second representation network is similar to that of the foregoing first representation network, and the structure of each second expert network is similar to that of the foregoing first expert network, which will not be elaborated here. It should be noted that the number of expert sub-networks arranged in parallel in the second expert network may be the same as or different from the number of expert sub-networks arranged in parallel in the first expert network.
[0079] In this step, for the second user features corresponding to each second feature group (taking the second user feature x as an example), based on the second user feature x, through its corresponding second representation network, the feature representation data corresponding to the second user feature x can be determined; based on the feature representation data corresponding to the second user feature x, through its corresponding gating sub-network, the weights corresponding to its corresponding expert sub-networks can be obtained; based on the feature representation data corresponding to the second user feature x, through its corresponding expert sub-networks, the result data corresponding to its corresponding expert sub-networks can be obtained; then, by combining the weights and result data corresponding to the expert sub-networks corresponding to the second user feature x, the fourth feature representation data corresponding to the second user feature x is determined. By analogy, a number of fourth feature representation data corresponding to a number of second user features are obtained.
[0080] Next, in step S380, the third feature representation data and a number of fourth feature representation data are aggregated through an aggregation network to obtain aggregated feature representation data. Among them, the aggregated feature representation data can be used to determine the discrimination result of the first user in a specified discrimination task.
[0081] In some possible examples, the aggregation network may be any network in the related art that can aggregate data. In some specific examples, the aggregation network may be an aggregation network based on an attention mechanism. As Figure 3A shown, which may include one or more sets of characterization networks (subsequently referred to as the third characterization networks) and their corresponding expert networks (subsequently referred to as the third expert networks). The third expert network includes its gating sub-network and multiple expert sub-networks arranged in parallel. Among them, Figure 3A shows a set of third characterization networks and their corresponding third expert networks. The structure of the third characterization network is similar to the structure of the foregoing first characterization network, and its structure can be referred to the structure of the first characterization network; the structure of the third expert network is similar to the structures of the foregoing first expert network and second expert network, and its structure can be referred to the structures of the first expert network and second expert network, which will not be elaborated here. As Figure 3A shown, after the third characterization network and its corresponding third expert network, the aggregation network may further include an expert network to achieve better aggregation of the feature representation data corresponding to the foregoing multiple sets of feature groups.
[0082] In some implementation manners, when the aggregation network is an aggregation network based on an attention mechanism, the third feature representation data and several fourth feature representation data may be concatenated first to obtain concatenated feature representation data. Then, the concatenated feature representation data is input into the aggregation network to process the concatenated feature representation data through the aggregation network to obtain aggregated feature representation data.
[0083] In the discrimination process based on the characterization model (i.e., the actual inference process) or the training process of the characterization model (i.e., the actual training process), for example, it may be possible that the user features corresponding to a certain or certain feature groups of the user are not collected, that is, there is a situation where the characterization network and the expert network corresponding to a certain or certain feature groups are not input with the corresponding user features; for another example, in the actual training process, it is necessary to control the input of the features of different feature groups to enhance the sample diversity. Subsequently, when performing the aggregation network, the input of the aggregation network during aggregation can be controlled by preset activation information. Correspondingly, in some possible examples, in step S280, it may specifically include: through the aggregation network, based on the preset activation information, aggregating the third feature representation data and several fourth feature representation data to obtain aggregated feature representation data.
[0084] Among them, the preset activation information can be manually set according to the actual situation, which is used to indicate which or which feature groups' corresponding representation networks and expert networks are activated, that is, the output data can be adopted, and which or which feature groups' corresponding representation networks and expert networks are not activated, that is, the output data is not adopted. In some examples, the preset activation information can be represented by a mask vector, where each bit corresponding to a feature group is included. If the value of the bit corresponding to the feature group is a specified value, it can indicate that the representation network and expert network corresponding to the feature group are activated, that is, its output participates in the aggregation of the aggregation network; if the value of the bit corresponding to the feature group is a non-specified value, it can indicate that the representation network and expert network corresponding to the feature group are not activated, that is, its output does not participate in the aggregation of the aggregation network.
[0085] In the above embodiments, through the preset activation information, that is, the Activation / Selection mechanism, the model can be allowed to dynamically select and activate different combinations of feature groups, that is, indicate which feature groups' corresponding feature representation data participate in the aggregation of the subsequent aggregation network, and which feature groups' corresponding feature representation data do not participate in the aggregation of the subsequent aggregation network, so as to more flexibly control the output of the aggregation network to better meet the actual inference requirements. It can not only achieve the progressive learning of the features of different feature groups during the training process, but also adjust the representation output according to the needs during inference, which provides a richer and more flexible feature representation for the model.
[0086] Next, the structure of the first representation network will be introduced.
[0087] In some possible examples, the aforementioned first representation network may include a first attention layer, and this first attention layer can achieve task perception in combination with the label representation vocabulary, and can be called an attention layer based on the task perception attention mechanism; as Figure 4 shown, in the aforementioned step S210, it may include the following steps S410 - S420:
[0088] In step S410, obtain the label representation vocabulary, where the label representation vocabulary includes label representation data corresponding to various labels under the specified discrimination task. Here, the label representation data can be a vector or a matrix.
[0089] In some implementation manners, the specified discrimination task may include but is not limited to any one of the following: various classification tasks about users, such as but not limited to: user risk classification, user intention classification, user rental classification, whether the user is a recommended user for a certain type of object (such as goods, services, etc.), user credit assessment, etc., and may also include the discrimination task of the user's affiliated gang and other regression tasks of specified dimensions.
[0090] Wherein, when the above-mentioned specified discrimination task includes various classification tasks about users, the classification task can be a binary classification task or a multi-classification task. When the classification task about users is a multi-classification task, it can correspond to multiple labels. For example, taking the classification task of user risk classification as an example, the multiple labels corresponding to it can be, for example: "low-risk user", "medium-risk user", "high-risk user". Correspondingly, the label representation vocabulary can include the label representation data corresponding to "low-risk user", "medium-risk user", and "high-risk user" respectively under the user's risk classification.
[0091] Taking the classification task about users as the user intention classification as an example again, the multiple labels corresponding to it can be, for example: standard questions such as "how to open the payment function", "how to change the bound mobile phone number", "how to make a complaint", etc. Correspondingly, the label representation vocabulary can include the label representation data corresponding to "how to open the payment function", "how to change the bound mobile phone number", and "how to make a complaint" respectively under the user's intention classification.
[0092] Taking the classification task about users as the user rental classification as an example again, the multiple labels corresponding to it can be, for example: "short-term rental user", "long-term rental user", "frequent rental user", etc. Correspondingly, the label representation vocabulary can include the label representation data corresponding to "short-term rental user", "long-term rental user", and "frequent rental user" respectively under the user's rental classification.
[0093] When the above-mentioned specified discrimination task includes discrimination tasks of other non-classification tasks, such as regression tasks in other specified dimensions, the label representation vocabulary can include the label representation data corresponding to various labels under the regression task.
[0094] It can be understood that the specified discrimination task can be one or more. When the specified discrimination task is multiple, the label representation vocabulary can include the label representation data corresponding to various labels under each specified discrimination task in the multiple specified discrimination tasks.
[0095] In some possible implementation manners, the representation model may further include a label representation model. As Figure 3A shown, the label representation model is used to vectorize the labels under the specified discrimination task; correspondingly, in step S410, it specifically includes: based on multiple labels under the specified discrimination task, obtaining the label representation vocabulary through the label representation model.
[0096] In some possible implementation manners, the label representation model may be implemented as any word embedding model, such as Word2Vec, GloVe, FastText, etc. Specifically, multiple labels under a specified discrimination task may be input into the label representation model, and each label is processed by the label representation model to obtain label representation data corresponding to each label under the specified discrimination task. Then, based on the label representation data corresponding to each label under the specified discrimination task, a label representation vocabulary is formed.
[0097] In some cases, when there are multiple specified discrimination tasks, when training the label representation model, it can be set that different labels under different specified discrimination tasks are all different, and during the training process of the representation model, at least the goal is to minimize the label representation data corresponding to different labels under different specified discrimination tasks, and the representation model is trained, that is, the feature representation model and the label representation model are trained. Correspondingly, multiple labels under each specified discrimination task may be input into the label representation model to process multiple labels under each specified discrimination task through the label representation model to obtain label representation data corresponding to each label under each specified discrimination task. Then, based on the label representation data corresponding to each label under each specified discrimination task, a label representation vocabulary is formed. Among them, the label representation data corresponding to each label under each specified discrimination task is different from each other.
[0098] In still some other possible implementation manners, there are multiple specified discrimination tasks; correspondingly, in step S410, it may specifically include: based on multiple labels under each specified discrimination task and the task identifier of the specified discrimination task corresponding to each label, a label representation matrix is obtained through the label representation model.
[0099] In the above implementation manner, for each label under each specified discrimination task (taking the label j under the specified discrimination task i as an example for illustration), the label j under the specified discrimination task i and the task identifier i of the specified discrimination task i corresponding to the label j are input into the label representation model to use the label representation model to perform word embedding processing on the label j and the task identifier i respectively to obtain representation data e1 and representation data e2. Then, the representation data e1 and the representation data e2 are subjected to a fusion process to obtain label representation data ij corresponding to the label j under the specified discrimination task i. Among them, the fusion process may be, for example, a stacking process, or may be a fusion process based on an attention mechanism, etc. By analogy, label representation data corresponding to each label under each specified discrimination task is obtained, and then a label representation vocabulary is obtained by using the label representation data corresponding to each label under each specified discrimination task.
[0100] In some other possible implementations, in the discrimination process based on the representation model, the representation model is a trained model, that is, the feature representation model and the label representation model are trained models. At this time, when the various labels under the specified discrimination task do not change, the corresponding label representation data will not change, and accordingly, the label representation vocabulary may not change. In view of this, in the discrimination process based on the representation model, the electronic device can store the label representation vocabulary in the specified space in advance, and accordingly, the electronic device can obtain the label representation vocabulary from the specified space.
[0101] Exemplarily, the label representation vocabulary may exist in the form of a matrix. For example, each row (or column) may represent label representation data corresponding to a label under a specified task; or each row (or column) may represent label representation data corresponding to multiple labels under a specified task.
[0102] In step S420, first feature representation data is determined through a first attention layer based on the first user feature and the tag representation vocabulary.
[0103] In some possible examples, the first representation network is a network based on a Transformer structure, and the first representation network may deploy one or more attention layers arranged in series, and one or more attention layers in the first representation network may be set as attention layers based on a task-aware attention mechanism (and / or an attention mechanism based on a dynamic masking mechanism mentioned later), such as the first attention layer provided in the embodiments of this specification. Exemplarily, the first attention layer may be the first attention layer in the first representation network, or a non-first attention layer, which is possible.
[0104] In some implementations, the first representation network may further include a processing layer corresponding to each attention layer after each attention layer. Exemplarily, the processing layer may be implemented as a feedforward neural network layer or an expert network layer. The structure of the expert network layer is similar to that of the aforementioned expert network, and will not be described in detail here.
[0105] In this step, the electronic device can obtain the output matrix corresponding to the first attention layer based on the first user feature and the label representation vocabulary through the first attention layer, and then determine the first feature representation data of the first user feature based on the output matrix corresponding to the first attention layer.
[0106] In some possible implementations, in step S420, as Figure 4 As shown, steps S11-S13 may be included:
[0107] In step S11, based on the first user feature, the first query matrix, the first key matrix, and the first value matrix are determined through the first attention layer.
[0108] In a possible implementation, when the first attention layer is the first attention layer of the first representation network, the electronic device may input the first user feature into the first attention layer, so as to process the first user feature through the Q weight matrix, the K weight matrix, and the V weight matrix (hereinafter simply referred to as the QKV weight matrix) of the first attention layer, to obtain the corresponding first query Q matrix, the first key K matrix, and the first value V matrix, as Figure 5 shown. Among them, in the process of training the representation model, the aforementioned QKV weight matrix of the first attention layer (as well as the QKV weight matrices of other attention layers and the parameters of the aforementioned embedding network) and the parameters in the first expert network need to be adjusted, and the parameters in the processing layer after the first attention layer (as well as the processing layers after other attention layers) also need to be adjusted.
[0109] Exemplarily, the process of processing the first user feature can be represented by the following formula (1):
[0110]
[0111] where Q represents the first query matrix, K represents the first key matrix, V represents the first value matrix, and W Q 、W K and W V respectively represent the Q weight matrix, the K weight matrix, and the V weight matrix of the first attention layer, and E represents the first user feature or the first processing result mentioned later.
[0112] In another implementation, when the first attention layer is a non-first attention layer of the first representation network, correspondingly, the electronic device may, based on the first user feature and the structure before the first attention layer in the first representation network (such as at least one group of attention layers and their corresponding processing layers), obtain the corresponding first processing result, and then input the first processing result into the first attention layer, so as to process the first processing result through the QKV weight matrix of the first attention layer, to obtain the corresponding first query Q matrix, the first key K matrix, and the first value V matrix.
[0113] In step S12, based on the label representation vocabulary and the first query matrix, the second query matrix is determined.
[0114] In this implementation manner, considering that the query Q matrix is mainly used for querying, which mainly affects the weight distribution among the features in the input (such as the word vectors in the user features) and does not affect the overall data change of the input. Correspondingly, in this implementation manner, the first query matrix is adjusted related to the task information. Specifically, based on the label representation vocabulary and the first query matrix, the second query matrix is determined.
[0115] In some specific examples, the aforementioned first representation network may further include: a vocabulary processing layer; correspondingly, in step S12, it may include steps 121-122:
[0116] In step 121, based on the label representation vocabulary, through the vocabulary processing layer, a label matrix is obtained, where the parameters in this vocabulary processing layer are trainable parameters during the training process of the representation model. In this step, the label representation vocabulary is input into the vocabulary processing layer to process the label representation vocabulary through the vocabulary processing layer to obtain the label matrix.
[0117] After that, in step 122, based on the label matrix and the first query matrix, the second query matrix is determined.
[0118] In some possible examples, as Figure 5 shown, the aforementioned vocabulary processing layer may include a first feedforward network, a second feedforward network, and a preset activation function; in some examples, the preset activation function may be, for example, the sigmoid activation function.
[0119] Correspondingly, in the aforementioned step 121, it may include steps 1211-1212:
[0120] In step 1211, based on the label representation vocabulary, through the first feedforward network and the preset activation function, a label weight matrix is obtained. In this step, the label representation vocabulary is input into the first feedforward network to process the label representation vocabulary through the first feedforward network to obtain an intermediate matrix. Among them, both the first feedforward network and the second feedforward network can be implemented as linear layers to linearly process the label representation vocabulary; then the intermediate matrix is processed using the preset activation function to obtain the label weight matrix. The parameters of both the first feedforward network and the second feedforward network are trainable parameters.
[0121] And, in step 1212, based on the label representation vocabulary, through the second feedforward network, a label bias matrix is obtained. In this step, the label representation vocabulary is input into the second feedforward network to process the label representation vocabulary through the second feedforward network to obtain the label bias matrix.
[0122] Afterwards, in one implementation, at step 122, the first query matrix is multiplied element-wise with the label weight matrix. Then, the result obtained by multiplying the first query matrix and the label weight matrix element-wise is added to the label bias matrix to obtain an added matrix. Subsequently, the added matrix can be used as the second query matrix to obtain a second query matrix embedded with task label information, which can help obtain feature representation data related to the task label information.
[0123] In yet another implementation, to avoid the occurrence of over-task perception, at step 122, a residual connection method can be adopted to determine the second query matrix based on the label matrix and the first query matrix. Specifically, the first query matrix is multiplied element-wise with the label weight matrix. Then, the result obtained by multiplying the first query matrix and the label weight matrix element-wise is added to the label bias matrix to obtain an added matrix. After that, the added matrix is added to the first query matrix again to obtain a second query matrix embedded with task label information, as Figure 5 shown. Among them, the second query matrix can help obtain feature representation data related to the task label information.
[0124] In the above implementation, the residual connection method is used to determine the second query matrix, which not only strengthens the task orientation but also retains the information of the original input, preventing the performance degradation of the representation model caused by excessive adjustment, and this realizes an effective strategy for balancing task guidance and original feature preservation.
[0125] In some other possible examples, the vocabulary processing layer can only include the aforementioned first feed-forward network and the preset activation function, or only include the aforementioned second feed-forward network, which are both acceptable. Exemplarily, when the vocabulary processing layer includes the aforementioned first feed-forward network and the preset activation function, the aforementioned label matrix can include the aforementioned label weight matrix. Correspondingly, the second query matrix embedded with task label information can be determined based on the label weight matrix and the first query matrix.
[0126] At step S13, based on the second query matrix, the first key matrix, and the first value matrix, through the first attention layer, the first feature representation data is determined. In this step, based on the second query matrix, the first key matrix, and the first value matrix, through the preset attention formula of the first attention layer, the output matrix corresponding to the first attention layer is obtained. Then, based on the output matrix corresponding to the first attention layer, the feature representation data of the user feature is determined, as Figure 5 shown. Among them, the preset attention formula can be expressed as the following formula (2):
[0127]
[0128] Among them, O represents the output matrix corresponding to the first attention layer, and Q ′ represents the second query matrix, K represents the first key matrix, V represents the first value matrix, and d k represents the dimension number of the first key matrix, and softmax(.) represents the activation function.
[0129] In some possible examples, the first feature representation network may further include a processing layer (which may be implemented as a feed-forward neural network layer or an expert network layer) disposed after the first attention layer. The process of determining the feature representation data of the user feature based on the output matrix corresponding to the first attention layer may include: inputting the output matrix corresponding to the first attention layer into the processing layer after the first attention layer to obtain the output matrix corresponding to the processing layer after the first attention layer; then, in a possible implementation manner, if the processing layer after the first attention layer is the last layer of the first feature representation network, then determining the output matrix corresponding to the processing layer after the first attention layer as the first feature representation data; in another possible implementation manner, if the first feature representation network further includes several attention layers and processing layers after the processing layer after the first attention layer, then based on the output matrix corresponding to the processing layer after the first attention layer, through several attention layers and processing layers after the processing layer after the first attention layer in the first feature representation network, the first feature representation data is obtained.
[0130] In some other possible examples, after determining the first query matrix, the first key matrix, and the first value matrix based on the first user feature through the first attention layer, the electronic device may further determine the second key matrix based on the first key matrix and the label representation vocabulary (or the label matrix determined by the vocabulary processing layer based on the label representation vocabulary); then, by combining the first query matrix, the second key matrix, and the first value matrix, through the foregoing preset attention formula of the first attention layer, determine the output matrix corresponding to the first attention layer, and then based on the output matrix corresponding to the first attention layer, determine the first feature representation data.
[0131] In the above process, using the label representation vocabulary, that is, using the label representation data corresponding to various labels under the specified discrimination task as an agent, to assist the first attention layer in processing its input (that is, the word vector sequence corresponding to the user feature), so that each feature in the input perceives the task label information of the specified discrimination task and its own influence, so as to adjust the final output of the first attention layer, so that the first attention layer has the ability of task perception. And, combined with the vocabulary processing layer, it can better implement embedding the task label information into the calculation process of the attention mechanism. Through the first attention layer in the feature representation model, not only can the interaction between features in the input be captured, but also the sensitivity of the feature representation model to features highly relevant to the specified discrimination task can be significantly improved.
[0132] Similarly, through the first attention layer of the first representation network, label representation vocabulary is injected into the first user features, that is, label representation data corresponding to various labels under the specified discrimination task is injected, so as to obtain feature representation data that not only pays attention to the relevance between the first user features but also pays attention to the relevance between the first user features and the specified discrimination task, enabling the first feature representation data to be directly related to the specified discrimination task. The third feature representation data determined based on such first feature representation data can better improve the accuracy of the prediction result of the subsequent specified discrimination task.
[0133] In the above process, based on the innovative integration of task label information directly into the attention mechanism (combining user features and label representation vocabulary, and determining the feature representation data of user features through the first attention layer), by introducing task label information, the representation model can consider the relevance between the input and the target task while paying attention to the internal correlation of the input content, thereby narrowing the gap between the traditional attention mechanism and the discrimination task target. Moreover, based on the label representation vocabulary and user features, the feature representation data of user features is determined, which can avoid the leakage of the first label corresponding to the user features during the training process and ensure the privacy information security of the sample users. In addition, due to the addition of the direct perception of task label information in the feature representation model, the generalization ability of the representation model in new scenarios and data can be improved to a certain extent.
[0134] In the above example, after processing the label representation vocabulary through the vocabulary processing layer, a label matrix is obtained, and based on the label matrix, the weight distribution of the original query matrix (i.e., the first query matrix) is dynamically adjusted, which can better enhance the attention of the representation model to the discrimination task orientation and also ensure the flexibility and pertinence during the adjustment process.
[0135] In some possible examples, the aforementioned first representation network may include a second attention layer and weight masking information. The second attention layer, in conjunction with the weight masking information, can achieve dynamic screening (or dynamic masking) of features more relevant to the specified discrimination task. Such a second attention layer can be referred to as an attention layer of an attention mechanism adopting a dynamic masking mechanism.
[0136] Among them, the "first" in the first attention layer and the "second" in the second attention layer are only for convenience of description to distinguish different attention layers and do not have other limiting meanings. In some implementations, the first representation network may also include an attention layer, which can determine the first feature representation data in conjunction with the aforementioned label representation vocabulary and weight masking information, and this is also possible.
[0137] As Figure 6 shown, in the aforementioned step S210, it may include the following steps S610 - S630:
[0138] In step S610, based on the first user feature, the attention weight matrix and the second value matrix are determined through the second attention layer.
[0139] In some possible examples, the first representation network is a network based on the Transformer structure. The first representation network may deploy one or more attention layers arranged in series. One or more attention layers in the first representation network may be set as attention layers based on the aforementioned task-aware attention mechanism and / or the attention mechanism based on the dynamic masking mechanism, such as the second attention layer provided in the embodiments of this specification. Exemplarily, the second attention layer may be the first attention layer in the first representation network, or a non-first attention layer, which is all possible.
[0140] In some implementation manners, the first representation network may further include a processing layer corresponding to each attention layer after each attention layer. Exemplarily, the processing layer may be implemented as a feed-forward neural network layer or an expert network layer. The structure of the expert network layer is similar to the structure of the aforementioned expert network, and will not be elaborated here. During the training process of the representation model, one or more attention layers (including the second attention layer) in the first representation network, their corresponding processing layers, and the weight masking information are all trainable.
[0141] In a possible implementation manner, based on the first user feature, through the second attention layer, a third query matrix, a third key matrix, and a second value matrix may be obtained. The implementation principle of obtaining the third query matrix, the third key matrix, and the second value matrix is similar to the implementation principle of obtaining the first query matrix, the first key matrix, and the first value matrix described above. The implementation process may refer to the implementation process of obtaining the first query matrix, the first key matrix, and the first value matrix described above, and will not be elaborated here.
[0142] After that, in combination with the aforementioned preset attention weight formula, based on the third query matrix and the third key matrix, the attention weight matrix Y is determined. The preset attention weight formula may be expressed as:
[0143]
[0144] where Q″ represents the third query matrix, K ′ represents the third key matrix, d k ′ represents the dimension number of the third key matrix, and softmax(.) represents the activation function.
[0145] Then, in step S620, based on the attention weight matrix and the weight masking information, a masking matrix is obtained, where the weight masking information is used to indicate screening features more relevant to the specified discrimination task.
[0146] Considering that in the input (the first user feature), there may be some features that have little contribution or impact on the discrimination result of the specified discrimination task. During the process of the feature representation model representing the input, if the final feature representation data is determined based on all the features of the input, this may cause the feature representation (such as the third feature representation data) learned by the feature representation model to be interfered by irrelevant information, that is, some features that have little contribution or impact on the discrimination result of the specified discrimination task. To a certain extent, this may reduce the accuracy and efficiency of determining the discrimination result under the specified discrimination task based on this feature representation data.
[0147] In view of the above situation, the electronic device adopts an attention mechanism based on a dynamic masking mechanism to adaptively mask features that are not highly important for the specified discrimination task. Specifically, in step S620, weight masking information can be obtained, and based on each attention weight value in the attention weight matrix and the weight masking information, a masking matrix is obtained. This weight masking information is used to indicate the screening of features that are more relevant to the specified discrimination task.
[0148] It can be understood that in the process of representing the user feature based on the representation model, when it is a subprocess of the discrimination process based on the representation model, both the feature representation model and the weight masking information have been trained. Correspondingly, the electronic device can directly obtain the trained weight masking information and then execute the subsequent process. In the process of representing the user feature based on the representation model, when it is a subprocess of the training process of the representation model, the weight masking information has been adjusted based on the previous iteration process of the current iteration process (taking the k-th training iteration process as an example), that is, the (k - 1)-th iteration process. Correspondingly, the electronic device can obtain the weight masking information corresponding to the current iteration process (that is, the weight masking information adjusted by the (k - 1)-th iteration process) and then execute the subsequent process.
[0149] In the embodiments of this specification, the unimportant features in the second value matrix can be dynamically masked according to the feature importance, further highlighting the value of the important features in the second value matrix, that is, the effective features for the discrimination result of the specified discrimination task. At the same time, in order to ensure that the masked features are neither too many nor too few, upper and lower limits of the masking quantity are also configured. In some implementation manners, the weight masking information may include: a weight masking threshold, a maximum masking quantity, and a minimum masking quantity; the maximum masking quantity is used to limit the maximum number of unmasked features, and the minimum masking quantity is used to limit the minimum number of unmasked features.
[0150] Among them, the aforementioned feature importance can be reflected by the magnitudes of the respective attention weight values in the aforementioned attention weight matrix. The larger the corresponding attention weight value, the greater the correlation between the corresponding feature value in the second value matrix and the specified discrimination task.
[0151] The above process of obtaining the masking matrix may include: comparing each attention weight value in the attention weight matrix with the weight masking threshold in sequence; if a certain attention weight value is not less than the weight masking threshold, it indicates that the eigenvalue corresponding to this attention weight value in the second value matrix needs to be retained (or it indicates that the feature corresponding to this attention weight value in the output matrix corresponding to the second attention layer needs to be retained); if a certain attention weight value is less than the weight masking threshold, it indicates that the eigenvalue corresponding to this attention weight value in the second value matrix needs to be masked (or it indicates that the feature corresponding to this attention weight value in the output matrix corresponding to the second attention layer needs to be masked).
[0152] In addition, in order to retain sufficient information for subsequent accurate prediction and try to mask out features with low importance as much as possible, it is also necessary to ensure that the minimum number of the finally unmasked features in the second value matrix (or the output matrix corresponding to the second attention layer) is not less than the minimum masking quantity, and the maximum number of the unmasked features is not greater than the maximum masking quantity.
[0153] For example, the minimum masking quantity is a, and the maximum masking quantity is b, where b is greater than a. Each attention weight value in the attention weight matrix is compared with the weight masking threshold in sequence to determine the attention weight values that are not less than the weight masking threshold. Theoretically, the features corresponding to the determined attention weight values that are not less than the weight masking threshold (such as the corresponding eigenvalues in the second value matrix or the corresponding features in the output matrix corresponding to the second attention layer) need to be retained, and the features corresponding to other attention weight values that are less than the weight masking threshold need to be masked.
[0154] In practice, in one case, if the number c of the attention weight values that are not less than the weight masking threshold is greater than the maximum masking quantity b, in order to try to mask out features with low importance as much as possible, subsequently, c - b attention weight values that are not less than the weight masking threshold need to be determined from the c attention weight values that are not less than the weight masking threshold, and the features corresponding to them are masked so that the maximum number of the finally unmasked features does not exceed the maximum masking quantity b.
[0155] Exemplarily, the process of determining c - b attention weight values that are not less than the weight masking threshold from the c attention weight values that are not less than the weight masking threshold may be to determine the c - b attention weight values with the smallest values from the c attention weight values that are not less than the weight masking threshold.
[0156] In another case, after successively comparing each attention weight value in the attention weight matrix with the weight masking threshold, if the number d of attention weight values not less than the weight masking threshold is less than the minimum masking quantity a, in order to retain sufficient information for subsequent accurate prediction, subsequently, (a - d) attention weight values need to be selected from multiple attention weight values less than the weight masking threshold to retain the corresponding features (such as the corresponding eigenvalue in the second value matrix or the corresponding feature in the output matrix corresponding to the second attention layer), so that the minimum number of unmasked features is not less than the minimum masking quantity a.
[0157] Exemplarily, the foregoing selection of (a - d) attention weight values from multiple attention weight values less than the weight masking threshold may be: selecting (a - d) attention weight values with the largest values from multiple attention weight values less than the weight masking threshold.
[0158] Through the above method, based on the attention weight matrix and the weight masking information, a masking matrix is obtained. In some possible examples, the size of the masking matrix is the same as that of the attention weight matrix, and each element therein corresponds one-to-one to each attention weight value in the attention weight matrix. Each element in the masking matrix may take a first value (such as 1) or a second value (such as 0), where the first value indicates that the feature corresponding to the corresponding attention weight value needs to be retained, and the second value indicates that the feature corresponding to the corresponding attention weight value needs to be masked.
[0159] In some other possible examples, it may be to set the attention weight values in the attention weight matrix that need to mask the corresponding features determined by the above method and the weight masking information to 0; and retain the attention weight values in the attention weight matrix that need to retain the corresponding features determined by the above method and the weight masking information to obtain the masking matrix. Correspondingly, the masking matrix includes the attention weight values that need to retain the corresponding features, and the attention weight values that need to mask the corresponding features are set to 0.
[0160] After that, in step S630, the first feature representation data is determined using the masking matrix and the second value matrix.
[0161] In some possible examples, when each element in the masking matrix takes a first value (such as 1) or a second value (such as 0), the output matrix corresponding to the second attention layer may be determined based on the attention weight matrix and the second value matrix through the foregoing preset attention formula, and then the output matrix corresponding to the second attention layer is multiplied element by element with the masking matrix to obtain the masked output matrix corresponding to the second attention layer, as Figure 7As shown. Among them, the features with low importance relative to the specified discrimination task in the output matrix are masked, for example, set to 0, and the features with high importance relative to the specified discrimination task are retained.
[0162] In some other possible examples, when the attention weight values for retaining the corresponding features are included in the masking matrix and the attention weight values for masking the corresponding features are set to 0, the second value matrix can be multiplied element-wise with the masking matrix to obtain the masked output matrix corresponding to the second attention layer.
[0163] Exemplarily, taking the process of multiplying the second value matrix and the masking matrix element-wise as an example, the element-wise multiplication can be introduced as multiplying the element (or eigenvalue) in the nth row and mth column of the second value matrix with the element in the nth row and mth column of the masking matrix respectively to obtain the element (feature) in the nth row and mth column of the masked output matrix corresponding to the second attention layer.
[0164] After obtaining the masked output matrix corresponding to the aforementioned second attention layer, based on the masked output matrix corresponding to the second attention layer, the first feature representation data is determined. The process of determining the first feature representation data based on the masked output matrix corresponding to the second attention layer can refer to the process of determining the first feature representation data based on the output matrix corresponding to the first attention layer mentioned above, which will not be elaborated here.
[0165] In the above process, the features most relevant to the specified discrimination task can be determined from the input, and the invalid features in the input, that is, the features not important enough for the specified discrimination task, can be adaptively and dynamically reduced, reducing the impact of the invalid features in the input on the feature representation model, and further improving the accuracy of the subsequent discrimination results.
[0166] For the aforementioned several second representation networks and third representation networks, their structures can refer to the structure of the following first representation network. In some examples, the attention mechanisms (such as including task-aware attention mechanisms, attention mechanisms with dynamic masking mechanisms, and any attention mechanisms in related technologies) based on the attention layers deployed between the several second representation networks, third representation networks, and the first representation network can be the same or different.
[0167] In the above example, a flexible and adaptive control mechanism is provided. The degree of feature masking is determined by trainable parameters. During the training process of the representation model, three parameters, namely the weight masking threshold, the minimum masking number, and the maximum masking number, can be initialized and trained, enabling the model to automatically learn relatively more appropriate masking rules, reducing the trouble of manual configuration, ensuring the rationality of the threshold and the upper and lower limits of masking, ensuring that each sample still retains sufficient information for accurate prediction after feature masking, and ultimately enabling each sample to have its dynamic output, achieving the purpose of highlighting important features. Correspondingly, during the actual inference process, the model can dynamically and adaptively mask unimportant features in the input user features, and try to retain features that are more useful for the specified discrimination task, ensuring that the obtained feature representation data can determine a more accurate prediction and discrimination result.
[0168] Moreover, by selectively masking less important features in the input, the processing of useless or interfering information by the model is effectively reduced, improving the computational efficiency and the generalization ability of the model. The dynamic masking mechanism encourages the representation model to focus more on features closely related to the task objective, which helps to improve the performance of the model when facing datasets with high noise and sparse key information, and is particularly suitable for the optimization requirements of the representation model in discrimination tasks. The above solution reduces the impact of invalid features on the model, enabling the model to learn features that are more useful for the specified discrimination task. In this way, not only can the accuracy of the discrimination result in the subsequent specified discrimination task be improved, but also the training efficiency of the model can be optimized, and the consumption of computing resources can be reduced, thereby reducing costs while improving the business efficiency of the model.
[0169] In addition, during the dynamic masking process of features through weight masking information above, the individual differences of each input can be fully considered to ensure that the masking decision can reflect the uniqueness of the input, strengthening the robustness and accuracy of the model under complex data distributions and achieving the enhancement of input features.
[0170] The above dynamic masking mechanism and task-aware attention mechanism solutions can be seamlessly integrated into existing models based on the Transformer structure, such as the BERT model, the GPT model, etc., without fundamentally changing the model architecture, improving the practicality and promotion potential of the technical solutions. And they jointly solve the problems of information redundancy, inaccurate feature selection, and insufficient task orientation in the discrimination task, providing strong support for promoting the progress of deep learning technology in applications with high-precision requirements and resource sensitivity.
[0171] In some possible implementation manners, the process of characterizing user features based on the characterization model is a subprocess of the discrimination process based on the characterization model. That is, after obtaining the third feature characterization data (or aggregated feature characterization data) based on the aforementioned process of characterizing user features based on the characterization model, the discrimination process is continued based on this third feature characterization data (or aggregated feature characterization data). Correspondingly, as Figure 8 shown, an exemplary flowchart of a discrimination process based on a characterization model is illustrated. Among them, the characterization model includes a feature characterization model and a label characterization model. The feature characterization model includes a first characterization network and a first expert network, and the first expert network includes a gating subnet and multiple expert subnets arranged in parallel. The label characterization model includes a label characterization vocabulary, and the label characterization vocabulary includes label characterization data corresponding to various labels under a specified discrimination task.
[0172] In the discrimination process based on the characterization model, both the feature characterization model and the label characterization model therein have been trained. After the label characterization model is trained, the final label characterization data of various labels under a specified discrimination task is actually obtained. Then, based on the final label characterization data of various labels under any specified discrimination task, the corresponding discrimination task can be executed for the newly received user features.
[0173] Among them, the discrimination process based on the characterization model may include steps S810 - S870:
[0174] In step S810, obtain the first user feature of the first user.
[0175] In step S820, based on the first user feature, determine the first feature characterization data through the first characterization network.
[0176] In step S830, based on the first feature characterization data, obtain the first weights corresponding to each expert subnet through the gating subnet.
[0177] In step S840, based on the first feature characterization data, obtain the second feature characterization data corresponding to each expert subnet through each expert subnet.
[0178] In step S850, combine the first weights and the second feature characterization data corresponding to each expert subnet to determine the third feature characterization data.
[0179] Among them, the implementation principles of steps S810 - S850 are similar to those of the aforementioned steps S210 - S250, and the implementation process can refer to the implementation process of the aforementioned steps S210 - S250, which will not be elaborated here.
[0180] In step S860, calculate the target similarity between the third feature representation data and the label representation data corresponding to the target label in the target discrimination task, where the target discrimination task is any discrimination task in the specified discrimination tasks. The label representation data corresponding to the target label in the target discrimination task can be obtained from the aforementioned label representation vocabulary.
[0181] In this step, the electronic device can determine the target discrimination task from the specified discrimination tasks based on the user's discrimination requirement, and then calculate the target similarity between the third feature representation data and the label representation data corresponding to each target label in the target discrimination task. Among them, the target similarity can be, for example, cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, and so on.
[0182] Taking cosine similarity as an example, it can be to perform an inner product on the third feature representation data and the label representation data corresponding to the target label, so as to obtain the above-mentioned target similarity.
[0183] In step S870, determine the predicted discrimination result for the first user according to the target similarity.
[0184] Exemplarily, the target similarity can include the respective target similarities between the feature representation data and the label representation data corresponding to multiple target labels in the target discrimination task; correspondingly, the target label corresponding to the target similarity with the largest value can be used as the predicted discrimination result of the user feature in the target discrimination task.
[0185] Another example is that in the case of a binary classification task of the target discrimination task, the target similarity can be the similarity between the feature representation data and the label representation data corresponding to any target label in the target discrimination task. Correspondingly, if the value of the target similarity exceeds the specified threshold, it can be determined that the predicted discrimination result of the user feature in the target discrimination task conforms to the target label (that is, it is determined that the user indicated by the user feature belongs to the category indicated by the target label in the target discrimination task); if the value of the target similarity does not exceed the specified threshold, it can be determined that the predicted discrimination result of the user feature in the target discrimination task does not conform to the target label (that is, it is determined that the user indicated by the user feature does not belong to the category indicated by the target label in the target discrimination task and belongs to another category in the target discrimination task).
[0186] In the above process, when performing discrimination based on the representation model for the user, only the similarity between the feature representation data corresponding to the user feature of the user and the label representation data corresponding to the target label in the target discrimination task needs to be calculated to achieve discrimination for the user, thereby greatly improving the discrimination efficiency for the user.
[0187] It can be understood that for the process of determining the predicted discrimination result of the first user in the target discrimination task based on the aggregated feature representation data, reference can be made to the process of determining the predicted discrimination result of the first user in the target discrimination task based on the third feature representation data, which will not be elaborated here.
[0188] In some possible implementation manners, the above process of representing user features based on the representation model can be a subprocess of the training process of the representation model. That is, after obtaining the third feature representation data (or aggregated feature representation data) based on the above process of representing user features based on the representation model, it is necessary to jointly use the third feature representation data (or aggregated feature representation data) and the label corresponding to the first user feature to continue training the representation model. Among them, as Figure 9A , a schematic diagram of the principle of training the representation model is shown. Among them, the representation model includes a feature representation model and a label representation model. During the training process, based on the user features (such as the first user features), the feature representation data corresponding to the user features (such as the third feature representation data) can be obtained through the feature representation model. Based on the first label corresponding to the user features, through the label representation model (such as from the label representation word list obtained based on the label representation model), the first label representation data corresponding to the first label is obtained. After that, at least the similarity between the feature representation data corresponding to the user features and the first label representation data is jointly determined to determine the representation loss, and the representation model is trained with the goal of minimizing the representation loss, that is, the parameters of the feature representation model and the label representation model are adjusted.
[0189] Correspondingly, as Figure 9B shown, schematic diagrams of a training process of a representation model are respectively exemplarily shown. Among them, the representation model includes a feature representation model, and the feature representation model includes a first representation network and a first expert network. The first expert network includes a gating subnet and a plurality of expert subnets arranged in parallel. The training process of this representation model can include steps S910 - S990:
[0190] In step S910, the first user features of the first user are obtained.
[0191] In step S920, based on the first user features, through the first representation network, the first feature representation data is determined.
[0192] In step S930, based on the first feature representation data, through the gating subnet, the first weights corresponding to each expert subnet are obtained.
[0193] In step S940, based on the first feature representation data, through each expert subnet, the second feature representation data corresponding to each expert subnet is obtained.
[0194] In step S950, the first weights corresponding to each expert sub-network and the second feature representation data are combined to determine the third feature representation data.
[0195] Among them, the implementation principles of steps S910 - S950 are similar to those of the aforementioned steps S210 - S250, and the implementation process can refer to the implementation process of the aforementioned steps S210 - S250, which will not be elaborated here.
[0196] In step S960, the first label under the specified discrimination task corresponding to the first user feature is obtained. In this step, the first label under the specified discrimination task corresponding to the user feature can be obtained from the training set used to train the representation model.
[0197] In step S970, the first label representation data corresponding to the first label is determined from the label representation vocabulary.
[0198] In this step, in the current iteration process (illustrated by the k-th training iteration process as an example), the label representation vocabulary has been obtained after adjusting the label representation model based on the previous k - 1 iteration processes. The label representation vocabulary includes the label representation data corresponding to various labels under the specified discrimination task. Correspondingly, the electronic device can determine the first label representation data corresponding to the first label from the label representation vocabulary. As Figure 9A shown, the first label representation data corresponding to the first label can also be said to be determined by the label representation model.
[0199] In step S980, based on the first similarity between the third feature representation data and the first label representation data, the representation loss is determined, where the representation loss is negatively correlated with the first similarity value.
[0200] In this step, the first similarity between the third feature representation data and the first label representation data is calculated. As shown before, the first similarity can be, for example, cosine similarity, Euclidean distance, Manhattan distance, and Pearson correlation coefficient, etc. Taking cosine similarity as an example, it can be the inner product of the feature representation data and the first label representation data to obtain the above first similarity. Then, based on this first similarity, the representation loss is determined.
[0201] In some embodiments, the process of determining the representation loss can be: using the first loss function, based on the first similarity, to determine the representation loss. Wherein, the representation loss is negatively correlated with the first similarity. Wherein, the first loss function and the subsequent second loss function and third loss function can be the binary cross-entropy loss function BCEWithLogitsLoss, or can also be the focal loss, etc.
[0202] In some examples, there may be multiple specified discrimination tasks. Correspondingly, the first label may include the labels of the first user under each specified discrimination task. Correspondingly, the first label characterization data may include the label characterization data corresponding to each label of the first user under each specified discrimination task. Calculating the first similarity between the third feature characterization data and the first label characterization data may include calculating the respective first similarities between the third feature characterization data and each of the first label characterization data. After that, a characterization loss may be determined based on each of the first similarities using a first loss function.
[0203] In some possible implementation manners, in this embodiment, the dissimilarity loss between the labels within each task may also be combined to determine the characterization loss, so as to train the feature characterization model and the label characterization model. Correspondingly, in step S980, it may include steps 21-22:
[0204] In step 21, calculate the second similarity between the label characterization data corresponding to each type of label under the specified discrimination task. In this step, for each specified discrimination task, calculate the second similarity between the label characterization data corresponding to every two labels among each type of label under the specified discrimination task to obtain a number of second similarities.
[0205] In step 22, determine the characterization loss based on the first similarity between the third feature characterization data and the first label characterization data and the second similarity, where the characterization loss is also positively correlated with the second similarity and negatively correlated with the first similarity.
[0206] In this step, the first loss function may be used to determine the first loss based on the first similarity, where the first loss is negatively correlated with the first similarity; a second loss function may be used to determine the second loss based on a number of second similarities, where the second loss is positively correlated with each of the number of second similarities; after that, the characterization loss is determined based on the sum or mean of the first loss and the second loss.
[0207] In some possible implementation manners, in this embodiment, the dissimilarity loss between the labels within different tasks may also be combined to determine the characterization loss, so as to train the feature characterization model and the label characterization model. Correspondingly, in step S980, it may include 31-32:
[0208] In step 31, calculate the third similarity between the label characterization data corresponding to every two labels under different specified discrimination tasks. The process of calculating the third similarity may refer to the process of calculating the first similarity described above and will not be elaborated here.
[0209] In step 32, based on the first similarity between the third feature representation data and the first label representation data, and the third similarity, determine the representation loss, where the representation loss is also positively correlated with the third similarity and negatively correlated with the first similarity.
[0210] In this step, the foregoing first loss function can be used to determine the first loss based on the first similarity, where the first loss is negatively correlated with the first similarity; the third loss function can be used to determine the third loss based on a number of third similarities, where the third loss is positively correlated with each of the number of third similarities; and then, based on the sum or mean of the first loss and the third loss, determine the representation loss.
[0211] In another implementation, the foregoing first similarity, second similarity, and third similarity can also be combined to determine the representation loss. The implementation method of this determination process can refer to the process of determining the representation loss based on the first similarity and the second similarity described above, and will not be elaborated here.
[0212] In step S990, train the representation model by minimizing the representation loss. In this step, by minimizing the representation loss, that is, with the goal of making the third feature representation data and the first label representation data more similar, and making the label representation data corresponding to different labels under the same specified discrimination task and the label representation data corresponding to different labels under different specified discrimination tasks more dissimilar, adjust the parameters of the representation model, that is, adjust the parameters of the feature representation model (and the label representation model).
[0213] When adjusting the parameters of the representation model through the representation loss determined based on the first similarity, it can be ensured that the feature representation data of samples with the same label are close to each other, so that the representation model can better learn the similarity between samples, thereby promoting the effective extraction of features.
[0214] When adjusting the parameters of the representation model through the representation loss determined based on the first similarity and the second similarity (and the third similarity), it can be achieved that while ensuring that the feature representation data of samples with the same label are close to each other, so that the representation model can better learn the similarity between samples, thereby promoting the effective extraction of features, it can also create a greater dissimilarity (i.e., difference) between different labels and / or tasks. In other words, it is to ensure that the label representation data corresponding to different labels are independent and dissimilar to each other, so that the model can clearly distinguish different labels under the same specified discrimination task (and different labels under different specified discrimination tasks), and can encourage different samples corresponding to the same specified discrimination task to be as dispersed as possible in the embedding space to enhance the discriminability of the representation.
[0215] Adjust the parameters of the feature representation model through the above-mentioned representation loss, enabling the model to learn more detailed and accurate label boundaries, promoting the consistency of representations within the same label (e.g., within the same category) and the distinctiveness of representations between different labels (e.g., between different categories). By combining the first loss, second loss, and third loss to determine the representation loss, the feature representation model can better refine and optimize the feature representation data, thereby effectively improving its performance in discrimination tasks, especially in scenarios with a large number of similar categories or unclear category boundaries.
[0216] The feature representation model provided in the above example can flexibly select different functional components according to the requirements of different stages. Each functional module in the feature representation model is modularized and customized, which is lacking in traditional models, and can better improve the broader applicability and higher flexibility of the feature representation model.
[0217] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments, and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible, or may be advantageous.
[0218] Corresponding to the above method embodiments, an embodiment of this specification provides a device 1000 for representing user features based on a representation model. The representation model includes a feature representation model, and the feature representation model includes a first representation network and a first expert network. The first expert network includes a gating sub-network and a plurality of expert sub-networks arranged in parallel. Its schematic block diagram is as Figure 10 shown, including:
[0219] A first acquisition module 1010, configured to acquire the first user feature of a first user;
[0220] A first determination module 1020, configured to determine first feature representation data based on the first user feature through the first representation network;
[0221] A first obtaining module 1030, configured to obtain first weights corresponding to each expert sub-network based on the first feature representation data through the gating sub-network;
[0222] A second obtaining module 1040, configured to obtain second feature representation data corresponding to each expert sub-network based on the first feature representation data through each expert sub-network;
[0223] The second determination module 1050 is configured to determine third feature representation data by combining the first weights corresponding to each expert sub-network and the second feature representation data.
[0224] In some possible examples, the first representation network includes a first attention layer;
[0225] The first determination module 1020 includes: an acquisition unit (not shown in the figure), configured to acquire a label representation vocabulary, where the label representation vocabulary includes label representation data corresponding to various labels under a specified discrimination task; a first determination unit (not shown in the figure), configured to determine the first feature representation data based on the first user feature and the label representation vocabulary through the first attention layer.
[0226] In some possible examples, the first determination unit (not shown in the figure) includes:
[0227] A first determination sub-module (not shown in the figure), configured to determine a first query matrix, a first key matrix, and a first value matrix based on the first user feature through the first attention layer;
[0228] A second determination sub-module (not shown in the figure), configured to determine a second query matrix based on the label representation vocabulary and the first query matrix;
[0229] A third determination sub-module (not shown in the figure), configured to determine the first feature representation data based on the second query matrix, the first key matrix, and the first value matrix through the first attention layer.
[0230] In some possible examples, the first representation network further includes: a vocabulary processing layer;
[0231] The second determination sub-module (not shown in the figure) includes: a first obtaining sub-unit (not shown in the figure), configured to obtain a label matrix based on the label representation vocabulary through the vocabulary processing layer; a first determination sub-unit (not shown in the figure), configured to determine a second query matrix based on the label matrix and the first query matrix.
[0232] In some possible examples, the vocabulary processing layer includes a first feed-forward network, a second feed-forward network, and a preset activation function;
[0233] The first obtaining sub-unit (not shown in the figure) is specifically configured to obtain a label weight matrix based on the label representation vocabulary through the first feed-forward network and the preset activation function; obtain a label bias matrix based on the label representation vocabulary through the second feed-forward network, so as to obtain the label matrix.
[0234] In some possible examples, the characterization model further includes a label characterization model, which is used to vectorize the labels under a specified discrimination task; the obtaining unit is specifically configured to: based on multiple labels under the specified discrimination task, through the label characterization model, obtain the label characterization vocabulary.
[0235] In some possible examples, there are multiple specified discrimination tasks;
[0236] The obtaining unit is specifically configured to: based on multiple labels under each specified discrimination task and the task identifier of the specified discrimination task corresponding to each label, through the label characterization model, obtain the label characterization matrix.
[0237] In some possible examples, it further includes: a second obtaining module (not shown in the figure), configured to obtain the first label under the specified discrimination task corresponding to the first user feature; a third determining module (not shown in the figure), configured to determine, from the label characterization vocabulary, the first label characterization data corresponding to the first label; a fourth determining module (not shown in the figure), configured to determine a characterization loss based on a first similarity between the third feature characterization data and the first label characterization data, where the characterization loss is negatively correlated with the first similarity value; a training module (not shown in the figure), configured to train the characterization model by minimizing the characterization loss.
[0238] In some possible examples, the fourth determining module (not shown in the figure) is specifically configured to calculate a second similarity between the label characterizations corresponding to various labels under the specified discrimination task;
[0239] Based on the first similarity and the second similarity, determine the characterization loss, where the characterization loss is also positively correlated with the second similarity.
[0240] In some possible examples, the fourth determining module (not shown in the figure) is specifically configured to calculate a third similarity between the label characterization data corresponding to two labels under different specified discrimination tasks;
[0241] Based on the first similarity and the third similarity, determine the characterization loss, where the characterization loss is also positively correlated with the third similarity.
[0242] In some possible examples, the first characterization network includes a second attention layer and weight masking information;
[0243] The first determining module 1020 is specifically configured to, based on the first user feature, through the second attention layer, determine an attention weight matrix and a second value matrix;
[0244] Based on the attention weight matrix and the weight masking information, a masking matrix is obtained, where the weight masking information is used to indicate screening features more relevant to a specified discrimination task;
[0245] Using the masking matrix and the second value matrix, the first feature representation data is obtained.
[0246] In some possible examples, the weight masking information includes: a weight masking threshold, a maximum masking number, and a minimum masking number; the maximum masking number is used to limit the maximum number of unmasked features, and the minimum masking number is used to limit the minimum number of unmasked features.
[0247] In some possible examples, the feature representation model further includes a sequence segmentation network;
[0248] The apparatus further includes: a third obtaining module (not shown in the figure), configured to, before determining the first feature representation data based on the first user feature through the first representation network, segment the first user feature through the sequence segmentation network to obtain a plurality of first user feature blocks;
[0249] The first determining module 1020 is specifically configured to respectively determine, based on each first user feature block, the feature representation data corresponding to each first user feature block through the first representation network, and further jointly determine the first feature representation data.
[0250] In some possible examples, the first user feature is the user feature corresponding to the first feature group of the first user; the feature representation model further includes second representation networks and second expert networks corresponding to a plurality of second feature groups; the feature representation model further includes: an aggregation network; the apparatus further includes: a third obtaining module (not shown in the figure), configured to obtain a plurality of second user features corresponding to a plurality of second feature groups of the first user;
[0251] A fourth obtaining module (not shown in the figure), configured to respectively process the second user features corresponding to each second feature group through the second representation network and the second expert network corresponding to each second feature group to obtain a plurality of fourth feature representation data;
[0252] An aggregation module (not shown in the figure), configured to aggregate the third feature representation data and the plurality of fourth feature representation data through the aggregation network to obtain aggregated feature representation data.
[0253] In some possible examples, the aggregation module (not shown in the figure) is specifically configured to aggregate the third feature representation data and the plurality of fourth feature representation data through the aggregation network based on preset activation information to obtain aggregated feature representation data.
[0254] In some possible examples, the characterization model further includes a label characterization model, the label characterization model includes a label characterization vocabulary, and the label characterization vocabulary includes label characterization data corresponding to various labels under a specified discrimination task;
[0255] The apparatus further includes: a calculation module (not shown in the figure), configured to calculate a target similarity between the third feature characterization data and the label characterization data corresponding to a target label under a target discrimination task, where the target discrimination task is any discrimination task in the specified discrimination tasks; a fifth determination module (not shown in the figure), configured to determine a predicted discrimination result for the first user according to the target similarity.
[0256] The above apparatus embodiments correspond to the method embodiments. For specific descriptions, reference may be made to the descriptions in the method embodiment section, which will not be elaborated here. The apparatus embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference may be made to the corresponding method embodiments.
[0257] An embodiment of this specification also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method for characterizing user features based on a characterization model provided in this specification.
[0258] An embodiment of this specification also provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method for characterizing user features based on a characterization model provided in this specification is implemented.
[0259] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the storage medium and the computing device, since they are basically similar to the method embodiments, the descriptions are relatively simple. For relevant parts, reference can be made to the partial descriptions of the method embodiments.
[0260] Those skilled in the art should be able to realize that, in the above one or more examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0261] The specific embodiments described above further elaborate in detail the objectives, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above description is only the specific embodiments of the embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for characterizing user features based on a characterization model, wherein the characterization model includes a feature characterization model, wherein the feature characterization model includes a first characterization network and a first expert network, wherein the first expert network includes a gated subnetwork and a plurality of expert subnetworks arranged in parallel, wherein the method includes: Acquire a first user feature of a first user; Based on the first user feature, determining first feature representation data through the first representation network; Based on the first feature representation data, obtaining first weights corresponding to each expert sub-network through the gating sub-network; Based on the first feature characterization data, obtaining second feature characterization data corresponding to each expert sub-network through each expert sub-network; The first weights and the second feature representation data corresponding to each expert sub-network are combined to determine the third feature representation data.
2. The method of claim 1, wherein: The first representation network includes a first attention layer; The determining, based on the first user feature, first feature characterization data through the first characterization network includes: Obtaining a label representation vocabulary, wherein the label representation vocabulary includes label representation data corresponding to various labels under the specified discrimination task; Based on the first user feature and the tag representation vocabulary, the first feature representation data is determined through the first attention layer.
3. The method of claim 2, wherein: The determining the first feature characterization data includes: Based on the first user feature, determining a first query matrix, a first key matrix, and a first value matrix through the first attention layer; Determine a second query matrix based on the label representation vocabulary and the first query matrix; Based on the second query matrix, the first key matrix and the first value matrix, the first feature representation data is determined through the first attention layer.
4. The method of claim 3, wherein: The first representation network also includes: a vocabulary processing layer; The determining a second query matrix based on the label representation vocabulary and the first query matrix includes: Based on the label representation vocabulary, a label matrix is obtained through the vocabulary processing layer; Based on the label matrix and the first query matrix, a second query matrix is determined.
5. The method of claim 4, wherein: The vocabulary processing layer includes a first feedforward network, a second feedforward network and a preset activation function; The obtained label matrix includes: Based on the label representation vocabulary, a label weight matrix is obtained through the first feedforward network and the preset activation function; Based on the label representation vocabulary, a label bias matrix is obtained through the second feedforward network, thereby obtaining the label matrix.
6. The method of claim 2, wherein: The representation model also includes a label representation model, and the label representation model is used to vectorize the labels under the specified discrimination task; The step of obtaining a label representation vocabulary includes: Based on the multiple labels under the specified discrimination task, the label representation vocabulary is obtained through the label representation model.
7. The method according to claim 6, wherein the designated identification tasks are multiple; The step of obtaining a label representation vocabulary includes: Based on a plurality of labels under each designated discrimination task and a task identifier of the designated discrimination task corresponding to each label, the label representation matrix is obtained through the label representation model.
8. The method of claim 2, further comprising: Obtaining a first label under the specified discrimination task corresponding to the first user feature; Determining first label representation data corresponding to the first label from the label representation vocabulary; Determining a representation loss based on a first similarity between the third feature representation data and the first label representation data, wherein the representation loss is negatively correlated with the first similarity value; The representation model is trained to minimize the representation loss.
9. The method of claim 8, wherein: The determining of the characterization loss comprises: Calculating the second similarity between the label representations corresponding to the various labels under the specified discrimination task; The representation loss is determined based on the first similarity and the second similarity, wherein the representation loss is also positively correlated with the second similarity.
10. The method of claim 8, wherein determining the characterization loss comprises: Calculate the third similarity between the label representation data corresponding to two labels under different specified discrimination tasks; The representation loss is determined based on the first similarity and the third similarity, wherein the representation loss is also positively correlated with the third similarity.
11. The method of claim 1, wherein: The first representation network includes a second attention layer and weight masking information; The determining, based on the first user feature, first feature characterization data through the first characterization network layer includes: Based on the first user feature, determining an attention weight matrix and a second value matrix through the second attention layer; Based on the attention weight matrix and the weight masking information, a masking matrix is obtained, wherein the weight masking information is used to indicate the selection of features that are more relevant to the specified discrimination task; The first feature characterization data is obtained by using the masking matrix and the second value matrix.
12. The method of claim 11, wherein: The weight masking information includes: a weight masking threshold, a maximum masking number and a minimum masking number; the maximum masking number is used to limit the maximum number of unmasked features, and the minimum masking number is used to limit the minimum number of unmasked features.
13. The method according to any one of claims 1 to 12, wherein: The feature representation model also includes a sequence segmentation network; Before determining the first feature characterization data based on the first user feature through the first characterization network, the method further includes: Segmenting the first user feature by the sequence segmentation network to obtain a plurality of first user feature blocks; The determining, based on the first user feature, first feature characterization data through the first characterization network includes: Based on each first user feature block respectively, feature representation data corresponding to each first user feature block is determined through the first representation network, and then the first feature representation data is jointly determined.
14. The method according to any one of claims 1 to 12, wherein: The first user feature is a user feature corresponding to a first feature group of the first user; the feature representation model further includes a second representation network and a second expert network corresponding to a plurality of second feature groups; The feature characterization model also includes: an aggregation network; The method further comprises: Acquire a plurality of second user features corresponding to a plurality of second feature groups of the first user; Processing the second user features corresponding to each second feature group through the second representation network and the second expert network corresponding to each second feature group, respectively, to obtain a plurality of fourth feature representation data; The third feature characterization data and the plurality of fourth feature characterization data are aggregated through the aggregation network to obtain aggregated feature characterization data.
15. The method according to claim 14, wherein aggregating the third feature characterization data and the plurality of fourth feature characterization data through the aggregation network comprises: The third feature characterization data and the plurality of fourth feature characterization data are aggregated through the aggregation network based on preset activation information to obtain aggregated feature characterization data.
16. The method according to any one of claims 1 to 12, wherein: The representation model further includes a label representation model, the label representation model includes a label representation vocabulary, and the label representation vocabulary includes label representation data corresponding to various labels under the specified discrimination task; The method further comprises: Calculating the target similarity between the third feature representation data and the label representation data corresponding to the target label under the target discrimination task, wherein the target discrimination task is any discrimination task in the specified discrimination task; A prediction result for the first user is determined according to the target similarity.
17. A user feature characterization device based on a characterization model, the characterization model comprising a feature characterization model, the feature characterization model comprising a first characterization network and a first expert network, the first expert network comprising a gated sub-network and a plurality of expert sub-networks arranged in parallel, the device comprising: A first acquisition module, configured to acquire a first user feature of a first user; A first determining module is configured to determine first feature characterization data based on the first user feature through the first characterization network; A first obtaining module is configured to obtain first weights corresponding to each expert sub-network through the gating sub-network based on the first feature representation data; A second obtaining module is configured to obtain second feature characterization data corresponding to each expert sub-network through each expert sub-network based on the first feature characterization data; The second determination module is configured to determine the third feature representation data by combining the first weights and the second feature representation data corresponding to each expert sub-network.
18. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 16 is implemented.