Method and device for representing user features based on representation model
By injecting label representation vocabulary into user characteristics, the task-perceptual attention mechanism of the representation model is used to solve the problem that existing models only focus on input internal logic when generating representations, and improve the prediction accuracy of the discriminant class model.
Patent Information
- Application Number
- CN202510243993.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-24
AI Technical Summary
When generating representations, existing deep learning-based models only focus on the internal logic and features of the input itself, resulting in the limitation of the accuracy of the prediction results of the discriminant class model.
A task-aware attention mechanism based on the representation model is adopted to inject label representation vocabulary into user characteristics to generate feature representation data more relevant to the discriminant task.
The accuracy of the prediction results of the discriminant class model is improved, so that the feature representation data can be more directly related to the designated discriminant task, thereby improving the prediction effect of the discriminant task.
Smart Images

Figure CN120197124A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of data processing, and in particular, to a method and apparatus for characterizing user features based on a characterization model. Background Art
[0002] With the rapid development of fields such as natural language processing (NLP) technology and computer vision technology, deep learning-based models have become the key technologies for solving complex problems in these fields. Currently, in the process of a deep learning-based model performing various tasks, when generating a characterization of a given input, it generally pays more attention to the internal logic and features of the given input itself to obtain an input characterization of the given input, and then determines prediction discrimination data based on the input characterization, and jointly trains the model with the prediction discrimination data and label data.
[0003] In the above process, for a discriminative model (i.e., a deep learning-based model for performing a discrimination task), its key task is to predict the relationship between a given input and a specific task target. For example, for a model that divides users into groups, that is, a user classification model, it pays more attention to the correlation between the input user features and the classification label. In the above process of generating a characterization of a given input, only paying attention to the internal logic and features of the given input itself has certain limitations in improving the accuracy of the prediction results of such models.
[0004] Then, how to provide an improved method for characterizing user features to improve the accuracy of the prediction results of discriminative models has become an urgent problem to be solved. Summary of the Invention
[0005] One or more embodiments of this specification provide a method and apparatus for characterizing user features based on a characterization model to implement a task-aware attention mechanism, so as to obtain feature characterization data that is more relevant to the task label information of the discrimination task, thereby improving the accuracy of subsequent discrimination prediction results.
[0006] According to a first aspect, there is provided a method for characterizing user features based on a characterization model, where the characterization model includes a feature characterization model, and the feature characterization model includes a first attention layer. The method includes:
[0007] Obtain user features;
[0008] Obtain a label characterization vocabulary, where the label characterization vocabulary includes label characterization data corresponding to various labels under a specified discrimination task;
[0009] Based on the user features and the label characterization vocabulary, determine the feature characterization data of the user features through the first attention layer.
[0010] According to a second aspect, there is provided an apparatus for characterizing user features based on a characterization model, the characterization model including a feature characterization model, the feature characterization model including a first attention layer, the apparatus including:
[0011] A first acquisition module configured to acquire user features;
[0012] A second acquisition module configured to acquire a label characterization vocabulary, wherein the label characterization vocabulary includes label characterization data corresponding to various labels under a specified discrimination task;
[0013] A first determination module configured to determine, based on the user features and the label characterization vocabulary and through the first attention layer, the feature characterization data of the user features.
[0014] According to a third aspect, there is provided a computer-readable storage medium having stored thereon a computer program which, when executed on a computer, causes the computer to execute the method described in the first aspect.
[0015] According to a fourth aspect, there is provided a computing device including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0016] According to the method and apparatus for characterizing user features based on a characterization model provided in the embodiments of the present specification, user features are acquired, and a label characterization vocabulary including label characterization data corresponding to various labels under a specified discrimination task is acquired. Then, based on the user features and the label characterization vocabulary, through the first attention layer included in the feature characterization model of the characterization model, the feature characterization data of the user features is determined. In the above process, through the first attention layer, the label characterization vocabulary is injected into the user features, that is, the label characterization data corresponding to various labels under the specified discrimination task is injected, so as to obtain feature characterization data that not only pays attention to the correlation between user features but also pays attention to the correlation between user features and the specified discrimination task, enabling the feature characterization data to be directly related to the specified discrimination task, thereby better improving the accuracy of the prediction result of the subsequent specified discrimination task. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic diagram of the implementation framework of an embodiment disclosed in this specification;
[0019] Figure 2 A flowchart of a method for characterizing user features based on a characterization model provided for an embodiment;
[0020] Figure 3A A flowchart of a process for determining feature characterization data provided for an embodiment;
[0021] Figure 3B Another flowchart of a process for determining feature characterization data provided for an embodiment;
[0022] Figure 4 Another flowchart of a method for characterizing user features based on a characterization model provided for an embodiment;
[0023] Figure 5 A flowchart of a discrimination process based on a characterization model provided for an embodiment;
[0024] Figure 6A A schematic diagram of a training process of a characterization model provided for an embodiment;
[0025] Figure 6B A flowchart of a training process of a characterization model provided for an embodiment;
[0026] Figure 7 Another flowchart of a training process of a characterization model provided for an embodiment;
[0027] Figure 8 A schematic block diagram of a device for characterizing user features based on a characterization model provided for an embodiment. Detailed implementation manners
[0028] Next, the technical solutions of the embodiments of the present specification will be described in detail with reference to the accompanying drawings.
[0029] The embodiments of the present specification disclose a method and a device for characterizing user features based on a characterization model. First, the application scenarios and technical concepts of the method will be introduced as follows:
[0030] Currently, in the process of performing various tasks, deep learning-based models generally pay more attention to the internal logic and features of the given input when generating the representation of the given input. Specifically, for example, in a model based on the Transformer architecture, the design intention of its core component, the Self-Attention mechanism, is to identify the correlation between elements in the input sequence, enabling the model to consider the context of the entire input sequence when generating each output. This Self-Attention mechanism has brought great success to models performing generation tasks such as language models (LMs) (such as BERT, GPT), etc., because their goal is to better understand the internal logic and structure of the input.
[0031] However, for some non-generation models, such as discriminative models, whose key task is to predict the relationship between a given input and a specific task target. For example, a user classification model pays more attention to the correlation between the user features of the input and the classification label. In the process of generating the representation of the given input as mentioned above, only focusing on the internal logic and features of the given input itself has certain limitations in improving the accuracy of the prediction results of such models.
[0032] To address the above problems and requirements, we propose a new Task Aware attention mechanism. It embeds task label information into the attention calculation process to bridge the gap between the focus of the attention mechanism in discriminative models and the discriminative task objectives of discriminative models. Through this mechanism, when a discriminative model generates the representation of the input, its attention is no longer just the self-correlation of the input content, but also simultaneously focuses on the correlation and relevance between the input content and its discriminative task objective, thereby improving its performance, that is, improving the prediction accuracy of its discriminative results.
[0033] In view of this, the inventors propose a method for representing user features based on a representation model, so as to obtain feature representation data that not only pays attention to the correlation of the user features themselves, but also pays attention to the correlation and relevance between the user features and their discriminative task objectives during the process of representing user features.
[0034] Figure 1 The schematic diagram of an implementation scenario according to an embodiment disclosed in this specification is shown. In this implementation scenario, the representation model includes a feature representation model, and the feature representation model includes a first attention layer. During the process of representing user features, user features are obtained; and a label representation vocabulary is obtained, where the label representation vocabulary includes label representation data corresponding to various labels under a specified discriminative task; then, based on the user features and the label representation vocabulary, through the first attention layer, the feature representation data of the user features is determined.
[0035] Exemplarily, the label representation vocabulary can be obtained through the label representation model included in the representation model.
[0036] In the above process, through the first attention layer of the feature representation model, the label representation vocabulary is injected into the user features, that is, the label representation data corresponding to various labels under the specified discrimination task is injected, so as to obtain feature representation data that not only pays attention to the correlation between user features but also pays attention to the correlation between user features and the specified discrimination task, making the feature representation data directly related to the specified discrimination task, thereby better improving the accuracy of the prediction result of the subsequent specified discrimination task.
[0037] Next, combined with specific embodiments, the method for representing user features based on the representation model provided in this specification will be elaborated in detail.
[0038] Figure 2 The flowchart of the method for representing user features based on the representation model in an embodiment of this specification is shown. This method is executed by an electronic device, and the electronic device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities.
[0039] Among them, the representation model includes a feature representation model, and the feature representation model includes a first attention layer. Exemplarily, the feature representation model can be understood as an encoder model for encoding the input, and its output can be used to execute the discrimination process under the subsequent specified discrimination task.
[0040] In some implementation manners, the feature representation model can be implemented as a model based on the Transformer structure, such as the Bert model, or can be implemented as other models that can deploy attention layers. Taking the feature representation model as a model based on the Transformer structure as an example, the feature representation model can deploy multiple attention layers arranged in series. One or more attention layers in the feature representation model can be set as attention layers based on the task-aware attention mechanism (and / or the subsequent mentioned weight-based dynamic masking mechanism), such as the first attention layer provided in the embodiments of this specification. Exemplarily, the first attention layer can be the first attention layer or a non-first attention layer in the feature representation model, which is all acceptable.
[0041] In some implementation manners, the feature representation model may further include a feed-forward network layer corresponding to each attention layer after each attention layer.
[0042] Next, the process of representing user features based on the representation model will be introduced. As Figure 2 shown, the method for representing user features based on the representation model includes the following steps S210 - S230:
[0043] In step S210, user features are obtained.
[0044] It can be understood that in one implementation, the process of representing user features based on the representation model can be a subprocess in the training process of the representation model. That is to say, the feature representation data corresponding to the user features obtained by the process of representing user features based on the representation model can continue to be jointly used with the labels corresponding to the user features to train the representation model. Correspondingly, the user features can be the features of any sample user in the training set used to train the representation model.
[0045] In another implementation, the process of representing user features based on the representation model can be a subprocess in the discrimination process based on the representation model. That is to say, the feature representation data corresponding to the user features obtained by the process of representing user features based on the representation model can continue to be used to determine the discrimination result of the user features under a specified discrimination task, that is, the discrimination result of the user indicated by the user features under the specified discrimination task. Correspondingly, the user features can be the features of any user who needs to be discriminated under a specified discrimination task.
[0046] Among them, the feature representation data here can refer to a vector or a matrix.
[0047] Exemplarily, the user features can include but are not limited to user portraits, transaction records, scene activity levels, historical risk information, user operation behaviors for a specified application, and user gang information, etc. The specified application can include but is not limited to electronic payment platforms, electronic trading platforms, and financial platforms, etc.
[0048] In step S220, a label representation vocabulary is obtained, where the label representation vocabulary includes label representation data corresponding to various labels under a specified discrimination task.
[0049] Among them, the label representation data here can refer to a vector or a matrix.
[0050] In some implementations, the specified discrimination task can include but is not limited to any one of the following: various classification tasks regarding users, such as but not limited to: user risk classification, user intention classification, user rental classification, whether the user is a recommended user for a certain type of object (such as goods, services, etc.), user credit assessment, etc., and can also include the discrimination task for the user's affiliated gang and other regression tasks for specified dimensions, etc.
[0051] Among them, when the above-mentioned specified discrimination task includes various classification tasks about users, the classification task can be a binary classification task or a multi-classification task. When the specified classification task is a multi-classification task, it can correspond to multiple labels. For example, taking the classification task of user risk classification as an example, the multiple labels corresponding to it can be, for example, "low-risk user", "medium-risk user", "high-risk user". Correspondingly, the label representation vocabulary can include the label representation data corresponding to "low-risk user", "medium-risk user", and "high-risk user" under the user's risk classification respectively.
[0052] Taking the specified classification task of user intention classification as an example again, the multiple labels corresponding to it can be, for example, standard questions such as "how to open the payment function", "how to change the bound mobile phone number", "how to make a complaint", etc. Correspondingly, the label representation vocabulary can include the label representation data corresponding to "how to open the payment function", "how to change the bound mobile phone number", and "how to make a complaint" under the user's intention classification respectively.
[0053] Taking the specified classification task of user rental classification as an example again, the multiple labels corresponding to it can be, for example, "short-term rental user", "long-term rental user", "frequent rental user", etc. Correspondingly, the label representation vocabulary can include the label representation data corresponding to "short-term rental user", "long-term rental user", and "frequent rental user" under the user's rental classification respectively.
[0054] When the above-mentioned specified discrimination task includes discrimination tasks of other non-classification tasks, such as regression tasks in other specified dimensions, the label representation vocabulary can include the label representation data corresponding to various labels under the regression task.
[0055] It can be understood that the specified discrimination task can be one or more. When the specified discrimination task is multiple, the label representation vocabulary can include the label representation data corresponding to various labels under each specified discrimination task among the multiple specified discrimination tasks.
[0056] It should be noted that Figure 2 In the shown process, the order of first executing step S210 and then executing step S220 is shown. In practice, it is also possible to first execute step S220 and then execute step S210, or execute step S210 and step S220 simultaneously. This specification does not make any limitations in this regard.
[0057] In some possible implementation manners, the representation model may further include a label representation model, and the label representation model is used to perform vectorization processing on the labels under the specified discrimination task; correspondingly, in step S220, it specifically includes: based on the multiple labels under the specified discrimination task, through the label representation model, obtaining the label representation vocabulary.
[0058] In some possible implementations, the label representation model can be implemented as any word embedding model, such as Word2Vec, GloVe, FastText, etc. Specifically, multiple labels under a specified discrimination task can be input into the label representation model, and each label can be processed by the label representation model to obtain label representation data corresponding to each label under the specified discrimination task. Then, based on the label representation data corresponding to each label under the specified discrimination task, a label representation vocabulary is formed.
[0059] In some cases, when there are multiple specified discrimination tasks, when training the label representation model, it can be set that different labels under different specified discrimination tasks are all different, and during the training process of the representation model, at least minimizing the label representation data corresponding to different labels under different specified discrimination tasks is used as the goal to train the representation model, that is, to train the feature representation model and the label representation model. Correspondingly, multiple labels under each specified discrimination task can be input into the label representation model to process multiple labels under each specified discrimination task through the label representation model to obtain label representation data corresponding to each label under each specified discrimination task. Then, based on the label representation data corresponding to each label under each specified discrimination task, a label representation vocabulary is formed. Among them, the label representation data corresponding to each label under each specified discrimination task is different from each other.
[0060] In still some other possible implementation manners, there are multiple specified discrimination tasks; correspondingly, in step S220, it may specifically include: based on multiple labels under each specified discrimination task and the task identifier of the specified discrimination task corresponding to each label, a label representation matrix is obtained through the label representation model.
[0061] In the above implementation manner, for each label under each specified discrimination task (taking the label j under the specified discrimination task i as an example), the label j under the specified discrimination task i and the task identifier i of the specified discrimination task i corresponding to the label j are input into the label representation model to respectively perform embedding processing on the label j and the task identifier i by using the label representation model to obtain representation data e1 and representation data e2. Then, the representation data e1 and the representation data e2 are fused to obtain the label representation data ij corresponding to the label j under the specified discrimination task i. Among them, the fusion processing can be, for example, stacking processing, or can be fusion processing based on an attention mechanism, etc. By analogy, the label representation data corresponding to each label under each specified discrimination task is obtained, and then the label representation vocabulary is obtained by using the label representation data corresponding to each label under each specified discrimination task.
[0062] In some other possible embodiments, in the discrimination process based on the representation model, the representation model is a trained model, that is, the feature representation model and the label representation model therein are trained models. At this time, when the various labels under the specified discrimination task do not change, the corresponding label representation data will not change either. Correspondingly, the label representation vocabulary can also remain unchanged. In view of this, in the discrimination process based on the representation model, the electronic device can pre-store the label representation vocabulary in a specified space. Correspondingly, the electronic device can obtain the label representation vocabulary from the specified space.
[0063] Exemplarily, the label representation vocabulary can exist in the form of a matrix. For example, each row (or each column) represents the label representation data corresponding to a label under a specified task; or, for another example, each row (or each column) represents the label representation data corresponding to multiple labels under a specified task.
[0064] After obtaining the user features and the label representation vocabulary, in step S230, based on the user features and the label representation vocabulary, through the first attention layer, determine the feature representation data of the user features.
[0065] In some possible examples, the feature representation model may further include an embedding layer (emmbedding layer), which is arranged before the first attention layer. In the embedding layer (emmbedding layer) of the feature representation model, word embedding processing can be performed on the user features to obtain the corresponding word vector sequence. Then, based on the obtained word vector sequence and the label representation vocabulary, through the first attention layer, obtain the output matrix corresponding to the first attention layer. Then, based on the output matrix corresponding to the first attention layer, determine the feature representation data of the user features.
[0066] In some possible embodiments, the embedding layer can be implemented as an embedding layer implemented based on any word embedding algorithm in the related art. In some cases, in the discrimination process based on the representation model, the parameters of the embedding layer have been trained based on a large amount of text corpora (including the features of sample users and their corresponding labels). Correspondingly, the word vectors of each word involved in the user features are obtained through training, that is, a trained word vector table is formed; or, during the training process of the representation model, the embedding layer has been trained based on the previous iteration process of the current iteration process (taking the kth training iteration process as an example), that is, the k - 1th iteration process, to obtain the corresponding word vectors of each word, that is, a trained word vector table corresponding to the current iteration process is formed. Thus, the word vector sequence corresponding to the user features can be determined by referring to the corresponding word vector table.
[0067] In some possible embodiments, as Figure 2 shown, in step S230, it may include steps S11 - S13:
[0068] In step S11, based on user features, a first query matrix, a first key matrix, and a first value matrix are determined through a first attention layer.
[0069] In some possible examples, the electronic device can determine its corresponding word vector sequence based on user features through the embedding layer of the feature representation model; then, in one implementation, when the first attention layer is the first attention layer of the feature representation model, the electronic device can input the word vector sequence corresponding to the user features into the first attention layer, so as to process the word vector sequence through the Q weight matrix, K weight matrix, and V weight matrix (hereinafter referred to as the QKV weight matrix) of the first attention layer, so as to obtain the corresponding first query Q matrix, first key K matrix, and first value V matrix, as Figure 3A shown.
[0070] Exemplarily, it can be represented by the following formula (1):
[0071]
[0072] where Q represents the first query matrix, K represents the first key matrix, V represents the first value matrix, and W Q 、W K and W V respectively represent the Q weight matrix, K weight matrix, and V weight matrix of the first attention layer, and E represents the word vector sequence or the intermediate processing result mentioned later.
[0073] In another implementation, when the first attention layer is not the first attention layer of the feature representation model, correspondingly, the electronic device can obtain the corresponding intermediate processing result based on the word vector sequence corresponding to the user features and the structure before the first attention layer in the feature representation model (such as at least one group of attention layers and their corresponding feed-forward neural networks), and then input the intermediate processing result into the first attention layer, so as to process the intermediate processing result through the QKV weight matrix of the first attention layer, so as to obtain the corresponding first query Q matrix, first key K matrix, and first value V matrix.
[0074] Among them, during the process of training the representation model, the aforementioned QKV weight matrix of the first attention layer (as well as the QKV weight matrices of other attention layers and the parameters of the aforementioned embedding layer) needs to be adjusted, and the parameters in the feed-forward neural network after the first attention layer (as well as the feed-forward neural networks after other attention layers) also need to be adjusted.
[0075] Next, in step S12, a second query matrix is determined based on the label representation word list and the first query matrix.
[0076] In this implementation manner, considering that the query Q matrix is mainly used for querying, which mainly affects the weight distribution among the features in the input (such as the word vectors corresponding to user features), and does not affect the overall data change of the input. Accordingly, in this implementation manner, the first query matrix is adjusted related to the task information. Specifically, based on the label representation vocabulary and the first query matrix, the second query matrix is determined.
[0077] In some specific examples, the foregoing feature representation model may further include: a vocabulary processing layer, where the parameters in the vocabulary processing layer are trainable parameters during the training process of the representation model; accordingly, in step S12, it may include steps 121-122:
[0078] In step 121, based on the label representation vocabulary, through the vocabulary processing layer, a label matrix is obtained. In this step, the label representation vocabulary is input into the vocabulary processing layer to process the label representation vocabulary through the vocabulary processing layer to obtain the label matrix.
[0079] After that, in step 122, based on the label matrix and the first query matrix, the second query matrix is determined.
[0080] In some possible examples, as Figure 3A shown, the foregoing vocabulary processing layer may include a first feedforward network, a second feedforward network, and a preset activation function; in some examples, the preset activation function may be, for example, a sigmoid activation function.
[0081] Accordingly, in the foregoing step 121, it may include steps 1211-1212:
[0082] In step 1211, based on the label representation vocabulary, through the first feedforward network and the preset activation function, a label weight matrix is obtained. In this step, the label representation vocabulary is input into the first feedforward network to process the label representation vocabulary through the first feedforward network to obtain an intermediate matrix. Among them, the first feedforward network may be a linear layer to linearly process the label representation vocabulary; then the intermediate matrix is processed using the preset activation function to obtain the label weight matrix.
[0083] And, in step 1212, based on the label representation vocabulary, through the second feedforward network, a label bias matrix is obtained. In this step, the label representation vocabulary is input into the second feedforward network to process the label representation vocabulary through the second feedforward network to obtain the label bias matrix.
[0084] After that, in one implementation, in step 122, the first query matrix is multiplied element by element with the label weight matrix. Then, the result obtained by multiplying the first query matrix and the label weight matrix element by element is added to the label bias matrix to obtain an added matrix. After that, the added matrix can be used as the second query matrix to obtain a second query matrix embedded with task label information, and this second query matrix can help obtain feature representation data related to the task label information.
[0085] In yet another implementation, to avoid the occurrence of over-task perception, in step 122, a residual connection method can be adopted to determine the second query matrix based on the label matrix and the first query matrix. Specifically, the first query matrix is multiplied element by element with the label weight matrix. Then, the result obtained by multiplying the first query matrix and the label weight matrix element by element is added to the label bias matrix to obtain an added matrix. After that, the added matrix is added to the first query matrix again to obtain a second query matrix embedded with task label information, as Figure 3A shown. Among them, this second query matrix can help obtain feature representation data related to the task label information.
[0086] In the above implementation, the residual connection method is used to determine the second query matrix, which not only strengthens the task orientation but also retains the information of the original input, preventing the performance degradation of the representation model caused by excessive adjustment, and this realizes an effective strategy for balancing task guidance and original feature preservation.
[0087] In some other possible examples, the vocabulary processing layer can only include the aforementioned first feed-forward network and the preset activation function, or only include the aforementioned second feed-forward network, which are both acceptable. Exemplarily, when the vocabulary processing layer includes the aforementioned first feed-forward network and the preset activation function, the aforementioned label matrix can include the aforementioned label weight matrix. Correspondingly, the second query matrix embedded with task label information can be determined based on the label weight matrix and the first query matrix.
[0088] After that, in step S13, based on the second query matrix, the first key matrix, and the first value matrix, through the first attention layer, the feature representation data of the user feature is determined. In this step, based on the second query matrix, the first key matrix, and the first value matrix, through the preset attention formula of the first attention layer, the output matrix corresponding to the first attention layer is obtained. Then, based on the output matrix corresponding to the first attention layer, the feature representation data of the user feature is determined, as Figure 3A shown. Among them, the preset attention formula can be expressed as the following formula (2):
[0089]
[0090] Among them, O represents the output matrix corresponding to the first attention layer, Q ′ represents the second query matrix, K represents the first key matrix, V represents the first value matrix, and d k represents the number of dimensions of the first key matrix, softmax(.) represents the activation function, represents the attention weight matrix.
[0091] In some possible examples, the feature representation model may further include a feed-forward neural network (subsequently referred to as the third feed-forward neural network) disposed after the first attention layer. The process of determining the feature representation data of the user feature based on the output matrix corresponding to the first attention layer may include: inputting the output matrix corresponding to the first attention layer into the third feed-forward neural network to obtain the output matrix corresponding to the third feed-forward neural network; then, in a possible implementation manner, if the third feed-forward neural network is the last layer of the feature representation model, the output matrix corresponding to the third feed-forward neural network is determined as the feature representation data of the user feature; in another possible implementation manner, if the feature representation model further includes several attention layers and feed-forward neural networks after the third feed-forward neural network, based on the output matrix corresponding to the third feed-forward neural network, through several attention layers and feed-forward neural networks in the feature representation model after the third feed-forward neural network, the feature representation data of the user feature is obtained.
[0092] In some other possible examples, after the electronic device determines the first query matrix, the first key matrix, and the first value matrix based on the user feature through the first attention layer, it may determine the second key matrix based on the first key matrix and the label representation vocabulary (or the label matrix determined by the vocabulary processing layer based on the label representation vocabulary); then, jointly with the first query matrix, the second key matrix, and the first value matrix, through the foregoing preset attention formula of the first attention layer, it determines the output matrix corresponding to the first attention layer, and then based on the output matrix corresponding to the first attention layer, it determines the feature representation data of the user feature.
[0093] In the above process, using the label representation vocabulary, that is, using the label representation data corresponding to various labels under the specified discrimination task as an agent, to assist the first attention layer in processing its input (that is, the word vector sequence corresponding to the user feature), so that each feature in the input perceives the task label information and its own influence of the specified discrimination task, to adjust the final output of the first attention layer, so that the first attention layer has the ability of task perception. And, combined with the vocabulary processing layer, it can better implement embedding the task label information into the calculation process of the attention mechanism. Through the first attention layer in the feature representation model, not only can it capture the interaction between features in the input, but also it can significantly improve the sensitivity of the feature representation model to features highly relevant to the specified discrimination task.
[0094] Similarly, through the first attention layer of the feature representation model, injecting the label representation vocabulary into the user features, that is, injecting the label representation data corresponding to various labels under the specified discrimination task, so as to obtain feature representation data that not only pays attention to the relevance between user features but also pays attention to the relevance between user features and the specified discrimination task, making the feature representation data directly related to the specified discrimination task, thereby better improving the accuracy of the prediction result of the subsequent specified discrimination task.
[0095] Corresponding to the method for representing user features based on the representation model provided in the above-mentioned specification embodiments, as Figure 4 shown, the embodiments of this specification also provide a method for representing user features based on the representation model. Specifically, in the process of representing user features based on the representation model, the method includes the following steps S410 - S470:
[0096] In step S410, obtain user features.
[0097] In step S420, obtain the label representation vocabulary, where the label representation vocabulary includes the label representation data corresponding to various labels under the specified discrimination task.
[0098] In step S430, based on the user features, through the first attention layer, determine the first query matrix, the first key matrix, and the first value matrix.
[0099] In step S440, based on the label representation vocabulary and the first query matrix, determine the second query matrix.
[0100] Among them, the implementation principles of steps S410 - S420 are similar to those of the foregoing steps S210 - S220, and the implementation process can refer to the implementation process of the foregoing steps S210 - S220. The implementation principles of steps S430 - S440 are the same as those of the foregoing steps S11 - S12, and the implementation process can refer to the implementation process of the foregoing steps S11 - S12.
[0101] In step S450, based on the second query matrix and the first key matrix, through the first attention layer, determine the attention weight matrix.
[0102] In this step, the attention weight matrix corresponding to the first attention layer can be jointly determined based on the second query matrix, the first key matrix, and the dimension number d of the foregoing first key matrix through the foregoing preset attention formula in the first attention layer k ,
[0103] After that, in step S460, based on the attention weight matrix and the weight masking information, obtain the masking matrix, where the weight masking information is used to indicate screening features more relevant to the specified discrimination task.
[0104] Considering that in the input (user features), some features may have little contribution or impact on the discrimination result of a specified discrimination task. During the process of the feature representation model representing the input, if the final feature representation data is determined based on all features of the input, this may cause the feature representation (such as feature representation data) learned by the feature representation model to be interfered by irrelevant information, that is, some features with little contribution or impact on the discrimination result of the specified discrimination task. To a certain extent, this may reduce the accuracy and efficiency of determining the discrimination result under the specified discrimination task based on this feature representation data.
[0105] In view of the above situation, the electronic device adopts a weight-based dynamic masking mechanism to adaptively mask features that are not highly important for the specified discrimination task. Specifically, in step S460, weight masking information can be obtained, and based on each attention weight value in the attention weight matrix and the weight masking information, a masking matrix is obtained. The weight masking information is used to indicate the screening of features more relevant to the specified discrimination task. The weight masking information is a trainable parameter.
[0106] It can be understood that in the process of representing user features based on the representation model, when it is a subprocess of the discrimination process based on the representation model, both the feature representation model and the weight masking information have been trained. Correspondingly, the electronic device can directly obtain the trained weight masking information and then execute the subsequent process. In the process of representing user features based on the representation model, when it is a subprocess of the training process of the representation model, the weight masking information has been adjusted based on the previous iteration process of the current iteration process (illustrated by the kth training iteration process), that is, the (k - 1)th iteration process. Correspondingly, the electronic device can obtain the weight masking information corresponding to the current iteration process (that is, the weight masking information adjusted in the (k - 1)th iteration process) and then execute the subsequent process.
[0107] In the embodiments of this specification, the unimportant features in the first value V matrix can be dynamically masked according to the feature importance in the first value V matrix, further highlighting the value of the important features in the first value V matrix, that is, the effective features for the discrimination result of the specified discrimination task. At the same time, to ensure that the masked features are neither too many nor too few, upper and lower limits of the masking quantity are also configured. In some implementation manners, the weight masking information may include: a weight masking threshold, a maximum masking quantity, and a minimum masking quantity; the maximum masking quantity is used to limit the maximum number of unmasked features, and the minimum masking quantity is used to limit the minimum number of unmasked features.
[0108] Among them, the aforementioned feature importance can be reflected by the magnitudes of the respective attention weight values in the aforementioned attention weight matrix. The larger the corresponding attention weight value, the greater the correlation between the corresponding feature V value in the first value V matrix and the specified discrimination task.
[0109] The aforementioned process of obtaining the masking matrix may include: sequentially comparing each attention weight value in the attention weight matrix with a weight masking threshold; if an attention weight value is not less than the weight masking threshold, it indicates that the feature V value corresponding to this attention weight value in the first value V matrix needs to be retained (or it indicates that the feature corresponding to this attention weight value in the output matrix corresponding to the first attention layer needs to be retained); if an attention weight value is less than the weight masking threshold, it indicates that the feature V value corresponding to this attention weight value in the first value V matrix needs to be masked (or it indicates that the feature corresponding to this attention weight value in the output matrix corresponding to the first attention layer needs to be masked).
[0110] In addition, in order to retain sufficient information for subsequent accurate prediction and to mask out features with relatively low importance as much as possible, it is also necessary to ensure that the minimum number of the finally unmasked features in the first value V matrix (or the output matrix corresponding to the first attention layer) is not less than a minimum masking quantity, and the maximum number of unmasked features is not greater than a maximum masking quantity.
[0111] For example, the minimum masking quantity is a, and the maximum masking quantity is b, where b is greater than a. Each attention weight value in the attention weight matrix is sequentially compared with the weight masking threshold to determine the attention weight values not less than the weight masking threshold. Theoretically, the features corresponding to the determined attention weight values not less than the weight masking threshold (such as the corresponding feature V values in the first value V matrix or the corresponding features in the output matrix corresponding to the first attention layer) need to be retained, and the features corresponding to other attention weight values less than the weight masking threshold need to be masked.
[0112] In practice, in one case, if the number c of attention weight values not less than the weight masking threshold is greater than the maximum masking quantity b, in order to mask out features with relatively low importance as much as possible, subsequently, c - b attention weight values not less than the weight masking threshold need to be determined from the c attention weight values not less than the weight masking threshold, and the features corresponding to them are masked so that the maximum number of finally unmasked features does not exceed the maximum masking quantity b.
[0113] Exemplarily, the process of determining c - b attention weight values not less than the weight masking threshold from the c attention weight values not less than the weight masking threshold may be to determine the c - b attention weight values with the smallest values from the c attention weight values not less than the weight masking threshold.
[0114] In another case, after successively comparing each attention weight value in the attention weight matrix with the weight masking threshold, if the number d of attention weight values not less than the weight masking threshold is less than the minimum masking quantity a, in order to retain sufficient information for subsequent accurate prediction, it is necessary to select (a - d) attention weight values from multiple attention weight values less than the weight masking threshold in the following, so as to retain the corresponding features (such as the corresponding feature V values in the first value V matrix or the corresponding features in the output matrix corresponding to the first attention layer), so that the minimum number of unmasked features is not less than the minimum masking quantity a.
[0115] Exemplarily, the foregoing selection of (a - d) attention weight values from multiple attention weight values less than the weight masking threshold may be: selecting (a - d) attention weight values with the largest values from multiple attention weight values less than the weight masking threshold.
[0116] Through the above method, based on the attention weight matrix and the weight masking information, a masking matrix is obtained. In some possible examples, the size of this masking matrix is the same as that of the attention weight matrix, and each element therein corresponds one-to-one with each attention weight value in the attention weight matrix. Each element in this masking matrix may take a first value (such as 1) or a second value (such as 0), where the first value indicates that the feature corresponding to the corresponding attention weight value needs to be retained, and the second value indicates that the feature corresponding to the corresponding attention weight value needs to be masked.
[0117] In some other possible examples, it may be to set the attention weight values in the attention weight matrix that need to mask the corresponding features determined by the above method and the weight masking information to 0; and retain the attention weight values in the attention weight matrix that need to retain the corresponding features determined by the above method and the weight masking information, so as to obtain a masking matrix. Correspondingly, this masking matrix includes the attention weight values that need to retain the corresponding features, and the attention weight values that need to mask the corresponding features are set to 0.
[0118] After that, in step S470, the feature representation data corresponding to the user features is determined by using the foregoing masking matrix and the first value matrix.
[0119] In some possible examples, in the case where each element in the masking matrix takes a first value (such as 1) or a second value (such as 0), the output matrix corresponding to the first attention layer may be determined based on the attention weight matrix and the first value matrix through the foregoing preset attention formula, and then the output matrix corresponding to the first attention layer is multiplied element by element with the masking matrix to obtain the masked output matrix corresponding to the first attention layer, as Figure 3BAs shown. Among them, the features with low importance relative to the specified discrimination task in the output matrix are masked, for example, set to 0, and the features with high importance relative to the specified discrimination task are retained.
[0120] In some other possible examples, when the attention weight values for retaining the corresponding features are included in the masking matrix and the attention weight values for masking the corresponding features are set to 0, the first value matrix can be multiplied element-wise with the masking matrix to obtain the masked output matrix corresponding to the first attention layer.
[0121] Exemplarily, taking the process of multiplying the first value matrix and the masking matrix element-wise as an example, the per-pixel multiplication is introduced. It can be to multiply the element (or called feature V value) in the nth row and mth column of the first value matrix with the element in the nth row and mth column of the masking matrix respectively to obtain the element (feature) in the nth row and mth column of the masked output matrix corresponding to the first attention layer.
[0122] After obtaining the masked output matrix corresponding to the aforementioned first attention layer, based on the masked output matrix corresponding to the first attention layer, the feature representation data corresponding to the user features is determined. Among them, the process of determining the feature representation data corresponding to the user features based on the masked output matrix corresponding to the first attention layer can refer to the process of determining the feature representation data corresponding to the user features based on the output matrix corresponding to the first attention layer described above, which will not be elaborated here.
[0123] In the above process, when perceiving the task label, the features most relevant to the specified discrimination task can also be determined from the input, and the invalid features in the input, that is, the features not important enough for the specified discrimination task, can be adaptively and dynamically reduced, reducing the influence of the invalid features in the input on the feature representation model, and further improving the accuracy of the subsequent discrimination result.
[0124] In some possible implementation manners, the above process of representing user features based on the representation model is a subprocess of the discrimination process based on the representation model. That is, after obtaining the feature representation data based on the above process of representing user features based on the representation model, the discrimination process is continued based on the feature representation data. Correspondingly, as Figure 5 shown, an exemplary flowchart of a discrimination process based on a representation model is shown, where the representation model includes a feature representation model, and the feature representation model includes a first attention layer.
[0125] In the discrimination process based on the representation model, the representation model, namely the feature representation model therein (including the aforementioned first attention layer, vocabulary processing layer, and weight masking information) and the label representation model, have all been trained. After the label representation model is trained, in fact, the final label representation data of various labels under the specified discrimination task is also obtained. Subsequently, based on the final label representation data of various labels under any specified discrimination task, the corresponding discrimination task can be performed on the newly received user features.
[0126] Among them, the discrimination process based on the representation model may include steps S510 - S550:
[0127] In step S510, user features are obtained.
[0128] In step S520, a label representation vocabulary is obtained, where the label representation vocabulary includes the label representation data corresponding to various labels under the specified discrimination task. Exemplarily, the label representation vocabulary can be determined through the trained label representation model.
[0129] In step S530, based on the user features and the label representation vocabulary, through the first attention layer, the feature representation data of the user features is determined.
[0130] Among them, the implementation principles of steps S510 - S530 are similar to those of the aforementioned steps S210 - S230, and the implementation process can refer to the implementation process of the aforementioned steps S210 - S230.
[0131] After that, in step S540, the target similarity between the feature representation data and the label representation data corresponding to the target label under the target discrimination task is calculated. The target discrimination task is any discrimination task among the aforementioned specified discrimination tasks. The label representation data corresponding to the target label under the target discrimination task can be obtained from the aforementioned label representation vocabulary.
[0132] In this step, the electronic device can determine the target discrimination task from the specified discrimination tasks based on the user's discrimination requirements, and then calculate the target similarity between the feature representation data and the label representation data corresponding to each target label under the target discrimination task. Among them, the target similarity can be, for example, cosine similarity, Euclidean distance, Manhattan distance, and Pearson correlation coefficient, etc.
[0133] Taking cosine similarity as an example, it can be the inner product of the feature representation data and the label representation data corresponding to the target label to obtain the above - mentioned target similarity.
[0134] In step S550, based on the target similarity, determine the predicted discrimination result of the user feature under the target discrimination task. Exemplarily, the target similarity may include the respective target similarities between the feature representation data and the label representation data corresponding to multiple target labels under the target discrimination task; correspondingly, the target label corresponding to the maximum target similarity value may be used as the predicted discrimination result of the user feature under the target discrimination task.
[0135] Also exemplarily, the target similarity may be the similarity between the feature representation data and the label representation data corresponding to any one target label under the target discrimination task. Correspondingly, if the value of the target similarity exceeds a specified threshold, it may be determined that the predicted discrimination result of the user feature under the target discrimination task conforms to the any one target label (for example, when the target label indicates a certain category, it may be determined that the user feature belongs to the category indicated by the target label under the target discrimination task); if the value of the target similarity does not exceed the specified threshold, it may be determined that the predicted discrimination result of the user feature under the target discrimination task does not conform to the target label (for example, when the target label indicates a certain category, it may be determined that the user feature does not belong to the category indicated by the target label under the target discrimination task).
[0136] In the above process, when performing discrimination based on the representation model for a user, only the similarity between the feature representation data corresponding to the user feature of the user and the label representation data corresponding to the target label under the target discrimination task needs to be calculated to achieve discrimination for the user. Thus, the discrimination efficiency for the user can be greatly improved. Moreover, through the first attention layer of the feature representation model, the label representation vocabulary is injected into the user feature, that is, the label representation data corresponding to various labels under the specified discrimination task is injected, so as to obtain feature representation data that not only pays attention to the relevance between user features but also pays attention to the relevance between user features and the specified discrimination task, making the feature representation data directly related to the specified discrimination task, thereby better improving the accuracy of the prediction result of the subsequent specified discrimination task.
[0137] In some possible implementation manners, the above process of representing the user feature based on the representation model may be a subprocess of the training process of the representation model, that is, after obtaining the feature representation data based on the above process of representing the user feature based on the representation model, it is necessary to jointly train the representation model with the feature representation data and the label corresponding to the user feature. Among them, as Figure 6A, which shows a schematic diagram of the principle of training a characterization model. Among them, the characterization model includes a feature characterization model and a label characterization model. During the training process, based on user features, the feature characterization data corresponding to the user features can be obtained through the feature characterization model. Based on the first label corresponding to the user features, the first label characterization data corresponding to the first label can be obtained through the label characterization model (for example, from the label characterization word list obtained based on the label characterization model). After that, at least combining the similarity between the feature characterization data corresponding to the user features and the first label characterization data, the characterization loss is determined, and the characterization model is trained with the goal of minimizing the characterization loss, that is, adjusting the parameters of the feature characterization model and the label characterization model.
[0138] Correspondingly, as Figure 6B shown, an exemplary flowchart of the training process of a characterization model is shown. Among them, the characterization model includes a feature characterization model, and the feature characterization model includes a first attention layer (as well as the aforementioned word list processing layer and weight masking information). As Figure 6A shown, the training process of this characterization model may include steps S610 - S670:
[0139] In step S610, user features are obtained.
[0140] In step S620, a label characterization word list is obtained, where the label characterization word list includes the label characterization data corresponding to various labels under a specified discrimination task.
[0141] In step S630, based on the user features and the label characterization word list, through the first attention layer, the feature characterization data of the user features is determined.
[0142] Among them, the implementation principles of steps S610 - S630 are similar to those of the aforementioned steps S210 - S230, and the implementation process can refer to the implementation process of the aforementioned steps S210 - S230.
[0143] In step S640, the first label under the specified discrimination task corresponding to the user features is obtained. In this step, the first label under the specified discrimination task corresponding to the user features can be obtained from the training set used to train the characterization model.
[0144] After that, in step S650, the first label characterization data corresponding to the first label is determined from the label characterization word list.
[0145] In this step, in the current iteration process (illustrated by taking the k-th training iteration process as an example), the label representation vocabulary has been obtained after adjusting the label representation model based on the previous k-1 iteration processes. The label representation vocabulary includes label representation data corresponding to various labels under the specified discrimination task. Correspondingly, the electronic device can determine the first label representation data corresponding to the first label from the label representation vocabulary based on the first label under the specified discrimination task. As Figure 6A shown, the first label representation data corresponding to the first label can also be referred to as being determined through the label representation model.
[0146] Next, in step S660, based on the first similarity between the feature representation data and the first label representation data, a representation loss is determined, where the representation loss is negatively correlated with the first similarity. In this step, the first similarity between the feature representation data and the first label representation data is calculated. As shown above, the first similarity can be, for example, cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, etc. Taking cosine similarity as an example, it can be the inner product of the feature representation data and the first label representation data to obtain the above first similarity. Then, based on this first similarity, the representation loss is determined.
[0147] In some embodiments, the process of determining the representation loss can be: using the first loss function, based on the first similarity, to determine the representation loss. Wherein, the representation loss is negatively correlated with the first similarity. Wherein, the first loss function and the subsequent second loss function and third loss function can be the binary cross-entropy loss function BCEWithLogitsLoss, or can also be the focal loss, etc.
[0148] In some examples, there can be multiple specified discrimination tasks. Correspondingly, the first label can include the labels of the user features under each specified discrimination task. Correspondingly, the first label representation data can include the label representation data corresponding to the labels of the user features under each specified discrimination task respectively. The aforementioned calculation of the first similarity between the feature representation data and the first label representation data can include calculating the respective first similarities between the feature representation data and each first label representation data. Then, the first loss function can be used to determine the representation loss based on each first similarity.
[0149] In some possible embodiments, in this embodiment, the dissimilarity loss between the labels within the task can also be combined to determine the representation loss, so as to train the feature representation model and the label representation model. Correspondingly, in step S660, it can include steps 21-22:
[0150] In step 21, calculate the second similarity between the label representation data corresponding to each type of label under a specified discrimination task. In this step, for each specified discrimination task, the second similarity between the label representation data corresponding to every two labels among the various types of labels under the specified discrimination task can be calculated to obtain a number of second similarities.
[0151] In step 22, determine the representation loss based on the first similarity between the feature representation data and the first label representation data, and the second similarity, where the representation loss is also positively correlated with the second similarity and negatively correlated with the first similarity.
[0152] In this step, the first loss function can be used to determine the first loss based on the first similarity, where the first loss is negatively correlated with the first similarity; the second loss function can be used to determine the second loss based on a number of second similarities, where the second loss is positively correlated with each of the number of second similarities; then, the representation loss is determined based on the sum or mean of the first loss and the second loss.
[0153] In some possible implementation manners, in this embodiment, the dissimilarity loss between the labels within different tasks can also be combined to determine the representation loss, so as to train the feature representation model and the label representation model. Correspondingly, in step S660, it may include 31 - 32:
[0154] In step 31, calculate the third similarity between the label representation data corresponding to every two labels under different specified discrimination tasks. The process of calculating the third similarity can refer to the process of calculating the first similarity described above, and will not be elaborated here.
[0155] In step 32, determine the representation loss based on the first similarity between the feature representation data and the first label representation data, and the third similarity, where the representation loss is also positively correlated with the third similarity and negatively correlated with the first similarity.
[0156] In this step, the first loss function can be used to determine the first loss based on the first similarity, where the first loss is negatively correlated with the first similarity; the third loss function can be used to determine the third loss based on a number of third similarities, where the third loss is positively correlated with each of the number of third similarities; then, the representation loss is determined based on the sum or mean of the first loss and the third loss.
[0157] In yet another implementation manner, the first similarity, the second similarity, and the third similarity described above can also be combined to determine the representation loss. The implementation manner of this determination process can refer to the process of determining the representation loss based on the first similarity and the second similarity described above, and will not be elaborated here.
[0158] After that, in step S670, the representation model is trained with the goal of minimizing the representation loss. In this step, with the goal of minimizing the representation loss, that is, making the feature representation data and the first label representation data more similar, and making the label representation data corresponding to different labels under the same specified discrimination task and the label representation data corresponding to different labels under different specified discrimination tasks less similar, the parameters of the representation model are adjusted, that is, the parameters of the feature representation model (and the label representation model) are adjusted.
[0159] When adjusting the parameters of the representation model through the representation loss determined based on the first similarity, it can be ensured that the feature representation data of samples with the same label are close to each other, enabling the representation model to better learn the similarity between samples, and thus promoting the effective extraction of features.
[0160] When adjusting the parameters of the representation model through the representation loss determined based on the first similarity and the second similarity (and the third similarity), it can be achieved that while ensuring that the feature representation data of samples with the same label are close to each other, enabling the representation model to better learn the similarity between samples, and thus promoting the effective extraction of features, it can also create a greater dissimilarity (i.e., difference) between different labels and / or tasks. In other words, it is to ensure that the label representation data corresponding to different labels remain independent and dissimilar to each other, so that the model can clearly distinguish different labels under the same specified discrimination task (and different labels under different specified discrimination tasks), and can encourage different samples corresponding to the same specified discrimination task to be as dispersed as possible in the embedding space to enhance the discriminability of the representation.
[0161] By adjusting the parameters of the representation model through the above-mentioned representation loss, the model learns more detailed and accurate label boundaries, promoting the consistency of representations within the same label (e.g., within the same category) and the distinctiveness of representations between different labels (e.g., between different categories). Through the representation loss determined by combining the aforementioned first loss, second loss, and third loss, the representation model can better refine and optimize the feature representation data, thereby effectively improving its performance in the discrimination task, especially in scenarios with a large number of similar categories or unclear category boundaries.
[0162] In the above process, by innovatively integrating task label information directly into the attention mechanism (combining user features and the label representation vocabulary, and determining the feature representation data of user features through the first attention layer), by introducing task label information, the representation model can consider the relevance between the input and the target task while paying attention to the internal relationships of the input content, thus narrowing the gap between the traditional attention mechanism and the discriminant task objective. Moreover, based on the label representation vocabulary and user features, the feature representation data of user features is determined, which can avoid the leakage of the first label corresponding to user features during the training process and ensure the privacy and security of the sample user's information. In addition, due to the addition of direct perception of task label information in the feature representation model, the generalization ability of the representation model in new scenarios and data can be improved to a certain extent.
[0163] In the above example, after processing the label representation vocabulary through the vocabulary processing layer, a label matrix is obtained, and the weight distribution of the original query matrix (i.e., the first query matrix) is dynamically adjusted based on the label matrix, which can better enhance the attention of the representation model to the discriminant task orientation and also ensure the flexibility and pertinence during the adjustment process.
[0164] The above solution also brings many beneficial effects in terms of commercial value, as follows:
[0165] 1. In terms of improving the generalization ability and accuracy of the model: Through the task-aware attention mechanism, the model can more accurately capture the features closely related to the target task, thus achieving a significant improvement in performance in various discriminant tasks (such as classification, regression), enhancing the practicality and competitiveness of the model. 2. In terms of enhancing the interpretability of the model: The task-aware attention mechanism, by explicitly integrating task label information into the model decision-making process, helps to improve the transparency and interpretability of the model decision-making, which has important value for industry applications that require high trust and supervision (such as financial risk control, medical diagnosis). 3. In terms of promoting cross-domain applications: Since the above solution can flexibly adapt to the needs of different discriminant tasks, it can be quickly applied to new fields or new tasks by adjusting the label representation data, reducing the cost of model migration, and accelerating the popularization and commercialization process of artificial intelligence technology in multiple industries. 4. In terms of optimizing resource efficiency: By finely adjusting the allocation of attention weights, unnecessary calculations are reduced, and the efficiency of model training and inference is improved, which is particularly important for scenarios with limited resources (such as mobile devices, edge computing), and is conducive to reducing operating costs and improving the user experience.
[0166] In some other possible implementation manners, such as Figure 7 shown, an exemplary flowchart of the training process of another representation model is illustrated, where the training process may include steps S710 - S770:
[0167] In step S710, user features are obtained.
[0168] In step S720, a label representation vocabulary is obtained, where the label representation vocabulary includes label representation data corresponding to various labels under a specified discrimination task.
[0169] In step S730, based on the user features, through a first attention layer, a first query matrix, a first key matrix, and a first value matrix are determined.
[0170] In step S740, based on the label representation vocabulary and the first query matrix, a second query matrix is determined.
[0171] In step S750, based on the second query matrix and the first key matrix, through the first attention layer, an attention weight matrix is determined.
[0172] In step S760, based on the attention weight matrix and weight masking information, a masked matrix is obtained, where the weight masking information is used to indicate screening features more relevant to the specified discrimination task. The weight masking information is information to be trained.
[0173] In step S770, using the masked matrix and the first value matrix, feature representation data is determined;
[0174] In step S780, a first label under the specified discrimination task corresponding to the user features is obtained;
[0175] In step S790, from the label representation vocabulary, the first label representation data corresponding to the first label is determined;
[0176] In step S7100, based on a first similarity between the feature representation data and the first label representation data, a representation loss is determined, where the representation loss is negatively correlated with the first similarity;
[0177] In step S7110, with the goal of minimizing the representation loss, a representation model is trained.
[0178] Among them, the implementation principles of steps S710 - S770 are similar to those of the foregoing steps S410 - S470, and the implementation process can refer to the implementation process of the foregoing steps S410 - S470. The implementation principles of steps S780 - S7110 are the same as those of the foregoing steps S640 - S670, and the implementation process can refer to the implementation process of the foregoing steps S640 - S670.
[0179] In this embodiment, under the beneficial effects of the above - mentioned embodiments, the following beneficial effects can also be brought:
[0180] The above process can not only focus the attention of the representation model on the features that are relatively most critical for the specified discrimination task, but also provides a flexible and adaptive control mechanism to determine the degree of feature masking through trainable parameters. During the training process, by initializing and training three parameters: the weight masking threshold, the minimum masking number, and the maximum masking number, the model can automatically learn relatively more appropriate masking rules, reducing the trouble of manual configuration, ensuring the rationality of the threshold and the upper and lower limits of masking, ensuring that each sample still retains sufficient information for accurate prediction after feature masking, and finally enabling each sample to have its dynamic output, achieving the purpose of highlighting important features.
[0181] Moreover, by selectively masking the less important features in the input, it effectively reduces the model's processing of useless or interfering information in them, improving the computational efficiency and the generalization ability of the model. The dynamic masking mechanism encourages the representation model to focus more on the features closely related to the task objective, which helps to improve the model's performance when facing datasets with high noise and sparse key information, and is especially suitable for the optimization requirements of the representation model in discrimination tasks. The above solution reduces the impact of invalid features on the model, enabling the model to learn features that are more useful for the specified discrimination task. In this way, it can not only improve the accuracy of the discrimination results in the subsequent specified discrimination task, but also optimize the training efficiency of the model and reduce the consumption of computing resources, thus reducing costs while improving the business efficiency of the model.
[0182] In addition, during the dynamic masking process of features through weight masking information, the individual differences of each sample can be fully considered to ensure that the masking decision can reflect the uniqueness of the sample, strengthening the robustness and accuracy of the model under complex data distributions and achieving the enhancement of sample features. The dynamic masking mechanism and the task-aware attention mechanism solution can be seamlessly integrated into existing models based on the Transformer architecture, such as the BERT model, the GPT model, etc., without fundamentally changing the model architecture, improving the practicality and promotion potential of the technical solution. And they jointly solve the problems of information redundancy, inaccurate feature selection, and insufficient task orientation in the discrimination task, providing strong support for promoting the progress of deep learning technology in applications with high-precision requirements and resource sensitivity.
[0183] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0184] Corresponding to the above method embodiments, an embodiment of this specification provides an apparatus 800 for characterizing user features based on a characterization model. The characterization model includes a feature characterization model, and the feature characterization model includes a first attention layer. Its schematic block diagram is as Figure 8 shown and includes:
[0185] A first acquisition module 810, configured to acquire user features;
[0186] A second acquisition module 820, configured to acquire a label characterization vocabulary. Among them, the label characterization vocabulary includes label characterization data corresponding to various labels under a specified discrimination task;
[0187] A first determination module 830, configured to determine the feature characterization data of the user features based on the user features and the label characterization vocabulary through the first attention layer.
[0188] In some possible examples, the first determination module 830 includes: a first determination unit (not shown in the figure), configured to determine a first query matrix, a first key matrix, and a first value matrix based on the user features through the first attention layer;
[0189] A second determination unit (not shown in the figure), configured to determine a second query matrix based on the label characterization vocabulary and the first query matrix;
[0190] A third determination unit (not shown in the figure), configured to determine the feature characterization data of the user features based on the second query matrix, the first key matrix, and the first value matrix through the first attention layer.
[0191] In some possible examples, the feature characterization model further includes: a vocabulary processing layer; the second determination unit (not shown in the figure) includes: a first obtaining sub-module (not shown in the figure), configured to obtain a label matrix based on the label characterization vocabulary through the vocabulary processing layer; a first determination sub-module (not shown in the figure), configured to determine a second query matrix based on the label matrix and the first query matrix.
[0192] In some possible examples, the vocabulary processing layer includes a first feedforward network, a second feedforward network, and a preset activation function; the first obtaining sub-module (not shown in the figure) is specifically configured to obtain a label weight matrix based on the label characterization vocabulary through the first feedforward network and the preset activation function;
[0193] Based on the label characterization vocabulary, obtain a label bias matrix through the second feedforward network.
[0194] In some possible examples, the third determination unit (not shown in the figure) is specifically configured to determine an attention weight matrix based on the second query matrix and the first key matrix through the first attention layer;
[0195] Based on the attention weight matrix and the weight masking information, a masking matrix is obtained, where the weight masking information is used to indicate screening features more relevant to the specified discrimination task;
[0196] The feature representation data is determined by using the masking matrix and the first value matrix.
[0197] In some possible examples, the weight masking information includes: a weight masking threshold, a maximum masking number, and a minimum masking number. The maximum masking number is used to limit the maximum number of unmasked features, and the minimum masking number is used to limit the minimum number of unmasked features.
[0198] In some possible examples, the representation model further includes a label representation model, and the label representation model is used to vectorize the labels under the specified discrimination task;
[0199] The second acquisition module 820 is specifically configured to obtain the label representation vocabulary based on multiple labels under the specified discrimination task through the label representation model.
[0200] In some possible examples, there are multiple specified discrimination tasks;
[0201] The second acquisition module 820 is specifically configured to obtain the label representation matrix based on multiple labels under each specified discrimination task and the task identifier of the specified discrimination task corresponding to each label through the label representation model.
[0202] In some possible examples, it further includes: a second acquisition module (not shown in the figure), configured to acquire a first label under the specified discrimination task corresponding to the user feature; a second determination module (not shown in the figure), configured to determine first label representation data corresponding to the first label from the label representation vocabulary; a third determination module (not shown in the figure), configured to determine a representation loss based on a first similarity between the feature representation data and the first label representation data, where the representation loss is negatively correlated with the first similarity;
[0203] A training module (not shown in the figure), configured to train the representation model with the goal of minimizing the representation loss.
[0204] In some possible examples, the third determination module (not shown in the figure) is specifically configured to calculate a second similarity between label representation data corresponding to various labels under the specified discrimination task;
[0205] Based on the first similarity and the second similarity, determine the representation loss, where the representation loss is also positively correlated with the second similarity.
[0206] In some possible examples, the third determination module (not shown in the figure) is specifically configured to calculate a third similarity between label representation data corresponding to every two labels under different specified discrimination tasks;
[0207] Based on the first similarity and the third similarity, determine the representation loss, where the representation loss is also positively correlated with the third similarity.
[0208] In some possible examples, the device further includes: a calculation module (not shown in the figure), configured to calculate a target similarity between the feature representation data and label representation data corresponding to a target label under a target discrimination task, where the target discrimination task is any discrimination task in the specified discrimination tasks; a fourth determination module (not shown in the figure), configured to determine a predicted discrimination result of the user feature under the target discrimination task according to the target similarity.
[0209] The above device embodiments correspond to the method embodiments. For specific descriptions, reference can be made to the descriptions in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference can be made to the corresponding method embodiments.
[0210] An embodiment of this specification also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method for representing user features based on a representation model provided in this specification.
[0211] An embodiment of this specification also provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method for representing user features based on a representation model provided in this specification is implemented.
[0212] The various embodiments in this specification are all described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the storage medium and the computing device, since they are basically similar to the method embodiments, the descriptions are relatively simple. For the relevant parts, reference can be made to the partial descriptions of the method embodiments.
[0213] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0214] The specific embodiments described above have further elaborated on the objectives, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above are only specific embodiments of the embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for characterizing user features based on a characterization model, wherein the characterization model includes a feature characterization model, wherein the feature characterization model includes a first attention layer, and wherein the method includes: Get user characteristics; Obtaining a label representation vocabulary, wherein the label representation vocabulary includes label representation data corresponding to various labels under the specified discrimination task; Based on the user features and the tag representation vocabulary, feature representation data of the user features are determined through the first attention layer.
2. The method of claim 1, wherein: The feature characterization data for determining the user feature includes: Based on the user features, determining a first query matrix, a first key matrix, and a first value matrix through the first attention layer; Determine a second query matrix based on the label representation vocabulary and the first query matrix; Based on the second query matrix, the first key matrix and the first value matrix, feature representation data of the user features are determined through the first attention layer.
3. The method of claim 2, wherein: The feature representation model also includes: a vocabulary processing layer; The determining a second query matrix based on the label representation vocabulary and the first query matrix includes: Based on the label representation vocabulary, a label matrix is obtained through the vocabulary processing layer; Based on the label matrix and the first query matrix, a second query matrix is determined.
4. The method of claim 3, wherein: The vocabulary processing layer includes a first feedforward network, a second feedforward network and a preset activation function; The obtained label matrix includes: Based on the label representation vocabulary, a label weight matrix is obtained through the first feedforward network and the preset activation function; Based on the label representation vocabulary, a label bias matrix is obtained through the second feed-forward network.
5. The method of claim 2, wherein: The feature characterization data for determining the user feature includes: Determining an attention weight matrix based on the second query matrix and the first key matrix through the first attention layer; Based on the attention weight matrix and the weight masking information, a masking matrix is obtained, wherein the weight masking information is used to indicate the selection of features that are more relevant to the specified discrimination task; The feature characterization data is determined using the masking matrix and the first value matrix.
6. The method of claim 5, wherein: The weight masking information includes: a weight masking threshold, a maximum masking number and a minimum masking number, wherein the maximum masking number is used to limit the maximum number of unmasked features, and the minimum masking number is used to limit the minimum number of unmasked features.
7. The method according to any one of claims 1 to 6, wherein the representation model further comprises a label representation model, and the label representation model is used to vectorize the labels under the specified discrimination task; The step of obtaining a label representation vocabulary includes: Based on the multiple labels under the specified discrimination task, the label representation vocabulary is obtained through the label representation model.
8. According to the method of claim 7, the number of the designated identification tasks is multiple; The step of obtaining a label representation vocabulary includes: Based on a plurality of labels under each designated discrimination task and a task identifier of the designated discrimination task corresponding to each label, the label representation matrix is obtained through the label representation model.
9. The method of claim 7, further comprising: Obtaining a first label under the specified discrimination task corresponding to the user feature; Determining first label representation data corresponding to the first label from the label representation vocabulary; Determining a representation loss based on a first similarity between the feature representation data and the first label representation data, wherein the representation loss is negatively correlated with the first similarity; The representation model is trained with the goal of minimizing the representation loss.
10. The method of claim 8, wherein: The determining of the characterization loss comprises: Calculating the second similarity between the label representation data corresponding to each type of label under the specified discrimination task; The representation loss is determined based on the first similarity and the second similarity, wherein the representation loss is also positively correlated with the second similarity.
11. The method of claim 8, wherein determining the characterization loss comprises: Calculate the third similarity between the label representation data corresponding to each two labels under different specified discrimination tasks; The representation loss is determined based on the first similarity and the third similarity, wherein the representation loss is also positively correlated with the third similarity.
12. The method according to any one of claims 1 to 6, further comprising: Calculating target similarity between the feature representation data and label representation data corresponding to a target label under a target discrimination task, wherein the target discrimination task is any discrimination task in the specified discrimination task; According to the target similarity, a prediction and discrimination result of the user feature under the target discrimination task is determined.
13. A device for characterizing user features based on a characterization model, the characterization model comprising a feature characterization model, the feature characterization model comprising a first attention layer, the device comprising: A first acquisition module, configured to acquire user characteristics; A second acquisition module is configured to acquire a label representation vocabulary, wherein the label representation vocabulary includes label representation data corresponding to various labels under a specified discrimination task; The first determination module is configured to determine the feature representation data of the user feature through the first attention layer based on the user feature and the tag representation vocabulary.
14. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 12 is implemented.