Model training method and device, electronic equipment, medium and program product

By encoding and fusing features of tabular sample data to generate a target feature extraction model, the problem of low generalization of tabular data models in the existing technology is solved, and stronger feature adaptation and generalization capabilities are achieved.

CN120744446APending Publication Date: 2025-10-03ZHEJIANG E COMMERCE BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510836311.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing technology, model training methods for structured tabular data are difficult to fully explore the correlation information between features, resulting in low generalization of the model.

Method used

By obtaining sample data from multiple tables and performing encoding processing to generate the initial feature vector, the fused feature vector is determined using multi-head attention calculation and gated weight calculation network, the loss function is constructed and the model parameters are updated to generate the target feature extraction model.

Benefits of technology

The model's adaptability and generalization performance to complex and diverse features are improved, and the performance of the target feature extraction model in various tasks is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744446A_ABST
    Figure CN120744446A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training method and device, electronic equipment, a medium and a program product. The method comprises the steps that multiple pieces of table sample data are obtained, the table sample data are coded, initial feature vectors corresponding to the table sample data are obtained, and each piece of table sample data in the multiple pieces of table sample data comprises multiple pieces of feature data of multiple feature types; determining a fusion feature vector corresponding to the table sample data based on the initial feature vector, and constructing a loss function according to the fusion feature vector; and finally, updating model parameters of the initial feature extraction model based on the loss function to obtain a target feature extraction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a model training method, device, electronic device, medium, and program product. Background Art

[0002] Tabular data is two-dimensional structured data consisting of multiple fields. With the widespread application of artificial intelligence in finance, e-commerce, government affairs and other fields, the demand for modeling based on large amounts of structured tabular data is growing.

[0003] In related technologies, model training methods for structured tabular data usually independently model individual features, making it difficult to fully explore the correlation information between features, resulting in low model generalization. Summary of the Invention

[0004] The embodiments of this specification provide a model training method, device, electronic device, medium, and program product that can improve the model's adaptability and generalization performance to complex and diverse features, thereby improving the performance of the final target feature extraction model in various tasks. The above technical solutions are as follows:

[0005] In a first aspect, the embodiments of this specification provide a model training method, including:

[0006] Acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types;

[0007] Encode the sample data of each table to obtain the initial feature vector corresponding to the sample data of the table;

[0008] Determine the fused feature vector corresponding to the sample data in the table based on the initial feature vector;

[0009] Construct a loss function based on the above fused feature vector;

[0010] Based on the above loss function, the model parameters of the initial feature extraction model are updated to obtain the target feature extraction model.

[0011] In a possible implementation, encoding the sample data of each table to obtain the initial feature vector corresponding to the sample data of the table includes:

[0012] For each of the above table sample data, masking is performed on at least part of the feature data in the above table sample data based on a preset masking strategy to obtain a masking result;

[0013] The masked feature data in the mask result is used as the first feature data, and each first feature data is encoded to obtain a first feature vector corresponding to each first feature data;

[0014] Using the unmasked feature data in the mask results as second feature data, and performing encoding on each of the second feature data to obtain a second feature vector corresponding to each of the second feature data;

[0015] The first feature vectors corresponding to the above-mentioned first feature data and the second feature vectors corresponding to the above-mentioned second feature data are combined as the initial feature vectors corresponding to the above-mentioned table sample data.

[0016] In a possible implementation, encoding each first feature data to obtain a first feature vector corresponding to each first feature data includes:

[0017] The feature column name information of each first feature data is encoded to obtain a first feature vector corresponding to each first feature data.

[0018] In a possible implementation, encoding the second feature data to obtain the second feature vector corresponding to the second feature data includes:

[0019] Determine a feature type of the second feature data, where the feature type is a numerical value or a categorical type;

[0020] In the case where the feature type is the numerical type, encoding the product of the feature column name embedding vector of the second feature data and the feature value to obtain the second feature vector corresponding to each second feature data; or

[0021] In the case where the feature type is the category type, encoding processing is performed on a combination of the feature column name information and the feature value of the second feature data to obtain a second feature vector corresponding to each of the second feature data.

[0022] In a possible implementation, determining the fused feature vector corresponding to the table sample data based on the initial feature vector includes:

[0023] Perform multi-head attention calculation based on the above initial feature vector to obtain the target feature vector corresponding to the above table sample data, and the above target feature vector is used to represent the correlation relationship between the feature data in the above table sample data;

[0024] The target feature vector is weighted by a preset gate weight calculation network to obtain feature gate weight information corresponding to the sample data in the table;

[0025] Based on the target feature vector and the feature gating weight information, a fusion feature vector corresponding to the sample data in the table is determined.

[0026] In one possible implementation, the multi-head attention calculation is performed based on the initial feature vector to obtain the target feature vector corresponding to the sample data in the table, including:

[0027] Generate a query vector based on a preset sequence of learnable clustering vectors;

[0028] Perform linear transformation on the above initial feature vector to generate key vector and value vector;

[0029] Performing a multi-head attention calculation based on the query vector, the key vector, and the value vector to project the initial feature vector into a preset cluster space to obtain a feature representation in the cluster space;

[0030] Based on the feature representation in the above cluster space, multi-head attention calculation is performed to obtain the cluster feature interaction result;

[0031] The above cluster feature interaction results are restored to the original space of the above initial feature vector to obtain the target feature vector corresponding to the above table sample data.

[0032] In a possible implementation, the feature gating weight information is used to indicate the importance of each feature data in the feature update process;

[0033] The above-mentioned determination of the fused feature vector corresponding to the sample data in the above-mentioned table based on the above-mentioned target feature vector and the above-mentioned feature gating weight information includes:

[0034] Based on the feature gating weight information, the target feature vector is weighted feature by feature to obtain a weighted feature vector;

[0035] The weighted feature vector and the target feature vector are fused to obtain a fused feature vector corresponding to the sample data in the table.

[0036] In one possible implementation, the above loss function includes mask prediction loss, contrast loss, and consistency loss;

[0037] The loss function constructed based on the above fused feature vector includes:

[0038] Determining a first mask loss corresponding to the first feature data and a second mask loss corresponding to the second feature data based on the fused feature vector;

[0039] Determining a mask prediction loss according to the first mask loss and the second mask loss;

[0040] Construct contrast loss based on multiple fused feature vectors generated under multiple different preset masking strategies based on the sample data in the above table;

[0041] Constructing consistency loss based on the multiple mask results corresponding to the multiple different preset mask strategies based on the sample data in the above table;

[0042] The above mask prediction loss, the above contrast loss and the above consistency loss are weighted and combined to obtain the loss function.

[0043] In one possible implementation, updating the model parameters of the initial feature extraction model based on the loss function to obtain the target feature extraction model includes:

[0044] Based on the above loss function, backpropagation calculates the gradient of the initial feature extraction model;

[0045] Based on the above gradient, the initial parameters of the above initial feature extraction model are updated until the above loss function converges or meets the preset stopping condition, thereby obtaining the above target feature extraction model.

[0046] In a second aspect, the embodiments of this specification provide a model training method, including:

[0047] Acquire multiple tabular sample data corresponding to multiple users, each tabular sample data including multiple feature data of multiple feature types, each tabular sample data including at least one of the following: basic information, behavior information, and device information of the user;

[0048] Encode the sample data of each table to obtain the initial feature vector corresponding to the sample data of the table;

[0049] Determine the fused feature vector corresponding to the sample data in the table based on the initial feature vector;

[0050] Construct a loss function based on the above fused feature vector;

[0051] Based on the above loss function, the model parameters of the initial feature extraction model are updated to obtain the target feature extraction model;

[0052] The target feature extraction model is generated by the method provided in the first aspect.

[0053] In a third aspect, the embodiments of this specification provide a label prediction method, including:

[0054] Input the table data to be processed corresponding to the user into the target feature extraction model to obtain the feature vector representation output by the target feature extraction model;

[0055] Generating a predicted scenario label corresponding to the table data to be processed according to the feature vector representation, wherein the predicted scenario label is used to indicate the state information of the user in the corresponding scenario;

[0056] The target feature extraction model is generated by the method provided in the first or second aspect of the claim.

[0057] In a fourth aspect, an embodiment of this specification provides a model training device, comprising:

[0058] A first acquisition module is configured to acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types;

[0059] A first encoding module is used to encode the sample data of each table to obtain the initial feature vector corresponding to the sample data of the table;

[0060] A first determining module is used to determine the fused feature vector corresponding to the table sample data based on the initial feature vector;

[0061] The first construction module is used to construct a loss function based on the above fused feature vector;

[0062] The first updating module is used to update the model parameters of the initial feature extraction model based on the above loss function to obtain the target feature extraction model.

[0063] In a fifth aspect, an embodiment of this specification provides a label prediction device, including:

[0064] An input module is used to input the table data to be processed corresponding to the user into the target feature extraction model to obtain the feature vector representation output by the target feature extraction model;

[0065] A generating module, configured to generate a predicted scenario label corresponding to the table data to be processed based on the feature vector representation, wherein the predicted scenario label is used to indicate the state information of the user in the corresponding scenario;

[0066] Among them, the above-mentioned target feature extraction model is a target feature extraction model generated by the method provided by the first aspect.

[0067] In a sixth aspect, an embodiment of this specification provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and when the processor executes the computer program, the method steps provided in the first aspect of the embodiment of this specification are implemented.

[0068] In a seventh aspect, an embodiment of this specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method steps provided in the first aspect of the embodiment of this specification.

[0069] In an eighth aspect, an embodiment of this specification provides a computer program product comprising instructions. When the above-mentioned computer program product runs on a computer or a processor, the above-mentioned computer or the above-mentioned processor executes the model training method provided in the first aspect or the second aspect of the embodiment of this specification, or the label prediction method provided in the third aspect.

[0070] The embodiments of this specification obtain multiple tabular sample data, wherein each tabular sample data includes multiple feature data of multiple feature types; encodes each tabular sample data to obtain the initial feature vector corresponding to the tabular sample data; determines the fused feature vector corresponding to the tabular sample data based on the initial feature vector; constructs a loss function based on the fused feature vector; updates the model parameters of the initial feature extraction model based on the loss function to obtain the target feature extraction model. Thus, by uniformly encoding different types of features in the tabular sample data, the effective expression and utilization of different types of features are achieved, and the expressive power of feature representation is improved through feature fusion, avoiding the information loss caused by independent feature processing, thereby significantly improving the model's adaptability and generalization performance to complex and diverse features, and thereby improving the performance of the final target feature extraction model in various tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0072] Figure 1 An exemplary system architecture diagram of a model training method provided in an embodiment of this specification;

[0073] Figure 2 A flowchart of a model training method provided in an embodiment of this specification;

[0074] Figure 3 A flowchart of a model training method provided in an embodiment of this specification;

[0075] Figure 4 A flowchart of a model training method provided in an embodiment of this specification;

[0076] Figure 5 A flowchart of a model training method provided in an embodiment of this specification;

[0077] Figure 6 A schematic diagram of a model pre-training architecture provided in an embodiment of this specification;

[0078] Figure 7 A schematic diagram of the specific structure of a feature interaction layer provided in an embodiment of this specification;

[0079] Figure 8 A flowchart of a tag prediction method provided in an embodiment of this specification;

[0080] Figure 9 A schematic diagram of the architecture of a tag prediction system provided in an embodiment of this specification;

[0081] Figure 10 A schematic diagram of the structure of a model training device provided in an embodiment of this specification;

[0082] Figure 11 A schematic diagram of the structure of a model training device provided in an embodiment of this specification;

[0083] Figure 12 A schematic diagram of the structure of a label prediction device provided in an embodiment of this specification;

[0084] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0085] To make the features and advantages of the embodiments of this specification more obvious and easy to understand, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the embodiments of this specification.

[0086] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of the embodiments of this specification, as detailed in the appended claims. And in the description of the embodiments of this specification, unless otherwise indicated, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a way to describe the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this specification, "multiple" refers to two or more than two.

[0087] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0088] In application platforms such as online business marketing, online shopping, government affairs / people's livelihood services, and human resources management, a large amount of relevant data exists in the form of tables. Traditional modeling methods usually require independent modeling of labels for different scenarios on each platform. Each scenario must go through a complete process including feature screening, algorithm modeling, model evaluation, and online deployment. However, there are many problems with independent modeling for each label. On the one hand, modeling based on local data leads to sparse data for long-tail scenarios or specific customer groups, which limits the model effect. On the other hand, models for different scenarios need to be developed and maintained independently, resulting in low efficiency in online deployment and iteration. Pre-training methods based on tabular data help to connect multiple scenarios and build a unified feature modeling base model. On the one hand, it can fully explore the inherent correlation between data between different scenarios and improve model accuracy and generalization capabilities. On the other hand, it supports rapid fine-tuning of each scenario based on a unified model, improving development and deployment efficiency. Unlike existing text and image pre-training models, tabular data has the structured characteristics of consisting of horizontal rows and vertical columns. Different tables are highly heterogeneous in structure, fields, and value types. Field values ​​include both numerical values ​​and categories, and may also contain text information. These characteristics bring significant challenges to the unified pre-training of tabular data. There is an urgent need to design an efficient and universal pre-training modeling solution based on the characteristics of tabular data.

[0089] Therefore, the embodiments of this specification provide a model training method to solve the technical problem of low generalization of the above-mentioned model.

[0090] See also Figure 1 , Figure 1An exemplary system architecture diagram of a model training method provided in an embodiment of this specification.

[0091] like Figure 1 As shown, the system architecture may include a terminal 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal 101 and the server 103. The network 102 may include various types of wired communication links or wireless communication links, for example, a wired communication link may include an optical fiber, a twisted pair, or a coaxial cable, and a wireless communication link may include a Bluetooth communication link, a Wireless-Fidelity (Wi-Fi) communication link, or a microwave communication link.

[0092] The terminal 101 can interact with the server 103 through the network 102 to receive messages from the server 103 or send messages to the server 103, or the terminal 101 can interact with the server 103 through the network 102 to receive messages or data sent by other users to the server 103. The terminal 101 can be hardware or software. When the terminal 101 is hardware, it can be various electronic devices, including but not limited to tablet computers, laptop portable computers and desktop computers. When the terminal 101 is software, it can be installed in the electronic devices listed above, which can be implemented as multiple software or software modules (for example: for providing distributed services), or it can be implemented as a single software or software module, which is not specifically limited here.

[0093] In the embodiments of this specification, terminal 101 first obtains an initial feature extraction model. During the training preparation phase, terminal 101 obtains multiple tabular sample data, each of which includes multiple feature data of multiple feature types. Once the samples and model architecture are prepared, terminal 101 encodes each tabular sample data to obtain an initial feature vector corresponding to the tabular sample data; determines a fused feature vector corresponding to the tabular sample data based on the initial feature vector; constructs a loss function based on the fused feature vector; and updates the model parameters of the initial feature extraction model based on the loss function to obtain a target feature extraction model.

[0094] The server 103 may be a server that provides various services. It should be noted that the server 103 may be hardware or software. When the server 103 is hardware, it may be implemented as a distributed server cluster consisting of multiple servers, or it may be implemented as a single server. When the server 103 is software, it may be implemented as multiple software or software modules (for example, for providing distributed services), or it may be implemented as a single software or software module, which is not specifically limited herein.

[0095] Alternatively, the system architecture may not include the server 103. In other words, the server 103 may be an optional device in the embodiments of this specification, that is, the method provided in the embodiments of this specification may be applied to a system structure that only includes the terminal 101, and the embodiments of this specification do not limit this.

[0096] It should be understood that Figure 1 The number of terminals, networks, and servers in the figure is only for illustration and any number of terminals, networks, and servers may be used according to implementation requirements.

[0097] See also Figure 2 , Figure 2 This is a flowchart of a model training method provided in an embodiment of this specification. The execution subject of an embodiment of this specification can be a terminal executing the model training method, a processor in the terminal executing the model training method, or a feature extraction model training service in the terminal executing the model training method. For ease of description, the specific execution process of the model training method is described below using the example of a processor in a terminal as the execution subject.

[0098] See also Figure 2 , Figure 2 This is a flow chart of a model training method provided in the embodiment of this specification. Figure 2 As shown, the model training method may at least include:

[0099] S202: Acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types.

[0100] Optionally, the table sample data may be a data sample with a table structure, specifically data organized in the form of a two-dimensional table. For example, the table sample data corresponds to a user one-to-one, that is, each table sample data represents the information record of a single user. The table sample data includes feature fields of multiple feature types, that is, feature data. Among them, the multiple feature types include numerical and categorical types. Numerical feature data such as age, transaction amount, service access time, device resolution, etc. have continuous or discrete numerical features; categorical feature data such as gender, city, membership level, device type, operating system, etc. have categorical features with limited enumerated values. It is worth noting that categorical feature data can include binary feature data, such as gender.

[0101] It is understandable that obtaining a plurality of tabular sample data may refer to obtaining a plurality of sets of tabular data samples from an actual application system or an experimental platform, and these data may be used as training data for the model.

[0102] S204: Encode the sample data of each table to obtain the initial feature vector corresponding to the sample data of the table.

[0103] The initial feature vector corresponding to the table sample data may be a numerical representation formed after encoding, for example, a low-dimensional or dense vector for subsequent model processing.

[0104] Optionally, each tabular sample data is encoded, i.e., the original discrete or text features are converted into vectors or numerical representations, such as embedding vectors. In other words, encoding the above-mentioned tabular sample data can convert the feature data of different feature types in the tabular sample data into a unified vector format that can be processed by a neural network or machine learning model, facilitating the execution of subsequent tasks.

[0105] S206: Determine a fused feature vector corresponding to the table sample data based on the initial feature vector.

[0106] Optionally, the fused feature vector can be a high-quality feature representation that is further extracted, learned, and optimized based on the initial feature vector. Specifically, the final feature representation can be obtained through multi-head attention mechanisms, weighted fusion operations, and other processing (such as weighted summation, gating mechanisms, etc.). Among them, the multi-head attention mechanism is a deep learning mechanism that calculates the relationship between feature data through multiple parallel attention heads, each head capturing different feature interaction patterns or focus points. Feature gating weight information can be a feature importance score or weight vector generated by a gated weight calculation network, which is used to guide feature weighting or selection.

[0107] S208. Construct a loss function based on the above fused feature vector.

[0108] Optionally, the loss function is a mathematical expression that measures the difference between the model output and the target, where the target can be one or more of feature restoration, contrastive learning, or consistency preservation. That is, the loss function can include at least one of the following losses: mask prediction loss, contrastive loss, and consistency loss. Mask prediction loss can be achieved by randomly masking (masking) part of the feature data, allowing the model to learn to restore the masked content; contrastive loss can be used to bring semantically similar sample data closer together and semantically dissimilar sample data further apart; consistency loss can be used to require the model to maintain consistent output under different input perturbations or different perspectives, thereby improving the model's robustness and stability.

[0109] S210. Update the model parameters of the initial feature extraction model based on the above loss function to obtain a target feature extraction model.

[0110] Optionally, the initial feature extraction model may be a basic model used to process tabular sample data at the beginning of training, and may be a structure such as a shallow network or a Transformer.

[0111] It can be understood that the model parameter update is based on the above-mentioned loss function to update the model parameters of the initial feature extraction model, which means adjusting the weight parameters in the model through optimization methods such as back propagation to gradually improve its performance on the training samples, and then obtain the target feature extraction model. Among them, the target feature extraction model can be the final model obtained after multiple rounds of training. The target feature extraction model can extract tabular feature data more accurately and more generally, and be used for classification, recommendation, search and other tasks in downstream scenarios.

[0112] In an embodiment of the present specification, a model training method is provided, which is obtained by obtaining a plurality of tabular sample data, wherein each tabular sample data includes a plurality of feature data of a plurality of feature types; encoding each tabular sample data to obtain an initial feature vector corresponding to the tabular sample data; determining a fused feature vector corresponding to the tabular sample data based on the initial feature vector; constructing a loss function according to the fused feature vector; and updating the model parameters of the initial feature extraction model based on the loss function to obtain a target feature extraction model. Thus, by uniformly encoding different types of features in the tabular sample data, effective expression and utilization of different types of features are achieved, and then the expressive power of feature representation is improved through feature fusion, avoiding information loss caused by independent feature processing, thereby significantly improving the model's adaptability and generalization performance to complex and diverse features, and thereby improving the performance of the final target feature extraction model in various tasks.

[0113] See also Figure 3 , Figure 3 This is a flow chart of a model training method provided in the embodiment of this specification. Figure 3 As shown, the model training method may at least include:

[0114] S302: Acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types.

[0115] Specifically, the above S302 is consistent with S202 and will not be repeated here.

[0116] S304 : For each of the above table sample data, mask at least part of the feature data in the above table sample data based on a preset masking strategy to obtain a masking result.

[0117] Optionally, the preset masking strategy can be a pre-defined feature masking rule, for example, a random masking strategy of 30% of features. Masking can involve hiding or replacing feature values ​​with null values ​​or specific flags to train the model's restoration capabilities. The masked result can refer to new tabular sample data formed after the masking process, in which some feature data is masked and some is not.

[0118] S306 : Using the masked feature data in the mask result as first feature data, and performing encoding processing on each first feature data to obtain a first feature vector corresponding to each first feature data.

[0119] That is, the first feature data may be feature data masked in the mask result; and the first feature vector is a vector representation obtained by converting the masked feature data into a vector representation that can be processed by the model.

[0120] Optionally, in S306, encoding each first feature data to obtain a first feature vector corresponding to each first feature data includes encoding feature column name information of each first feature data to obtain a first feature vector corresponding to each first feature data. The feature column name information of the first feature data may be a field name or label to which the masked feature data belongs, such as "gender," "age," or "city."

[0121] Specifically, for the first feature data whose feature type is categorical, the feature column name information of the first feature data is used as the model input, and the feature column name information is embedded to obtain the corresponding first feature vector. The first feature vector does not contain the actual feature value information, and only prompts the model to restore and predict the feature based on the column name information. In addition, for the first feature data whose feature type is numerical, when the first feature data is masked, the feature column name information of the first feature data is also used as the model input, and the value of the numerical feature is set to a constant 1 by default. That is, only the embedding process is performed based on the feature column name information, and the feature value is fixed to the default value 1, thereby prompting the model to restore and predict the feature.

[0122] In the embodiments of this specification, by uniformly utilizing feature column name information for embedded coding, the model can break away from its reliance on the actual value of the feature when processing masking tasks for different types of features, and focus on making predictions based on feature context and feature semantic position, thereby improving the modeling capabilities of table structures and feature associations. In particular, for categorical features, inputting only column name information can effectively reduce the adverse effects of feature sparsity on model training; for numerical features, fixing the value to a constant of 1 helps avoid interference with model learning caused by numerical scale, unifies the feature expression space, and improves the model's generalization capabilities and mask reconstruction effects.

[0123] S308 : Taking the feature data that has not been masked in the mask result as the second feature data, and performing encoding processing on each of the second feature data to obtain a second feature vector corresponding to each of the second feature data.

[0124] That is to say, the second feature data may be feature data that is not masked in the mask result; the second feature vector may be a vector representation obtained after encoding the second feature data, serving as supporting information to help the model restore the first feature data.

[0125] Optionally, in S308, each of the above-mentioned second feature data is encoded to obtain a second feature vector corresponding to each of the above-mentioned second feature data, including: determining the feature type of the above-mentioned second feature data, the above-mentioned feature type is a numerical type or a categorical type; when the above-mentioned feature type is the above-mentioned numerical type, encoding the product of the feature column name embedding vector of the above-mentioned second feature data and the feature value to obtain the second feature vector corresponding to each of the above-mentioned second feature data; or, when the above-mentioned feature type is the above-mentioned categorical type, encoding the combination formed by splicing the feature column name information and the feature value of the above-mentioned second feature data to obtain the second feature vector corresponding to each of the above-mentioned second feature data.

[0126] In one embodiment, the feature column name information can be the field name or label to which the feature data belongs, such as "gender" or "age"; the feature value can be the specific value of the feature data (feature field), such as "male" (gender) or "25" (age). Specifically, when the above-mentioned feature type is the above-mentioned numerical type, the feature column name embedding vector of the second feature data is multiplied by the actual numerical value, and the product is used as the model input. For example, for the feature "age = 25", the column name corresponding to "age" is first embedded into a vector, and then the vector is multiplied by the value 25 to obtain a feature vector that combines position and numerical weights. While retaining the numerical information, the semantic representation of the feature position is introduced, which can improve the expression effect of the numerical feature. In addition, when the above-mentioned feature type is the above-mentioned categorical type, the feature column name and feature value of the categorical feature are combined as a whole as the model input. For example, for the feature "gender: male", "gender-male" is used as the overall feature identifier and input into the embedding layer for embedding processing to obtain the corresponding second feature vector. In this way, the semantic information of feature position (gender) and value (male) can be captured simultaneously, improving the model's ability to express category features.

[0127] It is worth noting that, since the values ​​of numerical feature data themselves are usually continuous or discrete values ​​(such as 0, 1, 2, etc.), in the actual input process, it is difficult for the model to directly distinguish whether the value is the real value of the sample or a placeholder value replaced due to the masking strategy. For example, when the real value of the feature is 1, the input performance is consistent with the case where the mask placeholder uses the same value 1, resulting in the model being unable to correctly distinguish whether the feature is masked. Therefore, for numerical feature data, an additional segment coding vector can be introduced. The segment coding vector can be used to indicate whether each feature data is the first feature data that has been masked, thereby providing the model with explicit mask status information. Specifically, a segment coding value of 0 can indicate that the feature data is a real feature value that has not been masked, and a value of 1 indicates that the feature is a placeholder feature after masking. Through this segment coding, the model can effectively distinguish between the first feature data and the second feature data, thereby improving the model's discrimination ability and learning effect in masking tasks.

[0128] In the embodiments of this specification, the feature encoding method jointly models the feature's location information (column name) and value information (specific numerical value or category) according to the different feature types, so that the model can take into account both location perception and value expression when processing numerical and categorical features, thereby improving the feature representation capability, and further helping to improve the model's ability to understand the table structure, feature semantics and value distribution, thereby enhancing the model's generalization and expression capabilities.

[0129] S310: Combine the first feature vectors corresponding to the first feature data and the second feature vectors corresponding to the second feature data to serve as the initial feature vectors corresponding to the table sample data.

[0130] Optionally, a complete representation formed by concatenating or fusing the first feature vector and the second feature vector may be used as the initial feature vector.

[0131] Specifically, the initial eigenvector corresponding to the sample data in the above table can be expressed as:

[0132] H 0 =Encoder(x cat ,x num )

[0133] Among them, H 0 is the initial eigenvector, x cat Represents categorical feature data, x num Represents numerical feature data, and Encoder(·) represents encoding of each type of input feature data to obtain the corresponding feature vector, that is, the initial feature vector.

[0134] S312: Determine the fused feature vector corresponding to the table sample data based on the initial feature vector.

[0135] Specifically, the above S312 is consistent with S206 and will not be repeated here.

[0136] S314: Construct a loss function based on the above fused feature vector.

[0137] Specifically, the above S314 is consistent with S208 and will not be repeated here.

[0138] S316. Update the model parameters of the initial feature extraction model based on the above loss function to obtain the target feature extraction model.

[0139] Specifically, the above S316 is consistent with S210 and will not be repeated here.

[0140] The embodiments of this specification effectively improve the model's ability to express categorical and numerical feature data by combining adaptive encoding of feature types, joint modeling of column names and values, and explicit mask prompts (segment encoding). It also solves the problem of indistinguishable numerical feature masks, enabling the model to effectively perceive feature positions, values, and their mask states, thereby enhancing the modeling capabilities of table structures, feature semantics, and feature associations.

[0141] See also Figure 4 , Figure 4 This is a flow chart of a model training method provided in the embodiment of this specification. Figure 4 As shown, the model training method may at least include:

[0142] S402: Acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types.

[0143] Specifically, the above S402 is consistent with S202 and will not be repeated here.

[0144] S404: Encode the sample data of each table to obtain the initial feature vector corresponding to the sample data of the table.

[0145] S406. Perform multi-head attention calculation based on the above initial feature vector to obtain the target feature vector corresponding to the above table sample data. The above target feature vector is used to characterize the correlation relationship between each feature data in the above table sample data.

[0146] Among them, multi-head attention calculation (mechanism) is a deep learning mechanism that calculates the relationship between features through multiple parallel attention heads, and each head captures different feature interaction patterns or focus points.

[0147] In one embodiment, based on the initial feature vector, the dependency relationship between features can be modeled through multiple groups of attention calculation modules, and the results of multiple attention heads can be spliced ​​or aggregated to obtain a richer and more comprehensive target feature vector.

[0148] Optionally, in S406, the above-mentioned multi-head attention calculation is performed based on the above-mentioned initial feature vector to obtain the target feature vector corresponding to the above-mentioned table sample data, including: generating a query vector based on a preset learnable clustering vector sequence; performing a linear transformation on the above-mentioned initial feature vector to generate a key vector and a value vector; performing a multi-head attention calculation based on the above-mentioned query vector, the above-mentioned key vector and the above-mentioned value vector to project the above-mentioned initial feature vector to a preset clustering space to obtain the feature representation in the above-mentioned clustering space; performing a multi-head attention calculation based on the feature representation in the above-mentioned clustering space to obtain a clustering feature interaction result; restoring the above-mentioned clustering feature interaction result to the original space of the above-mentioned initial feature vector to obtain the target feature vector corresponding to the above-mentioned table sample data.

[0149] The preset learnable clustering vector sequence can be a set of predefined vectors that can be learned and optimized through training, representing different cluster centers or semantic groups. Specifically, the query vector (Query, Q) for attention calculation can be directly generated from the learnable clustering vector sequence. The query vector can be a query matrix. Specifically, the query vector Q can be determined by the following formula:

[0150]

[0151] Among them, E cluster represents a preset sequence of learnable clustering vectors, is a linear transformation matrix, which is a predetermined trainable parameter matrix used to transform the learnable clustering vector sequence E cluster Mapping (or projection) to the query vector space.

[0152] In one embodiment, the key vector (Key, K) and the value vector (Value, V) can be obtained by linearly transforming the initial feature vector, and are used to match the query vector and the weighted feature representation, respectively. Specifically, the key vector K and the value vector V can be determined by the following formula:

[0153] K=H 0 W K , V=H 0 W V

[0154] Furthermore, a multi-head attention calculation is performed based on the query vector Q, the key vector K, and the value vector V to project the initial feature vector into a preset cluster space to obtain a feature representation in the cluster space. This can be specifically implemented by the following formula:

[0155]

[0156] Among them, CrossAttention(Q,K,V) is the feature representation in the cluster space, d represents the feature dimension of the query vector Q and the key vector K, It can be used as a scaling factor to improve the numerical stability of attention weight calculation.

[0157] Furthermore, based on the feature representation in the above cluster space, multi-head attention calculation is performed to obtain the cluster feature interaction result, which can be specifically achieved by the following formula:

[0158]

[0159] Among them, the i-th head in the multi-head attention calculation i The calculation can be expressed as:

[0160]

[0161] Among them, H l Represents the feature vector of the l-th layer input, W o is the linear transformation weight matrix output after multi-head splicing, and Concat(·) represents the splicing operation of multiple attention head outputs.

[0162] Furthermore, the clustering feature interaction results are restored to the original space of the initial feature vector, that is, the target feature vector corresponding to the sample data in the table can be obtained.

[0163] In the embodiments of this specification, by introducing a multi-head attention mechanism based on learnable clustering vectors, the deep correlation between each feature data in the tabular sample data can be fully explored, which not only improves the modeling ability of feature context and semantic structure, but also realizes the two-way feature interaction and alignment of feature data between cluster space and original space, helping the model to capture richer and more diverse feature interaction patterns, enhancing the comprehensiveness and robustness of feature expression, thereby improving the model's ability to understand complex tabular data and generalization performance.

[0164] S408 , performing weight estimation on the target feature vector through a preset gated weight calculation network to obtain feature gated weight information corresponding to the sample data in the table.

[0165] The feature gating weight information is used to indicate the importance of each feature data in the feature update process, and can be used to highlight more important feature data and suppress less important feature data, thereby achieving selective feature update.

[0166] Optionally, the target feature vector can be input into a preset gated weight calculation network for linear transformation. The result obtained is the original representation of the weight vector, and then the result of the linear transformation is subjected to nonlinear activation processing. For example, the weight value is compressed to the [0,1] interval through a related function to control the influence of each feature.

[0167] Specifically, the feature gating weight information corresponding to the table sample data can be obtained by the following formula:

[0168]

[0169] in, is the target feature vector, w G is the weight matrix of the gating network (learnable parameters), σ(·) represents the nonlinear activation function, g l It is the feature gating weight information corresponding to the table sample data.

[0170] S410 : Based on the target feature vector and the feature gating weight information, determine the fused feature vector corresponding to the table sample data.

[0171] Optionally, in S410, the above-mentioned fused feature vector corresponding to the above-mentioned table sample data is determined based on the above-mentioned target feature vector and the above-mentioned feature gating weight information, including: based on the above-mentioned feature gating weight information, performing feature-by-feature weighting processing on the above-mentioned target feature vector to obtain a weighted feature vector; and fusing the above-mentioned weighted feature vector with the above-mentioned target feature vector to obtain a fused feature vector corresponding to the above-mentioned table sample data.

[0172] The target feature vector is weighted feature by feature, i.e., the feature gating weight is multiplied element by element by the target feature vector, and each feature dimension is weighted independently. The resulting weighted feature vector can reflect the performance of each feature data after adjustment based on its importance. The weighted feature vector is fused with the target feature vector, which can be combined with the original target feature vector by concatenation or element-by-element addition to obtain a fused feature vector corresponding to the table sample data for subsequent task processing.

[0173] Specifically, the fusion feature vector can be determined by the following formula:

[0174]

[0175] Among them, H l+1 represents the fused feature vector, g l is the feature gating weight information corresponding to the table sample data, ⊙ is the element-by-element multiplication sign, indicating that the target feature is weighted feature by feature according to the weight. Represents a fusion operation, which can be feature concatenation or element-by-element addition. A linear transformation operation, such as a fully connected layer, is used to map the fusion result to the specified output space. linear(·) represents a linear transformation operation.

[0176] In the embodiments of this specification, by introducing a gated weight calculation network to evaluate the feature importance of the target feature vector and combining the weight information to implement feature-by-feature weighting and fusion processing, it is possible to not only highlight the important features that have a greater impact on model decision-making, but also suppress the interference of redundant or irrelevant features, effectively improving feature utilization efficiency. In addition, the fusion of weighted features with original features further enriches the feature expression content, enhances the robustness and adaptability of feature representation, and helps improve the learning effect of subsequent tasks and the generalization ability of the model.

[0177] See also Figure 5 , Figure 5 This is a flow chart of a model training method provided in the embodiment of this specification. Figure 5 As shown, the model training method may at least include:

[0178] S502: Acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types.

[0179] Specifically, the above S502 is consistent with S202 and will not be repeated here.

[0180] S504: Encode the sample data of each table to obtain the initial feature vector corresponding to the sample data of the table.

[0181] Specifically, the above S504 is consistent with S202 and will not be repeated here.

[0182] S506: Determine a fused feature vector corresponding to the table sample data based on the initial feature vector.

[0183] Specifically, the above S506 is consistent with S206 and will not be repeated here.

[0184] S508 . Determine a first mask loss corresponding to the first feature data and a second mask loss corresponding to the second feature data based on the fused feature vector.

[0185] The first mask loss can be a cross-entropy loss (CE), which is used to evaluate the prediction accuracy of the fused feature vector for categorical feature data, specifically to maximize the prediction probability of the correct category label and reduce the category prediction error. The second mask loss can be a mean squared error loss (MSE), which can be used to evaluate the ability of the fused feature vector to restore the value of numerical feature data, specifically to minimize the square error between the predicted value and the true value.

[0186] S510: Determine a mask prediction loss according to the first mask loss and the second mask loss.

[0187] Optionally, the sum of the first mask loss and the second mask loss can be used as the mask prediction loss, which can be specifically expressed as:

[0188] L1=CE(y cat ,y′ cat )+MSE(y num ,y′ num )

[0189] Among them, CE(y cat ,y′ cat ) represents the first mask loss, y cat Represents the true value of the categorical feature data, y′ cat Indicates the predicted value of categorical feature data. MSE(y num,y′ num ) represents the second mask loss, y ntm Represents the true value of the numerical feature data, y′ num represents the predicted value of numerical feature data, CE(·) represents the cross entropy loss function, MSE(·) represents the mean square error loss function, and L1 represents the mask prediction loss.

[0190] S512 , constructing a contrast loss based on multiple fused feature vectors generated under multiple different preset mask strategies based on the sample data in the above table.

[0191] Specifically, based on multiple fused feature vectors generated from the same feature data under different masking strategies, a contrast loss is constructed to encourage the model to learn a robust and consistent representation. This can be constructed using the following formula:

[0192]

[0193] Where h′ i∈cat and h″ i∈cat Represents the fusion feature vector of category features under different masking strategies, y′ i∈num and y″ i∈num It represents the fusion feature vector of numerical features under different masking strategies, and L2 represents the contrast loss.

[0194] S514: Construct consistency loss based on the plurality of mask results corresponding to the plurality of different preset mask strategies based on the sample data in the table.

[0195] Optionally, a consistency loss can be constructed based on the prediction results of the masked first feature data under different masking strategies to maintain a stable prediction of the missing features. Specifically, it can be determined by the following formula:

[0196]

[0197] Among them, y′ i∈miss and y″ i∈miss It represents the predicted value of the second feature data under different masking strategies. The second feature data can be obtained from multiple masking results corresponding to different preset masking strategies. L3 represents the consistency loss.

[0198] S516: Weighted combination of the mask prediction loss, the contrast loss, and the consistency loss to obtain a loss function. Specifically, the loss function L can be expressed as:

[0199] L=L1+L2+L3

[0200] Among them, L1 is the mask prediction loss, L2 is the contrast loss, and L3 is the consistency loss.

[0201] S518. Based on the above loss function, back propagate and calculate the gradient of the initial feature extraction model.

[0202] Optionally, based on the loss function, a backpropagation algorithm is used to perform gradient calculations on the parameters of the initial feature extraction model to obtain gradient information of the model parameters with respect to the loss function. The gradient information is used to indicate the sensitivity of each parameter to the loss function in its current state.

[0203] S520. Based on the gradient, the initial parameters of the initial feature extraction model are updated until the loss function converges or a preset stop condition is met, thereby obtaining the target feature extraction model.

[0204] Furthermore, based on the gradient information, a gradient descent algorithm is used to iteratively update the initial parameters of the initial feature extraction model to reduce the value of the loss function. This parameter update process continues until the loss function converges to a preset threshold or satisfies a preset stopping condition (e.g., reaching a maximum number of iterations), thereby obtaining the target feature extraction model.

[0205] In the embodiments of this specification, by combining the cross-entropy loss of categorical features, the mean square error loss of numerical features, the contrast loss under different masking strategies, and the missing feature prediction consistency loss, the model's ability to restore and predict multiple types of features, robustness, and stability can be comprehensively improved, and the model optimization can be guided by a weighted combination of unified loss functions; and through the backpropagation and gradient update mechanism, the model parameters are continuously improved so that it can maintain accurate and effective representation and prediction under different masking strategies and feature missing conditions, thereby obtaining a high-quality feature extraction model with stronger generalization ability, better adaptability, and deeper understanding of features.

[0206] The following describes the architecture of the initial feature extraction model pre-training.

[0207] See also Figure 6 , Figure 6 This is a schematic diagram of a model pre-training architecture provided in the embodiments of this specification. Figure 6 As shown, the architecture may include at least: an input layer, a feature interaction layer, and an output layer.

[0208] Optionally, the input layer can be used to perform feature encoding on the table sample data. Specifically, the input layer includes at least: a feature encoding unit for encoding the unmasked second feature data, tokenizing and embedding the category-type (category, cat) feature data, embedding the numerical (number, num) feature data in combination with normalization or default values ​​to generate a corresponding embedded feature vector; a mask identification encoding unit for generating a segment code that identifies whether the feature is masked; a feature embedding unit for processing the unmasked second feature data, wherein CLS represents classification encoding. Encoding converts text into a fixed-length vector representation. emb1, emb2, emb3, emb4, and emb5 represent the complete feature embedding vectors (Embedding) of the feature data in the first to fifth columns, combining the column name and value information. The column name embedding unit is used to embed the masked first feature data based only on the feature column name information to generate the corresponding embedding vector. Zero represents an all-zero vector, which serves as a placeholder or feature missing mark. Cemb1, Cemb2, Cemb3, Cemb4, and Cemb5 represent the column name embedding vectors (Column Embedding) of the features in the first to fifth columns.

[0209] Furthermore, the feature interaction layer includes a gated transformer, which is specifically composed of multiple stacked transformer layers (×L blocks). It can be used to perform deep feature interaction modeling on the feature vector sequence of tabular data and capture the dependencies and associations between feature data. The output layer includes two projection heads, which can represent feature projection units under different tasks (such as comparison tasks and prediction tasks). Among them, under the prediction task, the loss type of the loss function corresponding to the categorical feature data is cross entropy loss, and the loss type of the loss function corresponding to the numerical feature data is mean square error loss; under the comparison task, the type of loss function used is contrast loss.

[0210] In the embodiments of this specification, it is possible to simultaneously process masked and unmasked multi-type feature data, and improve the model's comprehensive perception of feature positions and values ​​by combining feature encoding, mask identification, and column name embedding; based on the feature interaction mechanism of the gated transformer, it effectively models the deep correlation between features; and at the same time, combined with the dual-task output design, it accurately restores categorical and numerical features in prediction tasks, and improves the model's robustness to different masking strategies in comparison tasks, thereby achieving coordinated optimization of feature recovery and feature consistency learning, and comprehensively enhancing the model's modeling capabilities and generalization performance of table structures, feature semantics, and contextual relationships.

[0211] See also Figure 7 , Figure 7 This is a schematic diagram of the specific structure of a feature interaction layer provided in the embodiment of this specification. Figure 7 As shown, the feature interaction layer can be the above Figure 6 Specifically, the feature interaction layer may include: original feature encoding unit, initialization clustering unit, cross attention mechanism (Cross Attention) unit, gated transformer (Gated Transformer) clustering space feature interaction unit and dual projection (Project Head × 2) unit.

[0212] Optionally, the original feature encoding unit can be used to embed each feature of the table sample data to generate a sequence of original feature encoding vectors (emb1, emb2, ..., emb6), and introduce the global feature CLS vector at the beginning of the sequence to obtain the initial feature vector corresponding to each table sample data. The initialization clustering unit can be used to provide a preset learnable clustering vector sequence (emb1, emb2) as the query vector Q in the Cross Attention unit. The Cross Attention unit can be used to perform multi-head attention calculation based on the above query vector, the above key vector, and the above value vector to project the above initial feature vector into a preset cluster space to obtain the feature representation in the above cluster space. The Gated Transformer cluster space feature interaction unit includes a multi-head attention mechanism subunit (Multi-headAttention), a gated layer (Gated Layer), a linear layer (Linear Layer), a residual connection, and a normalization subunit (Norm). The dual projection head unit can include two parallel projection branches, each projection head can be used for feature projection and output of different tasks (such as comparison tasks and prediction tasks).

[0213] By combining raw feature encoding, clustering hints, and cross-spatial interactions with deep feature modeling, the embodiments of this specification are able to fully capture the structural relationships and semantic dependencies between table features, thereby enhancing the richness and expressiveness of feature representation. Furthermore, a deep interactive design combining multi-head attention, gating mechanisms, and residual normalization enhances the stability and generalization of feature modeling. Furthermore, dual projection heads enable parallel optimization of feature prediction and contrastive learning tasks, effectively improving the model's adaptability and overall performance across multiple tasks.

[0214] See also Figure 8 , Figure 8 This is a flow chart of a tag prediction method provided in the embodiment of this specification. Figure 8 As shown, the label prediction method at least includes:

[0215] S802: Input the table data to be processed corresponding to the user into the target feature extraction model to obtain the feature vector representation output by the target feature extraction model.

[0216] The target feature extraction model is generated by any of the model training methods provided in the embodiments of this specification. This target feature extraction model can encode, interact with, and aggregate multidimensional features in user table data, ultimately outputting a user feature vector representation that comprehensively reflects the user's behavioral patterns or attribute characteristics in a multi-feature space.

[0217] S804: Generate a predicted scenario label corresponding to the table data to be processed according to the feature vector representation, where the predicted scenario label is used to indicate the status information of the user in the corresponding scenario.

[0218] Optionally, the above scenario can be any of the following: financial service scenario, online shopping scenario, government affairs / people's livelihood service scenario, etc. Among them, in the financial service scenario, the corresponding prediction scenario label may include but is not limited to: whether the user has subscribed to the financial product, whether the relevant service has been processed, whether there is a risk of overdue payment or whether there is a risk of loss, etc. In the online shopping scenario, the corresponding prediction scenario label may include but is not limited to: whether the user clicks on the product, whether it is added to the shopping cart, whether the order is placed, whether the product is returned, or whether the activity is stopped, etc. In the government affairs / people's livelihood service scenario, the corresponding prediction scenario label may include but is not limited to: whether the user has made an appointment to process the service, whether the relevant fees have been paid, whether the information update has been completed, or whether there is a risk of service loss, etc.

[0219] Further, see Figure 9 , Figure 9 This is a schematic diagram of the architecture of a tag prediction system provided in the embodiment of this specification. Figure 9 As shown, this structure can be used to execute the above-mentioned label prediction method. Specifically, it can uniformly process feature data from different scenarios, including numerical feature data and categorical feature data. By inputting the tabular features into a pre-trained tabular data model base, feature representations with global expression capabilities are extracted, and specific tasks are connected based on embedding or lightweight adaptation (Low-Rank Adaptation, LoRA) to achieve label prediction for a variety of scenarios, such as whether the user purchases, renews, or churns. This method has strong generalization and high adaptability. It can support efficient prediction of multiple tasks and multiple scenarios while maintaining the consistency of the input feature structure. It is suitable for applications in multiple fields such as finance, e-commerce, travel, and government affairs.

[0220] It should be understood that, although the steps in the flowcharts of the above-mentioned embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above-mentioned embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0221] Based on the inventive concept of the above model training method, Figure 10 As shown, the embodiment of this specification also provides a model training device 1000 for implementing the above-mentioned model training method. The model training device 1000 includes:

[0222] A first acquisition module 1010 is configured to acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types;

[0223] A first encoding module 1020 is used to encode the sample data of each table to obtain an initial feature vector corresponding to the sample data of the table;

[0224] A first determining module 1030 is configured to determine a fused feature vector corresponding to the table sample data based on the initial feature vector;

[0225] A first construction module 1040 is configured to construct a loss function based on the fused feature vector;

[0226] The first updating module 1050 is used to update the model parameters of the initial feature extraction model based on the above loss function to obtain a target feature extraction model.

[0227] In a possible implementation, the first encoding module 1020 includes:

[0228] a masking unit configured to perform masking processing on at least part of the feature data in the above-mentioned table sample data based on a preset masking strategy to obtain a masking result;

[0229] a first encoding unit, configured to use the masked feature data in the mask result as first feature data, and perform encoding processing on each first feature data to obtain a first feature vector corresponding to each first feature data;

[0230] a second encoding unit, configured to use the unmasked feature data in the masking result as second feature data, and perform encoding processing on each of the second feature data to obtain a second feature vector corresponding to each of the second feature data;

[0231] The first determining unit is configured to combine the first feature vectors corresponding to the first feature data and the second feature vectors corresponding to the second feature data as the initial feature vectors corresponding to the table sample data.

[0232] In a possible implementation, the first encoding unit includes:

[0233] The first encoding subunit is used to encode the feature column name information of each first feature data to obtain a first feature vector corresponding to each first feature data.

[0234] In a possible implementation, the second encoding unit includes:

[0235] A first determining subunit is configured to determine a feature type of the second feature data, where the feature type is a numerical value or a categorical type;

[0236] The second encoding subunit is used to encode the product of the feature column name embedding vector and the feature value of the above-mentioned second feature data when the above-mentioned feature type is the above-mentioned numerical type, so as to obtain the second feature vector corresponding to the above-mentioned each second feature data; or, when the above-mentioned feature type is the above-mentioned categorical type, encode the combination formed by splicing the feature column name information and the feature value of the above-mentioned second feature data to obtain the second feature vector corresponding to the above-mentioned each second feature data.

[0237] In a possible implementation, the first determining module 1030 includes:

[0238] A first calculation unit is used to perform multi-head attention calculation based on the above-mentioned initial feature vector to obtain a target feature vector corresponding to the above-mentioned table sample data, and the above-mentioned target feature vector is used to represent the correlation relationship between each feature data in the above-mentioned table sample data;

[0239] A second determining unit is configured to perform weight estimation on the target feature vector through a preset gated weight calculation network to obtain feature gating weight information corresponding to the sample data in the table;

[0240] The third determining unit is used to determine the fused feature vector corresponding to the table sample data based on the target feature vector and the feature gating weight information.

[0241] In a possible implementation, the first computing unit includes:

[0242] A first generating subunit, configured to generate a query vector based on a preset learnable clustering vector sequence;

[0243] A second generating subunit is used to perform a linear transformation on the initial feature vector to generate a key vector and a value vector;

[0244] a second determining subunit, configured to perform a multi-head attention calculation based on the query vector, the key vector, and the value vector to project the initial feature vector into a preset clustering space to obtain a feature representation in the clustering space;

[0245] The computing subunit is used to perform multi-head attention calculation based on the feature representation in the above cluster space to obtain the cluster feature interaction result;

[0246] The third determining subunit is used to restore the above cluster feature interaction results to the original space of the above initial feature vector to obtain the target feature vector corresponding to the above table sample data.

[0247] In a possible implementation, the feature gating weight information is used to indicate the importance of each feature data in the feature update process;

[0248] The third determining unit includes:

[0249] A first processing subunit is configured to perform feature-by-feature weighting processing on the target feature vector based on the feature gating weight information to obtain a weighted feature vector;

[0250] The second processing sub-unit is used to fuse the weighted feature vector with the target feature vector to obtain a fused feature vector corresponding to the table sample data.

[0251] In one possible implementation, the above loss function includes mask prediction loss, contrast loss, and consistency loss;

[0252] The first building block 1040 includes:

[0253] a fourth determining unit, configured to determine a first mask loss corresponding to the first feature data and a second mask loss corresponding to the second feature data based on the fused feature vector;

[0254] a fifth determining unit, configured to determine a mask prediction loss according to the first mask loss and the second mask loss;

[0255] A first construction unit is configured to construct a contrast loss based on a plurality of fused feature vectors generated under a plurality of different preset masking strategies based on the sample data in the above table;

[0256] A second construction unit is configured to construct a consistency loss based on the plurality of mask results corresponding to the plurality of different preset mask strategies of the sample data in the table;

[0257] The sixth determining unit is configured to perform a weighted combination of the mask prediction loss, the contrast loss, and the consistency loss to obtain a loss function.

[0258] In a possible implementation, the first update module 1050 includes:

[0259] A second computing unit is configured to calculate the gradient of the initial feature extraction model by back propagation based on the loss function;

[0260] An updating unit is used to update the initial parameters of the initial feature extraction model based on the gradient until the loss function converges or satisfies a preset stopping condition, thereby obtaining the target feature extraction model.

[0261] Based on the inventive concept of the above model training method, Figure 11 As shown, the embodiment of this specification also provides a model training device 1100 for implementing the above-mentioned model training method. The model training device 1100 includes:

[0262] A second acquisition module 1110 is configured to acquire a plurality of tabular sample data corresponding to a plurality of users, wherein each tabular sample data includes a plurality of feature data of a plurality of feature types, and each tabular sample data includes at least one of the following: basic information, behavior information, and device information of the user;

[0263] The second encoding module 1120 is used to encode the sample data of each table to obtain the initial feature vector corresponding to the sample data of the table;

[0264] A second determining module 1130 is configured to determine a fused feature vector corresponding to the table sample data based on the initial feature vector;

[0265] A second construction module 1140 is configured to construct a loss function based on the fused feature vector.

[0266] The second updating module 1150 is used to update the model parameters of the initial feature extraction model based on the above-mentioned loss function to obtain a target feature extraction model; wherein the above-mentioned target feature extraction model is a target feature extraction model generated by any one of the above-mentioned model training methods in the embodiments of this specification.

[0267] The division of the modules in the above-mentioned model training device is only for illustration. In other embodiments, the model training device can be divided into different modules as needed to complete all or part of the functions of the above-mentioned model training device. The implementation of each module in the model training device provided in the embodiments of this specification can be in the form of a computer program. The computer program can be run on a terminal or a server. The program modules constituted by the computer program can be stored in the memory of the terminal or the server. When the computer program is executed by the processor, all or part of the steps of the model training method described in the embodiments of this specification are implemented.

[0268] Based on the inventive concept of the above label prediction method, Figure 12 As shown, the embodiment of this specification also provides a label prediction device 1200 for implementing the above-mentioned label prediction method. The model training device 1200 includes:

[0269] Input module 1210, used to input the table data to be processed corresponding to the user into the target feature extraction model, and obtain the feature vector representation output by the target feature extraction model;

[0270] Generation module 1220 is used to generate the predicted scene label corresponding to the above-mentioned table data to be processed based on the above-mentioned feature vector representation, and the above-mentioned predicted scene label is used to indicate the status information of the above-mentioned user in the corresponding scene; wherein the above-mentioned target feature extraction model is a target feature extraction model generated by any model training method provided in the embodiments of this specification.

[0271] The division of the modules in the above-mentioned label prediction device is for illustration only. In other embodiments, the label prediction device can be divided into different modules as needed to complete all or part of the functions of the above-mentioned label prediction device. The implementation of each module in the label prediction device provided in the embodiments of this specification can be in the form of a computer program. The computer program can be run on a terminal or a server. The program modules constituting the computer program can be stored in the memory of the terminal or the server. When the computer program is executed by the processor, all or part of the steps of the label prediction method described in the embodiments of this specification are implemented.

[0272] The embodiment of this specification also provides an electronic device, which can be a server, and its internal structure diagram can be as follows: Figure 13As shown. The electronic device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and an external device. The communication interface of the electronic device is used to communicate with an external terminal through a network connection. The processor of the electronic device executes a computer program to implement a model training method or a label prediction method.

[0273] Those skilled in the art will understand that Figure 13 The structure shown in the figure is merely a block diagram of a portion of the structure related to the embodiment scheme of this specification, and does not constitute a limitation on the electronic device to which the embodiment scheme of this specification is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0274] In one possible implementation, a computer storage medium is provided, storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps in the above-described embodiments. If the components of the electronic device described above are implemented as software functional units and sold or used as independent products, they may be stored in the computer-readable storage medium.

[0275] In a possible implementation, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0276] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The above-mentioned computer program product includes one or more computer instructions. When the above-mentioned computer program instructions are loaded and executed on a computer, the above-mentioned process or function according to the embodiment of this specification is generated in whole or in part. The above-mentioned computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above-mentioned computer instructions can be stored in a computer storage medium or transmitted by the above-mentioned computer storage medium. The above-mentioned computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The above-mentioned computer storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available media may be magnetic media (eg, floppy disks, hard disks, tapes), optical media (eg, digital versatile discs (DVDs)), or semiconductor media (eg, solid state disks (SSDs)).

[0277] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the account balance, loan balance, outstanding amount, user identification field, asset balance data, etc. involved in this specification are all obtained with full authorization.

[0278] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. The technical features of this embodiment and the implementation scheme can be combined in any manner unless they conflict.

[0279] The above embodiments are merely preferred embodiments of the embodiments of this specification and are not intended to limit the scope of the embodiments of this specification. Without departing from the design spirit of the embodiments of this specification, various modifications and improvements made to the technical solutions of the embodiments of this specification by ordinary technicians in this field should fall within the scope of protection determined by the claims.

[0280] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A model training method, comprising: Acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types; Encoding each of the table sample data to obtain an initial feature vector corresponding to the table sample data; Determine a fused feature vector corresponding to the table sample data based on the initial feature vector; Constructing a loss function based on the fused feature vector; The model parameters of the initial feature extraction model are updated based on the loss function to obtain a target feature extraction model.

2. The method according to claim 1, wherein encoding each of the table sample data to obtain the initial feature vector corresponding to the table sample data comprises: For each of the table sample data, masking is performed on at least part of the feature data in the table sample data based on a preset masking strategy to obtain a masking result; Using the masked feature data in the mask result as first feature data, and performing encoding processing on each first feature data to obtain a first feature vector corresponding to each first feature data; Taking the feature data that has not been masked in the mask result as second feature data, and performing encoding processing on each second feature data to obtain a second feature vector corresponding to each second feature data; The first feature vectors corresponding to the first feature data and the second feature vectors corresponding to the second feature data are combined to serve as the initial feature vectors corresponding to the table sample data.

3. The method according to claim 2, wherein encoding each first feature data to obtain a first feature vector corresponding to each first feature data comprises: The feature column name information of each first feature data is encoded to obtain a first feature vector corresponding to each first feature data.

4. The method according to claim 2, wherein encoding the second feature data to obtain the second feature vector corresponding to the second feature data comprises: Determine a feature type of the second feature data, where the feature type is a numerical type or a categorical type; In the case where the feature type is the numerical type, encoding the product of the feature column name embedding vector and the feature value of the second feature data to obtain the second feature vector corresponding to each second feature data; or In the case where the feature type is the category type, encoding processing is performed on a combination of feature column name information and feature values ​​of the second feature data to obtain a second feature vector corresponding to each second feature data.

5. The method according to claim 1, wherein determining the fused feature vector corresponding to the table sample data based on the initial feature vector comprises: Perform multi-head attention calculation based on the initial feature vector to obtain a target feature vector corresponding to the table sample data, wherein the target feature vector is used to represent the correlation relationship between each feature data in the table sample data; The target feature vector is weighted estimated by a preset gated weight calculation network to obtain feature gated weight information corresponding to the table sample data; Based on the target feature vector and the feature gating weight information, a fused feature vector corresponding to the table sample data is determined.

6. The method of claim 5, wherein performing multi-head attention calculation based on the initial feature vector to obtain the target feature vector corresponding to the table sample data comprises: Generate a query vector based on a preset sequence of learnable clustering vectors; Performing a linear transformation on the initial feature vector to generate a key vector and a value vector; Performing a multi-head attention calculation based on the query vector, the key vector, and the value vector to project the initial feature vector into a preset cluster space to obtain a feature representation in the cluster space; Performing multi-head attention calculation based on the feature representation in the cluster space to obtain a cluster feature interaction result; The cluster feature interaction result is restored to the original space of the initial feature vector to obtain the target feature vector corresponding to the table sample data.

7. The method according to claim 5, wherein the feature gating weight information is used to indicate the importance of each feature data in the feature updating process; The determining, based on the target feature vector and the feature gating weight information, a fused feature vector corresponding to the table sample data includes: Based on the feature gating weight information, performing feature-by-feature weighting processing on the target feature vector to obtain a weighted feature vector; The weighted feature vector is fused with the target feature vector to obtain a fused feature vector corresponding to the table sample data.

8. The method of claim 2, wherein the loss function comprises a mask prediction loss, a contrast loss, and a consistency loss; The constructing a loss function according to the fused feature vector includes: Determining a first mask loss corresponding to the first feature data and a second mask loss corresponding to the second feature data based on the fused feature vector; determining a mask prediction loss according to the first mask loss and the second mask loss; Constructing a contrast loss based on a plurality of fused feature vectors generated under a plurality of different preset masking strategies based on the tabular sample data; Constructing consistency loss based on the table sample data and the plurality of mask results corresponding to the plurality of different preset mask strategies; The mask prediction loss, the contrast loss, and the consistency loss are weighted and combined to obtain a loss function.

9. The method according to claim 1, wherein updating the model parameters of the initial feature extraction model based on the loss function to obtain the target feature extraction model comprises: Based on the loss function, back propagation is used to calculate the gradient of the initial feature extraction model; Based on the gradient, the initial parameters of the initial feature extraction model are updated until the loss function converges or meets a preset stopping condition, thereby obtaining the target feature extraction model.

10. A model training method comprising: Acquire multiple tabular sample data corresponding to multiple users, each tabular sample data in the multiple tabular sample data includes multiple feature data of multiple feature types, and each tabular sample data includes at least one of the following: basic information, behavior information, and device information of the user; Encoding each of the table sample data to obtain an initial feature vector corresponding to the table sample data; Determine a fused feature vector corresponding to the table sample data based on the initial feature vector; Constructing a loss function based on the fused feature vector; Updating the model parameters of the initial feature extraction model based on the loss function to obtain a target feature extraction model; The target feature extraction model is a target feature extraction model generated by the model training method according to any one of claims 1 to 9.

11. A label prediction method, comprising: Input the table data to be processed corresponding to the user into the target feature extraction model to obtain the feature vector representation output by the target feature extraction model; Generating a predicted scenario label corresponding to the table data to be processed according to the feature vector representation, wherein the predicted scenario label is used to indicate the state information of the user in the corresponding scenario; The target feature extraction model is a target feature extraction model generated by the model training method according to any one of claims 1 to 10.

12. A model training device comprising: A first acquisition module is configured to acquire a plurality of table sample data, wherein each of the plurality of table sample data includes a plurality of feature data of a plurality of feature types; A first encoding module is used to encode the sample data of each table to obtain an initial feature vector corresponding to the sample data of the table; A first determining module, configured to determine a fused feature vector corresponding to the table sample data based on the initial feature vector; A first construction module is used to construct a loss function according to the fused feature vector; The first updating module is used to update the model parameters of the initial feature extraction model based on the loss function to obtain a target feature extraction model.

13. A model training device comprising: A second acquisition module is configured to acquire a plurality of tabular sample data corresponding to a plurality of users, wherein each tabular sample data includes a plurality of feature data of a plurality of feature types, and each tabular sample data includes at least one of the following: basic information, behavior information, and device information of the user; A second encoding module is used to encode the sample data of each table to obtain an initial feature vector corresponding to the sample data of the table; A second determining module is used to determine the fused feature vector corresponding to the table sample data based on the initial feature vector; A second construction module is used to construct a loss function according to the fused feature vector; A second updating module is used to update the model parameters of the initial feature extraction model based on the loss function to obtain a target feature extraction model; The target feature extraction model is a target feature extraction model generated by the model training method according to any one of claims 1 to 9.

14. A label prediction device, comprising: An input module, configured to input the table data to be processed corresponding to the user into the target feature extraction model, and obtain a feature vector representation of the output of the target feature extraction model; a generating module, configured to generate a predicted scenario label corresponding to the table data to be processed according to the feature vector representation, wherein the predicted scenario label is used to indicate the state information of the user in the corresponding scenario; The target feature extraction model is a target feature extraction model generated by the model training method according to any one of claims 1 to 10.

15. An electronic device comprising: processor and memory; The processor is connected to the memory; The memory is used to store executable program code; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method according to any one of claims 1 to 11.

16. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 11.

17. A computer program product comprising instructions, which, when run on a computer or a processor, enables the computer or the processor to execute the model training method according to any one of claims 1 to 11.