Method and apparatus for training a predictive model with missing data
Patent Information
- Application Number
- CN202611207713.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-10
- Publication Date
- 2026-09-08
AI Technical Summary
然而,这类方法对缺失值所携带的结构性信息的利用往往较为有限,仍然难以获得理想的性能表现
[0046]本说明书实施例提出的利用缺失数据训练预测模型的方法及装置,方法将预测任务和重构任务结合,在原始用户数据的基础上,人工构造缺失数据,经过模型处理后,基于模型输出的结果同时进行预测任务和人工缺失数据重构任务,并联合两个任务的训练损失对模型进行训练。使模型的重构过程被下游预测目标(包括分类目标和重构目标)所引导,避免了传统两阶段方法中重构目标与预测目标不对齐的问题。通过在训练样本的可见维度上额外施加掩码处理,模型在训练阶段能够接触多种不同的数据缺失模式,增强了模型对不同缺失率和不同缺失机制的适应能力。预测和重构均基于解码网络的输出表征,使得模型在进行预测时已融合了可见信息和对缺失信息的推断结果,形成了更为完整的全局表征,从而在数据缺失的场景下能够获得更优的预测性能。
Smart Images

Figure CN122713367A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of machine learning, and more particularly to methods and apparatus for training predictive models using missing data. Background Technology
[0002] Tabular data, organized in rows and columns, is a common and valuable data format in many fields such as financial risk control, healthcare, customer analysis, and industrial manufacturing, supporting downstream prediction tasks such as classification and regression. However, in actual data collection, due to various factors such as collection cost limitations, sensor equipment failures, data entry omissions, and privacy compliance requirements, tabular data commonly suffers from missing values. Failure to effectively handle missing values will directly lead to a reduction in the usable sample size, biased statistical estimations, and consequently affect the stability and generalization ability of downstream prediction models.
[0003] For handling missing values in tabular data, there are two main approaches in related technologies. The first is a two-stage approach of "imputation followed by prediction," where missing values are first filled using an imputation algorithm to obtain a complete dataset, and then the imputed dataset is input into a downstream prediction model for prediction. However, since the optimization objectives of the imputation stage and the prediction stage are not entirely the same, the reconstructed data does not necessarily lead to improved prediction performance, resulting in these methods often failing to achieve ideal performance in downstream prediction tasks. The second approach is direct prediction, which involves building an end-to-end prediction model capable of inference directly on incomplete data, bypassing the separate imputation stage. However, this approach often has limited utilization of the structural information carried by missing values, still struggling to achieve ideal performance. Therefore, a method is needed to better train prediction models on missing datasets. Summary of the Invention
[0004] This specification describes one or more embodiments of a method and apparatus for training a prediction model using missing data, which trains the prediction model end-to-end by jointly reconstructing and predicting optimization objectives.
[0005] In a first aspect, a method is provided for training a prediction model using missing data, the prediction model comprising an encoding network, a decoding network, a prediction network, and several reconstruction networks, the method comprising:
[0006] Obtain sequence data corresponding to the target user, including predetermined special words, attribute values of the target user in several visible dimensions, and padding markers for multiple invisible dimensions; the multiple invisible dimensions include masking dimensions where the attribute values of the target user are masked and missing dimensions where attribute values are missing.
[0007] The sequence data is input into the encoding network to obtain the encoded representation of each position in the sequence, wherein each invisible dimension position shares the same masked word representation;
[0008] Each encoded representation is input into the decoding network to obtain the decoded representation of each position in the sequence;
[0009] The decoded representation of the specific word position is input into the prediction network to obtain the prediction result for the target user; and the decoded representation of any mask dimension position is input into the corresponding reconstruction network to obtain the corresponding reconstructed user attribute value.
[0010] Based on the prediction results and reconstructed user attribute values, the total training loss is determined, and then the trainable parameters of the prediction model are adjusted.
[0011] In some possible implementations, obtaining the sequence data corresponding to the target user includes:
[0012] Obtain the original sequence data corresponding to the target user, including the attribute values of the target user in several visible dimensions, as well as the missing dimensions where several attribute values are missing;
[0013] The attribute values of several visible dimensions are masked to obtain several mask dimensions;
[0014] Add padding markers to each mask dimension and missing dimension, and add special terms at the beginning of the sequence to obtain the sequence data.
[0015] In some possible implementations, the encoding network is based on an attention mechanism; the step of inputting the sequence data into the encoding network to obtain the encoded representation of each position in the sequence includes:
[0016] The sequence data is input into an encoding network for target self-attention processing to obtain the encoded representation of each position in the sequence; wherein, the target self-attention processing only calculates the attention scores between special words and each visible dimension.
[0017] In some possible implementations, the step of inputting the sequence data into the encoding network for target self-attention processing includes:
[0018] After removing the data from each padding marker in the sequence data, it is input into the encoding network for self-attention processing.
[0019] In some possible implementations, the step of inputting the sequence data into the encoding network for target self-attention processing includes:
[0020] The sequence data is input into the encoding network, and attention masks are added to each padding marker in the sequence data before self-attention processing is performed.
[0021] In some possible implementations, the visible dimension includes a numerical dimension and a categorical dimension; the step of inputting the sequence data into the encoding network includes:
[0022] The sequence data is embedded to obtain an embedded representation; the embedding process includes: inputting data of any numerical dimension into the corresponding linear layer, and inputting data of any categorical dimension into the corresponding embedding table;
[0023] The embedded representation is input into the encoding network.
[0024] In some possible implementations, determining the total training loss based on the prediction results and reconstructed user attribute values includes:
[0025] Based on the prediction results and the target user's tags, determine the prediction loss;
[0026] The reconstruction loss is determined based on the difference between the reconstructed user attribute values and the corresponding real user attribute values for each mask dimension.
[0027] The total training loss is determined based on the predicted loss and the reconstruction loss.
[0028] In some possible implementations, the prediction network is a classification network, the prediction result is the predicted probability distribution of the target user belonging to each user category, and the label is a label category; determining the prediction loss based on the prediction result and the target user's label includes:
[0029] The prediction loss is determined based on the cross-entropy between the predicted probability distribution and the label category.
[0030] In some possible implementations, the prediction network is a regression network, the prediction result is the predicted value of the target user on the target attribute dimension, and the label is the true value of the target user on the target attribute dimension; determining the prediction loss based on the prediction result and the target user's label includes:
[0031] The prediction loss is determined based on the mean square error between the predicted and actual values.
[0032] In some possible implementations, determining the reconstruction loss based on the difference between the reconstructed user attribute values and the corresponding real user attribute values for each mask dimension includes:
[0033] The first reconstruction loss is determined based on the mean square error between the reconstructed user attribute values and their actual user attribute values for each numerical dimension.
[0034] The second reconstruction loss is determined based on the cross-entropy between the reconstructed user attribute values and their actual user attribute values for each category dimension.
[0035] The reconstruction loss is determined based on the first reconstruction loss and the second reconstruction loss.
[0036] In some possible implementations, the masked lexical representation is a trainable representation; adjusting the trainable parameters of the prediction model includes:
[0037] Adjust the trainable parameters of the prediction model and the masked lexical representation.
[0038] Secondly, an apparatus is provided for training a prediction model using missing data, the prediction model including an encoding network, a decoding network, a prediction network, and several reconstruction networks, the apparatus comprising:
[0039] The acquisition unit is configured to acquire sequence data corresponding to the target user, including predetermined special words, attribute values of the target user in several visible dimensions, and padding markers for multiple invisible dimensions; the multiple invisible dimensions include masking dimensions where the attribute values of the target user are masked and missing dimensions where attribute values are missing.
[0040] The encoding unit is configured to input the sequence data into the encoding network to obtain the encoded representation of each position in the sequence, wherein each invisible dimension position shares the same masked word representation;
[0041] The decoding unit is configured to input each encoded representation into the decoding network to obtain the decoded representation of each position in the sequence;
[0042] The prediction and reconstruction unit is configured to input the decoded representation of the special word position into the prediction network to obtain the prediction result for the target user; and to input the decoded representation of any mask dimension position into the corresponding reconstruction network to obtain the corresponding reconstructed user attribute value.
[0043] The training unit is configured to determine the total training loss based on the prediction results and reconstructed user attribute values, and then adjust the trainable parameters of the prediction model.
[0044] Thirdly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of the first aspect.
[0045] Fourthly, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method of the first aspect.
[0046] The method and apparatus for training a prediction model using missing data, as proposed in the embodiments of this specification, combine prediction and reconstruction tasks. Based on original user data, missing data is manually constructed. After model processing, prediction and reconstruction tasks are performed simultaneously based on the model output, and the training losses of both tasks are combined to train the model. This ensures that the model's reconstruction process is guided by downstream prediction objectives (including classification and reconstruction objectives), avoiding the misalignment between reconstruction and prediction objectives in traditional two-stage methods. By applying additional masking processing to the visible dimension of training samples, the model can encounter various data missing patterns during the training phase, enhancing its adaptability to different missing rates and mechanisms. Both prediction and reconstruction are based on the output representation of the decoding network, allowing the model to integrate visible information and inferences about missing information during prediction, forming a more complete global representation and thus achieving better prediction performance in scenarios with missing data. Attached Figure Description
[0047] To more clearly illustrate the technical solutions of the various embodiments disclosed in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only a few embodiments disclosed in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 A schematic diagram of an architecture for training a prediction model using missing data, according to one embodiment, is shown. Figure 2 This diagram illustrates the architecture of the coding layer according to one embodiment. Figure 3 A flowchart illustrating a method for training a prediction model using missing data according to one embodiment; Figure 4 A schematic diagram of asymmetric self-attention masks is shown; Figure 5 A schematic block diagram of an apparatus for training a prediction model using missing data according to one embodiment is shown. Detailed Implementation
[0049] The solution provided in this specification will now be described with reference to the accompanying drawings.
[0050] As mentioned earlier, tabular data is a common and valuable data format in many fields such as financial risk control, online e-commerce, and customer analysis, supporting downstream prediction tasks such as classification and regression. In these scenarios, users can be financial customers, internet users, product consumers, etc., and the specific types of users are not limited in the embodiments of this specification. User samples typically have multiple attribute dimensions, and the composition of attribute dimensions and the division of user categories vary in different application scenarios. Several exemplary scenarios are listed below for illustration.
[0051] In financial risk control scenarios, users are customers applying for loans or holding credit products. Their attribute dimensions can include: account holding period, number of historical overdue payments, current debt ratio, number of transactions in the past six months, credit rating, industry category, and city level. Among these, credit rating, industry category, and city level are categorical dimensions with finite discrete values, while account holding period, number of historical overdue payments, current debt ratio, and number of transactions in the past six months are numerical dimensions. Corresponding user categories can include normal customers and risky customers, or be divided into low-risk, medium-risk, and high-risk users, etc. The classification task is to predict whether a user is a risky customer or their risk level based on their attribute values, while the regression task can be to predict the credit limit a user can afford. In actual data collection, due to some users refusing to provide certain information or data system upgrades at financial institutions resulting in incomplete historical data, some attribute dimensions of the user sample may be missing.
[0052] In e-commerce operations, users are registered users or active consumers on the platform. Their attribute dimensions can include: registration duration, cumulative spending, number of login days in the past 30 days, number of days since the last purchase, frequently used device type, membership level, preferred product categories, and province of delivery. Among these, frequently used device type, membership level, preferred product categories, and province of delivery are categorical dimensions, while registration duration, cumulative spending, number of login days in the past 30 days, and number of days since the last purchase are numerical dimensions. Corresponding user categories can include high-value users, active users, and users at risk of churn. The classification task is to predict the user group to which a user belongs based on their attribute values to support differentiated operational strategies. In real-world scenarios, attribute dimensions such as cumulative spending and preferred product categories for newly registered users may be missing due to insufficient accumulation of behavioral data.
[0053] In video recommendation scenarios, users can be registered users of the content platform, and their attribute dimensions can include: completion rate stratification, interaction behavior type distribution (likes / comments / shares), content tag dwell time, negative feedback trigger frequency, viewing time distribution, and creator attention concentration. Among these, completion rate stratification, interaction behavior type distribution, viewing time distribution, and creator attention concentration are categorical dimensions, while content tag dwell time and negative feedback trigger frequency are numerical dimensions. Corresponding user categories can include deeply immersed users, superficially browsing users, and actively searching users, etc. The classification task is to predict the user group to which a user belongs based on their attribute values to recommend videos to that user, increasing user stickiness and dwell time. The regression task can be to predict the user's daily online time. In real-world scenarios, insufficient historical viewing time for new users or their reluctance to interact may lead to missing data in relevant dimensions.
[0054] Each of the relevant techniques for handling missing tabular data has its limitations. For the two-stage "imputation followed by prediction" approach, the imputation stage aims to minimize reconstruction error, while the downstream prediction stage aims to minimize prediction error. These two independent optimization objectives prevent the imputation process from adaptively adjusting to the final prediction goal. Furthermore, excessively pursuing point-by-point approximation of missing values may blur or even eliminate discriminative structural features crucial for classification or regression decision boundaries, often resulting in poor performance in downstream prediction tasks.
[0055] In real-world scenarios, data loss is often not completely random but follows a systematic pattern. In such cases, the event of a "missing value" itself carries additional information about the sample. However, in related technologies, tree models use hard-coded default branching strategies, and neural network models use the same special labels for all missing locations. Neither approach effectively models the fine-grained locational information inherent in the missing patterns, thus failing to fully utilize the structural signals within the missing patterns to enhance predictive capabilities.
[0056] To overcome the above problems, this specification proposes a method for training a prediction model using missing data. This method combines the reconstruction of missing data with the optimization objective of the prediction part to perform end-to-end training of the prediction model, and enables the model to learn structural signals in the missing patterns.
[0057] Figure 1 This diagram illustrates an architecture for training a prediction model using missing data, according to one embodiment. Figure 1In the example, user attributes include five attribute dimensions, specifically numerical and categorical dimensions. Dimensions 1 and 5 are numerical dimensions, while dimensions 2, 3, and 4 are categorical dimensions. The attribute dimensions are sorted in a preset order (i.e., dimensions 1 to 5), and the specific attribute values of each user sample are arranged in the corresponding order to form sequence data.
[0058] Figure 1 The user sample shown as an example has missing values in dimension 3, but has data in the other four dimensions, which are respectively... , , , Dimension 3 represents the missing dimension of this user sample, while the other dimensions represent the visible dimensions. At this point, the original sequence data of this user sample is... , where NA represents missing data.
[0059] First, embedding processing is performed on each attribute dimension of the original sequence data, uniformly mapping them to embedded representations in a high-dimensional vector space. For example... Figure 1 As shown in the feature layer, the specific method of embedding is as follows: for numerical dimensions (such as dimension 1), its scalar value is... Input the independent linear layers corresponding to this dimension, and convert them into fixed-dimensional embedding vectors through linear mapping. For categorical dimensions (such as dimension 2), the category identifier is input into the independent embedding table corresponding to that dimension, and converted into a fixed-dimensional embedding vector through a table lookup operation. For missing attribute dimensions (such as dimension 3), a placeholder is inserted (e.g., ...). Figure 1 In After embedding, each attribute dimension is converted into a sequence of embedding vectors of a uniform dimension. Positional encoding can also be added to the embedding vectors at each position after embedding to preserve the order information of each attribute dimension in the sequence. The final embedded sequence is obtained.
[0060] After the embedding process is completed, the random masking stage begins. For example... Figure 1 As shown in the masking visible dimension step, a masking process is applied to the embedding vector of the visible dimension. Specifically, a portion of the visible dimension is randomly selected with a predetermined probability, and its embedding vector is marked with a mask (e.g., ...). Figure 1[R]). For example, in this specific sample, dimensions 4 and 5 are randomly selected for masking, and their embedding vectors are replaced with [M]. The dimensions that are randomly masked are called mask dimensions, and their true attribute values are temporarily hidden, but can still be used to calculate the subsequent reconstruction loss because their true attribute values are known in the training data. Simultaneously, a predefined special term [CLS] is added to the beginning of the sequence. This special term is a learnable embedding vector used to aggregate global sequence information during subsequent encoding and decoding. The [CLS] term, which aggregates global sequence information, can be used later for prediction tasks, including classification and regression tasks. The sequence data obtained after masking is... This includes special word elements. The network consists of three dimensions: visible dimensions (dimension 1 and dimension 2), missing dimensions (dimension 3), and masked dimensions (dimension 4 and dimension 5). The missing and masked dimensions can be collectively referred to as invisible dimensions, which are not encoded in subsequent coding processes.
[0061] Subsequently, the sequence data is input into the encoding network for encoding processing. For example... Figure 1 As shown in the lower left section of the encoding network, the encoding network consists of multiple ( The system consists of stacked coding layers with identical structures. Each coding layer can be built based on the Transformer encoder architecture and only applies to the visible dimensions of the input sequence data. Figure 1 The input sequence data is encoded using dimensions 1 and 2, while the remaining invisible dimensions are ignored. After being processed layer by layer through multiple encoding layers, the output is the encoded representation of each visible dimension position, i.e., the dimension 1 corresponding to... Corresponding to dimension 2 Dimensions 3 through 5, which are invisible dimensions, will have shared, learnable masked lexical representations added to them (corresponding to...). Figure 1 [M]), the masked word representation has the same dimension as the encoded representation at the visible dimension position, and is updated synchronously with each network layer of the model during training. The encoded representation sequence after passing through the encoding network is: .
[0062] Subsequently, the encoded representation sequence output by the encoding network is fed into the decoding network for decoding. For example... Figure 1 As shown in the middle section of the decoding network, the decoding network consists of multiple ( The decoding network consists of stacked decoding layers with identical structures. Each decoding layer can be built based on the Transformer decoder architecture and includes self-attention sub-layers and feedforward neural network sub-layers. Unlike the encoding network, which only processes the visible dimension, the decoding network processes all attribute dimensions and special terms, and uses global self-attention processing, allowing free attention interactions between different positions in the sequence. The decoded representation sequence after the decoding network is... ,exist Figure 1 The dashed box is used to represent this.
[0063] After obtaining the decoded representation sequence, prediction and reconstruction tasks can be performed simultaneously based on the result to achieve joint optimization of prediction and reconstruction. For the prediction task, the decoded representations at the positions of special lexical units [CLS] in the decoded representation sequence are extracted and fed into the prediction network for prediction, obtaining the prediction result for the user. When the prediction task is a classification task, the prediction network can be a classification network, consisting of one or more fully connected layers, and finally outputs the prediction results for each user category through the Softmax function. When the prediction task is a regression task, the prediction network can be a regression network, containing one or more linear layers, and finally outputs the predicted values for the user's attribute dimensions. Thus, the prediction loss can be calculated based on the prediction results and the user's labels.
[0064] For the data reconstruction task, extract the decoding representation of each mask dimension position, i.e., the position corresponding to position 4. Corresponding to position 5 The data is fed into the corresponding reconstruction network to reconstruct the attribute values, obtaining the reconstructed user attribute values for each mask dimension. The reconstruction network can be equipped with an independent linear layer or a small multilayer perceptron for each attribute dimension. The reconstruction loss can be calculated based on the difference between the reconstructed user attribute values and the corresponding true attribute values for each mask dimension. It should be noted that the reconstruction loss is calculated only on the mask dimensions, not on the missing dimensions, because there are no known true attribute values as supervision signals for the missing dimensions.
[0065] The prediction loss and reconstruction loss together constitute the total training loss, which is jointly updated through backpropagation of the parameters of the encoding network, decoding network, prediction network, reconstruction network, embedding layer, and learnable masked lexical representations. In this joint optimization framework, reducing the prediction loss directly optimizes the model's prediction performance, while reducing the reconstruction loss drives the model to learn the intrinsic distribution and correlations between various dimensions of the data. The gradient from the prediction loss is passed to the decoding and encoding networks through backpropagation, causing the reconstruction process to dynamically adjust in a direction favorable to prediction, rather than simply pursuing numerically accurate reconstruction.
[0066] Figure 1The architecture shown unifies the prediction and reconstruction tasks within an end-to-end joint training framework. By applying an additional random mask to each training sample, the model is exposed to various missing value patterns during training, enhancing its adaptability to different missing value rates and mechanisms. Both prediction and attribute value reconstruction are based on the output representation of the decoding network, allowing the model to integrate visible information and inferences about missing information during prediction, forming a more complete global representation. This results in superior prediction performance in scenarios where tabular data contains missing values.
[0067] The above describes the overall architecture of the prediction model trained in this manual. The specific architecture of the encoding layer is described below.
[0068] Figure 2 A schematic diagram of the coding layer architecture according to one embodiment is shown. Figure 2 As shown, a single coding layer in the coding network consists of a normalization layer, an attention layer, and a fully connected layer. The input to the coding layer is the output sequence of the previous coding layer (for the first encoder layer, the input is the embedded representation sequence after embedding and position encoding).
[0069] The input sequence first undergoes a layer normalization process, which normalizes the representation vectors at each position in the sequence to stabilize the distribution characteristics of the input at each layer and accelerate training convergence. The normalized sequence is then fed into the attention layer. In this layer, the sequence undergoes linear projection to obtain the query matrix, key matrix, and value matrix. Subsequently, attention scores between positions are calculated using an asymmetric attention mechanism, resulting in an attention-weighted output sequence. The core of the asymmetric attention mechanism lies in applying a customized masking rule to the attention score matrix, giving different types of positions different attention interaction permissions. This effectively isolates the interference of invisible dimensions on feature extraction of visible dimensions during the encoding process. This customized masking rule will be described in detail later in this specification.
[0070] The output of the attention layer is residually concatenated with the input of the coding layer to obtain an intermediate representation sequence. This intermediate representation sequence is then passed through a second normalization layer for layer normalization before being fed into a fully connected layer. The output of the fully connected layer is also residually concatenated with the intermediate representation sequence to obtain the final output sequence of the coding layer.
[0071] The encoding layer progressively extracts and integrates contextual information between various attribute dimensions through the stacking of normalization, attention computation, and feedforward transformation. The stacking of multiple encoding layers enables the model to capture deeper inter-dimensional correlation patterns, improving the quality of the final encoded representation.
[0072] The following describes the specific implementation steps of the method for training a prediction model using missing data, with reference to specific embodiments.
[0073] Figure 3 A flowchart illustrating a method for training a predictive model using missing data according to one embodiment is provided. The method can be executed by any platform, server, or device cluster with computing and processing capabilities. Figure 3 The prediction model described includes an encoding network, a decoding network, a prediction network, and several reconstruction networks. The method includes at least the following steps: Step S302, acquiring sequence data corresponding to the target user, including predetermined special lexical units, attribute values of the target user in several visible dimensions, and padding markers for multiple invisible dimensions; Step S304, inputting the sequence data into the encoding network to obtain encoded representations of each position in the sequence, wherein each position in an invisible dimension shares the same masked lexical representation; Step S306, inputting each encoded representation into the decoding network to obtain decoded representations of each position in the sequence; Step S308, inputting the decoded representations of the special lexical positions into the prediction network to obtain prediction results for the target user; and inputting the decoded representation of any masked dimension position into the corresponding reconstruction network to obtain the corresponding reconstructed user attribute value; Step S310, determining the total training loss based on the prediction results and the reconstructed user attribute values, and then adjusting the trainable parameters of the prediction model.
[0074] The specific execution process of each of the above steps is described below.
[0075] First, in step S302, sequence data corresponding to the target user is obtained, including special words, attribute values of the target user in several visible dimensions, and padding markers for multiple invisible dimensions; the multiple invisible dimensions include masking dimensions where the attribute values of the target user are masked and missing dimensions where attribute values are missing.
[0076] The target user can be any one of multiple user samples in the training sample set. The sequence data corresponding to the target user is a structured data sequence in which the various attribute dimensions of the target user are organized in a predetermined order.
[0077] In different embodiments, the order of attribute dimensions in the sequence data can employ various organizational strategies. One approach is to group the attribute dimensions according to their type, placing all numerical dimensions at the beginning of the sequence and all categorical dimensions at the end. This approach facilitates closer feature interactions between dimensions of the same type during encoding. Another approach is to sort the attribute dimensions according to their correlation, for example, by arranging highly correlated dimensions adjacently based on the correlation coefficients of attribute values in the training set. This allows the self-attention mechanism of the encoding network to more efficiently capture the correlation patterns between dimensions. Yet another approach is to use a predefined fixed arrangement order, which can be set through expert experience during model initialization or automatically determined after statistical analysis of the training data.
[0078] In this sequence data, the visible dimension refers to the dimension in which the target user has the actual attribute value that has been collected. The invisible dimension refers to the dimension in the sequence data that is not visible to the prediction model. It further includes two categories: one is the masked dimension, which means that the target user originally has the actual attribute value in this attribute dimension, but it is artificially masked during the training process, making its attribute value temporarily invisible; the other is the missing dimension, which means that the target user has not collected the attribute value in this attribute dimension, which is naturally missing.
[0079] By unifying the mask and missing dimensions into the sequence data and identifying them with padding markers, the model can simultaneously encounter invisible patterns generated by artificial masks and those generated by natural missing values during the training phase, thereby learning a more robust missing value handling strategy. The beginning of the sequence data contains a special term [CLS], which is a special marker used to aggregate global sequence information during subsequent encoding and decoding processes, providing a centralized feature representation for the prediction task.
[0080] In one embodiment, the sequence data is obtained by manually adding masks and special terms to the original sequence data. In this embodiment, step S302, obtaining the sequence data corresponding to the target user, may include steps 11 to 13.
[0081] In step 11, the raw sequence data corresponding to the target user is obtained, including the target user's attribute values in multiple visible dimensions, as well as missing dimensions where several attribute values are missing. The raw sequence data reflects the actual collection status of the target user's attribute values.
[0082] Subsequently, in step 12, the attribute values of at least one of the multiple visible dimensions are masked to obtain several masked dimensions. The masking process can involve randomly selecting a subset of visible dimensions with a predetermined probability and replacing their attribute values with masked states; the masked dimensions are then called the masked dimensions. The predetermined probability can be flexibly set according to training needs, for example, it can be set to 15%, 20%, or 30%.
[0083] Then, in step 13, padding markers are added to each mask dimension and missing dimension, and the special term is added at the beginning of the sequence to obtain the sequence data. The padding markers are used to identify the position as an invisible dimension, enabling the model to distinguish between visible and invisible positions.
[0084] In one different implementation, masking can employ a more refined strategy. Besides independently and randomly masking each visible dimension, a block masking strategy can be used. This involves dividing several related dimensions into dimension blocks according to a predetermined grouping rule, and applying masking to the entire dimension block. Each masking operation simultaneously hides a set of semantically related dimensions, forcing the model to learn to infer a set of related attributes from information in other dimensions, thus enhancing the model's ability to model deep relationships between dimensions. Furthermore, masking can complement missing dimensions. When a sample has many missing dimensions, the probability of masking can be reduced to ensure sufficient visible information during training; conversely, when there are few missing dimensions, the masking probability can be appropriately increased to increase training difficulty.
[0085] Step S302 organizes the target user's attribute dimensions into a unified sequence data structure, enabling different types of attribute dimensions (numerical and categorical) and dimensions with different visibility states (visible dimensions, masked dimensions, and missing dimensions) to be modeled within the same processing framework. The introduction of masked dimensions allows the model to encounter various missing patterns during training, enhancing its adaptability to different missing rates and mechanisms. By flexibly adjusting mask probabilities or employing block masking strategies, the model's ability to model inter-dimensional relationships and perceive missing cause information can be further improved, thus achieving a more robust representation learning foundation in scenarios where tabular data contains missing values.
[0086] Then, in step S304, the sequence data is input into the encoding network to obtain the encoding representation of each position in the sequence, wherein each invisible dimension position shares the same mask lexical representation.
[0087] Encoding networks are used to extract features and model the context of input sequence data, outputting encoded representations of each position in the sequence. An encoding network can contain multiple encoding layers, each built upon the encoding layer of a Transformer. During encoding, the input representations for visible dimensions in the sequence data come from the embedding vectors of the attribute values for that dimension, while invisible dimensions (including masked dimensions and missing dimensions) are not visible to the encoding network, thus reducing the data processing complexity of the encoding network.
[0088] In one embodiment, the encoding network is based on an attention mechanism, and the encoding process in step S304 may include: inputting sequence data into the encoding network for target self-attention processing to obtain encoded representations of each position in the sequence; wherein, the target self-attention processing only calculates attention scores between specific words and each visible dimension. In the target self-attention processing, the effective calculation range of attention scores is limited to between specific words and each visible dimension, so that the encoding process focuses on extracting contextual features from visible information. Specific words can aggregate information from all visible dimensions to form a global aggregated representation.
[0089] Different implementations can be used to achieve the aforementioned self-attention processing. One implementation removes the padding markers from the sequence data and inputs a short sequence containing only visible dimension data and special terms into the encoding network for self-attention processing. In other words, before entering the encoding network, all positions marked as invisible dimensions in the sequence data are directly removed, and the encoding network only receives special terms and visible dimension data as input. After the encoding network outputs the encoded representation, the masked term representations are then backfilled into the corresponding positions of each invisible dimension to restore the complete sequence length for subsequent decoding network processing.
[0090] In another implementation, the sequence data can be input into the encoding network, and attention masks can be added to each padding marker in the sequence data before self-attention processing. In this approach, the full length of the sequence data remains unchanged, but attention masks are applied to invisible dimensions during the self-attention computation. The specific rules for constructing the attention masks can be as follows: Figure 4 As shown.
[0091] Figure 4This diagram illustrates asymmetric self-attention masks. In the mask matrix M, rows represent query positions, columns represent key positions, and shaded areas represent positions where attention is masked. The specific attention rules are as follows: Special terms [CLS] can focus on themselves and all visible and masked dimensions (i.e., the reconstruction target), but cannot focus on naturally missing dimensions; visible dimensions (including numerical and categorical dimensions) can only focus on special terms [CLS] and other visible dimensions, and cannot focus on any masked positions (including masked and missing dimensions); query rows at all masked positions (including masked and missing dimensions) are completely masked, meaning these positions do not actively initiate any attention computation during the encoding phase and are in a passive state. Through these asymmetric attention rules, the encoding process can focus on extracting high-quality contextual features from visible information while avoiding interference from noise in invisible dimensions.
[0092] Attention masks can be implemented by adding negative infinity values to the rows and columns corresponding to the invisible dimensions of the attention score matrix. This makes the attention weights at these positions approach zero after Softmax normalization, thus allowing the model to essentially ignore the attention contribution at these positions. The advantage of this approach is that it maintains the structural consistency of the input sequence, eliminates the need for dynamic adjustment of the sequence length, and facilitates batch computation of rules on hardware accelerators.
[0093] In one different embodiment, the encoding network can employ a multi-head self-attention mechanism, which decomposes the attention mechanism into multiple parallel attention heads. Each head computes the attention interactions between queries, keys, and values in an independent linear projection space. Finally, the outputs of all heads are concatenated and subjected to a linear transformation to obtain the final attention output. The multi-head attention mechanism enables the model to capture diverse correlation patterns between dimensions from different representation subspaces. In multi-head attention computation, the attention mask matrix is broadcast to all attention heads, ensuring that each attention head follows the constraint rules of target self-attention, i.e., only calculating the attention scores between specific terms and each visible dimension. The number of layers in the encoding network can be flexibly set according to the number of attribute dimensions and the complexity of the data, for example, it can be set to 3, 6, or 12 layers. The more layers, the stronger the encoding network's ability to extract contextual features, but the computational cost also increases accordingly. The hidden layer dimensions of each encoder layer... It can be set to values such as 64, 128, 256, or 512, and this dimension is the same as the dimension of the embedding vector. They can be the same or different.
[0094] In other embodiments, the complete sequence data can also be directly input into the encoding network for conventional self-attention processing, which will not be elaborated here.
[0095] Before inputting sequence data into the encoding network, embedding processing can be performed on different types of attribute dimensions to uniformly map mixed-type tabular data into representations in a high-dimensional vector space. Visible dimensions include numerical and categorical dimensions. This embedding process includes: inputting data of any numerical dimension into the corresponding linear layer, and inputting data of any categorical dimension into the corresponding embedding table.
[0096] Specifically, for numerical dimensions, each dimension is equipped with an independent linear transformation layer, which converts the scalar values in that dimension into fixed-dimensional embedding vectors through linear mapping. For example, if the dimension of the embedding vector is set to... Then the linear layer maps the one-dimensional input value to... 3D vectors. For categorical dimensions, each dimension is equipped with an independent embedding table. The embedding table stores the embedding vectors corresponding to all possible categories for that dimension. Category identifiers are mapped to vectors through a lookup operation. Dimensional embedding vector.
[0097] Through the embedding process described above, both numerical and categorical dimensions are converted into embedding vectors of a uniform dimension, allowing for consistent computation in subsequent processing. After embedding, positional encoding can be added to the embedding vectors at each position to preserve the positional information of each attribute dimension within the sequence. Positional encoding can employ learnable positional embeddings or fixed sinusoidal positional encoding. The embedded representations obtained after embedding and positional encoding are then input into the encoding network for subsequent feature extraction.
[0098] In step S304, the encoding network extracts features from the sequence data and models the context, transforming the original values of each attribute dimension into high-dimensional representations containing rich contextual information. The target self-attention processing limits the attention computation to specific words and visible dimensions, avoiding noise interference from invisible dimensions, allowing the encoding process to focus on extracting high-quality contextual features from visible information. The removal of padding markers reduces the length of the input sequence to the encoding network, significantly lowering the computational complexity of the self-attention mechanism; while the addition of attention masks maintains the structural consistency of the input sequence, facilitating batch computation of rules on hardware accelerators.
[0099] Next, in step S306, each encoded representation is input into the decoding network to obtain the decoded representation of each position in the sequence.
[0100] The decoding network receives the encoded representations at each position from the output of the encoding network, decodes them, and outputs the decoded representations at each position of the sequence. The input to the decoding network is a full-length sequence composed of two parts: one part consists of the encoded representations of specific terms output by the encoding network and the encoded representations at each visible dimension position; the other part consists of shared learnable mask terms inserted at all invisible dimension positions (including mask dimensions and missing dimensions). Through a self-attention mechanism, the decoding network allows each mask term position to update its representation using contextual information from the visible dimensions and specific terms, thereby generating representations for each invisible dimension that incorporate global semantic information in the decoded output.
[0101] In one different embodiment, the decoding network can employ a different attention strategy than the encoding network when performing self-attention processing. The encoding network uses asymmetric target self-attention, limiting attention computation to specific terms and visible dimensions; while the decoding network can use global self-attention, allowing free attention interactions between all positions in the sequence (including visible, masked, and missing dimensions). This differentiated design stems from the fact that the encoding stage needs to extract high-quality contextual features from visible information to avoid noise interference from invisible positions; while the decoding stage's task is to utilize the contextual information provided by the encoding network to generate meaningful representations for invisible positions. In this case, the interactions between invisible positions help the model infer the joint distribution pattern between dimensions, thereby generating a more consistent reconstruction result.
[0102] Then, in step S308, the decoded representation of the special word position is input into the prediction network to obtain the prediction result for the target user; and the decoded representation of any mask dimension position is input into the corresponding reconstruction network to obtain the corresponding reconstructed user attribute value.
[0103] In step S308, the output of the decoding network is divided into two paths. The first path is geared towards the prediction task, feeding the decoded representation of the special word positions after processing by the decoding network into the prediction network and outputting the prediction result. After being processed layer by layer by the encoding and decoding networks, the decoded representation of the special word has integrated the attribute information of the visible dimension and the model's inference information of the invisible dimension, and can be used as a global compact representation of the target user for prediction.
[0104] The type of prediction result depends on the specific prediction task. When the prediction task is a classification task, the prediction network is a classification network, and the prediction result is the predicted probability distribution of the target user belonging to each user category. When the prediction task is a regression task, the prediction network is a regression network, and the prediction result is the predicted value of the target user on the target attribute dimension. The target attribute dimension is the attribute dimension to be predicted in the regression task, and it is different from the aforementioned visible dimensions, masked dimensions, and missing dimensions.
[0105] Classification networks can consist of one or more fully connected layers, ultimately outputting the predicted probability distribution for each preset user category through a Softmax function. Regression networks, on the other hand, can consist of one or more fully connected layers and output predicted values for the target attribute dimension.
[0106] The second path output by the decoding network is geared towards the reconstruction task. It inputs the decoded representation at each mask dimension position into the corresponding reconstruction network to obtain the reconstructed user attribute value for that dimension. The reconstruction network and the prediction network are independent of each other; each mask dimension can correspond to an independent reconstruction network. For numerical dimensions, the reconstruction network can use linear layers to output continuous numerical values for reconstruction; for categorical dimensions, the reconstruction network can use linear layers combined with a softmax function to output the probability distribution of each category. In the reconstruction task, by forcing the model to recover the masked attribute values from the visible information, the encoding and decoding networks are driven to learn the intrinsic relationships between the dimensions of the data, thereby improving the quality of the representation. It should be noted that the reconstruction loss is calculated only on the mask dimension, not on the missing dimension, because there are no known true attribute values for the missing dimension as supervision signals for training.
[0107] In step S308, prediction and reconstruction form two independent output paths, enabling the prediction task to make decisions using the global representation output by the decoding network, while the reconstruction task drives the model to learn the intrinsic relationships between the various dimensions of the data.
[0108] Finally, in step S310, the total training loss is determined based on the prediction results and the reconstructed user attribute values, and then the trainable parameters of the prediction model are adjusted.
[0109] The total training loss consists of two parts: prediction loss and reconstruction loss. Prediction loss is determined based on the prediction results and the target user's labels, reflecting the model's accuracy in the prediction task. Reconstruction loss is determined based on the differences between the reconstructed user attribute values in each mask dimension and the corresponding real user attribute values (i.e., the attribute values before initial masking), reflecting the model's ability to reconstruct attributes in a self-supervised task. Prediction loss and reconstruction loss together constitute the optimization objective of the training process, guiding the model to achieve a balance between prediction performance and data understanding ability.
[0110] Specifically, predicting losses The calculation method depends on the type of prediction task. In one embodiment, the prediction task is a classification task, and the labels used are the label categories to which the target user belongs. In this embodiment, the prediction loss can be determined using the cross-entropy loss function, as shown in formula (1):
[0111]
[0112] in, Total number of user categories For the tag category in the Indicator values for each category (when the label category is number 1) Class Time Select 1 otherwise select 0. The corresponding output of the classification network The predicted probability of each category.
[0113] In another embodiment, the prediction task is a regression task, and the labels used are the true values of the target user on the target attribute dimension. In this embodiment, the prediction loss can be determined using mean squared error, as shown in formula (2):
[0114]
[0115] in, The predicted value output by the regression network, and This represents the corresponding true value.
[0116] The reconstruction loss is calculated only on the mask dimension set. The mask dimensions can further include numerical and categorical dimensions, each with different loss calculation methods. For numerical dimensions, the mean squared error is used as the reconstruction loss, denoted as the first reconstruction loss. As shown in formula (3):
[0117]
[0118] in This is the set of indices for the numerical dimensions within the mask dimension. For the first Reconstructed values for each numerical dimension, For the first The true attribute values of each numerical dimension.
[0119] For categorical dimensions, cross-entropy is used as the reconstruction loss, denoted as the second reconstruction loss. As shown in formula (4):
[0120]
[0121] in This is the set of indices for the categorical dimensions within the mask dimension. For the first Reconstruction probability distribution of each categorical dimension For the first The one-hot encoded vector corresponding to the true category in each categorical dimension. The cross-entropy represents the two, and the specific calculation method can be found in formula (1).
[0122] The total reconstruction loss is the sum of the first reconstruction loss and the second reconstruction loss: .
[0123] After obtaining the prediction loss and reconstruction loss, the total training loss is the combination of the prediction loss and reconstruction loss, for example, by directly summing them, i.e. In another implementation, a weighted coefficient can also be introduced into the total training loss. Adjusting the reconstruction loss, i.e. ,in The hyperparameter is greater than zero and is used to balance the contribution of the prediction task and the reconstruction task to the model parameter update. In this joint optimization framework, the prediction loss is the primary objective, directly optimizing the model's prediction performance; the reconstruction loss is the secondary objective, driving the model to learn the intrinsic distribution and correlation between the various dimensions of the data. The gradient from the prediction loss is propagated to the decoding and encoding networks through backpropagation, thereby affecting the parameter updates of each module during the reconstruction process. This allows the reconstruction process to dynamically adjust in a direction favorable to prediction, rather than simply pursuing numerically accurate restoration.
[0124] After determining the total training loss according to step S310, the trainable parameters of the prediction model can be updated using the backpropagation algorithm. Trainable parameters include the weights and biases of each layer in the encoding network, the weights and biases of each layer in the decoding network, the weights and biases of the prediction network, the weights and biases of each reconstruction network, and the parameters of each linear layer and embedding table in the embedding process. Furthermore, the masked term representation itself is also a trainable parameter, updated synchronously with the model parameters during training. As training progresses, this shared masked term gradually evolves into a representation that effectively activates the reconstruction capabilities of the decoding network, helping the model generate more reasonable inference results in unseen dimensions. During training, optimizers such as Adam or AdamW can be used, along with a learning rate decay strategy, to iteratively update the model parameters until the model converges or reaches the preset training epochs.
[0125] According to another embodiment, an apparatus for training a prediction model using missing data is also provided. Figure 5A schematic block diagram of an apparatus for training a predictive model using missing data according to one embodiment is shown. This apparatus can be deployed in any device, platform, or cluster of devices with computing and processing capabilities. Figure 5 As shown, the prediction model includes an encoding network, a decoding network, a prediction network, and several reconstruction networks, and the device 500 includes:
[0126] The acquisition unit 502 is configured to acquire sequence data corresponding to the target user, including predetermined special words, attribute values of the target user in several visible dimensions, and padding marks for multiple invisible dimensions; the multiple invisible dimensions include masking dimensions where the attribute values of the target user are masked and missing dimensions where attribute values are missing.
[0127] The encoding unit 504 is configured to input the sequence data into the encoding network to obtain the encoding representation of each position in the sequence, wherein each invisible dimension position shares the same mask word representation;
[0128] Decoding unit 506 is configured to input each encoded representation into the decoding network to obtain the decoded representation of each position in the sequence;
[0129] The prediction and reconstruction unit 508 is configured to input the decoded representation of the special word position into the prediction network to obtain the prediction result for the target user; and to input the decoded representation of any mask dimension position into the corresponding reconstruction network to obtain the corresponding reconstructed user attribute value.
[0130] Training unit 510 is configured to determine the total training loss based on the prediction results and reconstructed user attribute values, and then adjust the trainable parameters of the prediction model.
[0131] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform the methods described in any of the above embodiments.
[0132] According to another embodiment, a computing device is also provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in any of the above embodiments.
[0133] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0134] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0135] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0136] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0137] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for training a prediction model using missing data, the prediction model comprising an encoding network, a decoding network, a prediction network, and several reconstruction networks, the method comprising: Obtain sequence data corresponding to the target user, including predetermined special words, attribute values of the target user in several visible dimensions, and padding markers for multiple invisible dimensions; the multiple invisible dimensions include masking dimensions where the attribute values of the target user are masked and missing dimensions where attribute values are missing. The sequence data is input into the encoding network to obtain the encoded representation of each position in the sequence, wherein each invisible dimension position shares the same masked word representation; Each encoded representation is input into the decoding network to obtain the decoded representation of each position in the sequence; The decoded representation of the specific word position is input into the prediction network to obtain the prediction result for the target user; Furthermore, the decoded representation at any mask dimension position is input into the corresponding reconstruction network to obtain the corresponding reconstructed user attribute value; Based on the prediction results and reconstructed user attribute values, the total training loss is determined, and then the trainable parameters of the prediction model are adjusted.
2. The method according to claim 1, wherein, The step of obtaining the sequence data corresponding to the target user includes: Obtain the raw sequence data corresponding to the target user, including the attribute values of the target user in multiple visible dimensions, as well as missing dimensions where several attribute values are missing; The attribute values of at least one of the multiple visible dimensions are masked to obtain several mask dimensions; Add padding markers to each mask dimension and missing dimension, and add the special term at the beginning of the sequence to obtain the sequence data.
3. The method according to claim 1, wherein, The encoding network is based on an attention mechanism; The step of inputting the sequence data into the encoding network to obtain the encoded representation of each position in the sequence includes: The sequence data is input into an encoding network for target self-attention processing to obtain the encoded representation of each position in the sequence; wherein, the target self-attention processing only calculates the attention scores between special words and each visible dimension.
4. The method according to claim 3, wherein, The step of inputting the sequence data into the encoding network for target self-attention processing includes: After removing the data from each padding marker in the sequence data, it is input into the encoding network for self-attention processing.
5. The method according to claim 3, wherein, The step of inputting the sequence data into the encoding network for target self-attention processing includes: The sequence data is input into the encoding network, and attention masks are added to each padding marker in the sequence data before self-attention processing is performed.
6. The method according to claim 1, wherein, The visible dimensions include numerical dimensions and categorical dimensions; The step of inputting the sequence data into the encoding network includes: The sequence data is embedded to obtain an embedded representation; the embedding process includes: inputting data of any numerical dimension into the corresponding linear layer, and inputting data of any categorical dimension into the corresponding embedding table; The embedded representation is input into the encoding network.
7. The method according to claim 1, wherein, The step of determining the total training loss based on the prediction results and reconstructed user attribute values includes: Based on the prediction results and the target user's tags, determine the prediction loss; The reconstruction loss is determined based on the difference between the reconstructed user attribute values and the corresponding real user attribute values for each mask dimension. The total training loss is determined based on the predicted loss and the reconstruction loss.
8. The method according to claim 7, wherein, The prediction network is a classification network, the prediction result is the predicted probability distribution of the target user belonging to each user category, and the label is a label category; determining the prediction loss based on the prediction result and the target user's label includes: The prediction loss is determined based on the cross-entropy between the predicted probability distribution and the label category.
9. The method according to claim 7, wherein, The prediction network is a regression network, the prediction result is the predicted value of the target user on the target attribute dimension, and the label is the true value of the target user on the target attribute dimension; determining the prediction loss based on the prediction result and the label of the target user includes: The prediction loss is determined based on the mean square error between the predicted and actual values.
10. The method according to claim 7, wherein, The determination of reconstruction loss based on the difference between the reconstructed user attribute values and the corresponding real user attribute values for each mask dimension includes: The first reconstruction loss is determined based on the mean square error between the reconstructed user attribute values and their actual user attribute values for each numerical dimension. The second reconstruction loss is determined based on the cross-entropy between the reconstructed user attribute values and their actual user attribute values for each category dimension. The reconstruction loss is determined based on the first reconstruction loss and the second reconstruction loss.
11. The method according to claim 1, wherein, The masked lexical representation is a trainable representation; adjusting the trainable parameters of the prediction model includes: Adjust the trainable parameters of the prediction model and the masked lexical representation.
12. An apparatus for training a prediction model using missing data, the prediction model comprising an encoding network, a decoding network, a prediction network, and several reconstruction networks, the apparatus comprising: The acquisition unit is configured to acquire sequence data corresponding to the target user, including predetermined special words, attribute values of the target user in several visible dimensions, and padding markers for multiple invisible dimensions; the multiple invisible dimensions include masking dimensions where the attribute values of the target user are masked and missing dimensions where attribute values are missing. The encoding unit is configured to input the sequence data into the encoding network to obtain the encoded representation of each position in the sequence, wherein each invisible dimension position shares the same masked word representation; The decoding unit is configured to input each encoded representation into the decoding network to obtain the decoded representation of each position in the sequence; The prediction and reconstruction unit is configured to input the decoded representation of the special word position into the prediction network to obtain the prediction result for the target user; Furthermore, the decoded representation at any mask dimension position is input into the corresponding reconstruction network to obtain the corresponding reconstructed user attribute value; The training unit is configured to determine the total training loss based on the prediction results and reconstructed user attribute values, and then adjust the trainable parameters of the prediction model.
13. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-11.
14. A computing device comprising a memory and a processor, wherein, The memory stores executable code, and when the processor executes the executable code, it implements the method of any one of claims 1-11.