Error Repair Method, Device, and Electronic Equipment Based on Differentiable Bilayer Optimization

CN122673473APending Publication Date: 2026-09-01ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610741813.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

该类方法虽然能够在一定程度上恢复数据表面的正确性,但其修复目标通常侧重于数据保真,而未充分考虑修复结果对下游任务性能的实际影响

Benefits of technology

由以上技术方案可知,本申请通过将表格错误修复过程构造成面向下游任务的可微双层优化过程,使修复模型的更新不仅受数据自身分布约束,而且直接受验证集上的任务损失引导,从而能够避免仅追求数据表面合理性而忽视下游任务性能的问题,提高修复结果对分类或回归任务的适配能力;通过对数值属性和类别属性分别编码,并在统一的连续嵌入空间中进行修复,能够有效处理混合类型表格数据,增强不同属性之间相关关系的建模能力,为复杂错误模式下的表格修复提供更加丰富和灵活的表示基础;在外层优化中采用超梯度近似方式更新修复模型参数,无需针对大量候选修复方案反复完整训练下游任务模型,降低了任务导向错误修复的计算代价,提高了方法的可实施性和效率;能够在仅已知错误位置或部分错误信息的条件下,对待修复单元格进行上下文感知修复,无需人工逐条制定修复规则或穷举有限候选修复策略,具有较强的自动化能力和实际应用价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122673473A_ABST
    Figure CN122673473A_ABST
Patent Text Reader

Abstract

This invention discloses an error repair method based on differentiable bilayer optimization: It acquires erroneous table data, error masks, and task labels; encodes numerical and category attributes as continuous embeddings, dividing them into repair targets and conditions according to the mask; adds noise to the target portion, and uses a conditional denoising model to progressively predict the noise; and decodes the data after repair. The inner layer optimization uses the repaired data to update the downstream task model, while the outer layer optimization combines the validation set task loss and diffusion self-supervised loss to update the repair model, forming a differentiable bilayer optimization. The trained model can directly repair tables. This method eliminates the need for manual rules or repeated training of the task model, can handle mixed-type tables, aligns repair with downstream tasks, and improves repair quality and task performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing and data quality management technology, specifically to an error repair method and apparatus, and electronic equipment based on differentiable bilayer optimization. Background Technology

[0002] Tabular data is one of the most common data organization formats in government, finance, healthcare, retail, and manufacturing, and is widely used for downstream analysis tasks such as classification, regression, risk assessment, and trend prediction. In practical applications, due to factors such as human input errors, system data collection anomalies, data migration and transformation deviations, and inconsistencies in the fusion of different data sources, tabular data often contains errors such as missing values, outliers, incorrect category entries, format conflicts, semantic inconsistencies, and cross-attribute logical violations. These errors disrupt data distribution and the relationships between attributes, thereby affecting the training effect and decision-making results of downstream models. Therefore, how to effectively correct errors in tabular data has become a key issue in data quality governance.

[0003] Most existing table error correction methods primarily optimize for the consistency, completeness, or statistical reasonableness of the data itself. They typically rely on predefined rules, constraints, similar sample searches, probabilistic inference, or single-step prediction-based correction strategies to generate correction results. While these methods can restore the apparent correctness of the data to a certain extent, their correction objectives usually focus on data fidelity without fully considering the actual impact of the correction results on the performance of downstream tasks. That is, some correction values ​​that seem statistically reasonable may not improve the performance of downstream classification or regression tasks, and may even weaken the model's generalization ability.

[0004] On the other hand, some task-oriented data cleaning methods attempt to adjust the data by combining feedback from downstream models, but existing solutions usually still have the following problems: First, the repair candidate space is limited, often relying on manually set repair rules, discrete operation sets, or fixed candidate values, which is difficult to adapt to complex, diverse, and continuously changing real error patterns; Second, when dealing with mixed-type tabular data that contains both numerical and categorical attributes, it is difficult to uniformly model the representation of different attributes and the higher-order correlations between attributes; Third, in order to evaluate the impact of different repair strategies on downstream tasks, it is usually necessary to repeatedly train or validate the downstream task model, which is computationally expensive and inefficient; Fourth, there is a lack of end-to-end differentiable connection between the repair process and the task optimization process, making it difficult to use the downstream task loss to directly guide the learning of the repair model.

[0005] Furthermore, in real-world error correction scenarios, tabular data often only provides indications of error locations, such as which cells might contain errors, rather than directly providing the true correct values. For such issues, relying solely on a one-time prediction for correction can easily lead to instability due to the complexity of the error context and the large candidate space. Conversely, relying entirely on manual verification of each cell would significantly increase labor costs, failing to meet the needs of large-scale data governance. Therefore, how to fully utilize correct data as context, gradually recover erroneous cells through learnable methods, and coordinate the optimization of the correction model with downstream tasks, when only the error location or partial error information is known, remains a pressing problem in current technologies.

[0006] Therefore, it is necessary to propose a new table error repair scheme that can establish a connection between the repair process and the task optimization process for downstream tasks, so as to improve the performance of downstream tasks while ensuring the rationality of the repair results, and take into account the expressive power and repair efficiency of mixed-type table data. Summary of the Invention

[0007] To overcome the above technical problems, embodiments of the present invention provide an error repair method, apparatus, and electronic device based on differentiable bilayer optimization, so as to improve the performance of downstream tasks and the table error repair effect while ensuring the rationality of the repair results.

[0008] According to a first aspect of the present invention, an error repair method based on differentiable bilayer optimization is provided, comprising: Obtain the table data containing errors, the corresponding error mask, and the downstream task labels, and divide the table data into training data and validation data; The numerical and categorical attributes in the table data are encoded to obtain the continuous embedding representation of the samples, and the continuous embedding representation of each sample is divided into the target part to be repaired and the condition part according to the error mask. Forward diffusion noise is applied to the target part to be repaired, and a conditional denoising and repair model is constructed with the noisy target part, conditional part and time step as input. The noise is predicted step by step and the repaired target embedding is restored. The repaired target embedding is combined with the condition part and then decoded to obtain the repaired table data; In the inner layer optimization, the conditional denoising and repair model is fixed, and the parameters of the downstream task model are updated using the repaired training data, thereby obtaining the updated downstream task model. In the outer layer optimization, the updated downstream task model is fixed, the task guidance loss is calculated using the repaired validation data, and the parameters of the conditional denoising and repair model are updated by combining the diffusion self-supervised loss, forming a differentiable bilayer optimization process. After training, the trained conditional denoising and repair model is used to correct errors in the table data to be repaired, and the repair results are output.

[0009] According to a second aspect of the embodiments of this application, an error repair apparatus based on differentiable bilayer optimization is provided, comprising: The data acquisition module is used to acquire table data containing errors, the corresponding error mask, and downstream task labels, and to divide the table data into training data and validation data. The embedding representation module is used to encode the numerical attributes and category attributes in the table data respectively to obtain the continuous embedding representation of the sample, and divide the continuous embedding representation of each sample into the target part to be repaired and the condition part according to the error mask. The diffusion repair module is used to perform forward diffusion noise addition on the target part to be repaired, and to construct a conditional denoising repair model with the noise target part, condition part and time step as input, to gradually predict the noise and restore the repaired target embedding. The data decoding module is used to combine and decode the repaired target embedding with the condition part to obtain the repaired table data; The inner optimization module is used to fix the conditional denoising and repair model in the inner optimization, and use the repaired training data to update the parameters of the downstream task model, thereby obtaining the updated downstream task model. The outer optimization module is used to fix the updated downstream task model in the outer optimization, calculate the task guidance loss using the repaired verification data, and update the parameters of the conditional denoising and repair model by combining the diffusion self-supervised loss, forming a differentiable bi-layer optimization process. The repair output module is used to repair errors in the table data to be repaired by the trained conditional denoising and repair model after training, and output the repair results.

[0010] According to a third aspect of the embodiments of this application, an electronic device is provided, comprising: One or more sensors; One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in the first aspect.

[0011] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method as described in the first aspect.

[0012] The technical solutions provided by the embodiments of this application may include the following beneficial effects: As can be seen from the above technical solutions, this application constructs the table error repair process as a differentiable bi-layer optimization process oriented towards downstream tasks. This ensures that the update of the repair model is not only constrained by the data's own distribution but also directly guided by the task loss on the validation set. This avoids the problem of pursuing only the surface rationality of the data while ignoring the performance of downstream tasks, and improves the adaptability of the repair results to classification or regression tasks. By encoding numerical and categorical attributes separately and performing repair in a unified continuous embedding space, it can effectively handle mixed-type table data, enhance the modeling ability of the correlation between different attributes, and provide a richer and more flexible representation basis for table repair under complex error patterns. In the outer optimization, the repair model parameters are updated using a super-gradient approximation method, eliminating the need to repeatedly train the downstream task model for a large number of candidate repair schemes. This reduces the computational cost of task-oriented error repair and improves the feasibility and efficiency of the method. It can perform context-aware repair of cells to be repaired when only the error location or partial error information is known, without the need for manually formulating repair rules one by one or exhaustively listing a limited number of candidate repair strategies. It has strong automation capabilities and practical application value.

[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0015] Figure 1 This is a flowchart of an error repair method based on differentiable bilayer optimization provided in an embodiment of this application.

[0016] Figure 2 This is a block diagram of an error repair device based on differentiable bilayer optimization provided in an embodiment of this application. Detailed Implementation

[0017] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0018] Existing table error correction methods mostly rely on predefined rules, integrity constraints, statistical inference, or finite candidate value enumeration to determine the correction result. While these methods can restore the reasonableness of the data surface in some scenarios, their correction goals are usually focused on data fidelity and are difficult to optimize directly for downstream classification or regression tasks. Ideally, error correction methods should not only utilize the statistical correlation and contextual information within the table data but also be able to dynamically adjust the correction strategy based on feedback from downstream task models, thereby obtaining correction results that are more beneficial to the performance of downstream tasks. However, current methods generally suffer from problems such as limited correction candidate space, difficulty in uniformly modeling the relationships between attributes, separation between the correction process and the task optimization process, and high computational cost when dealing with mixed-type table data that contains both numerical and categorical attributes.

[0019] To overcome the above technical problems, embodiments of this application provide an error repair method, apparatus, and electronic device based on differentiable bilayer optimization, so as to improve the quality of table error repair and the performance of downstream tasks.

[0020] This application provides an error correction method based on differentiable bilayer optimization, applicable to all scenarios involving tabular data quality governance and downstream predictive analysis, such as healthcare, financial risk control, retail operations, industrial monitoring, and government administration, offering a good solution for tabular error correction. In this embodiment, hospital information table data is used as an example. Each row in the hospital information table data can represent a hospital or a business record, and each column can represent attributes such as hospital number, region, level, number of beds, fee information, and disease classification code. Errors such as numerical anomalies, misclassified categories, garbled codes, and missing substitute values ​​may exist in the table due to manual entry, system synchronization, or historical migration.

[0021] In a financial risk control scenario, each row in the table can represent a loan application record or a customer business record, and each column can represent attributes such as customer number, region, income level, credit score, historical default status, and loan product type. The table may also contain errors such as abnormal values, misfilled categories, chaotic coding, and missing substitute values ​​due to manual entry, system synchronization, or historical migration.

[0022] The above are just two examples, and are not limited to these.

[0023] Figure 1 This is a flowchart illustrating an error repair method based on differentiable bilayer optimization, provided in an embodiment of this application. Figure 1 As shown, the method may include the following steps: S1: Obtain the table data containing errors, the corresponding error mask, and the downstream task labels, and divide the table data into training data and validation data; this step may include the following sub-steps: S11: Obtain the table data containing errors, the corresponding error mask, and the downstream task labels; Specifically, this embodiment uses hospital information table data as an example for illustration. Each row in the hospital information table data can represent a hospital or a business record, and each column can represent attributes such as hospital number, region, level, number of beds, charging information, and disease classification code. The table may contain errors such as abnormal values, incorrect category entries, garbled codes, and missing substitute values ​​due to manual input, system synchronization, or historical migration. The input table data containing errors is represented as a data set composed of multiple samples: in, Indicates the number of samples. This indicates the number of attributes contained in each sample. Indicates the first One sample. Corresponding to the table data, obtain the downstream task tag set: in, Indicates sample The corresponding downstream task tags, whereby the downstream tasks can be classification tasks or regression tasks, such as hospital level classification tags, fee risk classification tags, or disease classification code compliance tags.

[0024] Simultaneously, obtain the set of error masks corresponding to the table data: in, Indicates sample The The location of each attribute indicates the location of the error to be fixed. Indicates sample The The position of each attribute is a known correct position.

[0025] S12: Preprocess the table data and determine the data type of each attribute; Specifically, attributes such as the number of beds and fee information can be defined as numerical attributes, while attributes such as region, grade, and disease classification code can be defined as categorical attributes. For the original tabular data, field alignment, format standardization, and missing data marker standardization can be performed first. To facilitate subsequent coding and model training, the original tabular data undergoes field alignment, format standardization, abnormal input standardization, and missing data marker standardization, thereby improving the accuracy of the tabular data representation. Furthermore, all attribute indexes are recorded as an attribute set: The attribute set is then divided into a set of numerical attributes and a set of categorical attributes: in, This represents the set of indices corresponding to numeric attributes. This represents the set of indexes corresponding to the category attributes. By pre-identifying the types of each attribute, a basis is provided for subsequent numerical projection coding and category embedding coding. Through the above processing, the differences in the input format and field expression of hospital information table data can be reduced, providing a foundation for subsequent numerical attribute coding and category attribute coding.

[0026] S13: Determine the positions to be repaired and the positions to remain unchanged based on the error mask; Specifically, for any sample According to its corresponding error mask vector The set of attribute locations to be repaired and the set of conditional attribute locations are determined as follows: in, Indicates the first The set of target attribute locations that need error correction in a sample Indicates the first The set of correct attribute locations used as contextual inputs in a sample. In subsequent fixes, only... Repair the corresponding locations, and for The corresponding positions remain unchanged.

[0027] For example, if the billing information in a hospital record is clearly abnormal, or if there are errors or misfills in the disease classification code, the corresponding attribute position can be identified as the position to be repaired. During the subsequent repair process, only the attribute corresponding to the position to be repaired is fixed, while the attribute corresponding to the conditional position remains unchanged. That is, for any sample... any attribute position ,satisfy: in, Indicates sample The result obtained after repair.

[0028] S14: Divide the table data, the corresponding downstream task labels, and the error mask into training data and validation data according to the one-to-one correspondence of samples, and establish a training label set and a training error mask set corresponding to the training data, and a validation label set and a validation error mask set corresponding to the validation data, respectively. Specifically, the error table data, downstream task labels, and error masks are divided into a training part and a validation part according to a one-to-one correspondence among samples, as follows: in, Represents training data, This represents the verification data. and These represent the sets of downstream task labels corresponding to the training data and validation data, respectively. and Let represent the sets of error masks corresponding to the training data and validation data, respectively, and satisfy the following: In this embodiment, the samples can be divided according to the distribution of hospital records, so that both the training and validation data contain hospital records from different regions, levels, or business types. In some implementations, random partitioning, hierarchical partitioning, or time-order-based partitioning can also be used to ensure the rationality of the label distribution or business entity distribution of the training and validation data.

[0029] S2: Encode the numerical and categorical attributes in the table data to obtain continuous embedding representations of the samples, and divide the continuous embedding representation of each sample into a target part to be repaired and a conditional part according to the error mask; this step may include the following sub-steps: S21: Encode the numerical attributes in the sample to obtain the embedded representations corresponding to each numerical attribute; Specifically, for any sample If its first Each attribute belongs to the set of numerical attributes. Then, the numerical attribute is treated as a continuous scalar and mapped to a continuous embedding space through the projection network corresponding to the attribute, to obtain the first... Embedded representation of each attribute: in, Indicates the first Projection network corresponding to each numerical attribute Indicates sample The Embedding vectors corresponding to each attribute.

[0030] For example, attributes such as the number of beds, billing information, and outpatient volume in hospital information tables can all be processed as numerical attributes. Since these attributes usually have continuously changing characteristics, and different numerical ranges often correspond to different business meanings, mapping them to a continuous embedding space through a projection network can facilitate subsequent models to uniformly represent numerical anomalies, range offsets, and correlations between numerical values.

[0031] In some implementations, the projection network can employ a structure combining linear layers, normalization layers, and activation functions to enhance its ability to represent the nonlinear distribution characteristics of numerical attributes. For example, numerical attributes can be projected onto a preset dimension first through linear transformation, and then the projection result can be transformed through layer normalization and activation functions to obtain a numerical attribute embedding representation suitable for subsequent diffusion repair processing.

[0032] S22: Encode the category attributes in the sample to obtain the embedding representation corresponding to each category attribute; Specifically, if the sample The Each attribute belongs to the category attribute set. For attributes such as region, grade, and disease classification code in hospital information table data, a corresponding embedded table is created for that category attribute. And based on the value of this attribute, find its embedding vector in the embedding table to obtain the first... The embedding representation of each attribute facilitates the model's learning of the association features between different categories: in, Indicates the first Embedded table corresponding to each category attribute This indicates the cardinality of the values ​​for this category attribute. Indicates the embedding dimension. Indicates sample The Embedding vectors corresponding to each category attribute.

[0033] By using the above method, appropriate encoding methods can be adopted for numerical attributes and categorical attributes respectively, representing different types of attribute values ​​in a unified continuous space, thereby taking into account the semantic characteristics and statistical features of different attributes in mixed-type tabular data.

[0034] S23: Concatenate the embedding representations corresponding to each attribute to obtain the continuous embedding representation of the sample; Specifically, in obtaining samples After generating the embedding vectors corresponding to all attributes, the embedding vectors of each attribute are concatenated according to the attribute order. For example, the embedding vectors corresponding to each attribute are concatenated according to a preset order such as hospital number, region, level, number of beds, charging information, and disease classification code, to form a continuous representation of the corresponding hospital record, thus obtaining a continuous embedding representation of the sample. in, This represents a vector concatenation operation. Indicates sample The corresponding initial continuous embedding representation, This indicates the number of attributes contained in the sample. This represents the embedding dimension corresponding to each attribute.

[0035] In some implementations, the continuous embedding representation does not directly repair cell values ​​in the original data space. Instead, it maps the original table data to a low-dimensional continuous space for subsequent diffusion modeling and conditional denoising repair. Compared to operating directly in the original discrete or mixed data space, using continuous embedding representation is beneficial for constructing differentiable repair models and expanding the candidate repair space for the cells to be repaired.

[0036] S24: Divide the continuous embedding representation of the sample into a target part to be repaired and a condition part according to the error mask; Specifically, to ensure that the subsequent repair model only repairs the erroneous locations while retaining the correct locations as contextual input, it is necessary to partition the continuous embedding representations based on the error mask. First, the error mask vector corresponding to the sample... Extending along the embedding dimension yields a representation similar to the continuous embedding. A dimensionally consistent mask matrix is ​​used. If a hospital record contains anomalies in its billing information or incorrectly entered disease classification codes, while fields such as hospital number, region, and level are confirmed to be correct, then the embeddings corresponding to the billing information and disease classification codes can be divided into the target parts to be repaired, and the embeddings corresponding to the remaining correct fields can be divided into the conditional parts. Then, based on the expanded mask matrix, the continuous embedding representations are divided into the target parts to be repaired. and condition section Specifically, it is expressed as: in, This indicates element-wise multiplication. This represents a matrix of all ones with the same dimensions as the mask matrix. This represents the expanded error mask matrix. This represents the target portion to be repaired, consisting of the embedded parts corresponding to the error locations. This indicates the conditional part consisting of the correct positional embedding.

[0037] Furthermore, the target portion to be repaired is used for subsequent forward diffusion noise addition and conditional inverse diffusion repair. The conditional portion retains known contextual information throughout the repair process, providing constraints for the recovery of the target portion to be repaired. This partitioning method allows the repair model to focus on the embedding recovery of erroneous locations without altering the attribute representations of known correct locations, thereby improving the rationality and stability of the repair results.

[0038] S3: Perform forward diffusion noise addition on the target portion to be repaired, and construct a conditional denoising and repair model with the noisy target portion, conditional portion, and time step as inputs to progressively predict noise and restore the repaired target embedding; this step may include the following sub-steps: S31: Perform forward diffusion noise addition on the target part to be repaired to obtain the noise target embedding corresponding to each time step; Specifically, in step S2, the target part of the sample to be repaired is obtained. Subsequently, to enable the repair model to learn the recovery process of the error location through progressive denoising, forward diffusion is performed on the target part to be repaired to add noise. If there is an anomaly in the billing information or a misfilled classification code in a hospital record, the embedded representation of the corresponding data can be used as the target part to be repaired, and noise is progressively added during the forward diffusion process. Let the preset total number of diffusion steps be... The variance schedule during the diffusion process is and define: In some implementations, to make the noise injection process at different time steps smoother and more stable, a cosine noise schedule can be used, whose cumulative noise figure satisfies: in, in, This is a preset small offset used to improve numerical stability.

[0039] Based on the above definition, it can be embedded from the initial target to be repaired. At any time step The corresponding noise target embedding is obtained by direct sampling. Its form is: Equivalent land can also be expressed as: in, Indicates standard Gaussian noise. Represents the identity matrix.

[0040] Through the aforementioned forward diffusion process, the target part to be repaired can be gradually perturbed into representations under different noise intensities, thereby enabling the subsequent repair model to learn the ability to recover erroneous attributes under different degrees of damage.

[0041] S32: Construct a conditional denoising and restoration model with the noise target part, condition part, and time step as input; Specifically, in order to utilize the contextual information provided by the known correct attributes in the samples to guide the recovery process of the target part to be repaired, this application constructs a conditional denoising repair model. The input to the conditional denoising and restoration model includes: the noise target embedding at the current time step. (Including confirmed correct attribute information such as hospital number, region, level, number of beds, etc.) Conditions that remain unchanged (e.g., abnormal billing information or incorrect disease classification coding) and time steps The encoding result; the model will refer to other correct information in the same hospital record and output the prediction result with injected noise at the current time step. ,Right now: in, These represent the parameters of the conditional denoising and repair model.

[0042] In some implementations, the conditional denoising and restoration model can employ a U-Net-style encoder-decoder structure adapted for tabular data. Since tabular data differs from image data, this application does not use traditional convolution; instead, it concatenates the noise target portion with the conditional portion and performs linear projection, combining this with the temporal step encoding results to form the model input hidden representation. in, This represents a linear mapping operation. This indicates a splicing operation. Indicates time step Embedded representation.

[0043] Furthermore, the hidden representation is input into the encoder to obtain the bottleneck representation and multi-layer skip connection features: The decoder then combines the bottleneck representation and skip connection features to output the noise prediction result: in, The bottleneck layer representation of the encoder output. This represents multi-level skip connection features. Through this structure, multi-level attribute-related features can be extracted during the encoding stage, while local detail information can be preserved during the decoding stage. This allows for effective repair of complex error scenarios such as charging anomalies, category misfilling, and encoding chaos.

[0044] S33: Based on the noise prediction results output by the conditional denoising and repair model, perform conditional inverse diffusion to gradually restore the repaired target embedding; Specifically, in the inverse diffusion process, the target is embedded according to the noise target at the current time step. Conditions section and the noise prediction results output by the conditional denoising and repair model. For example, under constraints such as region, grade, and number of beds, the charging information data can be gradually restored to a lower noise representation from the previous time step. .

[0045] Based on the inverse process of the diffusion model, a conditional inverse diffusion transfer distribution can be constructed: in, The mean term of the conditional reverse diffusion transfer distribution is: Therefore, a low-noise target embedding can be obtained by sampling from the conditional inverse diffusion transfer distribution. This is then fed into the conditional denoising and repair model at the next time step, and the process is repeated until the time step is reduced to 1 or 0.

[0046] Throughout the entire conditional reverse diffusion process, the conditional part It remains unchanged and continues to participate in the noise prediction and target recovery process at each time step as known context information, thereby ensuring that the recovery result of the attribute to be repaired is consistent with the original correct attribute of the sample.

[0047] S34: Perform conditional inverse diffusion step by step in descending order of time steps to obtain the repaired target embedding; Specifically, after completing the construction of the forward diffusion denoising and conditional denoising repair model, starting from the maximum diffusion time step... Starting with the corresponding noise target embedding, conditional inverse diffusion is performed sequentially in descending order of time steps, so that the charging information, disease classification code, and pending repair attributes gradually approach a reasonable representation that matches the hospital record from a high-noise state: in, This represents the repaired target embedding obtained after stepwise conditional backdiffusion.

[0048] By employing the stepwise recovery method described above, the embedding corresponding to the original erroneous position can be gradually approximated from a high-noise state to a reasonable repaired representation, avoiding the instability caused by directly predicting the repaired value all at once. Furthermore, since the recovery process is always constrained by the conditional part, the final repaired target embedding not only reflects the reasonable value of the position to be repaired but also maintains semantic and statistical consistency with other known correct attributes in the same sample.

[0049] S4: Combine and decode the repaired target embedding with the conditional portion to obtain the repaired table data; this step may include the following sub-steps: S41: Fuse the repaired target embedding with the conditional part according to the position indicated by the error mask to obtain the repaired complete sample embedding; Specifically, in step S3, the repaired target embedding can be obtained through a conditional reverse diffusion process. Since the target embedding only corresponds to the position to be repaired indicated by the error mask in the original sample, while the other positions in the sample that are determined to be correct have already been divided into conditional parts in step S2. Therefore, it is necessary to fuse the repaired target embedding with the conditional part to restore the complete sample embedding representation.

[0050] Specifically, using an error mask matrix By performing position selection and fusion on the repaired target embedding and conditional parts, the repaired complete sample embedding representation can be obtained. : in, This indicates element-wise multiplication. This represents a matrix of all ones with the same dimensions as the mask matrix. This represents the complete sample embedding representation after fusion.

[0051] Due to the condition section Since the correct positions in the original samples are known and remain unchanged throughout the aforementioned repair process, the fusion method described above ensures that only the positions indicated by the error mask are replaced by the repaired target embedding, while the remaining correct positions retain their original contextual information. Therefore, the repaired complete sample embedding contains both the error position representations recovered by the model and retains the original semantics and statistical features of the correct positions.

[0052] In this embodiment, if the billing information in a hospital record is abnormal, or the disease classification code is incorrectly entered, while attributes such as hospital number, region, level, and number of beds are confirmed to be correct, the corrected target embedding corresponding to the billing information or disease classification code can be fused with the conditional parts corresponding to the other correct attributes to restore the complete embedded representation corresponding to that hospital record. Through this method, only the positions of the erroneous attributes can be replaced while retaining the original correct information.

[0053] S42: Decode the repaired complete sample embedding according to the attribute type to obtain the repair result corresponding to each attribute; Specifically, after obtaining the repaired complete sample embedding representation Next, it needs to be mapped back from the continuous embedding space to the original tabular data space to obtain the specific repair values ​​for each attribute. To this end, the repaired complete sample embedding is split into embedding vectors corresponding to each attribute along the attribute dimension, as follows: in, Indicates sample The Middle The repaired embedding vectors corresponding to each attribute.

[0054] Furthermore, different decoding methods are used depending on the attribute type: For numerical attributes, if the first Each attribute belongs to the set of numerical attributes. Then, it is mapped to a scalar repair value through the corresponding numerical decoding head, represented as: in, Indicates the first A decoding function corresponding to each numerical attribute. In some implementations, the numerical decoding function can be implemented using a linear layer or a regression head to embed and map the attribute into a corresponding continuous numerical value.

[0055] For category attributes, if the first Each attribute belongs to the category attribute set. Then, the confidence vector or logits of the attribute in the candidate class space is output through the corresponding category decoding header, as follows: in, Indicates the first The candidate category output vector corresponds to each category attribute. Indicates the first The classification decoding function corresponding to each category attribute.

[0056] Furthermore, the repair result for the category attribute can be determined based on the position corresponding to the maximum confidence in the candidate category output vector, i.e.: In some implementations, the candidate value range of the category attribute can be constituted by the set of categories in which the attribute has already appeared in the training data or dirty data. This allows the decoding process to be performed within a finite and learnable category space, improving the stability and computational efficiency of category repair. For example, attributes such as the number of beds and billing information in hospital information tables can be restored to their corresponding numerical results using a numerical decoding head, while attributes such as region, grade, and disease classification codes can have their corresponding category results determined using a category decoding head. By decoding separately according to attribute type, it is easier to recover the actual values ​​of different types of attributes.

[0057] S43: Generate the repaired table data based on the repair results corresponding to each attribute; Specifically, after obtaining the sample Repair results for each attribute Then, the repaired sample can be further assembled: For all samples, the repaired tabular data set can be obtained: in, This represents the set of table data after the repair.

[0058] Furthermore, since the method of this application only repairs the position indicated by the error mask, while keeping the original value unchanged for the correct position, therefore for any sample any attribute position If it satisfies Then there are still: Therefore, through step S4, the repaired target embedding obtained in step S3 can be effectively integrated with the original correct context, and the numerical attributes and category attributes can be decoded according to the attribute type, so as to obtain the repaired table data that is structurally complete, semantically reasonable and consistent with the context.

[0059] After the above processing, the repaired hospital information table data can be obtained. For example, for hospital records with abnormal billing information or incorrect disease classification codes in the original records, the corresponding erroneous fields can be repaired while keeping the correct fields such as hospital number, region, level, and number of beds unchanged, forming a repaired record with a complete structure and consistent attribute values.

[0060] S5: In the inner layer optimization, fix the conditional denoising and repair model, and use the repaired training data to update the parameters of the downstream task model; this step may include the following sub-steps: S51: Sample batches of samples from the training data, and under the conditional denoising and repair model parameters, perform encoding, target part and conditional part division, diffusion repair and decoding on the batches of samples to obtain repaired training data; Specifically, in the inner layer optimization process, the fixed-condition denoising and repair model... parameters The process remains unchanged, only using it to correct errors in the training data, and then using the corrected results to update the downstream task model. First, from the training dataset... A batch of samples is sampled, and the corresponding error mask and downstream task label are obtained, represented as: in, This represents the training batch samples obtained from the current sampling. This indicates the corresponding error mask. This indicates the corresponding downstream task label.

[0061] The training batch samples can be multiple hospital records from a hospital information table, some of which may contain errors such as abnormal billing information, incorrect disease classification codes, incorrect regional fields, or disordered grade fields. Subsequently, following the method in step S2, the training batch samples are encoded into a continuous embedding representation, and divided into a target part to be repaired and a conditional part based on the error mask, resulting in... Further, following the method in step S3, forward diffusion noise addition and conditional inverse diffusion repair are performed on the target part to be repaired to obtain the repaired target embedding. Then, following the method in step S4, the repaired target embedding is fused with the conditional part, and the repaired training sample is obtained by decoding. Its form can be expressed as: in, Indicates a decoding operation. This represents the repaired training batch samples.

[0062] Through the above processing, repaired training data tailored to the current downstream task can be obtained during the inner-layer optimization process, thus providing input for the training of models for subsequent downstream tasks. For example, for hospital records with abnormal billing information, reasonable values ​​corresponding to the billing information can be restored under the constraints of attributes such as hospital number, region, level, and number of beds; for hospital records with incorrect disease classification codes, the corresponding codes can be restored by combining other correct attributes in the record.

[0063] S52: Input the repaired training data and the corresponding downstream task labels into the downstream task model, and calculate the training loss; Specifically, after obtaining the repaired training batch samples Then, it is input into the downstream task model. And combined with the corresponding downstream task tags Calculate the training loss. The parameters of the downstream task model are denoted as... The training objective corresponding to the current inner layer optimization can be expressed as: Equivalent land, in batch form, can also be expressed as: in, Indicates batch size, Represents the single-sample loss function. Indicates the first in the batch One repaired training sample, This indicates the corresponding task tag.

[0064] When the downstream task is a hospital tier classification task, the downstream task label can represent the tier category to which the corresponding hospital record belongs; when the downstream task is a billing risk prediction task, the downstream task label can also represent the risk category or risk score of the corresponding hospital record. Accordingly, when the downstream task is a classification task, the training loss can be cross-entropy loss; when the downstream task is a regression task, the training loss can be mean squared error loss or mean absolute error loss. Therefore, the method of this application is not limited to a specific task form, but is applicable to various differentiable downstream task models.

[0065] Furthermore, from the perspective of two-layer optimization, with fixed repair model parameters... Under these conditions, the objective of inner layer optimization can be expressed as: in, Indicates the parameters of the current repair model. The parameters of the downstream task model that minimize training loss are then determined.

[0066] S53: Update the parameters of the downstream task model through backpropagation based on the training loss, and repeat the inner layer optimization a preset number of times to obtain the updated downstream task model. Specifically, upon obtaining training loss Then, gradient descent or other optimization methods are used to adjust the parameters of the downstream task model. The update is performed. The parameter update can be expressed as: in, This represents the learning rate of the downstream task model. The training loss is expressed with respect to the parameters. The gradient.

[0067] In some implementations, the inner layer optimization process can be repeated a preset number of times. This process involves repeatedly training the downstream task model on the repaired training data generated by the current repair model, allowing the downstream task model to gradually converge. Correspondingly, the inner-layer optimization process can be summarized as follows: under the condition of fixed repair model parameters, repeatedly train the downstream task model using the repaired training data, so that the downstream task model gradually adapts to the data distribution corresponding to the current repair result.

[0068] S6: In the outer layer optimization, fix the updated downstream task model, calculate the task guidance loss using the repaired validation data, and update the parameters of the conditional denoising and repair model by combining the diffusion self-supervised loss, forming a differentiable bilayer optimization process; this step may include the following sub-steps: S61: Sample batches from the verification data and construct an enhanced mask based on the original error mask; Specifically, after completing the inner-layer optimization in step S5, the updated downstream task model is fixed. and from the verification dataset A batch of samples is sampled, and the corresponding error mask and downstream task label are obtained simultaneously, represented as follows: in, This represents the validation batch samples obtained from the current sampling. This represents the corresponding original error mask. This indicates the corresponding downstream task label.

[0069] For example, the verification batch sample can be multiple hospital records from a hospital information table, some of which still contain errors such as abnormal billing information, incorrectly filled disease classification codes, incorrectly filled regional fields, or disordered grade fields. Furthermore, to enable the conditional denoising and repair model not only to learn and repair the original error locations but also to learn data recovery capabilities under different noise intensities and masking patterns in a self-supervised manner, an enhanced mask is constructed by randomly selecting some attribute locations from those marked as correct, based on the original error mask. The enhanced mask satisfies: in, This indicates that the enhancement mask is in the 1st... The first sample The value at each attribute position.

[0070] For a given hospital record, the original error mask might only mark the billing information field as an error location. However, when constructing the enhanced mask, additional perturbation locations such as the number of beds, region, or grade can be randomly selected from the originally marked correct locations. Subsequently, the original error mask and the enhanced mask are combined to obtain the joint mask used in the outer optimization layer. in, This represents a union operation of logical OR based on position. Based on the joint mask, the verification batch samples can be encoded and divided into the target part to be repaired and the condition part according to the method in step S2, resulting in: Through the above processing, the verification stage can simultaneously take into account the task-oriented repair of the original error location and the self-supervised denoising learning of the additional enhancement location, thereby improving the generalization ability and stability of the repair model.

[0071] S62: Calculate the diffusion self-supervised loss based on the joint mask; Specifically, in obtaining the target part to be repaired corresponding to the verification batch sample. and condition section Then, sample one diffusion time step. and standard Gaussian noise And construct the noise target embedding for the current time step according to the forward diffusion method in step S3: Then, the noise target is embedded Conditions section and time step Input conditional denoising and repair model The noise prediction results are obtained as follows: To ensure the model maintains its ability to model the data distribution during outer-layer optimization and to prevent the repair results from deviating from the true data manifold, a diffusion self-supervised loss is constructed by statistically analyzing the difference between the model's predicted noise and the actual injected noise only at the locations corresponding to the enhancement mask. in, This represents the matrix form after the enhanced mask is expanded. This indicates the spread of loss from supervision.

[0072] For hospital information table data, the augmentation mask can correspond to randomly selected attribute locations such as the number of beds, region, grade, or disease classification code. By calculating the difference between the model's predicted noise and the actual noise only at these locations, the conditional denoising and restoration model can continue to learn the inherent relationships between different attributes in the hospital records without relying on manually added true restoration values. In this way, the conditional denoising and restoration model can be trained solely using the restoration process with artificially added noise, without relying on manually added true restoration values, thus maintaining the model's constraint on data fidelity.

[0073] S63: Calculate the task guidance loss based on the repaired validation data, and use the supergradient approximation to characterize the impact of the repaired model parameters on the performance of downstream tasks; Specifically, in the outer-layer optimization, the conditional denoising and repair model not only needs to maintain data recovery capabilities but also needs to ensure that the repair results are beneficial to improving the performance of downstream tasks. Therefore, in the fixed-update downstream task model... In this case, a conditional denoising and repair model is used to repair the validation batch samples, resulting in repaired validation data. And calculate the verification loss.

[0074] If the downstream task is hospital tier classification, the validation loss can reflect the impact of the restored hospital records on the tier classification results. If the downstream task is billing risk prediction, the validation loss can reflect the impact of restored billing information, disease classification codes, and other attributes on the risk prediction results. By introducing downstream task losses during the validation phase, the update direction of the restoration model can consider not only the apparent rationality of the data but also the role of the restoration results in the hospital business prediction task.

[0075] From the perspective of two-layer optimization, repairing model parameters The goal is to minimize the task loss on the validation set, and its outer optimization objective can be expressed as: in, Indicates the parameters of the current repair model. The downstream task model parameters are obtained through inner-layer optimization in step S5.

[0076] Due to the repair of model parameters It will not only directly affect the repair results of the validation data, but also indirectly affect the parameters of downstream task models by influencing the repair results of the training data. Therefore, the total derivative of the verification loss with respect to the parameters of the repair model can be expressed as: The first term represents the direct impact of repairing the model parameters on the validation loss, while the second term represents the indirect impact of repairing the model parameters on the validation loss through changes in the downstream task model parameters after inner-layer optimization.

[0077] Furthermore, to avoid the high computational overhead of fully expanding and differentiating the entire inner training process, the implicit function theorem can be used to approximate the indirect influence term, which takes the following form: in, This represents the Hessian matrix representing the inner training loss with respect to the parameters of the downstream task model.

[0078] To efficiently solve the inverse Hessian and vector product terms mentioned above, a system of linear equations can be constructed: in, , This represents the intermediate vector being solved. The linear equations can be approximated by combining the conjugate gradient algorithm with the Hessian vector product.

[0079] After obtaining the intermediate vector Then, task-guided loss can be constructed: in, To characterize the impact of the current parameters of the repair model on the performance of downstream tasks, the feedback from downstream tasks on the validation set can be used to constrain the direction of parameter updates of the repair model. This will make the repaired hospital information table data not only more reasonable in terms of data distribution, but also better serve tasks such as hospital level classification and billing risk prediction.

[0080] S64: Construct an outer optimization objective based on the task-guided loss and the diffusion self-supervised loss, and update the parameters of the conditional denoising and repair model; Specifically, in obtaining mission guidance loss and diffusion of self-monitored loss Subsequently, to balance data fidelity and downstream task performance, a weighted combination of the two is used to construct the outer optimization objective: in, The weighting coefficient represents the diffusion of self-supervised loss, used to balance the task benefits of the repair results with the constraints of data authenticity.

[0081] Subsequently, the model parameters are repaired based on the outer optimization objective. An update can be performed, and its form can be represented as: in, This represents the learning rate of the conditional denoising and restoration model.

[0082] By incorporating both the task-guided loss and the diffusion-supervised loss into the outer optimization objective, the conditional denoising and repair model can, during the update process, maintain its ability to restore the original distribution characteristics of the hospital information table data while gradually optimizing in a direction that improves the performance of hospital business prediction tasks. Repeating this outer optimization process allows the repair model and the downstream task model to adjust collaboratively during training, thus forming a task-oriented, differentiable, two-layer optimization process.

[0083] S7: After training, use the trained conditional denoising and repair model to perform error repair on the table data to be repaired and output the repair results. Specifically, after completing steps S1 to S6 of the training process, the trained conditional denoising and insulation model, along with its corresponding encoder and decoder structures, can be obtained. In practical applications, only the table data to be repaired and its corresponding error mask need to be obtained, without needing to obtain the actual repair labels.

[0084] In this embodiment, the table data to be repaired can be multiple records to be processed in the hospital information table data. Some records may contain errors such as abnormal billing information, incorrect disease classification codes, incorrect regional fields, or disordered grade fields. After inputting the table data to be repaired and its corresponding error mask into the trained conditional denoising and repair model, the model can combine other correct attribute information in the same hospital record and perform stepwise denoising and recovery of the error location under conditional constraints based on the repair mechanism formed during the training phase. A conditional inverse diffusion distribution is constructed based on the inverse diffusion mean. This yields the lower-noise target embedding for the next time step. And continue iterating until the repaired target embedding is obtained. This process can be represented as: Throughout the reasoning and repair process, the conditional part It always participates in model inference as a known correct context, thereby ensuring that the restoration result of the location to be repaired is consistent with the original correct attributes in the sample, and finally outputs the repaired table data. Furthermore, since the method of this application only repairs the location indicated by the error mask, it is applicable to any sample. any attribute position If satisfied If so, the corresponding position will retain its original value, that is: By performing the above S7 step, the table data to be repaired can be directly inferred and repaired after training is completed, and the table error repair result that is consistent with the context, conforms to the data distribution, and is beneficial to the performance of downstream tasks can be output, thereby completing the error repair process based on differentiable bilayer optimization proposed in this application.

[0085] For example, for hospital records with abnormal billing information, the corrected billing results can be output under the constraints of attributes such as hospital number, region, level, and number of beds; for hospital records with incorrect disease classification codes, the corrected disease classification codes can be output by combining other correct attributes in the record. Thus, the corrected hospital information table data can be obtained.

[0086] As can be seen from the above technical solutions, this application constructs the table error repair process as a differentiable bi-layer optimization process oriented towards downstream tasks. This ensures that the update of the repair model is not only constrained by the data's own distribution but also directly guided by the task loss on the validation set. This avoids the problem of pursuing only the surface rationality of the data while ignoring the performance of downstream tasks, and improves the adaptability of the repair results to classification or regression tasks. By encoding numerical and categorical attributes separately and performing repair in a unified continuous embedding space, it can effectively handle mixed-type table data, enhance the modeling ability of the correlation between different attributes, and provide a richer and more flexible representation basis for table repair under complex error patterns. In the outer optimization, the repair model parameters are updated using a super-gradient approximation method, eliminating the need to repeatedly train the downstream task model for a large number of candidate repair schemes. This reduces the computational cost of task-oriented error repair and improves the feasibility and efficiency of the method. It can perform context-aware repair of cells to be repaired when only the error location or partial error information is known, without the need for manually formulating repair rules one by one or exhaustively listing a limited number of candidate repair strategies. It has strong automation capabilities and practical application value.

[0087] Corresponding to the aforementioned embodiments of the error repair method based on differentiable bilayer optimization, this application also provides embodiments of an error repair device based on differentiable bilayer optimization.

[0088] Figure 2 This is a structural block diagram of an error repair device based on differentiable bilayer optimization provided in an embodiment of this application. (Refer to...) Figure 2 The device may include: Data acquisition module 1 is used to acquire table data containing errors, corresponding error masks and downstream task labels, and divide the table data into training data and validation data; Embedding representation module 2 is used to encode the numerical attributes and category attributes in the table data respectively to obtain the continuous embedding representation of the sample, and divide the continuous embedding representation of each sample into the target part to be repaired and the condition part according to the error mask; The diffusion repair module 3 is used to perform forward diffusion noise addition on the target part to be repaired, and to construct a conditional denoising repair model with the noise target part, condition part and time step as input, to gradually predict the noise and restore the repaired target embedding. Data decoding module 4 is used to combine and decode the repaired target embedding with the condition part to obtain the repaired table data; The inner layer optimization module 5 is used to fix the conditional denoising and repair model in the inner layer optimization, and use the repaired training data to update the parameters of the downstream task model, thereby obtaining the updated downstream task model. The outer optimization module 6 is used to fix the updated downstream task model in the outer optimization, calculate the task guidance loss using the repaired verification data, and update the parameters of the conditional denoising and repair model by combining the diffusion self-supervised loss, forming a differentiable bilayer optimization process. Repair output module 7 is used to repair errors in the table data to be repaired by the trained conditional denoising and repair model after training is completed, and output the repair results.

[0089] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0090] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0091] Accordingly, this application also provides an electronic device, comprising: one or more sensors; one or more processors; and a memory for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in the first aspect.

[0092] Accordingly, this application also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the error repair method based on differentiable two-layer optimization as described above.

[0093] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0094] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. An error repair method based on differentiable bilayer optimization, characterized in that, include: Obtain the table data containing errors, the corresponding error mask, and the downstream task labels, and divide the table data into training data and validation data; The numerical and categorical attributes in the table data are encoded to obtain the continuous embedding representation of the samples, and the continuous embedding representation of each sample is divided into the target part to be repaired and the condition part according to the error mask. Forward diffusion noise is applied to the target part to be repaired, and a conditional denoising and repair model is constructed with the noisy target part, conditional part and time step as input. The noise is predicted step by step and the repaired target embedding is restored. The repaired target embedding is combined with the condition part and then decoded to obtain the repaired table data; In the inner layer optimization, the conditional denoising and repair model is fixed, and the parameters of the downstream task model are updated using the repaired training data, thereby obtaining the updated downstream task model. In the outer layer optimization, the updated downstream task model is fixed, the task guidance loss is calculated using the repaired validation data, and the parameters of the conditional denoising and repair model are updated by combining the diffusion self-supervised loss, forming a differentiable bilayer optimization process. After training, the trained conditional denoising and repair model is used to correct errors in the table data to be repaired, and the repair results are output.

2. The method according to claim 1, characterized in that, Obtain the erroneous table data, the corresponding error mask, and the downstream task labels, and divide the table data into training data and validation data, including: Retrieve the table data containing errors, the corresponding error mask, and the downstream task labels; The table data is preprocessed, and the data type of each attribute is determined; Based on the error mask, determine the locations to be repaired and the locations that remain unchanged; The table data, the corresponding downstream task labels, and the error masks are divided into training data and validation data according to the one-to-one correspondence between samples. Training label sets and training error mask sets corresponding to the training data, as well as validation label sets and validation error mask sets corresponding to the validation data, are established respectively.

3. The method according to claim 1, characterized in that, The numerical and categorical attributes in the table data are encoded to obtain continuous embedding representations of the samples. Based on the error mask, the continuous embedding representation of each sample is divided into a target part to be repaired and a conditional part, including: Encode the numerical attributes in the sample to obtain the embedded representation corresponding to each numerical attribute; The category attributes in the sample are encoded to obtain the embedding representation corresponding to each category attribute; The embedding representations corresponding to each attribute are concatenated to obtain the continuous embedding representation of the sample; The continuous embedding representation of the sample is divided into a target part to be repaired and a condition part based on the error mask.

4. The method according to claim 1, characterized in that, Forward diffusion noise is applied to the target portion to be repaired, and a conditional denoising and repair model is constructed with the noisy target portion, conditional portion, and time step as inputs. This model progressively predicts the noise and recovers the repaired target embedding, including: Forward diffusion is performed on the target part to be repaired according to a preset number of diffusion steps and noise schedule, and the noise target embedding is obtained at any time step; The noise target embedding, the conditional part, and the time step encoding result are input into the conditional denoising and repair model, and the noise prediction result of the current time step is output. Based on the noise prediction results, conditional inverse diffusion is performed in descending order of time steps to gradually restore the repaired target embedding. The conditional denoising and repair model adopts a U-Net encoder-decoder structure adapted to tabular data. It extracts multi-layer features at the encoder end, models global attribute dependencies at the bottleneck layer, and restores the target embedding by combining skip connections at the decoder end. Conditional inverse diffusion is performed step by step in descending order of time steps to obtain the repaired target embedding.

5. The method according to claim 1, characterized in that, The repaired target embedding is combined with the conditional portion and decoded to obtain the repaired table data, including: The repaired target embedding is fused with the conditional part at the position indicated by the error mask to obtain the repaired complete sample embedding; For numerical attributes, the corresponding embeddings are mapped to numerical results using a regression decoding head; for categorical attributes, the corresponding embeddings are mapped to the confidence scores of each candidate category using a classification decoding head. Based on the numerical results and confidence levels, the repair values ​​for each attribute are determined, and the repaired tabular data is generated.

6. The method according to claim 1, characterized in that, In the inner layer optimization, the conditional denoising and repair model is fixed, and the parameters of the downstream task model are updated using the repaired training data to obtain the updated downstream task model, including: Batch samples are sampled from the training data. With the parameters of the conditional denoising and repair model fixed, the batch samples are encoded, the target part and the conditional part are divided, forward diffusion, conditional back diffusion repair and decoding are performed to obtain the repaired training data. The repaired training data and its corresponding downstream task labels are input into the downstream task model to calculate the training loss. The downstream task model parameters are updated through backpropagation based on the training loss, and the inner layer optimization is repeated a preset number of times to obtain the updated downstream task model.

7. The method according to claim 1, characterized in that, In the outer layer optimization, the updated downstream task model is fixed, the task guidance loss is calculated using the repaired validation data, and the parameters of the conditional denoising and repair model are updated in conjunction with the diffusion self-supervised loss, including: Batch samples are sampled from the validation data, and an enhanced mask is constructed by randomly selecting some of the positions marked as correct based on the original error mask; Encoding and diffusion noise are performed based on the union of the original error mask and the enhancement mask. The diffusion self-supervised loss is calculated by predicting the artificially injected noise at the corresponding position of the enhancement mask. The repaired validation data is input into a fixed downstream task model, the task guidance loss on the validation set is calculated, and the hypergradient approximation of the task guidance loss with respect to the repaired model parameters is obtained by using the implicit function theorem approximation, the conjugate gradient algorithm, and the Hessian vector product. Based on the aforementioned diffusion self-supervised loss and task-guided loss, an outer optimization objective is constructed and the parameters of the conditional denoising and repair model are updated.

8. An error repair device based on differentiable bilayer optimization, characterized in that, include: The data acquisition module is used to acquire table data containing errors, the corresponding error mask, and downstream task labels, and to divide the table data into training data and validation data. The embedding representation module is used to encode the numerical attributes and category attributes in the table data respectively to obtain the continuous embedding representation of the sample, and divide the continuous embedding representation of each sample into the target part to be repaired and the condition part according to the error mask. The diffusion repair module is used to perform forward diffusion noise addition on the target part to be repaired, and to construct a conditional denoising repair model with the noise target part, condition part and time step as input, to gradually predict the noise and restore the repaired target embedding. The data decoding module is used to combine and decode the repaired target embedding with the condition part to obtain the repaired table data; The inner optimization module is used to fix the conditional denoising and repair model in the inner optimization, and use the repaired training data to update the parameters of the downstream task model, thereby obtaining the updated downstream task model. The outer optimization module is used to fix the updated downstream task model in the outer optimization, calculate the task guidance loss using the repaired verification data, and update the parameters of the conditional denoising and repair model by combining the diffusion self-supervised loss, forming a differentiable bi-layer optimization process. The repair output module is used to repair errors in the table data to be repaired by the trained conditional denoising and repair model after training, and output the repair results.

9. An electronic device, characterized in that, include: One or more sensors; One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-7.