Data quality enhancement method and device for data space entity analysis

Through adversarial positive sample difficulty enhancement and dynamic noise perturbation technology, the problems of data quality and generalization capabilities in entity analysis are solved, the robustness and accuracy of the model are improved, and it is suitable for multi-source heterogeneous data environments.

CN120387012APending Publication Date: 2025-07-29HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510401684.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing entity analytical methods have low efficiency, poor generalization ability, and strong dependence on data quality in multi-source heterogeneous data environments, making it difficult to effectively distinguish subtle differences, and data quality problems such as scarcity of positive samples and proportional imbalances affect the accuracy of the model.

Method used

Adversarial positive sample difficulty enhancement and dynamic noise perturbation technology, noise generation network injection can be used to learn noise to improve the difficulty of positive sample, combined with dynamic noise amplitude adjustment and information evaluation, high-quality candidate positive samples are generated, and the model is optimized through multi-objective loss function to ensure semantic consistency and classification accuracy.

Benefits of technology

It improves the generalization ability and robustness of the model, enhances the learning of fine-grained features, balances the proportion of positive and negative samples, improves the accuracy and diversity of entity resolution, and is suitable for multi-source heterogeneous environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387012A_ABST
    Figure CN120387012A_ABST
Patent Text Reader

Abstract

The invention discloses a data quality enhancement method and device for data space entity analysis, belongs to the field of data integration and cleaning, and particularly relates to a data quality enhancement method and device for data space entity analysis. The problems that an existing data enhancement and sampling method depends on manual design rules, the cross-domain migration ability is insufficient, and the diversity and accuracy of generated samples cannot be considered at the same time are solved. The method comprises the following steps: injecting adaptive noise into an original positive sample by adopting a learnable noise network, and constructing an adversarial training framework in combination with a difficulty loss function; the random noise amplitude is dynamically adjusted based on the embedding similarity of the entity pair, candidate positive samples are obtained, and semantic consistency and classification accuracy of the samples are generated by combining classification loss and comparison loss function constraints; and high-information-content samples are screened through similarity scores. The data quality enhancing method and device for entity analysis in the data space are suitable for enhancing the data quality of entity analysis in the data space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data integration and cleaning, and particularly to a data quality enhancement method and device for data space entity resolution. Background Art

[0002] In the process of digital transformation, the data space, as a dynamic aggregate of multi-source heterogeneous data, has become the core infrastructure to support intelligent decision-making and data-driven applications. Entity Resolution (ER) is one of the important technologies for data space governance, and its main task is to identify records pointing to the same real entity in different data sources, so as to eliminate data heterogeneity and construct a unified entity view. In industries such as healthcare, e-commerce, and finance, the accuracy of entity resolution directly affects the value density and credibility of data. For example, in the healthcare field, incorrect entity resolution may lead to confusion of patient information, thus affecting diagnosis and treatment results. Therefore, improving the accuracy and robustness of entity resolution is of great significance for optimizing intelligent decision-making and applications.

[0003] With the advent of the big data era, the exponential growth of data scale and complexity has put forward higher requirements for entity resolution technology. Traditional entity resolution methods perform well in dealing with small-scale and single data sources, but in the face of a large-scale data environment with multi-source heterogeneity and dynamic data changes, existing methods generally have problems such as low efficiency, poor generalization ability, and strong dependence on data quality. Therefore, how to improve the accuracy, efficiency, and adaptability of entity resolution in a multi-source heterogeneous environment has become a key challenge to be solved urgently.

[0004] Currently, entity resolution methods have mainly gone through different development stages such as rule-based, machine learning-based, and deep learning-based.

[0005] Rule-based entity resolution methods rely on manually defined matching rules and thresholds to calculate the similarity between attributes, such as Jaccard similarity, edit distance, etc. The matching results of such methods are interpretable and applicable to structured data scenarios. For example: The CME method uses domain constraints combined with a relaxed label algorithm to perform entity matching in an unsupervised manner. The GBF method models the entity resolution task using Boolean formulas and learns matching rules from positive and negative matching examples to improve interpretability. However, rule-based methods mainly have the following limitations: They require manual construction of rules, have poor generalization ability, and are difficult to adapt to cross-domain and multi-source heterogeneous data; they cannot effectively handle semantic heterogeneity. For example, the matching of "Peking Union Medical College Hospital" and "PUMCH" requires customized rules; the cost of formulating and maintaining rules is high, and it is difficult to scale to large-scale data sets.

[0006] To overcome the limitations of rule-based methods, researchers have proposed entity resolution methods based on machine learning. Such methods use classifiers such as support vector machines (SVMs) and random forests to predict whether entity pairs match by learning manually designed features. For example: The MARLIN method learns similarity measures for different fields and combines SVM for entity matching. The Magellan method provides an end-to-end entity resolution solution, automatically generates matching features, and uses multiple classifiers to train the model. Although ML methods reduce the dependence on manual rules and improve the automation of the model, there are still the following problems: high feature engineering complexity. For example, in an e-commerce product dataset with 50 attributes, the feature combination space can reach 10 15 species, resulting in the curse of dimensionality problem; poor adaptability to high-dimensional heterogeneous data, requiring a large number of manually designed features. The feature selection process depends on domain knowledge and is difficult to transfer to different datasets.

[0007] Deep learning methods can automatically extract text features and deep representations, significantly improving the matching effect of entity resolution and becoming a research hotspot in recent years. For example: The DeepER method first combines word embedding with LSTM to represent entity records in a high-dimensional vector form and calculates entity matching using vector similarity. The DeepMatcher method proposes four DL-based solutions, such as SIF, RNN, and Attention mechanism, to learn the similarity between entity records. The MCA method comprehensively captures the matching signals between entity record pairs through Self-Attention, Pairwise Attention, and Global Attention. The Seq2SeqMatcher method models the entity resolution task as a sequence matching problem by comparing words across attributes, thus improving the model's adaptability to pattern heterogeneity and noisy data.

[0008] In addition, with the breakthrough of the Transformer model in the field of natural language processing (NLP), this model has gradually been applied to the entity resolution task. For example: The Ditto method calculates the similarity of entity pairs through the BERT pre-trained model and combines optimization strategies such as data augmentation and domain knowledge injection to significantly improve the matching quality. HierGAT designs a hierarchical graph attention network to strengthen the modeling of the dependence relationship between entities and attributes and improve the discriminative ability of the model.

[0009] Although deep learning-based methods have made significant progress in entity resolution, these models typically rely on large-scale and high-quality datasets. However, when the data quality is low or the samples are scarce, it is often difficult for the models to effectively learn matching features, and the imbalance between positive and negative samples may cause the models to be biased towards negative samples, thus affecting the accuracy of the matching results. These data quality issues severely restrict the generalization ability and matching accuracy of the models. Especially when faced with large-scale heterogeneous data, it is difficult for the models to effectively distinguish subtle differences.

[0010] Currently, the main data challenges faced by the entity resolution task can be summarized into the following two major problems:

[0011] (1) Low-difficulty positive sample problem: The positive samples in the existing datasets have too high similarity, resulting in the models being prone to falling into local optimal solutions and unable to learn fine-grained discriminative features. Especially when faced with large-scale heterogeneous data, the generalization ability of the models is restricted and they cannot effectively distinguish subtle differences.

[0012] (2) Positive sample scarcity problem: Due to the inherent characteristics of entity resolution datasets, the ratio of positive to negative samples is severely imbalanced. For example, in the Amazon-Google dataset, positive samples only account for 10.2%, and this imbalance causes the models to be prone to being biased towards negative samples, thus affecting the accuracy of entity matching.

[0013] To address the above problems, some studies have attempted to alleviate the deficiencies in data quality through data augmentation and sampling methods. For example, the Ditto method improves the model performance by introducing domain knowledge and data augmentation techniques. However, the enhancement strategy of this method mainly relies on manually designed rules (such as attribute abbreviation, keyword replacement, etc.), and this approach lacks systematic theoretical support and is prone to problems of inconsistent data distribution when migrating across domains, restricting its application in the open domain.

[0014] Similarly, the application of oversampling methods such as SMOTE in entity resolution also has limitations. Although SMOTE alleviates the problem of positive sample scarcity to a certain extent, the generated synthetic samples often have high similarity to the original samples and lack sufficient diversity, resulting in the inability to effectively improve the model's learning ability for fine-grained features. In addition, the samples generated by SMOTE may cross the decision boundary, thus triggering false positive samples and affecting the robustness and generalization ability of the models.

[0015] Generative adversarial networks (GANs) can enhance data diversity, but their training process is complex and may generate noise or inaccurate samples. Especially when generating complex data patterns, GANs may cause the models to learn false data features, thereby affecting the parsing performance.

[0016] In addition, enhancement methods that rely on external knowledge bases (such as knowledge graphs or dictionaries) also face problems. The quality of the knowledge base, the difficulty of maintenance, and poor domain adaptability make it difficult for these methods to operate stably in a changing real - data environment.

[0017] Based on the above - mentioned analysis, there is an urgent need for a general, data - driven enhancement method in the current entity resolution field to overcome the limitations of existing technologies. Summary of the Invention

[0018] The present invention proposes a data - quality enhancement method and device for entity resolution in the data space, which solves the problems existing in existing data enhancement and sampling methods, such as relying on manually designed rules, insufficient cross - domain migration ability, and inability to balance the diversity and accuracy of generated samples.

[0019] The data - quality enhancement method for entity resolution in the data space according to the present invention includes the following steps:

[0020] Step of obtaining the original data set and feature extraction: Obtain the original data set and generate original positive samples by using the feature - extraction ability of the entity - resolution model;

[0021] Step of enhancing the difficulty of adversarial positive samples: Use a noise - generation network to generate learnable noise and inject it into the original positive samples to increase the discrimination difficulty of the positive samples and obtain enhanced positive samples; Define a difficulty loss function, and optimize the parameters of the entity - resolution model through the Minimax adversarial training framework. While minimizing the standard cross - entropy loss function of the entity - resolution model, maximize the difficulty loss function, where the difficulty loss function is used to maximize the misclassification probability of the enhanced positive samples;

[0022] Step of generating positive samples with dynamic noise perturbation: Generate random noise and inject it into the original positive samples, and dynamically adjust the amplitude of the random noise based on the similarity of the embedded representations of entity pairs to obtain newly generated candidate positive samples; Under the contrastive learning framework, jointly optimize and train the entity - resolution model through the classification loss function and the contrastive loss function including the candidate positive samples, and simultaneously constrain the semantic consistency and classification accuracy of the candidate positive samples;

[0023] Step of information - quantity evaluation and screening: Evaluate the information quantity of the candidate positive samples to obtain an information - quantity score; After sorting by the information - quantity score, select the candidate positive samples with higher information quantity to obtain the final new positive samples;

[0024] Step of overall optimization: Use a loss function to construct an overall optimization objective function for multi - objective joint optimization; The loss function includes the standard cross - entropy loss function, the difficulty loss function, the classification loss function including the candidate positive samples, and the contrastive loss function; By adjusting the weights of each loss function in the overall optimization objective function, ensure the balance of each loss term.

[0025] Further, a preferred embodiment is provided. In the step of obtaining and feature extracting the original data set, the original positive samples generated by using the feature extraction ability of the entity resolution model are as follows:

[0026] Let the entity pair in the original data set be (e, e′), and the embedding representation of each entity in the entity pair is generated through the entity resolution model:

[0027] v l = Encoder(e), v r = Encoder(e′);

[0028] The embedding representations of each entity are concatenated to form the original positive samples:

[0029] s p = [v l ; v r ∈ R 2d , where d is the single-entity embedding dimension of the entity resolution model.

[0030] Further, a preferred embodiment is provided. In the step of enhancing the difficulty of the adversarial positive samples:

[0031] The learnable noise is:

[0032] η = σ(Wη′ + b)

[0033] where W ∈ R 2d×2d and b ∈ R 2d×2d are learnable parameters; η′ is uniformly distributed noise randomly generated in the interval [0, 0.1]; σ(·) is the Sigmoid function, which maps the noise value to the interval (0, 1); η follows a uniform distribution in the interval [0, 1], and its dimension is the same as that of the original positive samples, ensuring the dimensional consistency of noise injection;

[0034] The enhanced positive samples are: <>

[0035] s p ′ = s p + η;

[0036] The difficulty loss function is

[0037]

[0038] where p is the matching probability of the entity resolution model for the enhanced positive samples, and θ h represents the parameter for processing the enhanced positive samples;

[0039] The standard cross-entropy loss function is To optimize the classification performance of the entity resolution model for the original positive samples;

[0040] While minimizing the standard cross-entropy loss function of the entity resolution model, maximizing the difficulty loss function is:

[0041]

[0042] Where γ is the hyperparameter that controls the weight of the difficulty loss function weight.

[0043] Furthermore, a preferred implementation is provided. In the step of generating positive samples with dynamic noise perturbation:

[0044] The newly generated candidate positive sample is:

[0045]

[0046] Where:

[0047] Δ is the generated random noise vector, which follows a uniform distribution in the interval [0,1], and its dimension is the same as that of the original positive sample to ensure the dimensional consistency of noise injection;

[0048] m is the amplitude of the random noise:

[0049] m = σ(v l ·v r )

[0050] When the similarity between these two entity embedding representations v l and v r is low, make m → 0;

[0051] When the similarity between these two entity embedding representations v l and v r is high, make m → 1;

[0052] The classification loss function including the candidate positive sample is used to ensure that the model correctly classifies the new candidate positive samples;

[0053] The contrastive loss function is used to constrain the semantic consistency of the new candidate positive samples:

[0054]

[0055] Where s r represents the embedding randomly sampled from the negative samples, and is used to constrain, through backpropagation, that the similarity between the generated candidate positive samples and the original positive samples is higher than that with the negative samples, ensuring the semantic consistency of the generated candidate positive samples.

[0056] Further, a preferred embodiment is provided. In the information quantity evaluation and screening step:

[0057] The information quantity score of each candidate positive sample is:

[0058]

[0059] Information quantity score The lower it is, the higher the information quantity contained in the corresponding candidate positive sample embedding representation is;

[0060] The final newly generated positive sample is:

[0061]

[0062] where τ is a dynamic threshold.

[0063] Further, a preferred embodiment is provided. In the total optimization step of the integrated loss function:

[0064] The total optimization objective function is:

[0065]

[0066] where γ, λ, and β are hyperparameters representing the weights of the loss terms;

[0067] Adjust the hyperparameters γ, λ, and β to ensure the balance of the loss terms during the optimization process.

[0068] The present invention also proposes a data quality enhancement device for data space entity resolution. The device includes the following modules:

[0069] Original dataset acquisition and feature extraction module: Acquire the original dataset and generate original positive samples by using the feature extraction ability of the entity resolution model;

[0070] Adversarial positive sample difficulty enhancement module: Use a noise generation network to generate learnable noise and inject it into the original positive samples to increase the discrimination difficulty of the positive samples and obtain enhanced positive samples; Define a difficulty loss function, and optimize the entity resolution model parameters through the Minimax adversarial training framework. While minimizing the standard cross-entropy loss function of the entity resolution model, maximize the difficulty loss function, and the difficulty loss function is used to maximize the misclassification probability of the enhanced positive samples;

[0071] Positive sample generation module with dynamic noise perturbation: Generate random noise and inject it into the original positive samples, and dynamically adjust the amplitude of the random noise based on the similarity of the embedded representations of the entity pairs to obtain newly generated candidate positive samples; in the contrastive learning framework, jointly optimize and train the entity resolution model through the classification loss function and the contrastive loss function that include the candidate positive samples, and synchronously constrain the semantic consistency and classification accuracy of the candidate positive samples;

[0072] Information quantity evaluation and screening module: Evaluate the information quantity of the candidate positive samples to obtain information quantity scores; after sorting by the information quantity scores, select the candidate positive samples with higher information quantity to obtain the final newly generated positive samples;

[0073] Total optimization module: Construct a total optimization objective function using a loss function for multi-objective joint optimization; the loss function includes a standard cross-entropy loss function, a difficulty loss function, a classification loss function that includes candidate positive samples, and a contrastive loss function; by adjusting the weights of each loss function in the total optimization objective function, ensure the balance of each loss term.

[0074] The present invention also provides a computer device, including: a processor and a memory, the memory is used to store executable instructions of the processor, and the processor is configured to execute the data quality enhancement method for data space entity resolution described in any one of the above via executing the executable instructions.

[0075] The present invention also provides a computer storage medium, in which a computer program is stored, and when the computer program runs, it executes the data quality enhancement method for data space entity resolution described in any one of the above.

[0076] The present invention also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the data quality enhancement method for data space entity resolution described in any one of the above are implemented.

[0077] The present invention has the following beneficial effects:

[0078] 1. The data quality enhancement method for data space entity resolution of the present invention uses an adversarial training mechanism combined with learnable noise to adaptively increase the difficulty of positive samples: by increasing the difficulty of training positive samples, the model can learn more fine-grained features and improve generalization ability.

[0079] 2. The data quality enhancement method for data space entity resolution of the present invention, through dynamic noise perturbation based on similarity combined with an information maximization sampling strategy, while maintaining semantic consistency, safely expands more high-quality positive samples, balances the positive and negative sample ratios, reduces model classification bias, alleviates the problem of scarce positive samples, and thus improves the quality and diversity of data.

[0080] 3. The data quality enhancement method for data space entity resolution according to the present invention is decoupled from the existing model and has the characteristics of plug-and-play. It can be seamlessly integrated into the existing entity resolution model to improve the robustness and matching accuracy of the model (entity resolution model), and is applicable to entity resolution tasks in multi-source heterogeneous environments.

[0081] 4. The data quality enhancement method for data space entity resolution according to the present invention effectively improves the quality of training samples in the data space entity resolution task through adversarial positive sample hardness enhancement and positive sample generation technology based on noise perturbation, thereby significantly improving the accuracy and robustness of entity resolution.

[0082] The data quality enhancement method and device for data space entity resolution according to the present invention are applicable to enhancing the data quality of entity resolution in the data space. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0084] Figure 1 It is a schematic flowchart of the data quality enhancement method for data space entity resolution in an embodiment of the present invention;

[0085] Figure 2 It is a schematic diagram of the adversarial positive sample hardness enhancement step and the positive sample generation step of dynamic noise perturbation in the data quality enhancement method for data space entity resolution in an embodiment of the present invention; wherein, Figure 2 (a) represents Adversarial Hardness Enhancement, that is, a schematic diagram of adversarial positive sample hardness enhancement; Figure 2 (b) represents Dynamic Positive Sample Generation, that is, a schematic diagram of the generation and sampling of positive samples with dynamic noise perturbation;

[0086] Figure 3An example diagram for the task of data space entity resolution; the figure shows the comparison of entity records from two data sources, where each entity record consists of multiple attributes and their corresponding values, and title and price represent the attributes in the entity record; the task objective of entity resolution is to identify records in different data sources that point to the same real entity. In the figure, a cross indicates that these two entity records do not match, i.e., they do not point to the same entity, and a tick indicates a match, i.e., they point to the same entity;

[0087] Figure 4 In an embodiment of the present invention, it is a schematic diagram of the results of hyperparameter experiments; among them, Figure (a) shows the influence of the γ parameter on the model performance; Figure (b) shows the influence of different λ values on the model performance; Figure (c) shows the influence of the β hyperparameter on the model performance. Detailed implementation manners

[0088] To make the technical solutions and advantages of the present invention more clearly expressed, the following will further describe in detail and completely the specific implementation manners of the present invention in conjunction with the accompanying drawings. The following described implementation manners are only some preferred solutions of the present invention, rather than all implementation solutions; the following described implementation manners are intended to explain the present invention and should not be construed as a limitation to the present invention; the reasonable combination of the technical features defined in each implementation manner of the present invention, and all other implementation manners obtained by those of ordinary skill in the art based on the implementation manners of the present invention without creative efforts belong to the scope of protection of the present invention.

[0089] Embodiment 1. Refer to Figure 1 To illustrate this embodiment, this embodiment provides a data quality enhancement method for data space entity resolution, including adversarial positive sample difficulty enhancement and dynamic noise perturbation positive sample generation, which are used to solve the problems of low-difficulty positive samples and positive sample scarcity in entity resolution data, thereby improving the accuracy and robustness of entity resolution.

[0090] A data quality enhancement method for data space entity resolution, the method comprising the following steps:

[0091] Original dataset acquisition and feature extraction step: Acquire the original dataset and generate original positive samples using the feature extraction ability of the entity resolution model;

[0092] Adversarial positive sample difficulty enhancement step: Use a noise generation network to generate learnable noise and inject it into the original positive samples to increase the discrimination difficulty of the positive samples and obtain enhanced positive samples; define a difficulty loss function, and optimize the entity resolution model parameters through the Minimax adversarial training framework. While minimizing the standard cross-entropy loss function of the entity resolution model, maximize the difficulty loss function, and the difficulty loss function is used to maximize the misclassification probability of the enhanced positive samples;

[0093] Positive sample generation steps with dynamic noise perturbation: Generate random noise and inject it into the original positive samples, and dynamically adjust the amplitude of the random noise based on the similarity of the embedded representations of entity pairs to obtain newly generated candidate positive samples; In the contrastive learning framework, jointly optimize and train the entity resolution model through the classification loss function and contrastive loss function that include the candidate positive samples, and simultaneously constrain the semantic consistency and classification accuracy of the candidate positive samples;

[0094] Information quantity evaluation and screening steps: Evaluate the information quantity of the candidate positive samples to obtain information quantity scores; After sorting by the information quantity scores, select the candidate positive samples with higher information quantity to obtain the final newly generated positive samples;

[0095] Total optimization steps: Construct a total optimization objective function using the loss function for multi-objective joint optimization; The loss function includes the standard cross-entropy loss function, difficulty loss function, classification loss function that includes the candidate positive samples, and contrastive loss function; Ensure the balance of each loss term by adjusting the weights of each loss function in the total optimization objective function.

[0096] In this embodiment, the data quality enhancement method for data space entity resolution is an enhancement framework or strategy, rather than an independent end-to-end model, nor an independent data quality enhancement model; The method collaboratively optimizes and trains with the existing entity resolution model through mechanisms such as dynamic noise injection, adversarial training, and contrastive learning to obtain enhanced positive samples and final newly generated positive samples, achieving data quality enhancement.

[0097] In this embodiment, the adversarial positive sample difficulty enhancement step realizes the adaptive improvement of positive sample difficulty through learnable noise generation and adversarial training, helps the entity resolution model learn more fine-grained features, and enhances the generalization ability.

[0098] In this embodiment, the positive sample generation steps with dynamic noise perturbation generate more high-quality positive samples under the premise of maintaining semantic consistency through dynamic noise perturbation based on similarity and combined with the information quantity maximization sampling strategy, balance the positive and negative sample ratios, reduce the model classification bias, and improve data diversity and model performance.

[0099] In this embodiment, the data quality enhancement method for data space entity resolution proposes an adversarial positive sample difficulty enhancement and positive sample generation technology based on dynamic noise perturbation for the data quality problems in the entity resolution task.

[0100] First, use the adversarial training mechanism to construct a positive sample difficulty enhancement method, dynamically increase the discrimination difficulty of positive samples through learnable noise, enable the model to learn more fine-grained features, and improve the generalization ability.

[0101] Secondly, the noise amplitude is dynamically adjusted based on the sample similarity to generate diverse candidate positive samples. Combining with the information maximization sampling strategy, the ratio of positive and negative samples is optimized to alleviate the problem of scarce positive samples, thereby improving the quality and diversity of the data.

[0102] Finally, this method is decoupled from existing models and has the characteristics of plug-and-play (i.e., flexible adaptation). It can be seamlessly integrated into existing entity resolution models to improve the robustness and matching accuracy of the models (entity resolution models), and is applicable to entity resolution tasks in multi-source heterogeneous environments.

[0103] Embodiment 2. Refer to Figure 2 To illustrate this embodiment, this embodiment describes the original dataset acquisition and feature extraction steps in the data quality enhancement method for data space entity resolution described in Embodiment 1:

[0104] In the original dataset acquisition and feature extraction steps, the original positive samples are generated using the feature extraction ability of the entity resolution model as follows:

[0105] Let the entity pair in the original dataset be (e, e′). The embedding representation (or embedding vector) of each entity in the entity pair is generated through the entity resolution model:

[0106] v l = Encoder(er), v r = Encoder(e′);

[0107] The embedding representations of each entity are concatenated to form the original positive sample:

[0108] s p = [v l ; v r ∈ R 2d , where d is the single-entity embedding dimension of the entity resolution model.

[0109] Embodiment 3. Refer to Figure 2 To illustrate this embodiment, this embodiment describes the adversarial positive sample difficulty enhancement step in the data quality enhancement method for data space entity resolution described above:

[0110] The adversarial positive sample difficulty enhancement step includes the following steps:

[0111] Step S2.1: Use a noise generation network to generate learnable noise:

[0112] η = σ(Wη′ + b)

[0113] where W ∈ R 2d×2d and b ∈ R 2d×2dis a learnable parameter; η′1 is uniformly distributed noise randomly generated within the interval [0, 0.1]; σ(·) is the Sigmoid function that maps the noise value and restricts it within the interval (0, 1); η is learnable noise that follows a uniform distribution within the interval [0, 1], and its dimension is the same as that of the original positive sample to ensure the dimensional consistency of noise injection;

[0114] Step S2.2: Inject the generated learnable noise into the original positive sample to obtain an enhanced positive sample:

[0115] s p ′ = s p + η;

[0116] Step S2.3: Define the difficulty loss function Optimize the entity resolution model parameters through the Minimax adversarial training framework. While minimizing the standard cross-entropy loss function of the entity resolution model maximize the difficulty loss function

[0117]

[0118] where γ is a hyperparameter that controls the weight of the difficulty loss function ;

[0119] The difficulty loss function is used to maximize the misclassification probability of the enhanced positive sample:

[0120]

[0121] where p is the matching probability of the entity resolution model for the enhanced positive sample, and θ h represents the parameter for processing the enhanced positive sample;

[0122] The standard cross-entropy loss function acts on the original positive sample and is used to optimize the classification performance of the entity resolution model for the original positive sample.

[0123] In this embodiment, by maximizing the difficulty loss function, the noise generation network can generate noise that makes the enhanced positive sample more likely to be misclassified.

[0124] In this embodiment, this adversarial optimization strategy (adversarial training) balances the complexity and stability of the model by ensuring that the introduction of noise remains at the optimal level, thereby improving the overall robustness of the entity resolution model.

[0125] Embodiment 4. Refer to Figure 2This embodiment describes the positive sample generation step of dynamic noise perturbation in the data quality enhancement method for data space entity parsing described in the above embodiment:

[0126] The positive sample generation step of the dynamic noise perturbation includes the following steps:

[0127] Step S3.1: Generate random noise and inject it into the original positive sample, and dynamically adjust the amplitude of the random noise based on the similarity of the embedding representations of the entity pairs to obtain the newly generated candidate positive samples:

[0128]

[0129] Where:

[0130] Δ is the generated random noise vector, which follows a uniform distribution in the interval [0,1], and its dimension is the same as that of the original positive sample to ensure the dimensional consistency of noise injection; is the newly generated candidate positive sample;

[0131] m is the amplitude of the random noise:

[0132] m = σ(v l ·v r )

[0133] When the similarity of the two entity embedding representations v l and v r is low, it indicates that the original positive sample already contains a lot of useful information, and injecting too much noise may produce incorrect positive samples. Therefore, a smaller random noise vector should be injected, that is, m → 0;

[0134] When the similarity of the two entity embedding representations v l and v r is high, a larger random noise vector should be injected, that is, m → 1, so as to improve the robustness of the model and increase the diversity of the data;

[0135] Step S3.2: In the contrastive learning framework, jointly optimize and train the entity parsing model through the classification loss function including the candidate positive samples and the contrastive loss function to synchronously constrain the semantic consistency and classification accuracy of the candidate positive samples;

[0136] The classification loss function including the candidate positive samples acts on the candidate positive samples to ensure that the model correctly classifies the candidate positive samples, compensates for the distribution shift that may be introduced by the generated samples, and enhances the generalization of the model;

[0137] The contrastive loss function To constrain the semantic consistency of candidate positive samples:

[0138]

[0139] Among them, s r represents the embedding randomly sampled from negative samples, which is used to constrain, through backpropagation, that the generated candidate positive samples are more similar to the original positive samples than to the negative samples, ensuring the semantic consistency of the generated candidate positive samples.

[0140] In this embodiment, the contrastive loss function can ensure that the distance between the newly generated candidate positive samples and the original positive samples is as small as possible, while the distance between them and the negative samples is as large as possible.

[0141] In this embodiment, m→0 means that m approaches 0 (such as taking values 0.2, 0.1, etc.), that is, a smaller random noise vector is injected.

[0142] In this embodiment, m→1 means that m approaches 1 (such as taking values 0.8, 0.9, etc.), that is, a larger random noise vector is injected.

[0143] Embodiment Five. Refer to Figure 2 To illustrate this embodiment, this embodiment is an illustration of the information quantity evaluation and screening step in the data quality enhancement method for data space entity parsing described in the above embodiment:

[0144] Step S4.1: Calculate the similarity between each candidate positive sample and the original positive sample to obtain its information quantity score:

[0145]

[0146] Information quantity score The lower it is, the higher the information content of the corresponding candidate positive sample contains;

[0147] Step S4.2: Select the candidate positive samples with low information quantity scores to construct the final newly generated positive samples:

[0148]

[0149] Among them, τ is a dynamic threshold, which is used to screen samples with low information quantity scores (i.e., high information content).

[0150] In this embodiment, sampling according to the information quantity score can ensure that the generated positive samples are diverse and challenging in each iteration, and finally obtain a set of diverse and high-quality positive samples (i.e., the final newly generated positive samples).

[0151] Embodiment Six. Refer to Figure 2This embodiment describes the overall optimization step in the data quality enhancement method for data space entity parsing described in the above embodiment:

[0152] The overall optimization step is as follows:

[0153] Construct the overall optimization objective function:

[0154]

[0155] Among them, is the standard cross-entropy loss function, is the positive sample difficulty loss function, is the classification loss including candidate positive samples, is the contrast loss function; γ, λ, and β are hyperparameters representing the weights of the loss terms;

[0156] Adjust the hyperparameters γ, λ, and β to ensure the balance of the loss terms during the optimization process.

[0157] Embodiment 7: To verify the effectiveness and generality of the above data quality enhancement method for data space entity parsing, this embodiment provides an experiment that combines the above data quality enhancement method for data space entity parsing with existing mainstream ER models:

[0158] The model parameters for executing the above data quality enhancement method for data space entity parsing are set as follows:

[0159] The overall model learning rate is default set to 1e-5, the batch size is default set to 16, dropout is set to 0.1, and the training period is set to 10.

[0160] The data of the datasets used for the experiment are shown in Table 1.

[0161] Table 1 Statistical data of the experimental datasets

[0162]

[0163] These datasets cover multiple fields. Each dataset consists of a set of candidate entity record pairs from two structured tables of the same schema, and all attributes in each dataset are used in the experiment. Among them, Walmart-Amazon2, DBLP-ACM2, and DBLP-Scholar2 are "dirty" datasets, which are derived from the Walmart-Amazon1, DBLP-ACM1, and DBLP-Scholar1 datasets by randomly moving the values of each entity attribute to other attributes in the same tuple with a probability of 50% to simulate data inaccuracies in the real world.

[0164] The entity-record pairs in each dataset are divided into a training set, a validation set, and a test set in a 3:1:1 ratio according to the standard of DeepMatcher.

[0165] The mainstream ER model methods participating in the experiment are as follows:

[0166] Seq2SeqMatcher model: It models the structural information in entity records by learning the embedding vectors of each attribute, models the ER task as a word-level sequence-to-sequence matching task, and effectively solves the problems of schema heterogeneity and noisy data by comparing words across attributes to achieve accurate ER decisions.

[0167] Ditto model: An ER method that utilizes a pre-trained language model based on the Transformer model. It converts the ER task into a sequence pair classification problem and performs ER decisions by fine-tuning the pre-trained language model.

[0168] HierGAT model: It combines the self-attention mechanism and the graph attention mechanism (Graph Attention Network, GAT) for the first time to identify distinctive words from attributes and find the most discriminative attributes, and models a hierarchical heterogeneous graph to describe entity records. It utilizes the interdependence between decisions of different entity-record pairs to obtain better matching performance and achieves leading results.

[0169] The results of this experiment are shown in Tables 2 to 8.

[0170] Table 2 represents the results of the model before and after combining this method on the Walmart-Amazon1 dataset

[0171]

[0172] Table 3 represents the results of the model before and after combining this method on the Amazon-Google dataset

[0173]

[0174] Table 4 represents the results of the model before and after combining this method on the DBLP-ACM1 dataset

[0175]

[0176]

[0177] Table 5 represents the results of the model before and after combining this method on the DBLP-Scholar1 dataset

[0178]

[0179] Table 6 represents the results of the model on the Walmart-Amazon2 dataset before and after incorporating this method.

[0180]

[0181] Table 7 represents the results of the model on the DBLP-ACM2 dataset before and after incorporating this method.

[0182]

[0183] Table 8 represents the results of the model on the DBLP-Scholar2 dataset before and after incorporating this method.

[0184]

[0185]

[0186] The experimental results show that after incorporating this method, all models exhibit significant improvements in accuracy, recall, and F1-score on almost all datasets. Meanwhile, it reflects the wide applicability and effectiveness of this method.

[0187] Meanwhile, in order to verify the influence of different hyperparameters γ, λ, and β values on the model performance in the data quality enhancement method for data space entity resolution described in the present invention, a hyperparameter experiment was carried out, and the experimental results are as Figure 4 shown, where:

[0188] Figure (a) shows the influence of the γ parameter on the model performance, and γ mainly controls the weight of the positive sample difficulty enhancement loss function (difficulty loss function) in the total loss function (or total optimization objective function). of the weight.

[0189] Figure (b) shows the influence of different λ values on the model performance. The λ parameter is mainly used to balance the classification loss of the newly generated positive samples (including the classification loss function of the new candidate positive samples ) and the weights of other loss terms in the total loss function.

[0190] Figure (c) shows the influence of the β hyperparameter on the model performance. β adjusts the weight of the contrast loss (contrast loss function ) in the total loss, thereby affecting the discrimination of positive and negative sample features.

[0191] The experimental results show that reasonably adjusting the hyperparameters γ, λ, and β is crucial for the model performance. As the value of γ increases, the model performance improves, verifying the effectiveness of the positive sample difficulty enhancement mechanism. However, when γ exceeds the threshold, the model may ignore other samples and reduce the classification accuracy, so it needs to be moderately controlled. λ affects the learning weight of new samples. When it is small, the model underutilizes new samples, and when it is too large, it will lead to excessive attention to new samples and neglect other tasks. Therefore, a moderate λ helps to balance between tasks and improve the overall performance. The β value adjusts the contrastive loss and plays a positive role in balancing the feature discrimination between newly generated samples and negative samples. However, if it is too large, the model will overly focus on the distance relationship between samples, affecting feature learning and classification effects. Therefore, reasonably selecting γ, λ, and β can ensure that the model takes into account new sample learning, task balance, and positive and negative sample discrimination, thereby optimizing the overall performance and robustness.

[0192] Embodiment 8. The data quality enhancement method for data space entity resolution described in the present invention can be entirely implemented by computer software. Therefore, correspondingly, the present invention also provides a data quality enhancement device for data space entity resolution, and the device includes the following modules:

[0193] Original dataset acquisition and feature extraction module: Acquire the original dataset and generate original positive samples by using the feature extraction ability of the entity resolution model;

[0194] Adversarial positive sample difficulty enhancement module: Use a noise generation network to generate learnable noise and inject it into the original positive samples to increase the discrimination difficulty of positive samples and obtain enhanced positive samples; Define a difficulty loss function, and optimize the parameters of the entity resolution model through the Minimax adversarial training framework. While minimizing the standard cross-entropy loss function of the entity resolution model, maximize the difficulty loss function, and the difficulty loss function is used to maximize the misclassification probability of the enhanced positive samples;

[0195] Dynamically noise-perturbed positive sample generation module: Generate random noise and inject it into the original positive samples, and dynamically adjust the amplitude of the random noise based on the similarity of the embedded representations of entity pairs to obtain newly generated candidate positive samples; Under the contrastive learning framework, jointly optimize and train the entity resolution model through the classification loss function and the contrastive loss function including candidate positive samples, and synchronously constrain the semantic consistency and classification accuracy of candidate positive samples;

[0196] Information quantity evaluation and screening module: Evaluate the information quantity of candidate positive samples to obtain an information quantity score; After sorting by the information quantity score, select candidate positive samples with higher information quantity to obtain the final newly generated positive samples;

[0197] Total optimization module: A total optimization objective function is constructed using a loss function for joint multi-objective optimization. The loss function includes a standard cross-entropy loss function, a difficulty loss function, a classification loss function including candidate positive samples, and a contrastive loss function. By adjusting the weights of the individual loss functions in the total optimization objective function, the balance of each loss term is ensured.

[0198] Embodiment Nine. This embodiment provides a computer-readable storage medium in which a computer program is stored. When the computer program runs, it executes the data quality enhancement method for data space entity parsing described in any one of the above.

[0199] Embodiment Ten. This embodiment provides a computer device that includes a memory and a processor. The memory is used to store executable instructions for the processor, and the processor is configured to execute the data quality enhancement method for data space entity parsing described in any one of the above by executing the executable instructions.

[0200] For the computer device or system provided in this embodiment, the hardware device in this part is of a general model and is not shown in the form of a diagram. The system includes a processor and a memory. The processor and the memory can be connected through a bus or other means. The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, as well as corresponding program instructions / modules. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, so as to implement the data quality enhancement method for data space entity parsing in the above method embodiments.

[0201] The memory can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, an enterprise intranet, a mobile communication network, and combinations thereof.

[0202] One or more modules are stored in the memory. When the processor executes, it performs the method steps in the embodiment. In this way, through the method, device, and process of the present invention, the invention purpose of the present invention can be achieved. The specific details of the above computer device can be understood by referring to the corresponding relevant descriptions and effects in the embodiment, and will not be elaborated here.

[0203] Those skilled in the art can understand that to implement all or part of the processes in the above-described embodiment methods, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0204] The above further describes the technical solutions provided by the present invention in several specific implementation manners to highlight the advantages and beneficial effects of the technical solutions provided by the present invention. However, the above-described several specific implementation manners are not used as a limitation to the present invention. Any reasonable modification and improvement of the present invention, reasonable combination and equivalent replacement of implementation manners, etc., within the spirit and principle of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for enhancing the data quality of data space entity parsing, characterized in that, The method includes the following steps: Original dataset acquisition and feature extraction step: Acquire the original dataset, and generate original positive samples by using the feature extraction ability of the entity resolution model; Adversarial positive sample difficulty enhancement step: Use a noise generation network to generate learnable noise and inject it into the original positive samples to increase the discrimination difficulty of the positive samples and obtain enhanced positive samples; Define a difficulty loss function, and optimize the parameters of the entity resolution model through the Minimax adversarial training framework. While minimizing the standard cross-entropy loss function of the entity resolution model, maximize the difficulty loss function, where the difficulty loss function is used to maximize the misclassification probability of the enhanced positive samples; Dynamically noisy perturbed positive sample generation step: Generate random noise and inject it into the original positive samples, and dynamically adjust the amplitude of the random noise based on the similarity of the embedded representations of entity pairs to obtain newly generated candidate positive samples; Under the contrastive learning framework, jointly optimize and train the entity resolution model through the classification loss function and the contrastive loss function including the candidate positive samples, and simultaneously constrain the semantic consistency and classification accuracy of the candidate positive samples; Information quantity evaluation and screening step: Evaluate the information quantity of the candidate positive samples to obtain information quantity scores; After sorting by the information quantity scores, select the candidate positive samples with higher information quantity to obtain the final newly generated positive samples; Total optimization step: Use a loss function to construct a total optimization objective function for multi-objective joint optimization; The loss function includes the standard cross-entropy loss function, the difficulty loss function, the classification loss function including the candidate positive samples, and the contrastive loss function; By adjusting the weights of each loss function in the total optimization objective function, ensure the balance of each loss term.

2. The data quality enhancement method for data space entity parsing according to claim 1, wherein In the original dataset acquisition and feature extraction step, generating the original positive samples by using the feature extraction ability of the entity resolution model is as follows: Let the entity pairs in the original dataset be (e, e'), and generate the embedded representations of each entity in the entity pair through the entity resolution model: v l = Encoder(e), v r = Encoder(e'); Concatenate the embedded representations of each entity to form the original positive samples: s p = [v l ; v r ∈ R 2d , where d is the single-entity embedding dimension of the entity resolution model.

3. The method for enhancing data quality in parsing data space entities according to claim 2, wherein In the adversarial positive sample difficulty enhancement step: The learnable noise is: η = σ(Wη′ + b) where \(W\in\mathbb{R}\) 2d×2d and \(b\in\mathbb{R}\) 2d×2d are learnable parameters; \(\eta'\) is uniformly distributed noise randomly generated in the interval \([0, 0.1]\); \(\sigma(\cdot)\) is the Sigmoid function that maps the noise value and restricts it within the interval \((0, 1)\); \(\eta\) follows a uniform distribution in the interval \([0, 1]\) and has the same dimension as the original positive samples to ensure dimensional consistency of noise injection; The enhanced positive samples are: s′ p = s p + η; The difficulty loss function is where p is the matching probability of the entity resolution model for the enhanced positive sample, and θ h represents the parameter for processing the enhanced positive sample; The standard cross-entropy loss function is used to optimize the classification performance of the entity resolution model for the original positive samples; The step of maximizing the difficulty loss function while minimizing the standard cross-entropy loss function of the entity resolution model is: Among them, γ is a hyperparameter that controls the weight of the difficulty loss function. The weight.

4. The data quality enhancement method for data space entity parsing according to claim 3, wherein In the dynamically noisy perturbed positive sample generation step: The newly generated candidate positive samples are: Where: Δ is the generated random noise vector, which follows a uniform distribution in the interval [0, 1], and its dimension is the same as that of the original positive samples to ensure the dimensional consistency of noise injection; m is the amplitude of the random noise: m = σ(v l ·v r ) When the similarity between the embedded representations v l and v r is low, let m → 0; When the similarity between the two entity embedding representations v l and v r is relatively high, let m → 1; The classification loss function including candidate positive samples is used to ensure that the model correctly classifies new candidate positive samples; The contrastive loss function is used to constrain the semantic consistency of the new candidate positive samples: Among them, s r represents the embedding randomly sampled from negative samples, which is used to constrain, through backpropagation, the similarity between the generated candidate positive samples and the original positive samples to be higher than that with negative samples, ensuring the semantic consistency of the generated candidate positive samples.

5. The method for enhancing data quality of data space entity parsing according to claim 4, wherein In the information quantity evaluation and screening step: The information quantity score of each candidate positive sample is: Information quantity score The lower it is, the higher the information quantity contained in the corresponding embedded representation of the candidate positive sample is; The final newly generated positive samples are: Where τ is a dynamic threshold.

6. The method for enhancing data quality of data space entity parsing according to claim 1, wherein In the integrated loss function total optimization step: The total optimization objective function is: Where γ, λ, and β are hyperparameters representing the weights of the loss terms; Adjust the hyperparameters γ, λ, and β to ensure the balance of the loss terms during the optimization process.

7. Data quality enhancement device for data space entity resolution, characterized in that The device includes the following modules: Original dataset acquisition and feature extraction module: Acquire the original dataset and generate original positive samples using the feature extraction ability of the entity resolution model; Adversarial positive sample difficulty enhancement module: Use a noise generation network to generate learnable noise and inject it into the original positive samples to increase the discrimination difficulty of the positive samples and obtain enhanced positive samples; Define a difficulty loss function, and optimize the entity resolution model parameters through the Minimax adversarial training framework. While minimizing the standard cross-entropy loss function of the entity resolution model, maximize the difficulty loss function, where the difficulty loss function is used to maximize the misclassification probability of the enhanced positive samples; Dynamically noise-perturbed positive sample generation module: Generate random noise and inject it into the original positive samples, and dynamically adjust the amplitude of the random noise based on the similarity of the embedded representations of entity pairs to obtain newly generated candidate positive samples; Under the contrastive learning framework, jointly optimize and train the entity resolution model through the classification loss function and the contrastive loss function including the candidate positive samples, and synchronously constrain the semantic consistency and classification accuracy of the candidate positive samples; Information quantity evaluation and screening module: Evaluate the information quantity of the candidate positive samples to obtain an information quantity score; After sorting by the information quantity score, select the candidate positive samples with higher information quantity to obtain the final new positive samples; Total optimization module: Construct a total optimization objective function using a loss function for multi-objective joint optimization; The loss function includes the standard cross-entropy loss function, the difficulty loss function, the classification loss function including the candidate positive samples, and the contrastive loss function; By adjusting the weights of each loss function in the total optimization objective function, ensure the balance of each loss term.

8. A computer device, comprising: A processor and a memory, characterized in that the memory is used to store executable instructions of the processor, and the processor is configured to execute the data quality enhancement method for data space entity resolution according to any one of claims 1-6 by executing the executable instructions.

9. A computer storage medium, characterized in that, A computer program is stored in the storage medium, and when the computer program runs, it executes the data quality enhancement method for data space entity resolution according to any one of claims 1-6.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the data quality enhancement method for data space entity resolution according to any one of claims 1-6 are implemented.