A Cross-Source Data Quality Defect Detection Method and System Based on Deep Learning
Through the cross-origin data quality defect detection method based on deep learning, the scanning and repair model is used to monitor, analyze and repair cross-origin data, data quality problems are solved, the accuracy and consistency of data integration are improved, and a reliable data foundation is laid for the smart government platform.
Patent Information
- Application Number
- CN202510510260.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-23
AI Technical Summary
There are data quality problems in the acquisition and integration of cross-origin data, including high redundancy rate, many missing values, inconsistent data standards and high proportion of unstructured data, resulting in difficulty in cross-departmental field mapping.
The cross-origin data quality defect detection method based on deep learning is adopted to scan and analyze the cross-origin data sets by building a scanning model, generate data quality problem reports, and use the data repair model to repair defective data, including precise repair using the rule engine and the generator-discriminator model.
It has achieved efficient classification and repair of cross-origin data quality problems, reduced decision-making errors, provided a reliable data foundation for smart government affairs platforms, and improved the accuracy and consistency of data integration.
Smart Images

Figure CN120067767B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a cross-source data quality defect detection method and system based on deep learning. Background Art
[0002] In the big data era, government affairs data often needs to integrate information from multiple different data sources to obtain more comprehensive and accurate insights.
[0003] The acquisition and integration of cross-source data face many challenges, among which the data quality problem is particularly prominent. Data from different data sources may have various defects due to inconsistent collection standards, data entry errors, etc.
[0004] The inventors found and summarized during the development of the intelligent government affairs platform that there is a relatively high redundancy rate and missing values in multi-source heterogeneous data; at the same time, the inconsistent data standards make it difficult to map fields across departments, and the proportion of unstructured data is also relatively high.
[0005] How to solve the above problems is a technical problem that those skilled in the art need to overcome. Summary of the Invention
[0006] To at least partially solve the above technical problems, this application provides a cross-source data quality defect detection method and system based on deep learning.
[0007] In a first aspect, a cross-source data quality defect detection method based on deep learning provided by this application adopts the following technical solution.
[0008] A cross-source data quality defect detection method based on deep learning includes:
[0009] Using the constructed scanning model to scan the cross-source data set; the scanning model monitors and analyzes the data in the cross-source data set during the scanning process; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems;
[0010] Classifying the data quality problems according to the scanning result of the scanning model and summarizing the data quality problems of the same type;
[0011] Generating a data quality problem report based on a preset template according to the classification and summary results; and,
[0012] Inputting the cross-source data with quality defects into the data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data.
[0013] Optionally, the method further includes:
[0014] Screen the first defective data and the second defective data based on the classification results of defective cross - source data; the repair difficulty of the first defective data is less than that of the second defective data; the first defective data can be repaired by a configured rule engine; the rule engine contains several data rules; the data rules include data integrity rules and data consistency rules;
[0015] Input the first defective data into the rule engine; the rule engine matches the first defective data one by one based on the included data rules; for the first defective data that matches the corresponding data rule, the rule engine repairs the first defective data based on the processing method defined by the corresponding data rule;
[0016] Input the second defective data into the data repair model.
[0017] Optionally, scan the cross - source data set using the constructed scanning model, including:
[0018] For the text data in the cross - source data set, use a tokenizer to convert the text into a vector representation;
[0019] Construct the vectorized data into a sequence in order; for time - series data, arrange it in chronological order; for non - time - series data, arrange it according to the correlation relationship of the data;
[0020] Use the sequence as input features and input them into the scanning model;
[0021] Among them, using a tokenizer to convert the text in the cross - source data set into a vector representation includes:
[0022] Based on the tokenizer, tokenize the text data in the cleaned cross - source data set to split it into several tokens;
[0023] Add a first marker at the start position of each token after tokenizing the text data, and add a second marker at the end position of the token; the first marker is used as an aggregated representation of the semantics of the entire sentence; the second marker is used to distinguish different sentences or text paragraphs;
[0024] Generate corresponding token - type embeddings and position embeddings for each token; the token - type embeddings are used to distinguish different text paragraphs or sentences; the position embeddings are used to represent the position information of the token in the sequence; add the token - type embeddings, position embeddings and the word embeddings of the token to obtain the final input representation; input the obtained input representation into the pre - trained language model; the pre - trained language model processes the input data through multiple layers of encoders for feature extraction.
[0025] Optionally, the pre-trained language model processes the input data through multiple layers of encoders for feature extraction, including:
[0026] S401. For each token in the token sequence input to each layer of the encoder, multiply its corresponding input representation by three learnable weight matrices respectively , , to generate a query vector Q, a key vector K, and a value vector V;
[0027] S402. Calculate the dot product of each query vector Q and all key vectors K to obtain the original attention scores and obtain the attention weights based on the original attention scores;
[0028] S403. Perform weighted summation on the value vectors based on the attention weights to obtain the updated token representation of each token ;
[0029] S404. Input the updated token representation into a feed-forward neural network; the feed-forward neural network includes two linear layers and a non-linear activation function; the updated token representation is multiplied by the weight matrix of the first linear layer and added with the first bias vector and the intermediate result is obtained through the non-linear activation function; the intermediate result is multiplied by the weight matrix of the second linear layer and added with the second bias vector to obtain the output of the feed-forward neural network ;
[0030] S405. Add the input representation to the token representation and add the output of the feed-forward neural network to obtain the output after residual connection and use it as the output vector of the current layer of the encoder;
[0031] Each layer of the multiple layers of encoders sequentially executes steps S401 - S405 on the input data; among them, the output of the previous layer of the encoder is used as the input of the next layer of the encoder; after the processing of each layer of the encoder is completed, record the output vector of each token at this layer ; the output vector is the hidden state vector of the token at this layer; after processing all the encoding layers, the pre-trained language model outputs the set of hidden state vectors of each token at all layers;
[0032] From the hidden state vectors of each token output by the pre-trained language model, locate the hidden state vector corresponding to the first token; during the pre-training process, the vector corresponding to the first token will continuously aggregate the semantic information of the entire sentence; select the hidden state vector corresponding to the first token as the semantic representation of the text data in the cross-source dataset.
[0033] Optionally, the method further includes:
[0034] After the data repair model completes the repair of the cross-source data with quality defects, collect the repaired cross-source data;
[0035] Input the collected repaired data into the scanning model again; the scanning model outputs the secondary scanning result of the repaired data; based on the secondary scanning result, determine whether there are still data quality problems; if there are still data quality problems, input them into the data repair model again.
[0036] Optionally, the generation method of the data repair model includes:
[0037] Construct an initial repair model including a generator and a discriminator; the generator is used to receive the cross-source input with quality defects as input and output the repaired data sample; the discriminator is used to receive the repaired sample data output by the generator and output the probability that the sample data has been repaired;
[0038] Initialize the generator and the discriminator;
[0039] Input the cross-source data with quality defects into the generator; the generator processes the input cross-source data with quality defects according to the current parameter settings to generate the repaired sample data;
[0040] Input the repaired sample data generated by the generator and the corresponding real data into the discriminator respectively;
[0041] The discriminator calculates the probability that the sample data and the real data have been repaired respectively; calculates the loss function of the discriminator according to the probabilities of the two; updates the parameters of the discriminator according to the loss function of the discriminator; repeat the training until the confidence of the initial repair model is greater than the preset value to obtain the repair model.
[0042] Optionally, the method further includes:
[0043] During the training process, the generator and the discriminator are alternately trained. The goal of the generator is to generate repaired data samples that can deceive the discriminator, and the goal of the discriminator is to accurately distinguish between real data and the repaired data samples generated by the generator;
[0044] Update the parameters of the generator and discriminator according to the output result of the discriminator to improve the authenticity of the repaired data samples generated by the generator and the discrimination ability of the discriminator.
[0045] In a second aspect, a cross-source data quality defect detection system based on deep learning provided by the present application adopts the following technical solutions.
[0046] A cross-source data quality defect detection system based on deep learning includes:
[0047] A first processing module, configured to: scan a cross-source data set using a constructed scanning model; monitor and analyze the data in the cross-source data set during the scanning process; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems.
[0048] A second processing module, configured to: classify the data quality problems according to the scanning result of the scanning model and summarize the data quality problems of the same type.
[0049] A third processing module, configured to: generate a data quality problem report based on a preset template according to the classified and summarized results.
[0050] A fourth processing module, configured to: input the cross-source data with quality defects into a data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data.
[0051] Compared with the prior art, the beneficial effects of the present invention include: first, the scanning model scans the cross-source data set, classifies and summarizes the data quality problems based on the scanning result, and then generates a data quality problem report based on a preset template; inputting the cross-source data with quality defects into the data repair model and using the learned repair strategy for repair forms a complete and efficient data quality defect detection and repair closed loop, which can effectively reduce decision-making errors caused by data quality problems and lay a data foundation for the development of the intelligent government affairs platform. Description of the Drawings
[0052] Figure 1 is a flowchart of a cross-source data quality defect detection method based on deep learning according to an embodiment of the present application;
[0053] Figure 2 is a system block diagram of a cross-source data quality defect detection system based on deep learning according to an embodiment of the present application;
[0054] In the figure, 201, the first processing module; 202, the second processing module; 203, the third processing module; 204, the fourth processing module. Detailed Embodiments
[0055] The following further describes the present application in conjunction with the accompanying Figure 1-2 drawings and specific embodiments:
[0056] An embodiment of the present application discloses a cross-source data quality defect detection method based on deep learning, including the following steps:
[0057] Step S101: Scan the cross-source data set using the constructed scanning model; the scanning model monitors and analyzes the data in the cross-source data set during the scanning process; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems. Specifically, the scanning model is a model constructed based on deep learning, which contains a multi-layer neural network structure inside. During the scanning process, it can identify possible quality defects in the data according to the learned feature patterns and output these problems in a specific form. The cross-source data set refers to a data set composed of data from multiple different data sources, and these data sources may come from different business systems, different databases, or different device acquisition ends. Due to the diversity of data sources, there are often differences in the data formats, data standards, and data meanings of the cross-source data set. Data quality problems refer to various situations where the data in the cross-source data set does not meet the expected quality standards. Common data quality problems include: data accuracy problems, that is, the data value does not match the true value; data integrity problems, such as missing data records and empty field values; data redundancy problems.
[0058] Step S102: Classify the data quality problems according to the scanning result of the scanning model and summarize the data quality problems of the same type.
[0059] Step S103: Generate a data quality problem report based on the classified and summarized results according to a preset template.
[0060] Step S104: Input the cross-source data with quality defects into the data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data. Specifically, the data repair model is a model used to repair the cross-source data with quality defects. It learns a large number of existing data repair cases, and repairs the cross-source data with defects for the data quality problems detected by the scanning model, so that the data reaches a state that meets the quality standards and outputs the repaired cross-source data.
[0061] Specifically, the first-pass scanning model scans cross-source datasets, classifies and summarizes data quality problems based on the scanning results, and then generates a data quality problem report based on a preset template, providing intuitive and standardized information display for data managers and helping them quickly understand the data quality status. Inputting cross-source data with quality defects into the data repair model and using the learned repair strategies to repair it forms a complete and efficient closed-loop for data quality defect detection and repair, which can effectively reduce decision-making errors caused by data quality problems and lay a data foundation for the development of the intelligent government affairs platform.
[0062] The following is an example: The intelligent government affairs platform needs to integrate data sources from multiple different departments, such as population information data of the civil affairs department, tax payment data of the tax department, household registration and public security data of the public security department, etc. The formats, standards, and update frequencies of these cross-source datasets are different. There may be inconsistencies between the resident marriage status data registered by the civil affairs department and the marriage status information in the household registration system of the public security department. The constructed scanning model scans the cross-source datasets in the intelligent government affairs platform. Taking tax data and enterprise industrial and commercial registration data as an example, the scanning model can monitor whether the enterprise's tax declaration data matches information such as the business scope and registered capital in the industrial and commercial registration. If the scanning finds that the declared tax amount of a certain enterprise is far from the tax amount payable according to its industrial and commercial registration information, this may be a data accuracy problem. According to the results of the scanning model, classify and summarize data quality problems. In the intelligent government affairs platform, data quality problems can be divided into information accuracy problems (such as the above-mentioned inconsistency between tax data and industrial and commercial registration information), data integrity problems (such as the lack of insurance records of some residents in the social security system), data consistency problems (such as different departments registering different addresses for the same enterprise), etc. Through classification and summarization, the distribution of various data quality problems can be clearly presented. Generate a data quality problem report based on a preset template. For example, the report details the quality problems existing in the data of different departments, such as the number and proportion of incorrect date formats of some residents' birth dates in the population information of the civil affairs department; the list of enterprises with abnormal declarations and the involved amounts in the enterprise tax declaration data of the tax department. Input cross-source data with quality defects into the data repair model, and the data repair model can repair the cross-source data according to the data repair model, such as repairing the date format.
[0063] As a specific implementation manner of a cross-source data quality defect detection method based on deep learning, the method further includes:
[0064] Screen the first defective data and the second defective data based on the classification results of the defective cross-source data; the repair difficulty of the first defective data is less than that of the second defective data; the first defective data can be repaired by a configured rule engine; the rule engine includes a number of data rules; the data rules include data integrity rules and data consistency rules;
[0065] Input the first defective data into the rule engine; the rule engine matches the first defective data one by one based on the included data rules; for the first defective data that matches the corresponding data rule, the rule engine repairs the first defective data based on the processing method defined by the corresponding data rule;
[0066] Input the second defective data into the data repair model.
[0067] Specifically, in the development of the intelligent government affairs platform, data sources from different departments are aggregated to form a cross-source data set, and data quality problems are complex and diverse. Based on the classification results of the defective cross-source data, the first defective data and the second defective data are obtained. Suppose in the population information data, some resident ID number records have simple format errors. This type of data belongs to the first defective data because its repair difficulty is relatively small and it can be repaired by a configured rule engine. The rule engine includes data integrity rules such as requiring the ID number to be 18 digits; and data consistency rules, for example, the date of birth corresponding to the ID number should be consistent with the date of birth registered in the household registration system. Input this type of first defective data into the rule engine, and the rule engine matches it one by one based on the included data rules. When it is found that the ID number format is incorrect, the rule engine completes the repair of the first defective data according to the processing method defined by the corresponding rule, such as automatically supplementing the correct numbers at the missing digits or correcting the format error. For some complex data quality problems that are difficult to repair by the rule engine, such as the semantic fuzzy differences in the descriptions of the business scope of enterprises by different departments, this type of data belongs to the second defective data. Input the second defective data into the data repair model for repair, so as to improve the quality of the cross-source data of the intelligent government affairs platform and provide a more reliable data basis for government affairs decision-making and services.
[0068] As a specific implementation manner of a cross-source data quality defect detection method based on deep learning, use the constructed scanning model to scan the cross-source data set, including:
[0069] For the text data in the cross-source data set, use a tokenizer to convert the text into a vector representation;
[0070] Construct the vectorized data into a sequence in order; for time series data, arrange it in chronological order; for non-time series data, arrange it according to the correlation relationship of the data;
[0071] Use the sequence as the input feature and input it into the scanning model;
[0072] Among them, for the text data in the cross-source dataset, use a tokenizer to convert the text into a vector representation, including:
[0073] Based on the tokenizer, tokenize the text data in the cleaned cross-source dataset to split it into several tokens;
[0074] Add a first marker at the start position of each token after tokenizing the text data, and add a second marker at the end position of the token; the first marker is used as an aggregated representation of the semantics of the entire sentence; the second marker is used to distinguish different sentences or text paragraphs;
[0075] Generate corresponding token type embeddings and position embeddings for each token; the token type embeddings are used to distinguish different text paragraphs or sentences; the position embeddings are used to represent the position information of the token in the sequence; add the token type embeddings, position embeddings and the word embeddings of the token to obtain the final input representation; input the obtained input representation into the pre-trained language model; the pre-trained language model processes the input data through multiple layers of encoders for feature extraction.
[0076] Specifically, for a large amount of text data in the intelligent government affairs platform, such as policy documents, service guides, resident feedback, etc., first use a tokenizer to split the cleaned text into tokens, which can deconstruct continuous text information into more easily processed units. Add specific markers at the start and end positions of the tokens, so that the model can clearly distinguish different sentences and paragraphs and accurately grasp the text structure. Generate type embeddings and position embeddings for the tokens and add them to the word embeddings to obtain the final input representation; input it into the pre-trained language model and process it through multiple layers of encoders, and use the general language knowledge learned by the model on a large amount of text to mine the key features in the text data. Thereby improving the understanding ability of the scanning model for text data, enabling it to accurately identify possible data quality defects in the text, such as semantic ambiguity and inconsistent expressions in policy documents.
[0077] As a specific implementation of a cross-source data quality defect detection method based on deep learning, the pre-trained language model processes the input data through multiple layers of encoders for feature extraction, including:
[0078] S401. For each token in the token sequence input to each layer of the encoder, its corresponding input representation is respectively multiplied by three learnable weight matrices , , Multiply them to generate a query vector Q, a key vector K, and a value vector V;
[0079] S402. Calculate the dot product of each query vector Q and all key vectors K to obtain the original attention scores, and obtain the attention weights based on the original attention scores;
[0080] S403. Perform a weighted sum of the value vectors based on the attention weights to obtain the updated token representation for each token ;
[0081] S404. Input the updated token representation into a feed-forward neural network; the feed-forward neural network includes two linear layers and a non-linear activation function; the updated token representation is multiplied by the weight matrix of the first linear layer and added with the first bias vector and passed through the non-linear activation function to obtain an intermediate result; the intermediate result is multiplied by the weight matrix of the second linear layer and added with the second bias vector to obtain the output of the feed-forward neural network ;
[0082] S405. Add the input representation to the token representation and add the result to the output of the feed-forward neural network to obtain the output after residual connection and use it as the output vector of the current layer encoder;
[0083] Each layer of the multi-layer encoder sequentially executes steps S401 - S405 on the input data; where, the output of the previous layer encoder is used as the input of the next layer encoder; after the processing of each layer encoder is completed, record the output vector of each token at this layer ; The output vector is the hidden state vector of the token at this layer; after processing all the encoding layers, the pre-trained language model outputs the set of hidden state vectors of each token at all layers;
[0084] Among the hidden state vectors of each token output by the pre-trained language model, locate the hidden state vector corresponding to the first token; during the pre-training process, the vector corresponding to the first token will continuously aggregate the semantic information of the entire sentence; select the hidden state vector corresponding to the first token as the semantic representation of the text data in the cross-source dataset.
[0085] Specifically, the token input representation is multiplied by three learnable weight matrices to generate a query vector Q, a key vector K, and a value vector V. In S402, the original attention scores are calculated to obtain attention weights, enabling the model to allocate attention based on the correlation between tokens. In S403, the value vectors are weighted and summed based on the attention weights to update the token representation. In S404, the feed-forward neural network processes the data, introducing non-linearity through two linear layers and a non-linear activation function, enhancing the model's expressive power and enabling it to capture more complex semantic patterns in the text. The residual connection in S405 addresses the problem of vanishing gradients in the training of deep networks, ensuring the effective transmission of information and enabling the model to construct deeper network structures to learn more advanced semantic features. The multi-layer encoder processes the input data sequentially, extracting features further on the basis of the previous layer and continuously deepening the understanding of the text. The hidden state vectors of each token at each layer are recorded, providing rich information for a comprehensive analysis of the text. Finally, the hidden state vector corresponding to the first token is located and selected as the semantic representation of the text data because this vector aggregates the semantic information of the entire sentence during the pre-training process and can accurately represent the core semantics of the text. In the application scenario of the intelligent government affairs platform, this helps the scanning model to more accurately identify quality defects in text data, such as expression ambiguities and logical contradictions in policy documents.
[0086] As one implementation of a cross-source data quality defect detection method based on deep learning, the method further includes:
[0087] After the data repair model completes the repair of cross-source data with quality defects, the repaired cross-source data is collected;
[0088] The collected repaired data is input into the scanning model again; the scanning model outputs the secondary scanning result of the repaired data; based on the secondary scanning result, it is determined whether there are still data quality problems; if there are still data quality problems, it is input into the data repair model again.
[0089] As one implementation of a cross-source data quality defect detection method based on deep learning, the method for generating a data repair model includes:
[0090] Construct an initial repair model including a generator and a discriminator; the generator is used to receive cross-source inputs with quality defects as input and output repaired data samples; the discriminator is used to receive the repaired sample data output by the generator and output the probability that the sample data has been repaired;
[0091] Initialize the generator and the discriminator;
[0092] Input cross-source data with quality defects into a generator; the generator processes the input cross-source data with quality defects according to the current parameter settings to generate repaired sample data;
[0093] Input the repaired sample data generated by the generator and the corresponding real data into a discriminator respectively;
[0094] The discriminator calculates the probabilities that the sample data and the real data have been repaired respectively; calculates the loss function of the discriminator according to the two probabilities; updates the parameters of the discriminator according to the loss function of the discriminator; repeats the training until the confidence of the initial repair model is greater than a preset value to obtain a repair model.
[0095] As one implementation manner of a cross-source data quality defect detection method based on deep learning, the method further includes:
[0096] During the training process, the generator and the discriminator are alternately trained. The goal of the generator is to generate repaired data samples that can deceive the discriminator, and the goal of the discriminator is to accurately distinguish real data from the repaired data samples generated by the generator;
[0097] According to the output result of the discriminator, update the parameters of the generator and the discriminator to improve the authenticity of the repaired data samples generated by the generator and the discrimination ability of the discriminator.
[0098] This application also provides a cross-source data quality defect detection system based on deep learning, including:
[0099] A first processing module 201, configured to: scan a cross-source data set using a constructed scanning model; the scanning model monitors and analyzes the data in the cross-source data set during the scanning process; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems;
[0100] A second processing module 202, configured to: classify the data quality problems according to the scanning result of the scanning model and summarize the data quality problems of the same type;
[0101] A third processing module 203, configured to: generate a data quality problem report based on a preset template according to the classification and summarization results;
[0102] A fourth processing module 204, configured to: input cross-source data with quality defects into a data repair model; the data repair model repairs the defective cross-source data and outputs the repaired cross-source data.
[0103] It should be noted that the above embodiments are only used to illustrate the present application and do not limit the technical solutions described in the present application. Although the present specification has described the present application in detail with reference to the above embodiments, those of ordinary skill in the art should understand that those skilled in the art can still modify the present application or make equivalent substitutions, and all technical solutions and their improvements that do not depart from the spirit and scope of the present application should be covered within the scope of the claims of the present application.
Claims
1. A cross-source data quality defect detection method based on deep learning, characterized in that, Including: Scanning a cross-source dataset using the constructed scanning model; During the scanning process, the scanning model monitors and analyzes the data in the cross-source dataset; When there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems; Classifying the data quality problems according to the scanning result of the scanning model and summarizing the data quality problems of the same type; Generating a data quality problem report based on the classified and summarized results and a preset template; And, Inputting the cross-source data with quality defects into the data repair model; the data repair model repairs the defective cross-source data and outputs the repaired cross-source data; Scanning a cross-source dataset using the constructed scanning model includes: For the text data in the cross-source dataset, using a tokenizer to convert the text into a vector representation; Constructing the vectorized data into a sequence in order; for time series data, arranging it in chronological order; for non-time series data, arranging it according to the correlation relationship of the data; Using the sequence as an input feature and inputting it into the scanning model; The method for generating the data repair model includes: constructing an initial repair model including a generator and a discriminator; The method further includes: Screening based on the classification result of the defective cross-source data to obtain first defective data and second defective data; the repair difficulty of the first defective data is less than that of the second defective data; the first defective data can be repaired by a configured rule engine; the rule engine includes several data rules; the data rules include data integrity rules and data consistency rules; Inputting the first defective data into the rule engine; the rule engine matches the first defective data one by one based on the included data rules; for the first defective data that matches the corresponding data rule, the rule engine repairs the first defective data based on the processing method defined by the corresponding data rule; Inputting the second defective data into the data repair model.
2. The cross-source data quality defect detection method based on deep learning according to claim 1, characterized in that, For the text data in the cross-source dataset, using a tokenizer to convert the text into a vector representation, including: Based on the tokenizer, tokenizing the text data in the cleaned cross-source dataset to split it into several tokens; Adding a first marker at the start position of each token after tokenizing the text data, and adding a second marker at the end position of the token; the first marker is used as an aggregated representation of the semantics of the entire sentence; the second marker is used to distinguish different sentences or text paragraphs; Generating corresponding token type embeddings and position embeddings for each token; the token type embeddings are used to distinguish different text paragraphs or sentences; the position embeddings are used to represent the position information of the token in the sequence; adding the token type embeddings, position embeddings and the word embeddings of the token to obtain the final input representation; inputting the obtained input representation into a pre-trained language model; the pre-trained language model processes the input data through multiple layers of encoders for feature extraction.
3. The cross-source data quality defect detection method based on deep learning according to claim 2, characterized in that The pre-trained language model processes the input data through multiple layers of encoders for feature extraction, including: S401. For each token in the token sequence input to each layer of the encoder, multiply its corresponding input representation separately by three learnable weight matrices , , to generate a query vector Q, a key vector K, and a value vector V; S402. Calculate the dot product of each query vector Q and all key vectors K to obtain the original attention scores, and obtain the attention weights based on the original attention scores; S403. Weighted-sum the value vectors based on the attention weights to obtain the updated token representations for each token ; S404. Input the updated token representation into a feedforward neural network; the feedforward neural network includes two linear layers and a non-linear activation function; the updated token representation is multiplied by the weight matrix of the first linear layer and added with the first bias vector and an intermediate result is obtained through the non-linear activation function; the intermediate result is multiplied by the weight matrix of the second linear layer and added with the second bias vector to obtain the output of the feedforward neural network ; S405. Add the said input representation to the said token representation and add the result to the output of the feedforward neural network to obtain the output after residual connection which is used as the output vector of the current layer encoder; Each layer of the multi-layer encoder sequentially performs steps S401 - S405 on the input data; among them, the output of the previous layer encoder is used as the input of the next layer encoder; after the processing of each layer encoder is completed, the output vector of each token at this layer is recorded ; The output vector is the hidden state vector of the token at this layer; after processing all the encoding layers, the pre-trained language model outputs the set of hidden state vectors of each token at all layers; Among the hidden state vectors of each token output by the pre-trained language model, locate the hidden state vector corresponding to the first token; during the pre-training process, the vector corresponding to the first token will continuously aggregate the semantic information of the entire sentence; select the hidden state vector corresponding to the first token as the semantic representation of the text data in the cross-source dataset.
4. A cross-source data quality defect detection method based on deep learning according to claim 3, characterized in that The method further includes: After the data repair model completes the repair of the cross-source data with quality defects, collect the repaired cross-source data; Input the collected repaired data into the scanning model again; the scanning model outputs the secondary scanning result of the repaired data; based on the secondary scanning result, determine whether there are still data quality problems; if there are still data quality problems, input them into the data repair model again.
5. A cross-source data quality defect detection method based on deep learning according to claim 4, characterized in that The generator is used to receive the cross-source input with quality defects as input and output the repaired data sample; the discriminator is used to receive the repaired sample data output by the generator and output the probability that the sample data has been repaired; The method for generating the data repair model further includes: Initialize the generator and the discriminator; Input the cross-source data with quality defects into the generator; the generator processes the input cross-source data with quality defects according to the current parameter settings to generate the repaired sample data; Input the repaired sample data generated by the generator and the corresponding real data into the discriminator respectively; The discriminator calculates the probabilities that the sample data and the real data have been repaired respectively; calculate the loss function of the discriminator according to the probabilities of the two; update the parameters of the discriminator according to the loss function of the discriminator; repeat the training until the confidence of the initial repair model is greater than the preset value to obtain the repair model.
6. A cross-source data quality defect detection method based on deep learning according to claim 5, characterized in that, The method further includes: During the training process, the generator and the discriminator are alternately trained. The goal of the generator is to generate repaired data samples that can deceive the discriminator, and the goal of the discriminator is to accurately distinguish between real data and the repaired data samples generated by the generator; According to the output result of the discriminator, update the parameters of the generator and the discriminator to improve the authenticity of the repaired data samples generated by the generator and the discrimination ability of the discriminator.
7. A cross-source data quality defect detection system based on deep learning, which is used to implement the method described in any one of claims 1-6, characterized in that It includes: The first processing module is used to: scan the cross-source dataset using the constructed scanning model; The scanning model monitors and analyzes the data in the cross-source dataset during the scanning process; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems; The second processing module is used to: classify the data quality problems according to the scanning result of the scanning model and summarize the data quality problems of the same type; The third processing module is used to: generate a data quality problem report based on the classification and summary results and a preset template; The fourth processing module is used to: input the cross-source data with quality defects into the data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data.
Citation Information
Patent Citations
Abnormal data detection and restoration method for electric power measurement system
CN115238563A
Pipeline interior detection method and system based on multi-sensor fusion
CN119150175A