Cross-source data quality defect detection method and system based on deep learning
Through the cross-origin data quality defect detection method based on deep learning, quality problems in cross-origin data are identified and repaired, and the problem of poor data quality is solved, effectively improving data quality and a reliable data foundation of the smart government platform are achieved.
Patent Information
- Application Number
- CN202510510260.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The acquisition and integration of cross-origin data faces data quality problems, including data redundancy, missing values, inconsistent standards and high proportion of unstructured data, making it difficult to ensure data quality.
The cross-origin data quality defect detection method based on deep learning is adopted to monitor and analyze cross-origin data by building a scanning model, identify data quality problems, and classify, summarize and generate reports. At the same time, the data repair model is used to repair defective data, forming a closed loop of data quality defect detection and repair.
Effectively identify and classify data quality problems, generate detailed reports, and improve data quality through data repair models, reduce decision-making errors caused by data quality problems, and provide a reliable data foundation for smart government affairs platforms.
Smart Images

Figure CN120067767A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a cross-source data quality defect detection method and system based on deep learning. Background Art
[0002] In the era of big data, government affairs data often needs to integrate information from multiple different data sources to obtain more comprehensive and accurate insights.
[0003] The acquisition and integration of cross-source data face many challenges, among which the data quality problem is particularly prominent. Data from different data sources may have various defects due to inconsistent collection standards, data entry errors, and other reasons. The inventors found and summarized during the development of the intelligent government affairs platform that there is a relatively high redundancy rate and missing values in multi-source heterogeneous data; at the same time, the inconsistent data standards lead to difficulties in cross-departmental field mapping, and the proportion of unstructured data is also relatively high. How to solve the above problems is a technical problem that those skilled in the art need to overcome. Summary of the Invention
[0004] To at least partially solve the above technical problems, this application provides a cross-source data quality defect detection method and system based on deep learning.
[0005] In the first aspect, a cross-source data quality defect detection method based on deep learning provided by this application adopts the following technical solutions.
[0006] A cross-source data quality defect detection method based on deep learning includes: Scanning a cross-source data set using a constructed scanning model; the scanning model monitors and analyzes the data in the cross-source data set during the scanning process; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems; Classifying the data quality problems according to the scanning result of the scanning model and summarizing the data quality problems of the same type; Generating a data quality problem report based on a preset template according to the classification and summary results; and, Inputting the cross-source data with quality defects into a data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data.
[0007] Optionally, the method further includes: Screen the first defective data and the second defective data based on the classification results of defective cross-source data; the repair difficulty of the first defective data is less than that of the second defective data; the first defective data can be repaired by a configured rule engine; the rule engine includes several data rules; the data rules include data integrity rules and data consistency rules; Input the first defective data into the rule engine; the rule engine matches the first defective data one by one based on the included data rules; for the first defective data that matches the corresponding data rule, the rule engine repairs the first defective data based on the processing method defined by the corresponding data rule; Input the second defective data into the data repair model.
[0008] Optionally, use the constructed scanning model to scan the cross-source data set, including: For the text data in the cross-source data set, use a tokenizer to convert the text into a vector representation; Construct the vectorized data into a sequence in order; for time series data, arrange it in chronological order; for non-time series data, arrange it according to the correlation relationship of the data; Use the sequence as input features and input them into the scanning model; Among them, for the text data in the cross-source data set, using a tokenizer to convert the text into a vector representation includes: Based on the tokenizer, tokenize the text data in the cleaned cross-source data set to split it into several tokens; Add a first marker at the start position of each token after tokenizing the text data, and add a second marker at the end position of the token; the first marker is used as an aggregated representation of the semantics of the entire sentence; the second marker is used to distinguish different sentences or text paragraphs; Generate corresponding token type embeddings and position embeddings for each token; the token type embeddings are used to distinguish different text paragraphs or sentences; the position embeddings are used to represent the position information of the token in the sequence; add the token type embeddings, position embeddings and the word embeddings of the token to obtain the final input representation; input the obtained input representation into the pre-trained language model; the pre-trained language model processes the input data through multiple layers of encoders for feature extraction.
[0009] Optionally, the pre-trained language model processes the input data through multiple layers of encoders for feature extraction, including: S401. For each token in the token sequence input to each layer of the encoder, its corresponding input representation is respectively Multiplied by three learnable weight matrices 、 , Multiply them to generate a query vector Q, a key vector K, and a value vector V; S402. Calculate the dot product of each query vector Q and all key vectors K to obtain the original attention scores, and obtain attention weights based on the original attention scores; S403. Perform weighted summation on the value vectors based on the attention weights to obtain the updated token representations for each token ; S404. Input the updated token representations into a feed-forward neural network; the feed-forward neural network includes two linear layers and a non-linear activation function; the updated token representations are multiplied by the weight matrix of the first linear layer and added with the first bias vector and the intermediate result is obtained through the non-linear activation function; the intermediate result is multiplied by the weight matrix of the second linear layer and added with the second bias vector to obtain the output of the feed-forward neural network ; S405. Add the input representation to the token representation and add it to the output of the feed-forward neural network to obtain the output after residual connection and use it as the output vector of the current layer encoder; Each layer of the multi-layer encoder sequentially executes steps S401 - S405 on the input data; wherein, the output of the previous layer encoder is used as the input of the next layer encoder; after the processing of each layer encoder is completed, record the output vector of each token at this layer ; The output vector is the hidden state vector of the token at this layer; after processing all encoding layers, the pre-trained language model outputs the set of hidden state vectors of each token at all layers; Among the hidden state vectors of each token output by the pre-trained language model, locate the hidden state vector corresponding to the first token; during the pre-training process, the vector corresponding to the first token will continuously aggregate the semantic information of the entire sentence; select the hidden state vector corresponding to the first token as the semantic representation of the text data in the cross-source dataset.
[0010] Optionally, the method further includes: After the data repair model completes the repair of the cross-source data with quality defects, collect the repaired cross-source data; Input the repaired data collected again into the scanning model; the scanning model outputs the secondary scanning result of the repaired data; determine whether there are still data quality problems based on the secondary scanning result; if there are still data quality problems, input them into the data repair model again.
[0011] Optionally, the method for generating the data repair model includes: Construct an initial repair model including a generator and a discriminator; the generator is used to receive cross-source inputs with quality defects as inputs and output repaired data samples; the discriminator is used to receive the repaired sample data output by the generator and output the probability that the sample data has been repaired; Initialize the generator and the discriminator; Input the cross-source data with quality defects into the generator; the generator processes the input cross-source data with quality defects according to the current parameter settings to generate repaired sample data; Input the repaired sample data generated by the generator and the corresponding real data into the discriminator respectively; The discriminator calculates the probability that the sample data and the real data have been repaired respectively; calculates the loss function of the discriminator according to the probabilities of the two; updates the parameters of the discriminator according to the loss function of the discriminator; repeat the training until the confidence of the initial repair model is greater than the preset value to obtain the repair model.
[0012] Optionally, the method further includes: During the training process, the generator and the discriminator are alternately trained. The goal of the generator is to generate repaired data samples that can deceive the discriminator, and the goal of the discriminator is to accurately distinguish between real data and the repaired data samples generated by the generator; Update the parameters of the generator and the discriminator according to the output result of the discriminator to improve the authenticity of the repaired data samples generated by the generator and the discrimination ability of the discriminator.
[0013] In a second aspect, a cross-source data quality defect detection system based on deep learning provided by the present application adopts the following technical solutions.
[0014] A cross-source data quality defect detection system based on deep learning includes: A first processing module, configured to: scan a cross-source data set using the constructed scanning model; the scanning model monitors and analyzes the data in the cross-source data set during the scanning process; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems; A second processing module, configured to: classify the data quality problems according to the scanning result of the scanning model and summarize the data quality problems of the same type; A third processing module, configured to: generate a data quality problem report based on a preset template according to the classified and summarized results; A fourth processing module, configured to: input cross-source data with quality defects into a data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data.
[0015] Compared with the prior art, the beneficial effects of the present invention include: the first-pass scanning model scans the cross-source data set, classifies and summarizes data quality problems based on the scanning results, and then generates a data quality problem report based on a preset template; inputting the cross-source data with quality defects into the data repair model and using the learned repair strategy for repair forms a complete and efficient closed-loop for data quality defect detection and repair, which can effectively reduce decision-making errors caused by data quality problems and lay a data foundation for the development of the intelligent government affairs platform. Description of the Drawings
[0016] Figure 1 is a flowchart of a method for detecting cross-source data quality defects based on deep learning according to an embodiment of the present application; Figure 2 is a system block diagram of a system for detecting cross-source data quality defects based on deep learning according to an embodiment of the present application; In the figure, 201 is a first processing module; 202 is a second processing module; 203 is a third processing module; 204 is a fourth processing module. Detailed Embodiments
[0017] The following further describes the present application in conjunction with the attached Figure 1-2 drawings and specific embodiments: An embodiment of the present application discloses a method for detecting cross-source data quality defects based on deep learning, including the following steps: Step S101: Use the constructed scanning model to scan the cross-source dataset; during the scanning process, the scanning model monitors and analyzes the data in the cross-source dataset; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems. Specifically, the scanning model is a model constructed based on deep learning, which contains a multi-layer neural network structure inside. During the scanning process, it can identify possible quality defects in the data according to the learned feature patterns and output these problems in a specific form. The cross-source dataset refers to a dataset formed by data from multiple different data sources, and these data sources may come from different business systems, different databases, or different device acquisition terminals. Due to the diversity of data sources, there are often differences in the data formats, data standards, and data meanings of the cross-source dataset. Data quality problems refer to various situations where the data in the cross-source dataset does not meet the expected quality standards. Common data quality problems include: data accuracy problems, that is, the data value does not match the true value; data integrity problems, such as missing data records and empty field values; data redundancy problems.
[0018] Step S102: Classify the data quality problems according to the scanning result of the scanning model and summarize the data quality problems of the same type.
[0019] Step S103: Generate a data quality problem report based on the classified and summarized results according to a preset template.
[0020] Step S104: Input the cross-source data with quality defects into the data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data. Specifically, the data repair model is a model used to repair the cross-source data with quality defects. It learns a large number of existing data repair cases and repairs the cross-source data with defects for the data quality problems detected by the scanning model, so that the data reaches a state that meets the quality standards and outputs the repaired cross-source data.
[0021] Specifically, the first-pass scanning model scans the cross-source dataset, classifies and summarizes the data quality problems based on the scanning results, and then generates a data quality problem report according to a preset template, providing intuitive and standardized information display for data managers, which helps them quickly understand the data quality status. Inputting the cross-source data with quality defects into the data repair model and using the learned repair strategies to repair it forms a complete and efficient closed-loop for data quality defect detection and repair, which can effectively reduce decision-making mistakes caused by data quality problems and lay a data foundation for the development of the smart government affairs platform.
[0022] The following is an example: The intelligent government affairs platform needs to integrate data sources from multiple different departments, such as population information data from the civil affairs department, tax payment data from the tax department, household registration and public security data from the public security department, etc. The formats, standards, and update frequencies of these cross-source data sets are different. The marital status data of residents registered by the civil affairs department may be inconsistent with the marital status information in the household registration system of the public security department. The constructed scanning model scans the cross-source data sets in the intelligent government affairs platform. Taking tax data and enterprise industrial and commercial registration data as an example, the scanning model can monitor whether the enterprise's tax declaration data matches information such as the business scope and registered capital registered in the industrial and commercial registration. If it is found through scanning that the declared tax amount of a certain enterprise is very different from the tax amount payable according to its industrial and commercial registration information, this may be a data accuracy problem. According to the results of the scanning model, the data quality problems are classified and summarized. In the intelligent government affairs platform, data quality problems can be divided into information accuracy problems (such as the above-mentioned inconsistency between tax data and industrial and commercial registration information), data integrity problems (such as the lack of insurance records of some residents in the social security system), data consistency problems (such as different departments registering different addresses for the same enterprise), etc. Through classification and summarization, the distribution of various data quality problems can be clearly presented. Generate a data quality problem report based on a preset template. For example, the report details the quality problems existing in the data of different departments, such as the quantity and proportion of incorrect date formats of some residents' birth dates in the population information of the civil affairs department; the list of enterprises with abnormal declarations and the amounts involved in the enterprise tax declaration data of the tax department. Input the cross-source data with quality defects into the data repair model, and the data repair model can repair the cross-source data according to the data repair model, such as repairing the date format, etc.
[0023] As a specific implementation manner of a cross-source data quality defect detection method based on deep learning, the method further includes: Screen the first defective data and the second defective data based on the classification results of the defective cross-source data; the repair difficulty of the first defective data is less than that of the second defective data; the first defective data can be repaired through a configured rule engine; the rule engine includes several data rules; the data rules include data integrity rules and data consistency rules; Input the first defective data into the rule engine; the rule engine matches the first defective data one by one based on the included data rules; for the first defective data that matches the corresponding data rule, the rule engine repairs the first defective data based on the processing method defined by the corresponding data rule; Input the second defective data into the data repair model.
[0024] Specifically, in the development of the intelligent government affairs platform, data sources from different departments are aggregated to form a cross-source dataset, and data quality problems are complex and diverse. Based on the classification results of the defective cross-source data, the first defective data and the second defective data are obtained. Suppose in the population information data, some records of resident ID numbers have simple format errors. Such data belongs to the first defective data because its repair difficulty is relatively small and it can be repaired through the configured rule engine. The rule engine includes data integrity rules such as requiring the ID number to be 18 digits; and data consistency rules such as the date of birth corresponding to the ID number should be consistent with the date of birth registered in the household registration system. Input this type of first defective data into the rule engine, and the rule engine matches it one by one according to the included data rules. When a format error in the ID number is found, the rule engine follows the processing method defined by the corresponding rule, such as automatically supplementing the correct numbers at the missing positions or correcting the format error, to complete the repair of the first defective data. For some complex data quality problems that are difficult to repair through the rule engine, such as the semantic fuzzy differences in the descriptions of the business scope of enterprises by different departments, such data belongs to the second defective data. Input the second defective data into the data repair model for repair, so as to improve the quality of the cross-source data of the intelligent government affairs platform and provide a more reliable data basis for government affairs decision-making and services.
[0025] As a specific implementation of a cross-source data quality defect detection method based on deep learning, use the constructed scanning model to scan the cross-source dataset, including: For the text data in the cross-source dataset, use a tokenizer to convert the text into a vector representation; Construct the vectorized data into a sequence in order; for time series data, arrange it in chronological order; for non-time series data, arrange it according to the correlation relationship of the data; Input the sequence as input features into the scanning model; Among them, for the text data in the cross-source dataset, using a tokenizer to convert the text into a vector representation includes: Based on the tokenizer, tokenize the text data in the cleaned cross-source dataset to split it into several tokens; Add a first marker at the start position of each token after tokenizing the text data, and add a second marker at the end position of the token; the first marker is used as an aggregated representation of the semantics of the entire sentence; the second marker is used to distinguish different sentences or text paragraphs; Generate corresponding token type embeddings and position embeddings for each token; the token type embeddings are used to distinguish different text paragraphs or sentences; the position embeddings are used to represent the position information of the token in the sequence; add the token type embeddings, position embeddings and the word embeddings of the tokens to obtain the final input representation; input the obtained input representation into the pre-trained language model; the pre-trained language model processes the input data through multiple layers of encoders for feature extraction.
[0026] Specifically, for a large amount of text data in the intelligent government affairs platform, such as policy documents, service guides, resident feedback, etc., first use a tokenizer to split the cleaned text into tokens, which can deconstruct continuous text information into more easily processed units. Add specific markers at the start and end positions of the tokens so that the model can clearly distinguish different sentences and paragraphs and accurately grasp the text structure. Generate type embeddings and position embeddings for the tokens and add them to the word embeddings to obtain the final input representation; input it into the pre-trained language model and process it through multiple layers of encoders, and use the general language knowledge learned by the model on a large amount of text to mine the key features in the text data. Thus, the understanding ability of the scanning model for text data is improved, enabling it to accurately identify possible data quality defects in the text, such as semantic ambiguity and inconsistent expressions in policy documents.
[0027] As a specific implementation of a cross-source data quality defect detection method based on deep learning, the pre-trained language model processes the input data through multiple layers of encoders for feature extraction, including: S401. For each token in the token sequence input to each layer of the encoder, multiply its corresponding input representation respectively by three learnable weight matrices , , to generate a query vector Q, a key vector K, and a value vector V; S402. Calculate the dot product of each query vector Q and all key vectors K to obtain the original attention scores and obtain attention weights based on the original attention scores; S403. Perform weighted summation on the value vectors based on the attention weights to obtain the updated token representation of each token ; S404. Input the updated token representation into a feed-forward neural network; the feed-forward neural network includes two linear layers and a non-linear activation function; the updated token representation is multiplied by the weight matrix of the first linear layer and added with the first bias vector and passed through the non-linear activation function to obtain an intermediate result; the intermediate result is multiplied by the weight matrix of the second linear layer Multiply and add the second bias vector to obtain the output of the feedforward neural network ; S405. Represent the input and add it to the token representation and add it to the output of the feedforward neural network to obtain the output after residual connection and use it as the output vector of the current layer encoder; Each layer of the multi-layer encoder sequentially performs steps S401 - S405 on the input data; wherein, the output of the previous layer encoder is used as the input of the next layer encoder; after each layer of encoder processing, record the output vector of each token at that layer ; the output vector is the hidden state vector of the token at that layer; after processing all encoding layers, the pre-trained language model outputs the set of hidden state vectors of each token at all layers; Among the hidden state vectors of each token output by the pre-trained language model, locate the hidden state vector corresponding to the first token; during the pre-training process, the vector corresponding to the first token will continuously aggregate the semantic information of the entire sentence; select the hidden state vector corresponding to the first token as the semantic representation of the text data in the cross-source dataset.
[0028] Specifically, multiply the token input representation by three learnable weight matrices to generate the query vector Q, the key vector K, and the value vector V; calculate the original attention scores and obtain the attention weights in S402, enabling the model to allocate attention according to the correlation between tokens. S403 performs weighted summation on the value vectors based on the attention weights to update the token representation. The processing of the feedforward neural network in S404 uses two linear layers and a non-linear activation function, introducing non-linearity and enhancing the model's expressive power to capture more complex semantic patterns in the text. The residual connection in S405 solves the problem of gradient vanishing in the training of deep networks, ensuring the effective transmission of information, enabling the model to construct deeper network structures, and thus learning more advanced semantic features. The multi-layer encoder processes the input data sequentially, and each layer further extracts features based on the previous layer, continuously deepening the understanding of the text. Recording the hidden state vectors of each token at each layer provides rich information for a comprehensive analysis of the text; finally, locate and select the hidden state vector corresponding to the first token as the semantic representation of the text data because this vector aggregates the semantic information of the entire sentence during the pre-training process and can accurately represent the core semantics of the text. In the application scenario of the intelligent government affairs platform, this helps the scanning model to more accurately identify quality defects in text data, such as expression ambiguities and logical contradictions in policy documents.
[0029] As one of the implementation manners of a cross-source data quality defect detection method based on deep learning, the method further includes: After the data repair model completes the repair of the cross-source data with quality defects, collect the repaired cross-source data; Input the collected repaired data into the scanning model again; the scanning model outputs a secondary scanning result of the repaired data; based on the secondary scanning result, determine whether there are still data quality problems; if there are still data quality problems, input them into the data repair model again.
[0030] As one of the implementation manners of a cross-source data quality defect detection method based on deep learning, the method for generating a data repair model includes: Construct an initial repair model including a generator and a discriminator; the generator is used to receive cross-source inputs with quality defects as inputs and output repaired data samples; the discriminator is used to receive the repaired sample data output by the generator and output the probability that the sample data has been repaired; Initialize the generator and the discriminator; Input the cross-source data with quality defects into the generator; the generator processes the input cross-source data with quality defects according to the current parameter settings to generate repaired sample data; Input the repaired sample data generated by the generator and the corresponding real data into the discriminator respectively; The discriminator calculates the probabilities that the sample data and the real data have been repaired respectively; calculate the loss function of the discriminator according to the probabilities of the two; update the parameters of the discriminator according to the loss function of the discriminator; repeat the training until the confidence of the initial repair model is greater than a preset value to obtain a repair model.
[0031] As one of the implementation manners of a cross-source data quality defect detection method based on deep learning, the method further includes: During the training process, the generator and the discriminator are alternately trained. The goal of the generator is to generate repaired data samples that can deceive the discriminator, and the goal of the discriminator is to accurately distinguish between real data and the repaired data samples generated by the generator; According to the output result of the discriminator, update the parameters of the generator and the discriminator to improve the authenticity of the repaired data samples generated by the generator and the discrimination ability of the discriminator.
[0032] This application also provides a cross-source data quality defect detection system based on deep learning, including: The first processing module 201 is configured to: scan the cross-source dataset using the constructed scanning model; the scanning model monitors and analyzes the data in the cross-source dataset during the scanning process; when there are quality defects in the cross-source data, the scanning result output by the scanning model includes data quality problems; The second processing module 202 is configured to: classify the data quality problems according to the scanning result of the scanning model and summarize the data quality problems of the same type; The third processing module 203 is configured to: generate a data quality problem report based on the classified and summarized results and a preset template; The fourth processing module 204 is configured to: input the cross-source data with quality defects into the data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data.
[0033] It should be noted that: the above embodiments are only used to illustrate the present application and do not limit the technical solutions described in the present application. Although this specification has described the present application in detail with reference to the above embodiments, those of ordinary skill in the art should understand that those skilled in the technical field can still modify the present application or make equivalent replacements, and all technical solutions and their improvements that do not depart from the spirit and scope of the present application should be covered within the scope of the claims of the present application.
Claims
1. A cross-source data quality defect detection method based on deep learning, characterized in that: include: Use the constructed scanning model to scan the cross-source dataset; The scanning model monitors and analyzes the data in the cross-source data set during the scanning process; When there are quality defects in the cross-source data, the scanning results output by the scanning model contain data quality problems; Classifying the data quality issues according to the scanning results of the scanning model and aggregating the data quality issues of the same type; Generate a data quality problem report based on the preset template according to the classification and summary results; as well as, The cross-source data with quality defects is input into a data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data.
2. According to a method for detecting cross-source data quality defects based on deep learning according to claim 1, it is characterized in that: The method further comprises: The first defective data and the second defective data are obtained by screening based on the classification results of the defective cross-source data; the first defective data is less difficult to repair than the second defective data; the first defective data can be repaired by a configured rule engine; the rule engine includes a number of data rules; the data rules include data integrity rules and data consistency rules; The first defect data is input into the rule engine; the rule engine matches the first defect data one by one based on the included data rules; for the first defect data matched with the corresponding data rules, the rule engine repairs the first defect data based on the processing method defined by the corresponding data rules; The second defect data is input into the data repair model.
3. A cross-source data quality defect detection method based on deep learning according to claim 2, characterized in that: Use the built scanning model to scan cross-source datasets, including: For text data in cross-source datasets, use a tokenizer to convert the text into a vector representation; The vectorized data is constructed into a sequence in order; for time series data, it is arranged in chronological order; for non-time series data, it is arranged according to the correlation relationship of the data; inputting the sequence as an input feature into the scanning model; Among them, for the text data in the cross-source dataset, a word segmenter is used to convert the text into a vector representation, including: Based on the word segmenter, the text data in the cleaned cross-source data set is segmented to split into a plurality of word units; A first marker is added to the starting position of each word unit after word segmentation of each text data, and a second marker is added to the ending position of the word unit; the first marker is used as an aggregate representation of the semantics of the entire sentence; the second marker is used to distinguish different sentences or text paragraphs; Generate a corresponding word-gram type embedding and position embedding for each word-gram; the word-gram type embedding is used to distinguish different text paragraphs or sentences; the position embedding is used to represent the position information of the word-gram in the sequence; the word-gram type embedding, the position embedding and the word embedding of the word-gram are added to obtain a final input representation; the obtained input representation is input into a pre-trained language model; the pre-trained language model processes the input data through a multi-layer encoder to extract features.
4. A cross-source data quality defect detection method based on deep learning according to claim 3, characterized in that: The pre-trained language model processes the input data through a multi-layer encoder to extract features, including: S401, for each word in the word-word sequence input to each layer of encoder, the corresponding input representation is With three learnable weight matrices , , Multiply to generate query vector Q, key vector K and value vector V; S402, calculating the dot product of each query vector Q and all key vectors K to obtain an original attention score and obtaining an attention weight based on the original attention score; S403: Perform weighted summation on the value vector based on the attention weight to obtain the updated word unit representation of each word unit ; S404: Update the word element representation Input into a feedforward neural network; the feedforward neural network includes two linear layers and a nonlinear activation function; the updated word unit representation The weight matrix of the first linear layer Multiply and add the first bias vector And obtain an intermediate result through the nonlinear activation function; and add the intermediate result to the weight matrix of the second linear layer. Multiply and add the second bias vector Get the output of the feedforward neural network ; S405: the input representation With the word element representation Add and combine with the output of the feedforward neural network Add the residual connection output And used as the output vector of the current layer encoder; Each layer of the multi-layer encoder performs steps S401-S405 on the input data in turn; the output of the previous layer of encoder is used as the input of the next layer of encoder; after each layer of encoder processing is completed, the output vector of each word in that layer is recorded. ; Output vector is the hidden state vector of the word at this layer; after processing all encoding layers, the pre-trained language model outputs the set of hidden state vectors of each word in all layers; The hidden state vector corresponding to the first tag is located from the hidden state vectors of each word unit output by the pre-trained language model; during the pre-training process, the vector corresponding to the first tag will continuously aggregate the semantic information of the entire sentence; the hidden state vector corresponding to the first tag is selected as the semantic representation of the text data in the cross-source dataset.
5. A cross-source data quality defect detection method based on deep learning according to claim 4, characterized in that: The method further comprises: After the data repair model completes the repair of the cross-source data with quality defects, the repaired cross-source data is collected; The collected repaired data is input into the scanning model again; the scanning model outputs the secondary scanning result of the repaired data; based on the secondary scanning result, it is determined whether there are still data quality problems; if there are still data quality problems, it is input into the data repair model again.
6. A method for detecting cross-source data quality defects based on deep learning according to claim 5, characterized in that: The method for generating the data repair model includes: Constructing an initial repair model including a generator and a discriminator; the generator is used to receive a cross-source input with quality defects as input and output a repaired data sample; the discriminator is used to receive the repaired sample data output by the generator and output the probability that the sample data has been repaired; Initialize the generator and the discriminator; Inputting the cross-source data with quality defects into the generator; the generator processes the input cross-source data with quality defects according to the current parameter settings to generate repaired sample data; The repaired sample data generated by the generator and the corresponding real data are input into the discriminator respectively; The discriminator calculates the probabilities that the sample data and the real data have been repaired respectively; calculates the loss function of the discriminator according to the probabilities of the two; updates the parameters of the discriminator according to the loss function of the discriminator; and repeats the training until the confidence of the initial repair model is greater than a preset value to obtain a repair model.
7. A method for detecting cross-source data quality defects based on deep learning according to claim 6, characterized in that: The method further comprises: During the training process, the generator and the discriminator are trained alternately, the goal of the generator is to generate repaired data samples that can deceive the discriminator, and the goal of the discriminator is to accurately distinguish between real data and repaired data samples generated by the generator; According to the output result of the discriminator, the parameters of the generator and the discriminator are updated to improve the authenticity of the repaired data samples generated by the generator and the discrimination ability of the discriminator.
8. A cross-source data quality defect detection system based on deep learning, characterized in that: include: The first processing module is used to: scan the cross-source data set using the constructed scanning model; The scanning model monitors and analyzes the data in the cross-source data set during the scanning process; when there are quality defects in the cross-source data, the scanning results output by the scanning model contain data quality problems; A second processing module is used to: classify the data quality problems according to the scanning results of the scanning model and summarize the data quality problems of the same type; The third processing module is used to: generate a data quality problem report based on a preset template according to the classification and summary results; The fourth processing module is used to: input the cross-source data with quality defects into the data repair model; the data repair model repairs the cross-source data with defects and outputs the repaired cross-source data.
Citation Information
Patent Citations
Abnormal data detection and restoration method for electric power measurement system
CN115238563A
Pipeline interior detection method and system based on multi-sensor fusion
CN119150175A
Method, device and equipment for determining abnormal data of power distribution network, medium and product
CN119167254A
Data consistency check method and system based on icc
WO2022021849A1