Information collection method for auditing big data

By identifying and classifying data sources in audit big data processing, formulating acquisition strategies, and using variational autoencoder and graph neural network algorithms for data quality evaluation, data acquisition inadaptability and quality control problems in the existing technology are solved, and efficient and intelligent data acquisition and quality evaluation are achieved.

CN120123401AInactive Publication Date: 2025-06-10YITONG FINANCIAL SERVICES PAYMENT CO LTD

Patent Information

Application Number
CN202510187488.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing audit big data processing methods are difficult to effectively identify and classify different data sources, resulting in the acquisition strategy not meeting actual needs, resulting in data redundancy or omission, and it is difficult to effectively control data quality, which cannot ensure the timeliness and accuracy of data.

Method used

Through the audit system, connect to multiple external data sources, identify and classify various data sources, formulate corresponding collection strategies, and use variational autoencoder and graph neural network algorithm to evaluate the data quality to ensure the integrity, abnormality and consistency of the data.

Benefits of technology

It realizes accurate identification and classification of different data sources, avoids data omissions or redundancy, ensures the timeliness and integrity of data, improves the intelligence level of data quality control, and reduces the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123401A_ABST
    Figure CN120123401A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing, in particular to an auditing big data information collection method, which comprises the following steps of S1, identifying and classifying various data sources; s2, making a corresponding acquisition strategy according to the features of the various data sources identified in the S1; s3, acquiring data from each data source in real time according to the acquisition strategy formulated in S2; s4, storing the data acquired in the S3 into a data warehouse according to a preset data structure; s5, performing quality evaluation on the data stored in the data warehouse; s6, information in the data warehouse is periodically updated through a data synchronization mechanism, efficient collection, standardized processing and intelligent quality control of multi-source data are achieved through a collection strategy and an innovative data quality evaluation method, and the automation degree of data processing and the reliability of auditing analysis are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and particularly to a method for collecting information of audit big data. Background Art

[0002] With the advent of the information age, enterprises and institutions generate a large amount of business data in their daily operations. Especially in the fields of audit, finance, and taxation, these data are not only huge but also of various types. To support data-driven decision-making and audit analysis, enterprises need to effectively collect, store, and process data from different sources. Existing audit big data processing methods usually rely on traditional manual operations or simple automation tools. In the process of data collection, storage, and analysis, these methods often face problems such as inconsistent data formats, untimely data updates, and unstable data quality. In addition, due to the different interfaces and storage structures of different data sources, data integration and synchronization also become a technical problem.

[0003] In the existing technology, during the data collection and storage process, it is difficult to effectively identify and classify various data sources, resulting in the collection strategy not adapting to the actual needs, causing data redundancy or omission. At the same time, there are relatively large defects in the existing methods in terms of data quality control. They cannot effectively identify data integrity problems, abnormal data, and data consistency problems, and it is difficult to ensure the timeliness and accuracy of data. Therefore, how to collect, synchronize, clean, and evaluate large-scale data sets in real time and efficiently through automated means has become a major challenge in the current technology. Summary of the Invention

[0004] Based on the above purpose, the present invention provides a method for collecting information of audit big data.

[0005] A method for collecting information of audit big data includes the following steps:

[0006] S1: Connect to multiple external data sources through an audit system, identify and classify various data sources, including financial data sources, transaction data sources, and tax data sources;

[0007] S2: According to the characteristics of various data sources identified in S1, formulate corresponding collection strategies, including data access methods, collection frequencies, and data format requirements;

[0008] S3: Obtain data from each data source in real time according to the collection strategy formulated in S2;

[0009] S4: Store the data collected in S3 into a data warehouse according to a preset data structure, and perform preliminary data cleaning and sorting to ensure the uniformity of data formats;

[0010] S5: Apply the variational autoencoder and graph neural network algorithms to evaluate the quality of the data stored in the data warehouse, including specifically data integrity check, abnormal data detection, and data consistency verification;

[0011] S6: Through the data synchronization mechanism, regularly update the information in the data warehouse to ensure the timeliness and accuracy of the data.

[0012] Optionally, the specific steps of S1 include:

[0013] S11: The audit system connects to the external data source through the API interface and automatically obtains the metadata of the external data source, including the data source name, data format, and access protocol;

[0014] S12: Based on the identified metadata, use natural language processing technology to perform semantic analysis on the data source name and description, and match them through the set keyword library to identify the type of data source;

[0015] S13: According to the data source type identified in S12, classify the data sources into financial data sources, transaction data sources, and tax data sources, and add labels to each type of data source to mark its specific characteristics.

[0016] Optionally, the specific steps of S12 include

[0017] S121: Convert the data source name and description into a word vector representation, and use the Word2Vec or TF-IDF method to encode the text to form a vectorized representation; assume the data source description is D = {d 1 , d 2 ,..., d n}, where each d i is a word of the data source;

[0018] S122: Construct a keyword library K = {k 1 , k 2 ,..., k m}, where each k i represents a term of a category, and the specific categories include finance, transaction, and tax;

[0019] S123: Calculate the similarity between the data source description D and each term k i in the keyword library, specifically using the cosine similarity formula for calculation. The formula is: where D and k i are word vector representations, ∥D∥ and ∥k i ∥ are the norms of the vectors respectively, and · represents the dot product of the vectors;

[0020] S124: According to the similarity score sim(D, k i) Match the data source description with relevant categories in the keyword library; if the similarity score of a certain category exceeds a preset threshold, classify the data source into the corresponding category, including finance, transactions, and taxation.

[0021] Optionally, S2 specifically includes:

[0022] S21: Based on the data source types identified in S1, extract the feature information of each type of data source, including data format, data structure, update frequency, and data access method;

[0023] S22: According to the feature information extracted in S21, use a rule engine to formulate specific collection strategies for different data source categories, specifically including data access method, collection frequency, and data format requirements;

[0024] Select the corresponding data access method according to the type of data source and the provided access protocol; for financial data sources, select to access data through an SQL database connection; for transaction data sources, select to obtain data through a RESTful API; for tax data sources, select a file access protocol based on XML / JSON format;

[0025] Set different collection frequencies according to the update frequency and timeliness requirements of the data source; for financial data sources, set to collect regularly every day; for transaction data sources, set to collect every hour; for tax data sources, set to collect weekly;

[0026] Formulate corresponding data format requirements according to the storage formats of different data sources; for financial data sources, require the collected format to be structured data; for transaction data sources, require JSON format; for tax data sources, require XML format.

[0027] Optionally, S3 specifically includes:

[0028] S31: According to the collection strategies formulated in S2, select the corresponding connection method to establish a connection with the data source; for financial data sources of the SQL database type, connect through the JDBC protocol; for transaction data sources of the RESTful API type, connect through an HTTP request interface; for tax data sources in XML / JSON format, access files through HTTP or FTP protocols;

[0029] S32: Send data requests in real time according to the collection frequency defined in S2;

[0030] S33: Obtain data from various data sources according to the set access method and collection frequency.

[0031] Optionally, S4 specifically includes:

[0032] S41: After receiving the data collected in S3, perform format conversion on the data according to the data format requirements specified in S2;

[0033] S42: Clean the converted data, including removing duplicate data, filling in missing values, and standardizing the data format; specifically, for missing numerical data, use the mean value filling or interpolation method to fill it; for duplicate records, remove the duplicates by comparing the unique identifiers; for date or amount data with inconsistent formats, use regular expressions to format the data to ensure consistency;

[0034] S43: Store the cleaned data in the data warehouse according to the corresponding data structure requirements; for financial data sources, store the data in a relational database; for transaction data sources, store it in a NoSQL database; for tax data sources, store the data in a big data format.

[0035] Optionally, the specific steps of S5 include:

[0036] S51: Use the variational autoencoder algorithm to calculate the reconstruction error by mapping the data to the latent space and reconstructing the original data; if the error exceeds the set threshold, it is considered that there is an integrity problem with the data, and the data is marked as missing or incomplete;

[0037] S52: Based on the graph neural network algorithm, regard the data as a graph structure, calculate the similarity between the data points and their adjacent nodes by learning the relationships between the nodes, and when the similarity is lower than the preset threshold, mark it as abnormal data;

[0038] S53: Reconstruct and compare the same type of data from different data sources through the variational autoencoder, calculate the differences between the data, and if the differences exceed the preset range, judge the data as inconsistent and mark it as a consistency problem.

[0039] Optionally, the specific steps of S51 include:

[0040] S511: Represent the data in the data warehouse as an input matrix where n represents the number of samples and m represents the feature dimension of each sample; and map the input data X to the latent space Z through the encoder part of the variational autoencoder, and the expression is: Z i = f enc (X i ; θ enc ), where, is the representation of the i-th sample in the latent space, d is the dimension of the latent space; f enc is the encoding function; θ enc is the parameter of the encoder; X i represents the input data of the i-th sample;

[0041] S512: Assume that the latent space Z follows a Gaussian distribution N(μ, σ 2 ), where μ i and σ i represent the mean and standard deviation of the latent space of the i-th sample respectively; the specific calculation formula is: μ i = g enc (X i ; θ μ ); σ i = h enc (X i ; θ σ ), where g enc and h enc are functions in the encoder network used to calculate the mean and standard deviation respectively, and θ μ and θ σ are the parameters for mean and standard deviation calculation respectively;

[0042] S513: Reconstruct the data Z in the latent space into an approximation of the original data i through the decoder part

[0043] S514: Evaluate the integrity of the data by calculating the error between the original data X i and the reconstructed data . Use the mean squared error MSE as the evaluation metric, and the calculation formula is: where MSE i is the mean squared error of the i-th sample, and X ij and are the original data and the reconstructed data of the j-th feature of the i-th sample respectively;

[0044] S515: Set a threshold. If exceeds this threshold, it is considered that there is an integrity problem with the data of the corresponding i-th sample, and it is marked as missing or incomplete data.

[0045] Optionally, the S52 specifically includes:

[0046] S521: Represent each data point in the data warehouse as a node in the graph, and each node represents a data sample; and construct a graph structure according to the feature similarity between data points. For each pair of data points, if the similarity between their features exceeds the preset threshold, an edge is established in the graph for these two data points;

[0047] S522: Calculate the similarity between a data point and its adjacent nodes through the graph neural network algorithm as an indicator to measure its abnormality;

[0048] S523: When the calculated similarity is lower than the set threshold, it indicates that the features of the corresponding data points are abnormal; and the data points are marked as abnormal data.

[0049] Optionally, the S6 specifically includes:

[0050] S61: Configure different data synchronization strategies according to the type of data source and the data update frequency; for financial data sources, set to automatically synchronize daily; for transaction data sources, set to synchronize hourly; for tax data sources, set to synchronize weekly;

[0051] S62: Trigger the data synchronization operation through a scheduled task scheduler, and automatically start the data synchronization process at a predetermined time point according to the preset synchronization period;

[0052] S63: In each synchronization process, first obtain the latest data from the external data source and compare it with the existing data; check whether the newly obtained data meets the format requirements and perform deduplication to ensure that the combined data is consistent and non-duplicate.

[0053] Advantages of the present invention:

[0054] In the present invention, through an accurate acquisition strategy, it is possible to perform regular synchronization according to the characteristics and update frequency of the data source, avoiding data omission or redundancy in the traditional method, and ensuring the timeliness and integrity of the data; in addition, through the data cleaning and sorting process, the formats of data from different sources are effectively unified, providing a stable and standardized basis for subsequent data analysis and quality assessment.

[0055] In the present invention, by introducing the variational autoencoder and the graph neural network algorithm, the data is innovatively evaluated for quality, and it can automatically detect data integrity problems, abnormal data, and data consistency problems, greatly improving the intelligent level of data quality control; the variational autoencoder can accurately judge the missing or damaged situation of the data, while the graph neural network accurately identifies abnormal data by learning the correlation relationship between data points, thus ensuring the accuracy and consistency of the data; through the combination of these technologies, the efficiency of big data audit analysis is effectively improved, the need for manual intervention is reduced, and the automation degree and reliability of data processing are improved. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only those of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1Schematic diagram of the information collection method according to an embodiment of the present invention;

[0058] Figure 2 Schematic diagram of the method for quality assessment of data according to an embodiment of the present invention. Detailed implementation manners

[0059] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.

[0060] It should be noted that in the specification, references to "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. indicate that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes the specific feature, structure, or characteristic. Additionally, when combining embodiments to describe a specific feature, structure, or characteristic, implementing such feature, structure, or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0061] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather, at least in part depending on the context, can allow for the existence of other factors that may not be explicitly described.

[0062] As Figure 1 - Figure 2 shown, an information collection method for auditing big data includes the following steps:

[0063] S1: Connect to multiple external data sources through an audit system, identify and classify various data sources, including financial data sources, transaction data sources, and tax data sources;

[0064] S2: According to the characteristics of the various data sources identified in S1, formulate corresponding collection strategies, including data access methods, collection frequencies, and data format requirements;

[0065] S3: According to the collection strategies formulated in S2, obtain data from each data source in real time to ensure the integrity and timeliness of the collection of various data;

[0066] S4: Store the data collected in S3 into the data warehouse according to a preset data structure, and perform preliminary data cleaning and sorting to ensure the uniformity of data formats;

[0067] S5: Apply the Variational Autoencoder (VAE) and the Graph Neural Network algorithm (GNN) to evaluate the quality of the data stored in the data warehouse, specifically including data integrity checking, abnormal data detection, and data consistency verification; among them, VAE is used for abnormal data detection and data integrity checking, and GNN is used for data consistency verification. The combination of the two can effectively improve the intelligent level of data quality evaluation;

[0068] S6: Through the data synchronization mechanism, regularly update the information in the data warehouse to ensure the timeliness and accuracy of the data.

[0069] S1 specifically includes:

[0070] S11: The audit system connects to external data sources through the API interface and automatically obtains the metadata of the external data sources, including data source names, data formats, and access protocols;

[0071] S12: Based on the identified metadata, use Natural Language Processing (NLP) technology to perform semantic analysis on the data source names and descriptions, match them through a set keyword library, and identify the types of data sources to ensure the accurate classification of data source types;

[0072] S13: According to the data source types identified in S12, classify the data sources into financial data sources, transaction data sources, and tax data sources, and add labels to each type of data source to mark its specific features, such as data structure, format requirements, update frequency, etc.; the above steps use the API interface, natural language processing technology, and a preset keyword library to automatically realize the type analysis and classification of data sources, effectively improving the accuracy and efficiency of data source identification.

[0073] The steps of matching through the set keyword library in S12 specifically include

[0074] S121: Convert the data source name and description into a word vector representation, and use methods such as Word2Vec or TF-IDF to encode the text to form a vectorized representation; assume the data source description is D = {d 1 ,d 2 ,...,d n}, where each d i is a word of the data source, and the word vector is w i ;

[0075] S122: Construct a keyword library K = {k 1 ,k 2 ,...,km}, where each k i represents a term of a category, and the specific categories include finance, transactions, and taxation, and the terms are converted into corresponding vectors w k ;

[0076] S123: Calculate the similarity between the data source description D and each term k i in the keyword library. Specifically, the cosine similarity formula is used for calculation, and the formula is: where D and k i are the word vector representations, ∥D∥ and ∥k i ∥ are the norms of the vectors respectively, and · represents the dot product of the vectors;

[0077] S124: According to the similarity score sim(D, k i ), match the data source description with the relevant categories in the keyword library; if the similarity score of a certain category exceeds the preset threshold, classify the data source into the corresponding category, including finance, transactions, and taxation.

[0078] S2 specifically includes:

[0079] S21: Based on the data source types identified in S1 (such as financial data sources, transaction data sources, tax data sources), extract the feature information of each type of data source, including data format, data structure, update frequency, and data access method;

[0080] S22: According to the feature information extracted in S21, use a rule engine to formulate specific collection strategies for different data source categories, including data access methods, collection frequencies, and data format requirements;

[0081] According to the type of data source and the provided access protocol, select the corresponding data access method; for financial data sources, select to access data through an SQL database connection (such as JDBC); for transaction data sources, select to obtain data through RESTful APIs; for tax data sources, select a file access protocol based on XML / JSON format;

[0082] According to the update frequency and timeliness requirements of the data source, set different collection frequencies; for financial data sources, set to collect data regularly every day; for transaction data sources, set to collect data every hour; for tax data sources, set to collect data weekly;

[0083] According to the storage formats of different data sources, corresponding data format requirements are formulated; for financial data sources, the required format for collection is structured data (such as CSV, SQL tables); for transaction data sources, the required format is JSON; for tax data sources, the required format is XML; the above steps can select the best access method, collection frequency and data format according to the different characteristics of the data sources, ensuring the efficiency, accuracy and flexibility of the data collection process.

[0084] S3 specifically includes:

[0085] S31: According to the collection strategy formulated in S2, select the corresponding connection method to establish a connection with the data source; for financial data sources of the SQL database type, connect through the JDBC protocol; for transaction data sources of the RESTful API type, connect through the HTTP request interface; for tax data sources in XML / JSON format, access the file through the HTTP or FTP protocol;

[0086] S32: According to the collection frequency defined in S2, send data requests in real time;

[0087] S33: According to the set access method and collection frequency, obtain data from various data sources; through the implementation of the above steps, the most suitable connection method and collection frequency can be selected according to the characteristics of different data sources, and data can be obtained efficiently, providing stable and reliable support for subsequent data storage and analysis.

[0088] S4 specifically includes:

[0089] S41: After receiving the data collected in S3, perform format conversion on the data according to the data format requirements formulated in S2;

[0090] S42: Clean the converted data, including removing duplicate data, filling in missing values, and standardizing the data format; specifically for missing numerical data, use the mean value filling or interpolation method to fill in; for duplicate records, remove duplicates by comparing unique identifiers (such as transaction numbers, tax numbers, etc.); for date or amount data with inconsistent formats, use regular expressions to format the data to ensure consistency;

[0091] S43: Store the cleaned data into the data warehouse according to the requirements of the corresponding data structure; for financial data sources, store the data in a relational database (such as MySQL or PostgreSQL); for transaction data sources, store it in a NoSQL database (such as MongoDB); for tax data sources, store the data in a big data format (such as Parquet or ORC); through the above steps, the format conversion and cleaning of the data can be completed in an efficient and standardized manner, ensuring the unity and availability of the data; the uniformly formatted data after cleaning not only improves the efficiency of data processing, but also lays a foundation for subsequent data analysis and quality assessment, ensuring data quality and accuracy.

[0092] S5 specifically includes:

[0093] S51: Use the Variational Autoencoder (VAE) algorithm to calculate the reconstruction error by mapping the data to the latent space and reconstructing the original data; if the error exceeds the set threshold, it is considered that there is an integrity problem with the data, and the data is marked as missing or incomplete.

[0094] S52: Based on the Graph Neural Network (GNN) algorithm, regard the data as a graph structure, calculate the similarity between the data point and its adjacent nodes by learning the relationships between the nodes, and mark it as abnormal data when the similarity is lower than the preset threshold.

[0095] S53: Reconstruct and compare the same type of data from different data sources through the variational autoencoder, calculate the differences between the data, and if the differences exceed the preset range, judge the data as inconsistent and mark it as a consistency problem; the above steps realize a comprehensive evaluation of data quality through the variational autoencoder and graph neural network algorithms, can effectively check the integrity, abnormality and consistency of the data, improve the accuracy and reliability of data processing, and contribute to ensuring the high quality and consistency of the data.

[0096] S51 specifically includes:

[0097] S511: Represent the data in the data warehouse as an input matrix where n represents the number of samples and m represents the feature dimension of each sample; and map the input data X to the latent space Z through the encoder part of the variational autoencoder, and the expression is: Z i = f enc (X i ; θ enc ), where, is the representation of the i-th sample in the latent space, d is the dimension of the latent space; f enc is the encoding function; θ enc is the parameter of the encoder; X i represents the input data of the i-th sample;

[0098] S512: Assume that the latent space Z follows a Gaussian distribution N(μ, σ 2 ), where μ i and σ i represent the mean and standard deviation of the i-th sample latent space respectively; the specific calculation formula is: μ i = g enc (X i ; θ μ ); σ i = h enc (X i ; θ σ ), where g enc and h enc are functions in the encoder network used to calculate the mean and standard deviation respectively, and θ μ and θ σ are the parameters for mean and standard deviation calculation respectively;

[0099] S513: Reconstruct the data Z in the latent space into an approximation of the original data i through the decoder part, and the calculation formula is: where is the reconstructed data of the i-th sample, f is the decoder function, and θ dec is the parameter of the decoder; dec is the parameter of the decoder;

[0100] S514: Evaluate the integrity of the data by calculating the error between the original data X i and the reconstructed data . Use the mean squared error MSE as the evaluation index, and the calculation formula is: where MSE i is the mean squared error of the i-th sample, X ij and are the original data and the reconstructed data of the j-th feature of the i-th sample respectively;

[0101] S515: Set a threshold. If exceeds this threshold, it is considered that there is an integrity problem with the data of the corresponding i-th sample, and it is marked as missing or incomplete data; by reconstructing the data through the variational autoencoder algorithm and calculating the reconstruction error, the integrity problem of the data can be efficiently detected; when the reconstruction error exceeds the set threshold, the system can accurately mark the missing or incomplete data, thereby improving the quality and consistency of the data, ensuring the reliability of the data, and providing guarantee for subsequent data analysis.

[0102] The specific abnormal data detection in S52 includes:

[0103] S521: Represent each data point in the data warehouse as a node in a graph, where each node represents a data sample; and construct a graph structure based on the feature similarity between data points. For each pair of data points, if the similarity between their features exceeds a preset threshold, an edge is established between these two data points in the graph; the structure of the graph reflects the mutual relationship between data points, where edges are connected between similar data points, indicating a strong association between them, while dissimilar data points have no connected edges;

[0104] S522: Calculate the similarity between a data point and its adjacent nodes through a graph neural network algorithm. For each data point, information is transmitted and aggregated through its adjacent nodes to learn the similarity relationship between the data point and its adjacent nodes; the similarity between nodes is judged through the connection relationship between nodes, and the similarity between each data point and its adjacent nodes is calculated as an index to measure its abnormality;

[0105] S523: When the calculated similarity is lower than the set threshold, it is considered that the relationship between the data point and its adjacent nodes is weak, indicating that there is an abnormality in the features of the corresponding data point; and the data point is marked as an abnormal data; abnormal data points usually show deviation from other data points in the feature space, and their similarity with adjacent nodes is significantly lower than that of other nodes.

[0106] The specific steps for abnormal data detection are as follows:

[0107] Graph structure construction: Represent the data as a graph structure, where each data point d in the graph i is used as a node, and the similarity between data points is connected by edges; for each pair of data points d i and d j , if the similarity between their features is higher than the set threshold, an edge is added to the graph; the adjacency matrix of the graph is defined as:

[0108]

[0109] where similarity(d i ,d j ) represents the similarity between data points d i and d j , τ is the similarity threshold, and if the similarity between two nodes is greater than τ, an edge is established in the graph;

[0110] Graph neural network learning node representation: Learn the relationship between nodes through a graph neural network, use a graph convolutional network to embed nodes, and update the node representation h i , and the calculation formula is: where, is the representation of the i-th node at the k-th layer, and (i) represents the set of neighbor nodes of node i, A ij is an element in the adjacency matrix, W (k) and b (k) are the weight and bias of the k-th layer respectively, and σ is the activation function; the representation of each node is updated through multi-layer graph convolution operations to learn the relationships between nodes;

[0111] Calculate node similarity: After multi-layer convolution of the graph neural network, the final representation of the node is used to calculate the similarity between node i and its adjacent node j; the node similarity S ij is defined by calculating the cosine similarity between node representations: where, and are the final representations of node i and node j after K-layer graph convolution, and represent the norms of the node representations respectively;

[0112] Anomaly data marking: Compare the calculated node similarity S ij with a preset similarity threshold. When S ij is lower than this threshold, it is considered that the relationship between the data point d i and its adjacent nodes is weak, and this data point is marked as abnormal data; through the anomaly data detection method based on the graph neural network algorithm, the audit system can effectively learn the potential relationships between nodes, accurately calculate the similarity of data points, and thus accurately identify the abnormal data that is significantly different from most data points. This method can improve the accuracy and efficiency of data quality detection, optimize the data analysis process, help discover potential abnormal patterns, and enhance the reliability of data analysis.

[0113] S6 specifically includes:

[0114] S61: Configure different data synchronization strategies according to the type of data source and data update frequency; for financial data sources, set it to automatic synchronization daily; for transaction data sources, set it to synchronization hourly; for tax data sources, set it to synchronization weekly;

[0115] S62: Trigger the data synchronization operation through a scheduled task scheduler (such as a Cron scheduler), and automatically start the data synchronization process at a predetermined time point according to the preset synchronization period to ensure that the information in the data warehouse is consistent with the external data source and the timeliness of the data is guaranteed;

[0116] S63: During each synchronization process, first obtain the latest data from the external data source and compare it with the existing data; check whether the newly obtained data meets the format requirements and perform deduplication to ensure that the merged data is consistent and non-duplicate; through the regular data synchronization mechanism, the above steps can ensure that the information in the data warehouse is always consistent with the external data source, guarantee the timeliness and accuracy of the data, reduce manual intervention, and optimize the data management efficiency.

[0117] The present invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0118] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for collecting information of audit big data, characterized in that: The following steps are involved: S1: Connect the audit system with multiple external data sources to identify and classify various data sources, including financial data sources, transaction data sources, and tax data sources; S2: Based on the characteristics of various data sources identified in S1, formulate corresponding collection strategies, including data access methods, collection frequency and data format requirements; S3: Obtain data from various data sources in real time according to the collection strategy formulated in S2; S4: Store the data collected in S3 into the data warehouse according to the preset data structure, and perform preliminary data cleaning and sorting to ensure the uniformity of data format; S5: Apply variational autoencoders and graph neural network algorithms to evaluate the quality of data stored in the data warehouse, including data integrity check, abnormal data detection, and data consistency verification; S6: Through the data synchronization mechanism, regularly update the information in the data warehouse to ensure the timeliness and accuracy of the data.

2. According to the information collection method of audit big data according to claim 1, it is characterized in that: The S1 specifically includes: S11: The audit system connects to the external data source through the API interface and automatically obtains the metadata of the external data source, including the data source name, data format and access protocol; S12: Based on the identified metadata, use natural language processing technology to perform semantic analysis on the data source name and description, match them through the set keyword library, and identify the type of data source; S13: According to the data source types identified in S12, the data sources are classified into financial data sources, transaction data sources, and tax data sources, and labels are added to each type of data source to mark its specific characteristics.

3. The method for collecting information of audit big data according to claim 1, characterized in that: The S12 specifically includes S121: Convert the data source name and description into word vector representation, use Word2Vec or TF-IDF method to encode the text to form a vector representation; suppose the data source description is D = {d1, d2, ..., d n }, where each d i A word for the data source; S122: Construct keyword library K = {k1, k2, ..., k m }, where each k i A term representing a category, including finance, transactions, and tax; S123: Calculate the data source description D and each term k in the keyword library i The similarity is calculated using the cosine similarity formula, which is: Among them, D and k i is the word vector representation, ∥D∥ and ∥k i ∥ is the modulus of the vector, · represents the dot product of the vector; S124: Based on the similarity score sim(D, k i ), matches the data source description with the relevant categories in the keyword library; if the similarity score of a category exceeds the preset threshold, the data source is classified into the corresponding category, including finance, transaction and tax.

4. The method for collecting information of audit big data according to claim 1, characterized in that: The S2 specifically includes: S21: Based on the data source types identified in S1, extract characteristic information of each type of data source, including data format, data structure, update frequency and data access method; S22: Based on the feature information extracted in S21, a rule engine is used to formulate specific collection strategies for different data source categories, including data access methods, collection frequency, and data format requirements; Select the corresponding data access method based on the type of data source and the access protocol provided; for financial data sources, choose to access data through SQL database connection; for transaction data sources, choose to obtain data through RESTful API; for tax data sources, choose a file access protocol based on XML / JSON format; Set different collection frequencies based on the update frequency and timeliness requirements of the data source; for financial data sources, set it to daily scheduled collection; for transaction data sources, set it to hourly collection; for tax data sources, set it to weekly collection; According to the storage format of different data sources, corresponding data format requirements are formulated; for financial data sources, the collected format is required to be structured data; for transaction data sources, the format is required to be JSON; for tax data sources, the format is required to be XML.

5. The method for collecting information of audit big data according to claim 4 is characterized in that: The S3 specifically includes: S31: According to the collection strategy formulated in S2, select the corresponding connection method to establish a connection with the data source; for financial data sources of SQL database type, connect through the JDBC protocol; for transaction data sources of RESTful API type, connect through the HTTP request interface; for tax data sources in XML / JSON format, access files through HTTP or FTP protocol; S32: Send data request in real time according to the collection frequency defined in S2; S33: Obtain data from various data sources according to the set access mode and collection frequency.

6. The method for collecting information of audit big data according to claim 1, characterized in that: The S4 specifically includes: S41: after receiving the data collected in S3, convert the data format according to the data format requirements set in S2; S42: Clean the converted data, including removing duplicate data, filling missing values, and standardizing data formats; specifically, for missing numerical data, use mean filling or interpolation to fill in; for duplicate records, remove duplicates by comparing unique identifiers; for date or amount data with inconsistent formats, use regular expressions to format the data to ensure data consistency; S43: According to the corresponding data structure requirements, the cleaned data is stored in the data warehouse; for financial data sources, the data is stored in the relational database; for transaction data sources, the data is stored in the NoSQL database; for tax data sources, the data is stored in the big data format.

7. The method for collecting information of audit big data according to claim 1, characterized in that: The S5 specifically includes: S51: Using the variational autoencoder algorithm, the reconstruction error is calculated by mapping the data into the latent space and reconstructing the original data. If the error exceeds the set threshold, the data is considered to have integrity problems and marked as missing or incomplete data. S52: Based on the graph neural network algorithm, the data is regarded as a graph structure, and the similarity between the data point and its adjacent nodes is calculated by learning the relationship between the nodes. When the similarity is lower than the preset threshold, it is marked as abnormal data; S53: Reconstruct and compare the same type of data from different data sources through the variational autoencoder, calculate the difference between the data, and if the difference exceeds the preset range, judge the data as inconsistent and mark it as a consistency problem.

8. The method for collecting information of audit big data according to claim 7 is characterized in that: The S51 specifically includes: S511: Representing data in the data warehouse as an input matrix Where n represents the number of samples, m represents the feature dimension of each sample; and the input data X is mapped to the latent space Z through the encoder part of the variational autoencoder, expressed as: Z i =f enc (X i θ enc ),in, is the representation of the i-th sample in the latent space, d is the dimension of the latent space; f enc is the encoding function; θ enc is the encoder parameter; X i Represents the input data of the i-th sample; S512: Assume that the latent space Z follows a Gaussian distribution N(μ, σ 2 ), where μ i and σ i Respectively represent the mean and standard deviation of the i-th sample potential space; the specific calculation formula is: μ i =g enc (X i θ μ );σ i =h enc (X i θ σ ), where g enc and h enc are functions used in the encoder network to calculate the mean and standard deviation, respectively, θ μ and θ σ are the parameters for calculating the mean and standard deviation respectively; S513: The data Z in the latent space is transformed through the decoder part i Reconstruct an approximation of the original data S514: By calculating the original data X i and reconstruction data The error between them is used to evaluate the integrity of the data, and the mean square error MSE is used as the evaluation indicator. The calculation formula is: Among them, MSE i is the mean square error of the ith sample, X ij and are the original data and reconstructed data of the jth feature of the i-th sample respectively; S515: Set the threshold. If MSE i If the threshold is exceeded, the data of the corresponding i-th sample is considered to have integrity issues and is marked as missing or incomplete data.

9. The method for collecting information of audit big data according to claim 8, characterized in that: The S52 specifically includes: S521: Represent each data point in the data warehouse as a node in a graph, where each node represents a data sample; and construct a graph structure based on the feature similarity between the data points. For each pair of data points, if the similarity between their features exceeds a preset threshold, an edge is established for the two data points in the graph; S522: Calculate the similarity between the data point and its adjacent nodes through the graph neural network algorithm as an indicator to measure its abnormality; S523: When the calculated similarity is lower than the set threshold, it indicates that the feature of the corresponding data point is abnormal; and the data point is marked as abnormal data.

10. The method for collecting information of audit big data according to claim 1, characterized in that: The S6 specifically includes: S61: Configure different data synchronization strategies according to the type of data source and data update frequency; for financial data sources, set it to automatic synchronization every day; for transaction data sources, set it to synchronization every hour; for tax data sources, set it to synchronization every week; S62: triggering a data synchronization operation through a scheduled task scheduler, and automatically starting the data synchronization process at a predetermined time point according to a preset synchronization cycle; S63: During each synchronization process, the latest data is first obtained from the external data source and compared with the existing data; the newly obtained data is checked to see if it meets the format requirements, and deduplication is performed to ensure that the merged data is consistent and non-repetitive.

Citation Information

Patent Citations

  • Power plant equipment defect visualization system and method, computer equipment and storage medium

    CN114153914A

  • Audit data acquisition and fusion method and device, equipment and storage medium

    CN116596683A

  • Adversarial sample generation method and system based on double-feature selection

    CN116701910A

  • Novel multi-source real-time transaction quotation data receiving and processing method

    CN117575791A

  • Emergency data standardization system and method based on data virtualization

    CN118673182A

Cited By

  • Hotspot data auditing method and device, equipment and storage medium

    CN121117017A

  • Multi-data-source information collection method and system based on RPA technology

    CN122412716A

  • A multi-data-source information collection method and system based on RPA technology

    CN122412716B