An audit method based on big data
Through the audit method based on big data, preprocessing and analyzing different types of data sets, calculating the similarity and matching of the matter nodes, determining the associated matters for auditing, solving the problem of low quality of existing audit data, improving audit quality and efficiency, ensuring compliance and reducing risks.
Patent Information
- Application Number
- CN202510205822.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Among the existing audit methods, the data used for auditing is of low quality, which affects the quality and efficiency of audits, and is difficult to effectively reduce compliance risks.
The audit method based on big data is adopted, and different types of data sets are obtained by acquiring and preprocessing different types of data sets, and the initial position characteristics and importance characteristics of the event node are obtained using Gaussian distribution random initialization. The similarity function is calculated to determine the target importance characteristics, and the associated matters are determined based on the matching degree for auditing.
Improves the quality and accuracy of audit data, helping organizations conduct audits more efficiently, improve audit quality, ensure compliance and reduce risks.
Smart Images

Figure CN119691470B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of auditing methods, and more specifically, it relates to an auditing method based on big data. Background Art
[0002] Auditing is the review and inspection of the economic activities, financial revenues and expenditures, and implementation of financial and economic regulations of a certain industry. It is carried out from top to bottom in an organized, leadership-based, and planned manner by auditing organs and auditors, expanding from the auditing of unit economic activities to the auditing of the economic activities of the entire industry, and shifting from the auditing of micro-economic activities to the auditing of meso-economic activities or macro-economic activities; data, as the basis of auditing, the quality of data will directly affect the quality of auditing, and currently the quality of data used for auditing still needs to be improved. Summary of the Invention
[0003] The purpose of the present invention is to provide an auditing method based on big data, which can improve the quality of data and materials used for auditing work, help organizations and institutions conduct auditing work more efficiently, improve the quality of auditing, and reduce risks while ensuring compliance.
[0004] The above technical purpose of the present invention is achieved through the following technical solutions:
[0005] In a first aspect, the present application provides an auditing method based on big data, including the following specific steps:
[0006] Obtain at least two data sets to be audited of different types, and preprocess the data sets. The preprocessing includes data cleaning and abnormal data detection;
[0007] Based on each preprocessed data set, use Gaussian distribution to randomly initialize the initial position features and initial importance features of each matter node in the data set;
[0008] Use the initial position features and initial importance features to calculate the support given by adjacent nodes of each matter node, and calculate the aggregation features of each matter node that fuse the features of adjacent nodes according to the normalized support;
[0009] Construct a similarity function for each matter node using the initial position features and aggregation features, and determine the target importance features of each matter node as the initial importance features when the similarity function is maximized;
[0010] According to the target importance features of each matter node in at least two data sets, calculate the matching degree between each matter node in different types of data sets, determine two matter nodes with a matching degree not less than the threshold as associated matters, and audit each data set based on each associated matter.
[0011] The beneficial effects of the present invention are as follows: In this solution, first, preprocessing such as data cleaning and abnormal data detection is performed on the dataset to be audited, and duplicate data and abnormal data in each dataset are removed, making the data in the dataset more concise and accurate. Then, the initial position features and initial importance features of each event node in the dataset are obtained through Gaussian distribution random initialization. Secondly, the support degree given by adjacent nodes of each event node is calculated using the initial position features and initial importance features, and the aggregated features of each event node that integrate adjacent node features are calculated based on the normalized support degree. Then, a similarity function for each event node is constructed using the initial position features and aggregated features, and the initial importance features when the similarity function is maximized are determined as the target importance features of each event node. Finally, according to the target importance features of each event node in the two datasets, the matching degree between each event node in different types of datasets is calculated, and two event nodes with a matching degree not less than the threshold are determined as associated events. The associated events indicate a relatively high degree of association between the two times and can provide relatively strong reference for each other, which has a promoting effect on the audit of both datasets. Finally, each dataset is audited based on each associated event.
[0012] In this solution, when auditing the datasets on economic activities and financial revenues and expenditures together, for example, if the expenditure of a certain event in economic activities is a certain value, and the expenditure of a certain item in financial revenues and expenditures is also this value, without knowing the detailed activity content, these two events in these two datasets show a relatively high matching degree. Therefore, when conducting a joint audit of different types of datasets, by finding the connections between each item in the two different types of datasets, it provides a reference for the audit process, which has a high improvement for the audit work. It not only improves the quality of the data materials used for the audit work, helps organizations and institutions conduct audit work more efficiently, but also improves the audit quality and reduces risks while ensuring compliance.
[0013] Based on the above technical solution, the present invention can also be improved as follows.
[0014] Further, the above data cleaning is specifically as follows:
[0015] Discretize the attribute items of each original data in the dataset after normalization processing, and transform each obtained discrete attribute value into a preset integer interval according to the size.
[0016] Calculate the information gain rate of the attribute items of each original data based on the transformed attribute values, and construct an attribute set through the information gain rates not less than the preset value.
[0017] Insert each piece of data corresponding to the attribute set into a preset prefix tree, traverse each leaf node in the prefix tree to delete duplicate data, and obtain a data set after data cleaning.
[0018] The beneficial effect of adopting the above further solution is that the process of data cleaning by calculating the information gain rate and inserting into the prefix tree can reduce the time complexity of the detection process while ensuring the accuracy of the data set.
[0019] Furthermore, the above abnormal data detection is specifically as follows:
[0020] Use the K-Means clustering algorithm to cluster the data set after normalization processing, and obtain multiple data clusters composed of each original data in the data set;
[0021] Based on multiple data clusters, calculate the first Euclidean distance between each data in each data cluster and other data in the same cluster, and the second Euclidean distance between each data in each data cluster and each data in other data clusters;
[0022] Based on the first Euclidean distance, calculate the outlier factor of each data in each data cluster, and determine the data with an outlier factor not less than the first threshold as local isolated data;
[0023] Determine the data with the second Euclidean distance not less than the second threshold as global isolated data, and determine the original data corresponding to the local isolated data and the global isolated data as abnormal data. After deleting the abnormal data, obtain a data set after abnormal data detection.
[0024] The beneficial effect of adopting the above further solution is that since the data to be audited is generally complex and the data volume is large, density-based outliers are not ideal in terms of algorithm execution efficiency and identifying global outliers. Therefore, combining the idea of the clustering algorithm can reduce the algorithm execution time while also identifying global outliers.
[0025] Furthermore, the information gain rate of the attribute items of each of the above original data is specifically as follows:
[0026] , where: ;
[0027] In the formula, represents the information gain rate of the attribute item D in the data set A , represents the information gain of the attribute item D in the data set A , represents the number of samples when the attribute item A takes the value of i . Denotes the total number of samples in the data set D in the n Denotes the number of values of the attribute item A of
[0028] Furthermore, the outlier factor of each data in the above data clusters is specifically:
[0029] ;
[0030] In the formula, Denotes the outlier factor of the data of Denotes the reachable distance of the data of The reachable density of the data of Denotes the set composed of the nearest k data to the data
[0031] Furthermore, the support given by the adjacent nodes of each of the above event nodes is specifically:
[0032] ; In the formula, Denotes the support given by the adjacent node n to the event node m of Denotes the initial importance feature of the event node m of Denotes the initial importance feature of the adjacent node n of
[0033] The aggregated feature of each event node that fuses the features of adjacent nodes is specifically:
[0034] ; In the formula, Denotes the aggregated feature of the event node m of Denotes the set of adjacent nodes of the event node m of Denotes the support after normalization by the softmax function Denotes the initial position feature of the adjacent node n of
[0035] Furthermore, the above similarity function is specifically:
[0036] ; In the formula, Denotes the objective function value Denotes the aggregated feature of the event node m of Indicate the event node m The initial position feature, subscript T Indicates the vector transpose operation;
[0037] The matching degree between each event node is specifically:
[0038] , where:
[0039] , ;
[0040] In the formula, Indicates the event node m And the event node n The matching degree between them, subscript T Indicates the vector transpose operation, Indicates the target importance feature of the event node m , Event node n The target importance feature, Respectively represent the weight matrix and bias term corresponding to the event node m respectively, Respectively represent the weight matrix and bias term corresponding to the event node n respectively.
[0041] In the second aspect, the present application provides a big data-based audit system, which is applied to a big data-based audit method according to any one of the first aspect, and includes:
[0042] The first module is used to obtain at least two datasets to be audited of different types, and preprocess the datasets, and the preprocessing includes data cleaning and abnormal data detection;
[0043] The second module is used to randomly initialize the initial position feature and initial importance feature of each event node in the dataset based on each preprocessed dataset by using the Gaussian distribution;
[0044] The third module is used to calculate the support degree given by adjacent nodes of each event node by using the initial position feature and initial importance feature, and calculate the aggregation feature of each event node integrating adjacent node features according to the normalized support degree;
[0045] The fourth module is used to construct a similarity function of each event node by using the initial position feature and the aggregation feature, and determine the initial importance feature when the similarity function is maximized as the target importance feature of each event node;
[0046] The fifth module is used to calculate the matching degrees between the item nodes in different types of datasets according to the target importance characteristics of the item nodes in at least two datasets, determine two item nodes with a matching degree not less than the threshold as associated items, and audit each dataset based on each associated item.
[0047] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method according to any one of the first aspects is implemented.
[0048] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause a computer to execute the method according to any one of the first aspects.
[0049] Compared with the prior art, the present invention has at least the following beneficial effects:
[0050] In the present application, first, preprocessing such as data cleaning and abnormal data detection is performed on the datasets to be audited, and the duplicate data and abnormal data in each dataset are removed, making the data in the datasets more concise and accurate. Then, the initial position characteristics and initial importance characteristics of each item node in the dataset are obtained by random initialization with a Gaussian distribution. Secondly, the support given by adjacent nodes to each item node is calculated using the initial position characteristics and initial importance characteristics, and the aggregated characteristics of each item node that fuse the characteristics of adjacent nodes are calculated according to the normalized support. Then, a similarity function for each item node is constructed using the initial position characteristics and the aggregated characteristics, and the initial importance characteristics when the similarity function is maximized are determined as the target importance characteristics of each item node. Finally, according to the target importance characteristics of the item nodes in two datasets, the matching degrees between the item nodes in different types of datasets are calculated, and two item nodes with a matching degree not less than the threshold are determined as associated items. The associated items indicate a relatively high degree of association between the two times and a relatively strong reference that can be provided to each other, which is helpful for the audit of the two datasets. Finally, each dataset is audited based on each associated item.
[0051] In this application, when jointly auditing different types of datasets, by finding the connections between various matters in two different types of datasets, it provides a reference for the auditing process, which has a high improvement for the auditing work. It not only improves the quality of the data materials used for auditing work, helps organizations and institutions conduct auditing work more efficiently, but also improves the auditing quality and reduces risks while ensuring compliance; and the process of data cleaning by calculating the information gain rate and inserting into the prefix tree can reduce the time complexity of the detection process while ensuring the accuracy of the dataset; at the same time, since the data to be audited is generally quite complex and the data volume is large, the density-based outlier detection is not ideal in terms of algorithm execution efficiency and identifying global outliers. Therefore, combining the clustering algorithm idea can identify global outliers while reducing the algorithm execution time. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:
[0053] Figure 1 is the flowchart of the auditing method in the embodiment of the present invention;
[0054] Figure 2 is the connection schematic diagram of the auditing system in the embodiment of the present invention;
[0055] Figure 3 is the connection schematic diagram of the electronic device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0057] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0058] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0059] In the description of the embodiments of the present invention, "a plurality of" means at least two.
[0060] Embodiment 1: In order to improve the quality of data for auditing work, help organizations and institutions conduct auditing work more efficiently, and reduce risks while improving audit quality and ensuring compliance, this embodiment provides an auditing method based on big data, as Figure 1 shown, including the following specific steps:
[0061] S1. Obtain at least two data sets to be audited of different types, and preprocess the data sets. The preprocessing includes data cleaning and abnormal data detection.
[0062] Optionally, the above data cleaning is specifically:
[0063] S11. Discretize the attribute items of each original data in the normalized data set, and transform each obtained discrete attribute value into a preset integer interval according to its size.
[0064] S12. Calculate the information gain rate of the attribute items of each original data based on the transformed attribute values, and construct an attribute set through the information gain rates not less than a preset value.
[0065] Among them, the information gain rate of the attribute items of each original data is specifically:
[0066] , where: ;
[0067] In the formula, represents the information gain rate of the attribute item D in the data set A , represents the information gain of the attribute item D in the data set A , represents the number of samples with the value of the attribute item A being i , represents the total number of samples in the data set D , n represents the number of value ranges of the attribute item A .
[0068] S13. Insert the corresponding data in the attribute set into a preset prefix tree, traverse each leaf node in the prefix tree to delete duplicate data, and obtain the data set after data cleaning.
[0069] Specifically, the improvements in the characteristics and structure of the preset prefix tree are as follows:
[0070] 1) Non-leaf nodes serve as index splitting items and do not store data information.
[0071] 2) Only leaf nodes store data, and multiple data items can be stored.
[0072] 3) For data with n attributes as index items, there are n + 2 layers. The first layer is the root node, and the last layer stores sample data information.
[0073] Among them, when detecting duplicate data in the prefix tree, first, traverse the data in each leaf node in turn; second, compare each sample point with the data in its corresponding leaf node. If the similarity between two pieces of data is greater than the given threshold, output it to the similar data set. After comparing with other data, mark this data as the data that has been compared and do not compare it again during the next traversal. For example, if A and B are in the same leaf node, compare A and B. If their similarity is greater than the given threshold, output A to the similar data set and delete A from its corresponding leaf node. While for C, D, and G, since there is only one piece of data in their corresponding leaf nodes, no comparison is needed, that is, these single-existing data are not duplicate data.
[0074] Specifically, after improving the prefix tree, similar data can be quickly aggregated in the same leaf node, reducing the operation process. Then traverse the leaf nodes and calculate the similarity between data within the leaf nodes to complete the detection of duplicate data, improving the efficiency of duplicate data detection.
[0075] Optionally, the above abnormal data detection is specifically as follows:
[0076] S14. Use the K-Means clustering algorithm to cluster the dataset after normalization processing and obtain multiple data clusters composed of each original data in the dataset.
[0077] S15. Based on multiple data clusters, calculate the first Euclidean distance between each data in each data cluster and other data in the same cluster, and the second Euclidean distance between each data in each data cluster and each data in other data clusters.
[0078] S16. Calculate the outlier factor of each data in each data cluster based on the first Euclidean distance, and determine the data with an outlier factor not less than the first threshold as local isolated data.
[0079] Among them, the outlier factor of each data in each of the above data clusters is specifically:
[0080] ;
[0081] In the formula, represents the data The outlier factor of represents the data of the reachable distance, data of the reachable density, represents the distance from the data nearest k data points that make up the set.
[0082] S17, determine each data with the second Euclidean distance not less than the second threshold as global isolated data, and determine the original data corresponding to the local isolated data and the global isolated data as abnormal data, and obtain the dataset after abnormal data detection by deleting the abnormal data.
[0083] Specifically, since the data to be audited is generally complex and the data volume is large, the density-based outlier detection is not ideal in terms of algorithm execution efficiency and identifying global outliers. Therefore, combining the clustering algorithm idea can reduce the algorithm execution time while also identifying global outliers.
[0084] S2, based on each preprocessed dataset, use Gaussian distribution to randomly initialize the initial position features and initial importance features of each event node in the dataset.
[0085] Among them, the support degree of adjacent nodes to the current node is reflected by the importance degree feature, and the support degree of the same node to other adjacent nodes is also different; therefore, the position feature reflects the environmental information of adjacent nodes, while the importance degree feature reflects the unique support relationship. Compared with the position feature, the importance degree feature has stronger discrimination.
[0086] S3, use the initial position features and initial importance features to calculate the support degree given by the adjacent nodes of each event node, and calculate the aggregated feature of each event node that fuses the features of adjacent nodes according to the support degree after normalization processing.
[0087] Among them, the support degree given by the adjacent nodes of each event node is specifically:
[0088] ; in the formula, represents the adjacent node n giving the support degree to the event node m , represents the initial importance feature of the event node m , represents the adjacent node n 's initial importance feature;
[0089] Furthermore, the aggregated feature of each event node that fuses the features of adjacent nodes is specifically:
[0090] ; In the formula, represents the aggregation feature of the event node m , represents the set of adjacent nodes of the event node m , represents the support degree after normalization by the softmax function, represents the initial position feature of the adjacent node n .
[0091] S4. Use the initial position feature and the aggregation feature to construct a similarity function for each event node, and determine the initial importance feature when the similarity function is maximized as the target importance feature of each event node; the maximum value of the similarity function is 100%, that is, 1.
[0092] Specifically, the above similarity function is specifically:
[0093] ; In the formula, represents the value of the objective function, represents the aggregation feature of the event node m , represents the initial position feature of the event node m , and the subscript T represents the vector transpose operation.
[0094] S5. According to the target importance features of each event node in at least two data sets, calculate the matching degree between each event node in different types of data sets, determine two event nodes with a matching degree not less than the threshold as associated events, and audit each data set based on each associated event.
[0095] Specifically, when jointly auditing different types of data sets, by finding the connections between each event in two different types of data sets, it provides a reference for the audit process, which has a high improvement for the audit work. It not only improves the quality of the data materials used for the audit work, helps organizations and institutions conduct audit work more efficiently, but also improves the audit quality and reduces risks while ensuring compliance.
[0096] Among them, the matching degree between each event node is specifically:
[0097] , where:
[0098] , ;
[0099] In the formula, represents the event node m and the event node nThe matching degree between, subscript T represents the vector transpose operation, represents the event node m of the target importance feature, event node n of the target importance feature, respectively represent the event node m respectively corresponding weight matrix and bias term, respectively represent the event node n respectively corresponding weight matrix and bias term.
[0100] Embodiment 2: The embodiment of the present application provides a big data-based audit system, which is applied to a big data-based audit method in any one of Embodiment 1, as Figure 2 shown, including:
[0101] The first module is used to obtain at least two datasets to be audited of different types, and preprocess the datasets, and the preprocessing includes data cleaning and abnormal data detection.
[0102] The second module is used to randomly initialize the initial position feature and the initial importance feature of each event node in the dataset based on each preprocessed dataset by using the Gaussian distribution.
[0103] The third module is used to calculate the support degree given by the adjacent nodes of each event node by using the initial position feature and the initial importance feature, and calculate the aggregation feature of each event node that fuses the adjacent node features according to the normalized support degree.
[0104] The fourth module is used to construct the similarity function of each event node by using the initial position feature and the aggregation feature, and determine the initial importance feature when the similarity function is maximized as the target importance feature of each event node.
[0105] The fifth module is used to calculate the matching degree between each event node in at least two datasets according to the target importance feature of each event node, determine two event nodes with a matching degree not less than the threshold as associated events, and audit each dataset based on each associated event.
[0106] Embodiment 3: The embodiment of the present application provides an electronic device, as Figure 3 shown, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method in any one of Embodiment 1 is implemented.
[0107] Example 4: An embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions that cause a computer to execute the method according to any one of the embodiments 1.
[0108] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An audit method based on big data, characterized in that: The specific steps include: Acquire at least two data sets to be audited of different types, and preprocess the data sets, wherein the preprocessing includes data cleaning and abnormal data detection; Based on each preprocessed data set, the initial position features and initial importance features of each event node in the data set are obtained by random initialization using Gaussian distribution; Using the initial position feature and the initial importance feature, calculate the support given by the adjacent nodes of each event node, and calculate the aggregated feature of each event node by fusing the features of the adjacent nodes according to the normalized support; Constructing a similarity function for each event node using the initial position feature and the aggregation feature, and determining the initial importance feature when the similarity function is maximized as the target importance feature of each event node; According to the target importance characteristics of each item node in at least two of the data sets, the matching degree between each item node in different types of data sets is calculated, two item nodes whose matching degree is not less than a threshold are determined as related items, and each data set is audited based on each of the related items; The support given by the adjacent nodes of each event node is as follows: In the formula, α mn It represents the support given by the adjacent node n to the event node m. represents the initial importance feature of the event node m, and the subscript T represents the vector transposition operation. Represents the initial importance feature of the adjacent node n; Each event node integrates the aggregate features of adjacent nodes, specifically: In the formula, represents the aggregated features of event node m, N(v m ) represents the set of adjacent nodes of event node m, β mn represents the support after normalization by the softmax function, Represents the initial position features of the adjacent node n; The similarity function is specifically: Where P s represents the objective function value, represents the aggregation features of event node m, Indicates the initial position feature of event node m, and the subscript T indicates the vector transposition operation; The matching degree between each event node is as follows: q = tanh((h m ×h n ) T ·(h m ×h n )),in: In the formula, q represents the matching degree between event node m and event node n, and the subscript T represents the vector transposition operation. represents the target importance feature of event node m, The target importance feature of event node n, w i 、b i Respectively represent the weight matrix and bias term corresponding to the event node m, w j 、b j They respectively represent the weight matrix and bias term corresponding to the event node n.
2. According to the big data audit method described in claim 1, it is characterized in that: The data cleaning is specifically as follows: Discretize the attribute items of each original data in the normalized data set, and transform each attribute value obtained after discretization into a preset integer range according to the size; The information gain rate of each attribute item of the original data is calculated based on the transformed attribute value, and the attribute set is constructed through each information gain rate that is not less than a preset value; Insert the corresponding data in the attribute set into the preset prefix tree, traverse each leaf node in the prefix tree to delete duplicate data, and obtain a data set after data cleaning.
3. According to the big data audit method described in claim 1, it is characterized in that: The abnormal data detection is specifically as follows: The K-Means clustering algorithm is used to cluster the normalized data set, and multiple data clusters consisting of the original data in the data set are obtained; Based on the plurality of data clusters, a first Euclidean distance between each data in each data cluster and other data in the same cluster is calculated, as well as a second Euclidean distance between each data in each data cluster and each data in other data clusters; Calculate the outlier factor of each data in each data cluster based on the first Euclidean distance, and determine the data whose outlier factor is not less than a first threshold as local isolated data; Each data whose second Euclidean distance is not less than a second threshold is determined as global isolated data, and the original data corresponding to the local isolated data and the global isolated data is determined as abnormal data, and the abnormal data is deleted to obtain a data set after abnormal data detection.
4. According to the big data audit method described in claim 2, it is characterized in that: The information gain rate of the attribute items of each original data is as follows: in: In the formula, g r (D, A) represents the information gain rate of attribute item A in data set D, g(D, A) represents the information gain of attribute item A in data set D, |D i | represents the number of samples whose attribute item A takes the value i, |D| represents the total number of samples in data set D, and n represents the number of values of attribute item A.
5. The audit method based on big data according to claim 3 is characterized in that: The outlier factor of each data in each data cluster is: In the formula, L(X i ) represents data X i The outlier factor, B(Z i ) represents data Z i The reachable distance, ρ(X i )DataX i The accessible density, A ik Represents distance data X i The set of the most recent k data.
6. An audit system based on big data, applied to an audit method based on big data as claimed in any one of claims 1 to 5, characterized in that: include: The first module is used to obtain at least two data sets to be audited of different types and preprocess the data sets, wherein the preprocessing includes data cleaning and abnormal data detection; The second module is used to obtain the initial position features and initial importance features of each event node in the data set based on each preprocessed data set by using Gaussian distribution random initialization; The third module is used to calculate the support given by the adjacent nodes of each event node by using the initial position feature and the initial importance feature, and obtain the aggregated feature of each event node fused with the features of the adjacent nodes according to the normalized support calculation; The fourth module is used to construct a similarity function of each event node using the initial position feature and the aggregation feature, and determine the initial importance feature when the similarity function is maximized as the target importance feature of each event node; A fifth module is used to calculate the matching degree between each item node in different types of data sets according to the target importance characteristics of each item node in at least two of the data sets, determine two item nodes whose matching degree is not less than a threshold as related items, and audit each data set based on each of the related items; In this system, the support given by the adjacent nodes of each event node is as follows: In the formula, α mn It represents the support given by the adjacent node n to the event node m. represents the initial importance feature of the event node m, and the subscript T represents the vector transposition operation. Represents the initial importance feature of the adjacent node n; Each event node integrates the aggregate features of adjacent nodes, specifically: In the formula, represents the aggregated features of event node m, N(v m ) represents the set of adjacent nodes of event node m, β mn represents the support after normalization by the softmax function, Represents the initial position features of the adjacent node n; The similarity function is specifically: Where P s represents the objective function value, represents the aggregation features of event node m, Indicates the initial position feature of event node m, and the subscript T indicates the vector transposition operation; The matching degree between each event node is as follows: q = tanh((h m ×h n ) T ·(h m ×h n )),in: In the formula, q represents the matching degree between event node m and event node n, and the subscript T represents the vector transposition operation. Represents the target importance feature of event node m, Z n H1 The target importance feature of event node n, w i 、b i Respectively represent the weight matrix and bias term corresponding to the event node m, w j 、b j They respectively represent the weight matrix and bias term corresponding to the event node n.
7. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the method according to any one of claims 1 to 5 is implemented when the processor executes the computer program.
8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which enable a computer to execute the method of any one of claims 1-5.
Citation Information
Patent Citations
Auditing method, electronic equipment and storage medium
CN113076352A
Medical data auditing method, device and equipment and storage medium
CN113657549A