Knowledge graph-based audit risk measurement method

CN115545468BActive Publication Date: 2026-10-09STATE GRID SHANDONG ELECTRIC POWER CO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211212383.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2026-10-09
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

[0002]目前为止,国内外工程项目审计风险度量主要是靠审计管理人员设置评价指标和权重,这需要经验丰富的专业人员根据进行判定和实施,没有充分利用企业审计信息系统中的数据,且科学性和客观性受限

Benefits of technology

[0050] 1. By utilizing the text data of engineering project audit reports, an audit risk knowledge graph can be automatically generated;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115545468B_ABST
    Figure CN115545468B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of audit management, and particularly relates to a kind of audit risk measurement method based on knowledge graph, comprising the following steps: step one: collect project audit report corpus, and construct operation set U;Step two: the operation set U is preprocessed, and obtains audit risk pending knowledge node;Step three: the relationship link between risk knowledge points is constructed, and constitutes audit risk knowledge graph;Step four: the engineering project audit risk calculation method is constructed;Step five: the engineering project audit risk value in the set time period is sorted, and grade division is carried out, provides a kind of audit risk measurement method based on knowledge graph, which can identify engineering project audit risk and does not depend on professional auditors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audit management technology, specifically to an audit risk measurement method based on knowledge graphs. Background Technology

[0002] Currently, the measurement of audit risks in domestic and international engineering projects mainly relies on audit managers setting evaluation indicators and weights. This requires experienced professionals to make judgments and implement them, failing to fully utilize data from enterprise audit information systems and limiting its scientific rigor and objectivity. How to discover audit risk knowledge from textual data without relying on professional auditors, and improve the efficiency of risk calculation, has become an urgent audit management problem. Solving this problem will enhance the enterprise's data mining algorithm capabilities and provide scientific support for audit departments and enterprise decision-making. Summary of the Invention

[0003] The technical problem to be solved by this invention is to overcome the shortcomings of the prior art and provide a knowledge graph-based audit risk measurement method that can identify audit risks in engineering projects without relying on professional auditors.

[0004] The technical solution adopted by this invention to solve its technical problem is: a knowledge graph-based audit risk measurement method, comprising the following steps:

[0005] Step 1: Collect project audit report corpus and construct operation set U;

[0006] Step 2: Preprocess the operation set U to obtain knowledge nodes with pending audit risks;

[0007] Step 3: Construct the relationships between risk knowledge points to form an audit risk knowledge graph;

[0008] Step 4: Construct a method for calculating audit risks in engineering projects; provide a method for quickly calculating and ranking audit risks in engineering projects without relying on professional auditors.

[0009] Step 5: Sort the audit risk values ​​of engineering projects within the set time period and classify them into levels.

[0010] The formula for calculating the operation set U in step one is as follows:

[0011]

[0012] In the public statement (1), n ​​is the total number of engineering projects in audit management, U represents the operation set, P1, P2...P n These are the 1st, 2nd, ... nth engineering projects.

[0013] In step two, text mining technology is used to process the data in the operation set U to obtain named entities and their corresponding relationships.

[0014] Step two includes the following sub-steps:

[0015] 2-1: Divide the operation set U into segments according to the whole sentence, then perform word segmentation and remove stop words;

[0016] 2-2: Using Spacy to identify named entities whose corresponding word is k i Let i = 1, 2, ..., n, where n is an integer, forming the set KG = {k1, k2, ..., kn}. n}, as an alternative head node or result node in audit risk events, is segmented into the word ck according to the content of the National Audit Data Dictionary. j j = 1, 2, ..., p, where p is an integer, forming the reference entity set CH_KG = {ck1, ck2, ..., ck3}. p};

[0017] 2-3: Using the Spacy and zh_web_lg models, word vectors are trained and assigned to named entities and reference entities in the set KG and set CH_KG, resulting in two sets of word counts with a dimension of 300;

[0018] 2-4: Select any candidate entity k from set KG i Compare the word vector cosine similarity with all reference entities in CH_KG. If the similarity is greater than 50%, then k is... i As audit risk event nodes, they are added to set V. V = {v1, v2, ... v} m} contains m nodes.

[0019] In steps 2-3, the calculation formulas for the two sets of term counts with a dimension of 300 are as follows:

[0020] KG′={k1[w 1,1 w 1,2 ...w 1,300 ], k2[w 2,1 w 2,2 ...w 2,300 ...k n [w n,1 w n,2 ...w n,300 ]};

[0021] CH_KG′={ck1[cw 1,1 cw 1,2 ...cw 1,300 ], ck2[cw 2,1 cw2,2 ...cw 2,300 ...ck p [cw p,1 cw p,2 ...cw p,300 ]};

[0022] In the formula, KG′ is the set of word vectors corresponding to the named entities, and the named entity k i The word vector is denoted as

[0023] k i [w i,1 w i,2 ...w i,l ...w i,300 ], w i,l The value of the l-th dimension, w i,l The range of values ​​is [0, 1], l = 1, 2, ... 300.

[0024] CH_KG′ is a set of word vectors consisting of word vectors corresponding to named entities, where the named entity ck j The word vector is denoted as

[0025] ck j [cw j,1 cw j,2 ...cw j,l ...cw j,300 ], cw j,l It is the value of the l-th dimension, cw j,l The range of values ​​is [0, 1], l = 1, 2, ... 300.

[0026] In steps 2-4, the cosine similarity calculation formula is as follows:

[0027]

[0028] Where k i [w i,1 w i,2 ...w i,300 ] is the named entity k i Word vectors, ck j [cw j,1 cw j,2 ...cw j,300 ] is a reference to the named entity ck j Word vectors, w il It is a named entity k i The value of the l-th dimension of the word vector, cw j,l It is the value of the l-th dimension.

[0029] Step three involves constructing relational edges between risk knowledge points using semantic role labeling. Step three includes the following sub-steps:

[0030] 3-1: Initialize G(V, L) as a knowledge graph network, where V is generated in step two, and input the set of all audit report texts U again; V is not mentioned in step two, please supplement how V is generated and the meaning of each letter in G(V, L).

[0031] 3-2: The operation set U is segmented according to the whole sentence, retaining the syntactic analysis structures of noun + noun, noun + verb, verb + noun, adjective + noun, and noun + adjective. The words in the sentence are labeled with semantic roles using the Spacy semantic role labeling toolkit. The words corresponding to the agent of the action (A0), the core predicate (Vb), and the patient (A1) are selected. If a word satisfies A0∈V and A1∈V, then an edge l is added to the knowledge graph network G. ij (v i wl ij TR i v j ), v i v j Corresponding to A0 and A1 respectively, wl ij It is the weight of the connected edges, which is updated based on the frequency of the connected edges. (TR) i Enter Vb;

[0032] 3-3: Delete isolated nodes in G(V,L).

[0033] Step four includes the following sub-steps:

[0034] 4-1: Input the audit report corpus of the project to be evaluated and generate a keyword set;

[0035] 4-2: Traverse the knowledge graph G(V, L) and generate a subgraph G'(V', L');

[0036] 4-3: Compute node v′ i degree and Representation degree is The probability of a node appearing in the knowledge graph G;

[0037] 4-4: Calculate the mutual information between interconnected nodes in the subgraph. The formula is as follows:

[0038]

[0039] In the formula, Represents node v′ i and v′ j The joint probability distribution;

[0040] 4-5: Calculate the audit risk value (risk) of project m. m The larger the calculated value, the greater the risk.

[0041] The audit risk value m The calculation formula is as follows:

[0042]

[0043] Where, Length k It is the path length of the k-th path in subgraph G'. This integrates the risks inherent in the relationships between multiple knowledge nodes. Represents node v′ i and v′ j The joint probability distribution, Representation degree is The probability of a node appearing in the knowledge graph G.

[0044] Step five, which uses the quartile method to classify the audit risk value of the engineering project, includes the following sub-steps:

[0045] 5-1: Calculate the audit risk value for all engineering projects. m Sort by size from smallest to largest;

[0046] 5-2: Find the quartiles using the interquartile method: Q1 (25%), Q2 (50%), Q3 (75%), and Q4 (100%).

[0047] 5-3: risk m Values ​​between 0 and Q1 are classified as class R1, values ​​between Q1 and Q2 as class R2, values ​​between Q2 and Q3 as class R3, values ​​between Q3 and Q4 as class R, and values ​​exceeding Q4 as class R5.

[0048] R5 represents major risk, with many risk factors and high management difficulty, which will cause huge economic losses or major accidents; R4 represents significant risk, with many risk factors and high management difficulty, which will cause large economic losses or major accidents; R3 represents general risk, which is within the controllable range and will cause significant economic losses; R2 represents low risk, which is within the controllable range and will cause general economic losses; R1 represents low risk, which is within the controllable range and will cause minor economic losses.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] 1. By utilizing the text data of engineering project audit reports, an audit risk knowledge graph can be automatically generated;

[0051] 2. This method can structure and make explicit the risk management experience and knowledge in internal auditing; by using text processing technology to structure the audit report corpus data, an audit risk knowledge graph is generated, providing an effective enterprise knowledge management method.

[0052] 3. Projecting the audit corpus of engineering projects onto a knowledge graph network and using the network structure to calculate audit risks can simultaneously consider the persistence and significance of risks, enabling objective internal audit risk management. Attached Figure Description

[0053] Figure 1 This is the distribution chart of audit risk values ​​for engineering projects in Example 2.

[0054] Figure 2 This is the quartile distribution of audit risk values ​​for engineering projects in Example 2. Detailed Implementation

[0055] The embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0056] Example 1

[0057] The knowledge graph-based audit risk measurement method includes the following steps:

[0058] Step 1: Collect project audit report corpus and construct operation set U;

[0059] The formula for calculating the operation set U in step one is as follows:

[0060]

[0061] In the public statement (1), n ​​is the total number of engineering projects in audit management, U represents the operation set, P1, P2...P n These are the 1st, 2nd, ... nth engineering projects.

[0062] Step 2: Preprocess the operation set U to obtain knowledge nodes with pending audit risks; in Step 2, text mining technology is used to process the data in the operation set U to obtain named entities and their correspondence.

[0063] Step two includes the following sub-steps:

[0064] 2-1: Divide the operation set U into segments according to the whole sentence, then perform word segmentation and remove stop words;

[0065] 2-2: Using Spacy to identify named entities whose corresponding word is k i Let i = 1, 2, ..., n, where n is an integer, forming the set KG = {k1, k2, ..., kn}.n}, as an alternative head node or result node in audit risk events, is segmented into the word ck according to the content of the National Audit Data Dictionary. j j = 1, 2, ..., p, where p is an integer, forming the reference entity set CH_KG = {ck1, ck2, ..., ck3}. p};

[0066] 2-3: Using the Spacy and zh_web_lg models, word vectors are trained and assigned to named entities and reference entities in the KG and CH_KG sets, resulting in two sets of term counts with a dimension of 300. The calculation formula for the two sets of term counts with a dimension of 300 in step 2-3 is as follows:

[0067] KG′={k1[w 1,1 w 1,2 ...w 1,300 ], k2[w 2,1 w 2,2 ...w 2,300 ...k n [w n,1 w n,2 ...w n,300 ]};

[0068] CH_KG′={ck1[cw 1,1 cw 1,2 ...Cw 1,300 ], ck2[cw 2,1 cw 2,2 ...cw 2,300 ...ck p [cw p,1 cw p,2 ...cw p,300 ]};

[0069] In the formula, KG′ is the set of word vectors corresponding to the named entities, and the named entity k i The word vector is denoted as

[0070] k i [w i,1 w i,2 ...w i,l ...w i,300 ], w i,l The value of the l-th dimension, w i,l The range of values ​​is [0, 1], l = 1, 2, ... 300.

[0071] CH_KG′ is a set of word vectors consisting of word vectors corresponding to named entities, where the named entity ck j The word vector is denoted as ckj [cw j,1 cw j,2 ...cw j,l ...cw j,300 ], cw j,l It is the value of the l-th dimension, cw j,l The range of values ​​is [0, 1], l = 1, 2, ... 300.

[0072] 2-4: Select any candidate entity k from set KG i Compare the word vector cosine similarity with all reference entities in CH_KG. If the similarity is greater than 50%, then k is... i As audit risk event nodes, they are added to set V. V = {v1, v2, ... v} m} contains m nodes.

[0073] In steps 2-4, the cosine similarity calculation formula is as follows:

[0074]

[0075] Where k i [w i,1 w i,2 ...w i,300 ] is the named entity k i Word vectors, ck j [cw j,1 cw j,2 ...cw j,300 ] is a reference to the named entity ck j Word vectors, w il It is a named entity k i The value of the l-th dimension of the word vector, cw j,l It is the value of the l-th dimension.

[0076] Step 3: Constructing the relationships between risk knowledge points to form an audit risk knowledge graph; Step 3 involves constructing the relationships between risk knowledge points using semantic role labeling. Step 3 includes the following sub-steps:

[0077] 3-1: Initialize G(V, L) as a knowledge graph network, where V is generated in step two, and input the set U of all audit report texts again;

[0078] 3-2: The operation set U is segmented according to the whole sentence, retaining the syntactic analysis structures of noun + noun, noun + verb, verb + noun, adjective + noun, and noun + adjective. The words in the sentence are labeled with semantic roles using the Spacy semantic role labeling toolkit. The words corresponding to the agent of the action (A0), the core predicate (Vb), and the patient (A1) are selected. If a word satisfies A0∈V and A1∈V, then an edge l is added to the knowledge graph network G. ij (v i wl ij TR i v j ), v i v j Corresponding to A0 and A1 respectively, wl ij It is the weight of the connected edges, which is updated based on the frequency of the connected edges. (TR) i Enter Vb;

[0079] 3-3: Delete isolated nodes in G(V,L).

[0080] The semantic role labeling toolkit comes from the paper Rs A, Mj A. The possibility of improving automated calculation of measures of lexical richness for EFL writing: A comparison of the LCA, NLTK and SpaCy tools. 2022.

[0081] Step 4: Construct a method for calculating audit risks in engineering projects; provide a method for quickly calculating and ranking audit risks in engineering projects without relying on professional auditors.

[0082] Step four includes the following sub-steps:

[0083] 4-1: Input the audit report corpus of the project to be evaluated and generate a keyword set;

[0084] 4-2: Traverse the knowledge graph G(V, L) and generate a subgraph G'(V', L');

[0085] 4-3: Compute node v′ i degree and Representation degree is The probability of a node appearing in the knowledge graph G;

[0086] 4-4: Calculate the mutual information between interconnected nodes in the subgraph. The formula is as follows:

[0087]

[0088] In the formula, Represents node v′ i and v′ j The joint probability distribution;

[0089] 4-5: Calculate the audit risk value (risk) of project m. m The larger the calculated value, the greater the risk.

[0090] The audit risk value m The calculation formula is as follows:

[0091]

[0092] Where, Length k It is the path length of the k-th path in subgraph G'. This integrates the risks inherent in the relationships between multiple knowledge nodes. Represents node v′ i and v′ j The joint probability distribution, Representation degree is The probability of a node appearing in the knowledge graph G.

[0093] Step 5: Generate a level determination result from the calculation results of the audit risk of the engineering project, sort the audit risk values ​​of the engineering projects within the set time period, and classify them into levels.

[0094] Step five, which uses the quartile method to classify the audit risk value of the engineering project, includes the following sub-steps:

[0095] 5-1: Calculate the audit risk value for all engineering projects. m Sort by size from smallest to largest;

[0096] 5-2: Find the quartiles using the interquartile method: Q1 (25%), Q2 (50%), Q3 (75%), and Q4 (100%).

[0097] 5-3: risk m Values ​​between 0 and Q1 are classified as class R1, values ​​between Q1 and Q2 as class R2, values ​​between Q2 and Q3 as class R3, values ​​between Q3 and Q4 as class R, and values ​​exceeding Q4 as class R5.

[0098] R5 represents major risk, with many risk factors and high management difficulty, which will cause huge economic losses or major accidents; R4 represents significant risk, with many risk factors and high management difficulty, which will cause large economic losses or major accidents; R3 represents general risk, which is within the controllable range and will cause significant economic losses; R2 represents low risk, which is within the controllable range and will cause general economic losses; R1 represents low risk, which is within the controllable range and will cause minor economic losses.

[0099] Example 2

[0100] Reference Figure 1 and Figure 2 Using the method in Example 1, if the project audit report corpus is as shown in the table below:

[0101]

[0102] Subsequently, after steps two and three, an audit risk knowledge graph is generated. After traversing the knowledge graph, the degree of each node and the probability of each node appearing are calculated, as shown in the table below:

[0103]

[0104] The mutual information between interconnected nodes in the subgraph is calculated, as shown in the table below:

[0105]

[0106] The audit risk value of project m m Calculation results

[0107] The final distribution of audit risk values ​​for all engineering projects is as follows: Figure 1 As shown.

[0108] All project audit risk values ​​were sorted from smallest to largest. Using the quartile method, the quartiles were determined, and the project risk levels were divided into five categories. Values ​​[6.5176, 12.274] resulted in 0-25% belonging to category R1; values ​​[12.274, 17.988] resulted in 25%-50% belonging to category R2; values ​​[17.988, 25.575] resulted in 50%-75% belonging to category R3; values ​​[25.575, 42.672] resulted in 75%-100% belonging to category R4; and values ​​greater than 45.672 and deviating from the quartile range belonged to category R5. The visualization results of the quartile method are shown below. Figure 2 As shown.

Claims

1. A knowledge graph-based method for measuring audit risk, characterized in that, Includes the following steps: Step 1: Collect project audit report corpus and construct operation set U; the calculation formula for constructing operation set U in Step 1 is as follows: ; Where, n is the total number of engineering projects in audit management, Represents the set of operations. These are the 1st, 2nd, ..., nth engineering projects; Step Two: Preprocess the operation set U to obtain risk knowledge nodes; Step Two uses text mining technology to process the data in the operation set U to obtain named entities and their correspondences; Step Two includes the following sub-steps: 2-1: Divide the operation set U into segments according to the whole sentence, then perform word segmentation and remove stop words; 2-2: Identifying the corresponding words of named entities using Spacy Let i = 1, 2, ..., n, where n is an integer, forming the set. ={ }, as an alternative head node or result node in audit risk events, is segmented into words according to the content of the National Audit Data Dictionary. j = 1, 2, ..., p, where p is an integer, constitute the set of reference entities. ={ }; 2-3: Using the Spacy and zh_web_lg models, as a set ,gather The named entities and reference entities in the training word vectors are assigned values ​​to obtain two sets of word vectors with a dimension of 300; 2-4: Set Any candidate entity in and All reference entities are compared using word vector cosine similarity. If the similarity is greater than 50%, then... Added to the set as a risk knowledge node middle; Step 3: Constructing the relationships between risk knowledge nodes to form an audit risk knowledge graph; Step 3 involves constructing the relationships between risk knowledge nodes using semantic role labeling. Step 3 includes the following sub-steps: 3-1: Initialization For a knowledge graph network, V is generated in step two and then input into the operation set again. ; 3-2: Segment the operation set U according to the entire sentence, retaining the syntactic analysis structures of noun + noun, noun + verb, verb + noun, adjective + noun, and noun + adjective. Use the Spacy semantic role labeling toolkit to label the words in the sentence with semantic roles, filtering out the word A0 corresponding to the agent label, the word Vb corresponding to the core predicate label, and the word A1 corresponding to the patient label. If the words satisfy... In knowledge graph networks Add edges , Corresponding to A0 and A1 respectively, It represents the weight of the edges, which is updated based on the frequency of each edge. Enter Vb; 3-3: Delete Isolated nodes in; Step Four: Constructing a method for calculating audit risks in engineering projects; Step Four includes the following sub-steps: 4-1: Input the audit report corpus of the project to be evaluated and generate a keyword set; 4-2: Traversing the Knowledge Graph Generate subgraph ; 4-3: Computation Nodes degree and , Representation degree is The nodes in the knowledge graph The probability of its occurrence; 4-4: Calculate the mutual information between interconnected nodes in the subgraph. The formula is as follows: ; In the formula, Representation degree is and The joint probability distribution of the nodes; 4-5: Calculate the audit risk value of project m. The larger the calculated value, the greater the risk. The audit risk value The calculation formula is as follows: ; in, Subgraph The path length of the k-th path in the equation. This integrates the risks inherent in the relationships between multiple knowledge nodes. Representation degree is and The joint probability distribution of the nodes. , Representation degree is , The nodes in the knowledge graph The probability of its occurrence; Step 5: Classify the audit risk values ​​of engineering projects within the set time period into different levels.

2. The audit risk measurement method based on knowledge graphs according to claim 1, characterized in that, In steps 2-3, the calculation formula for the two word vector sets with a dimension of 300 is as follows: ={ }; { } ; In the formula, It is a set of word vectors consisting of the word vectors corresponding to named entities. The word vector is denoted as , It is the first The value of dimension, The value range is [0,1]. ; It is a set of word vectors consisting of the word vectors corresponding to the reference entities, where the reference entities... The word vector is denoted as It is the first The value of dimension, The value range is [0,1]. .

3. The audit risk measurement method based on knowledge graphs according to claim 2, characterized in that, In steps 2-4, the cosine similarity calculation formula is as follows: ; in Named entities Word vectors, Reference entity Word vectors, Named entities Word vectors The value of dimension, Reference entity Word vectors The value that a dimension can take.

4. The audit risk measurement method based on knowledge graphs according to claim 1, characterized in that, Step five, which uses the quartile method to classify the audit risk value of the engineering project, includes the following sub-steps: 5-1: Calculate the audit risk value for all engineering projects. Sort by size from smallest to largest; 5-2: Find the quartiles using the quartile method: Q1, Q2, Q3, Q4; 5-3: Values ​​between 0 and Q1 are classified as class R1, values ​​between Q1 and Q2 as class R2, values ​​between Q2 and Q3 as class R3, values ​​between Q3 and Q4 as class R4, and values ​​exceeding Q4 as class R5. R5 represents major risk, with many risk factors and great difficulty in management, which will cause huge economic losses or major accidents; R4 represents relatively high risk, with many dangerous factors and great difficulty in management, which will cause large economic losses or major accidents. R3 represents general risk, which is within a controllable range but will cause significant economic losses. R2 represents low risk, which is within a controllable range and will cause general economic loss; R1 represents low risk, meaning the risk is within a controllable range and will cause minor economic losses.

Citation Information

Patent Citations

  • Intelligent question and answer intention recognition method based on knowledge graph

    CN114579709A