Enterprise audit information extraction method based on artificial intelligence
By constructing a multimodal conversion processing and dynamic knowledge graph of enterprise audit information, the problem of poor correlation of multimodal data in enterprise audits is solved, high-precision audit clue discovery and legal terms adaptation are achieved, and audit efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510764404.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing technology has problems in enterprise audits with poor cross-modal correlation of multimodal data and misjudgment of audit clues caused by static legal clause mapping, especially in cross-modal correlation and semantic understanding of unstructured data.
Using an artificial intelligence-based method, a standard data set is generated through multi-modal transformation processing, an audit knowledge graph is constructed, cross-modal alignment is performed, and feature fusion is combined with graph neural network and attention mechanism, dynamically verify and update the knowledge base, and the final audit document is generated.
It significantly improves the accuracy of audit clue discovery and the adaptability of legal clauses, improves the efficiency of cross-modal clue discovery, reduces character breakage and speech transcription error rates, and realizes dynamic correlation analysis of financial abnormalities and business flow fluctuations.
Smart Images

Figure CN120340042A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of enterprise audit information processing, and in particular to a method for extracting enterprise audit information based on artificial intelligence. Background Art
[0002] In recent years, the informatization technology of enterprise audit has developed rapidly, and the multi-modal data parsing method based on optical character recognition (OCR) and natural language processing has become the mainstream. The existing technology usually adopts a modular processing flow: extracting text information through traditional image processing, combining a rule engine or a statistical model to match audit features, and finally relying on manual review to associate legal provisions. Such methods perform stably in structured data processing, but have significant limitations in cross-modal association and semantic understanding of unstructured data. Traditional OCR relies on fixed threshold segmentation, which is prone to character breakage or adhesion errors for low-quality scanned documents; while entity recognition based on keyword matching is difficult to capture complex semantic combinations such as "fictitious revenue increase", resulting in a relatively high missed detection rate of audit clues. This is mainly reflected in the aspects of dynamic association and knowledge fusion. On the one hand, most systems lack a unified representation method for multi-source data (text, table, time series), resulting in the breakage of cross-modal evidence chains. For example, the abnormal financial statements and the time series fluctuations of business logs cannot be associated through simple rules. On the other hand, the mapping of legal provisions mostly adopts hard-coded rules, without considering the impact of the enterprise organizational structure on the applicability of the provisions. For example, the same "internal control defect" may correspond to different regulatory provisions in different departments, but the existing technology does not introduce an attention mechanism to quantify the dynamic relevance between the provisions and organizational characteristics, resulting in insufficient pertinence of audit suggestions. Summary of the Invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides a method for extracting enterprise audit information based on artificial intelligence to solve the problems of poor cross-modal relevance of multi-modal audit data and missed detection and misjudgment of audit clues caused by static mapping of legal provisions.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for extracting enterprise audit information based on artificial intelligence, which includes obtaining original audit documents from an enterprise and performing multi-modal conversion processing to generate a standard data set; Performing in-depth semantic parsing on the standard data set to construct an audit knowledge graph; Performing cross-modal alignment on the audit knowledge graph with image features, table features and time series features to construct a joint audit analysis model; Dynamically verifying the audit report generated by the joint audit analysis model with a historical case library to form a dynamically evolving knowledge base; Generate the final audit document based on the updated knowledge base.
[0006] As a preferred solution of the enterprise audit information extraction method based on artificial intelligence according to the present invention, wherein: the multimodal conversion process includes the following steps, Perform grayscale processing on the PDF file of the contract text and the JPG file of the scanned financial statement, calculate the adaptive threshold using the Otsu algorithm, generate a black-and-white binary image, and perform noise elimination; Perform optical character recognition on the denoised binary image, segment it by text line area, and output a JSON-formatted text with character coordinates; Perform spectral subtraction noise reduction and voice endpoint detection on the WAV file of the meeting minutes recording, segment the effective voice segments, and transcribe them into text; Perform sentence splitting and entity tagging on the CSV file of the business system log, and generate a corrected text after dependency syntactic analysis.
[0007] As a preferred solution of the enterprise audit information extraction method based on artificial intelligence according to the present invention, wherein: the construction of the structured audit knowledge graph includes the following steps, Based on the text modal data in the standard dataset, extract audit suspicion entities through entity recognition to generate an intermediate dataset; Input the audit suspicion entities in the intermediate dataset into a sequence labeling model driven by a bidirectional long short-term memory network, match the pre-constructed laws and regulations clause database, and generate a mapping data table of clause index codes; Perform integrity verification and conflict detection on the mapping data table, and dynamically complete the missing entity attribute fields; Based on the completed mapping data table, use the BERT-base model to identify the organizational entity in the text, and construct an association matrix between the audit suspicion entity and the organizational structure; Combine the association matrix with the clause index data, calculate the clause applicability score through the multi-head attention mechanism, and generate a verified structured audit knowledge graph.
[0008] As a preferred solution of the enterprise audit information extraction method based on artificial intelligence according to the present invention, wherein: the construction of the joint audit analysis model includes the following steps, Extract text feature vectors, image feature vectors, and time series feature vectors from the audit knowledge graph; Perform cross-modal alignment on the text feature vectors, image feature vectors, and time series feature vectors in the shared embedding space, and dynamically adjust the feature weights; Fuse the table structure features and audit text features through the multi-head attention mechanism to generate a joint representation; Train a fully-connected network branch based on the joint representation, and parallelly output the classification of audit questions, the prediction of risk levels, and the results of clause matching to generate a structured audit report.
[0009] As a preferred solution of the enterprise audit information extraction method based on artificial intelligence according to the present invention, wherein: the formation of the dynamically evolving knowledge base includes the following steps. Convert the structured audit report into an XML data packet, call a pre-trained embedding model to generate semantic vectors, retrieve similar cases in the historical case library, and generate a verification report. Based on the difference annotation results in the verification report, receive a correction feedback file, parse the clause index changes, and generate a difference comparison table. Use the proximal policy optimization algorithm to fine-tune the joint audit analysis model, update the training data set, and reconstruct the audit knowledge graph. Associate the updated audit suspicion entities with the clause index nodes, automatically inherit the cross-modal positioning information, and form a dynamically evolving knowledge base.
[0010] As a preferred solution of the enterprise audit information extraction method based on artificial intelligence according to the present invention, wherein: the generation of the final audit document includes the following steps. Locate the original voucher screenshot according to the audit suspicion entity code in the dynamically evolving knowledge base and generate a visual image with a red annotation box. Retrieve the legal clause content through hash indexing and construct a paragraph framework of the clause content. Generate a rectification process node based on the risk fan parameters in the latest version of the audit suspicion entity node set. Input the visual image, clause content, and rectification process into a natural language generation system, fill in the semantic template, and reconstruct the sentences to generate a structured document. Associate the text entity with the multi-modal evidence chain through the coordinate binding algorithm, trigger the interactive display of original vouchers, data tables, and chronological graphs, and generate the final audit document.
[0011] As a preferred solution of the enterprise audit information extraction method based on artificial intelligence according to the present invention, wherein: the further processing of the grayscale scanned financial statement includes the following steps. Use a table recognition algorithm based on deep learning to locate the table bounding box, divide the cells by projection analysis, and extract the row numbers, column numbers, and content. Map the table header and data rows to generate two-dimensional structured table data, and integrate it into a PNG format screenshot. Associate the two-dimensional structured table data with the chronological data of the business system log to generate chronological data.
[0012] As a preferred solution of the enterprise audit information extraction method based on artificial intelligence according to the present invention, wherein: the step of dynamically completing the missing entity attribute fields includes the following steps, Call the extractor through coordinate positioning of the original audit document to complete the amount, time and subject fields; Compare the difference in the clause mapping confidence of the same entity to generate a completion request instruction; Calculate the relationship strength between the audit suspicious entity and the organizational structure based on the co-occurrence frequency and syntactic distance.
[0013] In a second aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and wherein: when the computer program is executed by the processor, any step of the enterprise audit information extraction method based on artificial intelligence as described in the first aspect of the present invention is implemented.
[0014] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program is executed by the processor, any step of the enterprise audit information extraction method based on artificial intelligence as described in the first aspect of the present invention is implemented.
[0015] The beneficial effects of the present invention are as follows: through multi-modal feature alignment and dynamic knowledge graph construction, the discovery accuracy of audit clues and the adaptability of legal clauses are significantly improved. In the data preprocessing stage, the Otsu algorithm is used for adaptive binarization and median filtering for noise reduction, solving the problem of character breakage in scanned documents and improving the OCR accuracy; spectral subtraction for noise reduction and double-threshold endpoint detection technology reduce the speech transcription error rate; through cross-modal feature vector alignment and attention-driven clause applicability scoring, dynamic correlation analysis of financial anomalies and business transaction fluctuations is realized, improving the cross-modal clue discovery efficiency; the quality verification mechanism based on graph neural network and clause matching fine-tuning optimized by proximal policy ensure the dynamic evolution ability of the knowledge graph. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0017] Figure 1 It is a flowchart of the enterprise audit information extraction method based on artificial intelligence.
[0018] Figure 2 It is a schematic diagram of multi-modal conversion processing.
[0019] Figure 3Schematic diagram for constructing an audit knowledge graph.
[0020] Figure 4 Schematic diagram for constructing a joint audit analysis model. Specific implementation manners
[0021] To make the above objects, features and advantages of the present invention more obvious and understandable, the specific implementation manners of the present invention will be described in detail below with reference to the accompanying drawings of the specification.
[0022] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0023] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that are mutually exclusive with other embodiments.
[0024] Referring to Figures 1 to 4 , this embodiment provides an enterprise audit information extraction method based on artificial intelligence, including the following steps: S1. Obtain the PDF file of the contract text, the JPG file of the scanned financial statement, the CSV file of the business system log, and the WAV file of the meeting minutes recording from the enterprise information system.
[0025] Perform grayscale processing on the PDF file of the contract text and the JPG file of the scanned financial statement.
[0026] Specifically, use the weighted average method to convert the RGB images of the PDF file of the contract text and the JPG file of the scanned financial statement into grayscale images.
[0027] Perform binarization processing on the grayscale images.
[0028] Specifically, apply the Otsu algorithm to calculate the adaptive threshold, count the grayscale histogram, iteratively calculate the maximum between-class variance, determine the segmentation threshold, and generate a black-and-white binary image.
[0029] Further illustrate the pixel value mapping rule: set the area with pixel values lower than the segmentation threshold to black (0), and the area higher than the segmentation threshold to white (255), and generate the black-and-white binary images of the contract text and the financial statement.
[0030] Perform noise elimination operations on the binary images.
[0031] Specifically, a 3×3 pixel sliding window is created using median filtering, the nine pixel values within the window are sorted to obtain the median, and the reflection filling method is used to expand the image edges to remove noise.
[0032] Optical character recognition is performed on the denoised binary image.
[0033] Specifically, it is divided into continuous text line regions in the order from left to right and top to bottom, and vertical projection analysis is performed to determine the line spacing (such as a line height of 32 pixels). The text lines are segmented by horizontal projection (such as 28 lines are segmented on page 5). Sequence recognition is performed on each text line region, the image is cut into [32×600] pixel blocks by line, and the character content and its coordinates in the original binary image are output in sequence to generate a JSON-formatted digital text with character coordinates.
[0034] Spectral subtraction noise reduction processing is performed on the WAV file of the meeting minutes recording to eliminate ambient background noise.
[0035] Specifically, through frame processing, background noise with a duration exceeding M is eliminated to generate a denoised audio file.
[0036] Voice activity detection is performed on the denoised recording file, and effective speech segments are segmented based on the double-threshold method of short-time energy and zero-crossing rate; the effective speech segments are input into a speech-to-text engine based on connectionist temporal classification to generate a text record with timestamps; the text record generated by the speech-to-text engine and the CSV file of the business system log are clause-processed, and regular expressions are used to match full stops, question marks, exclamation marks, etc. as sentence boundary identifiers; entity tagging is performed on the clause-processed text record and business system log text, and exact string matching is performed using a predefined audit entity dictionary.
[0037] Dependency syntactic analysis is performed on the text record and business system log text with tagged entities.
[0038] Specifically, a syntax tree with "discovery" as the root node and "abnormal growth" as the object component is constructed; for sentences without a subject that require immediate verification procedures to be initiated, a default subject "audit group" is inserted to generate a corrected text ("The audit group needs to immediately initiate verification procedures").
[0039] Table area detection is performed on the JPG file of the scanned financial statement.
[0040] Specifically, a table recognition algorithm based on deep learning is adopted to locate the table bounding box; cell segmentation is performed on the detected table area, and the projection analysis method is applied to identify the intersection points of row and column lines to generate cell coordinates; the text content of the header cells and the numerical values of the data cells are extracted, and the mapping relationship between the column headers and the data rows is established to generate two-dimensional structured table data containing row numbers, column numbers, and cell contents.
[0041] The JSON-formatted digital text, the timestamped text record text generated by the speech-to-text engine, the two-dimensional structured table data, and the CSV time-series data of the business system log after dependency syntax analysis are unified in format and converted into a TXT text file encoded in UTF-8, a PNG-formatted table screenshot, and TIMESTAMP-VALUE (timestamp-numerical value) format time-series data, forming a standard dataset containing text modality, image modality, table modality, and time-series modality; the text modality data in the standard dataset is classified according to audit problem types such as contract disputes, financial anomalies, and process violations, and stored in a folder structure named according to the problem categories.
[0042] S2. Perform audit problem feature extraction on the text modality data in the standard dataset through entity recognition. By comparing with the pre-built audit feature dictionary library through a pattern matching algorithm, identify the audit doubt entities containing semantic feature markers, and generate an intermediate dataset containing entity type codes and text position coordinates.
[0043] Specifically, load the pre-built audit feature dictionary library and traverse the text data using an exact string matching algorithm.
[0044] Furthermore, perform multi-pattern scanning on each sentence. When a keyword is matched, generate triple data containing entity type codes, start positions, and end positions; for text segments that do not hit an exact match, use a rule-based feature extension module to identify verb-object structures through dependency syntax (e.g., "inflated revenue" → verb "inflated" + noun "revenue").
[0045] Based on the intermediate dataset, perform legal clause mapping on the semantic units where the audit doubt entities are located. Use a sequence labeling model driven by a bidirectional long short-term memory network to match the clause number format features in the pre-built database of laws and regulations, and generate a mapping data table containing clause index codes and the associated relationships of the corresponding audit doubt entities.
[0046] Specifically, input the sentence where the entity is located into a sequence labeling model driven by a bidirectional long short-term memory network. The model structure includes an input layer, a bidirectional long short-term memory network layer, and a conditional random field layer.
[0047] Among them, the input layer receives legal term vectors (such as pre-trained on 500,000 legal texts), the bidirectional long short-term memory network captures context dependencies through 128 hidden units, and the conditional random field layer decodes and outputs the legal clause number features.
[0048] Interactive verification is performed on the extracted audit suspicion entities and the mapped clause index data. Structured completion request instructions are automatically generated for the situation of missing entity attributes, and the entity attribute fields are dynamically updated by parsing the feedback data.
[0049] Specifically, a verification rule engine is established for the extracted audit suspicion entities and the clause index data in the mapped data table.
[0050] Furthermore, the verification rule engine includes integrity check and conflict detection.
[0051] Among them, integrity check refers to checking whether the mandatory attributes of the entity (such as amount, time, subject) exist, and conflict detection refers to comparing the mapping results of multiple clauses of the same entity (such as triggering an alarm when the confidence difference > 0.2).
[0052] Automatically generating structured completion request instructions for the situation of missing entity attributes means locating the original data file by coordinates and calling a specific extractor to complete the fields.
[0053] Perform organizational structure feature extraction on the audit suspicion entities with complete attributes, identify the audit management organization entity of the target enterprise from the text semantic units, and establish an association matrix between the audit suspicion entities and the organizational management structure.
[0054] Specifically, identifying the audit management organization entity of the target enterprise from the text semantic units means using the BERT-base fine-tuning model to identify the organizational entities in the text (such as the audit committee, finance department, etc.), and setting entity relationship extraction rules. For example, a subordination relationship and a supervision relationship are established within a range of <200 characters between entities.
[0055] Establishing an association matrix between the audit suspicion entities and the organizational management structure means setting the row dimension as the audit suspicion entity ID, the column dimension as the organizational entity code, and the matrix value as the relationship strength (0-1 normalized value) calculated based on the co-occurrence frequency and syntactic distance.
[0056] Perform association analysis on the clause index data and the extracted organizational structure features, calculate the semantic association degree between a specific legal clause library and the audit decision records of the target enterprise using the attention mechanism, and generate a quantitative evaluation parameter for clause applicability.
[0057] Specifically, when calculating the semantic association degree between a specific legal clause library and the audit decision records of the target enterprise using the attention mechanism, first construct an attention calculation module, and then perform multi-head attention calculation.
[0058] Further explanation: The construction of the attention calculation module includes queries, keys, and values.
[0059] Among them, the query refers to the BERT embedding vector of the clause text, the key refers to the TF-IDF feature vector of the enterprise audit decision record, and the value refers to the association tag between the decision record and the clause.
[0060] Multi-head attention calculation is to calculate the attention distribution between clause features and enterprise records separately for each head, and fuse all heads to generate the clause applicability score. The expression is: ; Among them, is the clause applicability score, is the total number of attention heads, used to average the multi-head output, is the attention head index, represents normalizing the similarity to generate attention weights (the sum is 1), highlighting key matching items, is the th query matrix of the attention head, representing the audit doubt features to be concerned (such as inflated revenue), is the th key matrix of the attention head, representing the enterprise audit decision record features to be matched (such as historical violation cases), is the transpose of the key matrix, converting the row vector into a column vector for matrix multiplication, is the dimension of the key vector, preventing the dot product value from being too large, is the dimension index, is the th value matrix of the attention head, carrying the specific content of the enterprise audit decision record (such as case amount, time, etc.).
[0061] Combining the audit doubt entity features, clause index data, and organizational structure features, construct multi-dimensional associated nodes according to the graph structure data specification. Among them, the audit doubt entity node contains the original data file location information, and the clause index node contains the globally unique identifier of the regulation library and the clause text features; perform quality verification on the constructed graph structure data, specifically as follows: Based on the analysis of the graph structure data of historical audit cases, set the association path complexity threshold by statistically calculating the average complexity of the effective association paths in the compliance knowledge graph. When the entity association path complexity exceeds the association path complexity threshold, trigger the automatic review mechanism, update the verified association relationship confidence parameter to the graph structure data, and generate a structured audit knowledge graph containing multi-dimensional feature nodes, verified clause index association parameters, and cross-modal data location information.
[0062] Further explanation, the construction of the audit knowledge graph includes node creation and edge relationship definition.
[0063] Specifically, node creation includes audit doubt nodes, clause nodes, and organization nodes, and edge relationship definition includes entity-clause edges, entity-organization edges, and clause-organization edges. Quality verification refers to path complexity detection.
[0064] Among them, the audit doubt node includes an ID, a type code, and the original file path. The clause node includes a GUID, a regulatory library index number, and a clause text summary. The organization node includes a department code and a hierarchical path (such as / headquarters / audit department).
[0065] S3. Extract the semantic confidence index output by the semantic analysis of the audit text for the audit doubt entity nodes in the structured audit knowledge graph to form a text feature vector containing the entity type code and confidence weight; extract the table structure features from the scanned financial statement through a convolutional neural network and convert them into an image feature vector of a fixed dimension; convert the financial data association relationship modeled by the graph neural network into a table feature vector in the form of an adjacency matrix; perform time window slicing on the business flow fluctuation pattern processed by the long short-term memory network to generate a time series feature vector.
[0066] Perform cosine similarity calculation on the text feature vector, image feature vector, table feature vector, and time series feature vector in the shared embedding space to establish a cross-modal feature alignment matrix. Among them, the initial alignment weight of the text feature vector is weighted by the semantic confidence index of the audit text semantic analysis, and the alignment weight between the image feature vector and the text feature vector is dynamically adjusted according to the matching result of the table number field in the scanned financial statement in the text knowledge graph.
[0067] Based on the cross-modal alignment matrix, calculate the cross-attention score between the text feature vector and the image feature vector through the multi-head attention mechanism to generate a joint representation that fuses the table structure feature and the audit text feature. At the same time, calculate the self-attention distribution between the table feature vector and the time series feature vector to form an association pattern graph of financial data fluctuations and business flow changes.
[0068] Hierarchically fuse the joint representation of the table structure feature and the audit text feature with the association pattern graph, and use a gating mechanism to control the proportion of each modal feature. The initial value of the text modal gating weight uses the normalized semantic confidence index, and the image modal weight is dynamically adjusted according to the recognition accuracy of the financial statement table structure.
[0069] On the basis of the joint representation that fuses the table structure feature and the audit text feature, parallelly construct three fully connected network branches of the joint audit analysis model: The first branch receives the fused features and outputs the classification probability distribution of audit problem types such as financial fraud and internal control deficiencies. The second branch receives the same input and generates the probability prediction values for high-risk, medium-risk, and low-risk levels. The third branch receives the concatenated input of the fused features and the legal clause text features and generates a cosine similarity sorted list of the best-matching clauses.
[0070] Alternately train the output results of the three branches, sequentially optimize the cross-entropy loss of the audit problem classifier, the mean squared error loss of risk prediction, and the contrastive loss of clause matching, while keeping the feature extraction parameters frozen, to obtain the structured audit report output by the trained joint audit analysis model.
[0071] Furthermore, the structured audit report includes the audit doubt entity codes in the problem list, the clause index numbers of legal suggestions, and the level parameters of risk warnings.
[0072] S4. Convert the structured audit report output by the joint audit analysis model into an XML format data packet compatible with the audit work platform; after importing the XML format data packet into the audit work platform, call the pre-trained embedding model to generate 768-dimensional semantic vectors for the audit doubt entity description text, and use the cosine similarity algorithm to retrieve the most similar cases in the historical case library to generate a verification report document containing the similar case numbers, matching scores, and difference point annotations.
[0073] Furthermore, the XML format data packet compatible with the audit work platform includes the file path of the audit doubt entity in the original multimodal dataset, the global unique identifier link of the clause index database, and the risk level probability distribution vector.
[0074] Based on the difference point annotation results in the verification report document, mark the audit doubt entities with legal clause reference differences as items to be verified; receive the corrected feedback file through the online interface of the audit work platform, parse the audit doubt entity codes in the corrected feedback file and perform field-level matching with the original structured audit report, and extract the entity pairs with changed clause index numbers to form a difference comparison table.
[0075] Furthermore, the corrected feedback file includes a CSV file of the audit doubt entity codes confirmed manually, the corrected clause index numbers, and the risk level adjustment values.
[0076] Combine the audit doubt entity codes and the corrected clause index numbers in the difference comparison table into training sample pairs containing entity codes and new clause indexes, and use the proximal policy optimization algorithm to calculate the policy gradient of the legal clause matching branch in the joint audit analysis model, specifically as follows: Under the premise of keeping the text feature extraction parameters frozen, the weight matrix of the fully connected layer is dynamically fine-tuned (such as setting the learning rate to 3e-5 and the batch size to 32). During the iterative update process, the cosine similarity loss function between the clause index number and the audit suspicious entity is optimized first. After the update, cross-validation is performed using similar cases in the verification report document.
[0077] The verified correct audit suspicion entity code and clause index number data pairs are converted into structured records that conform to the format of the training dataset.
[0078] Specifically, the structured record contains the original text fragment, entity location coordinates, clause index number, and risk level label.
[0079] Perform deduplication verification on the newly added data, append the verified data records to the end of the TXT text file of the training data set, and update the corresponding PNG format table screenshots and the index mapping table of TIMESTAMP-VALUE time series data; re-execute the audit knowledge graph construction based on the updated training data set, perform dependency syntax analysis and entity relationship verification on the newly added audit suspicious entities, establish bidirectional association edges between the verified entity nodes and the clause index nodes (edge attributes include manual confirmation flags and automatic matching confidence), and automatically inherit cross-modal data positioning information when inserting new nodes into the original graph structure data to form a dynamically evolving knowledge base.
[0080] S5. Extract the latest version of the audit doubt entity node set from the dynamically evolving knowledge base, each node contains the audit doubt entity code, the associated clause index number, the risk level parameter and the cross-modal data location information; locate the original voucher screenshot path in the standard data set according to the audit doubt entity code, load the PDF file of the contract text and take a screenshot of the image area after grayscale processing, and generate a visualization image of the audit evidence with a red rectangular annotation box based on the coordinate parameters determined by the character coordinate positioning algorithm.
[0081] The clause index numbers in the dynamically evolving knowledge base are accurately matched with the legal and regulatory clause database, the complete text content of the corresponding clauses is retrieved through the hash index mechanism, and the paragraph framework of the clause content is constructed in accordance with the qualitative format requirements in the standard audit document specifications.
[0082] Based on the risk level probability distribution vector, a corrective measures process is automatically generated for high-risk audit suspicious entities.
[0083] Among them, the process nodes include problem confirmation, responsibility allocation, solution design, execution monitoring and effect verification. The connecting lines of each process node are marked with time constraints and work deliverable standards.
[0084] Input the visualized image of audit evidence, clause content, and rectification process into the natural language generation system. Fill in the semantic template according to the entity attributes verified by the dynamically evolving knowledge base. Reconstruct the order of sentence components through the dependency syntax analysis engine, and generate a structured document in combination with the standard audit document format specifications.
[0085] Establish the spatial mapping relationship between the structured document and the multi-modal evidence chain. Intelligently associate text entities with visual elements through the coordinate binding algorithm. Trigger the multi-modal evidence display protocol when clicking on the audit suspicion entity, and synchronously display the original voucher screenshot optimized by the image processing algorithm, the associated data table, and the time series analysis graph on the interactive interface to generate the final audit document.
[0086] This embodiment also provides a computer device applicable to the situation of the enterprise audit information extraction method based on artificial intelligence, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the enterprise audit information extraction method based on artificial intelligence as proposed in the above embodiment.
[0087] This computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0088] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for extracting enterprise audit information based on artificial intelligence as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0089] In summary, through multi-modal feature alignment and dynamic knowledge graph construction, the present invention significantly improves the discovery accuracy of audit clues and the adaptability of legal provisions. In the data preprocessing stage, the Otsu algorithm is used for adaptive binarization and median filtering for noise reduction, solving the problem of character breakage in scanned documents and improving the OCR accuracy; spectral subtraction for noise reduction and double-threshold endpoint detection technology reduce the speech transcription error rate; through cross-modal feature vector alignment and attention-driven clause applicability scoring, dynamic correlation analysis of financial anomalies and business transaction fluctuations is realized, improving the cross-modal clue discovery efficiency; the quality verification mechanism based on graph neural network and clause matching fine-tuning based on proximal policy optimization ensure the dynamic evolution ability of the knowledge graph.
[0090] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An enterprise audit information extraction method based on artificial intelligence, characterized in that: including obtaining original audit documents from an enterprise and performing multimodal conversion processing to generate a standard data set; performing in-depth semantic parsing on the standard data set to construct an audit knowledge graph; performing cross-modal alignment on the audit knowledge graph with image features, table features, and time series features to construct a joint audit analysis model; dynamically validating the audit report generated by the joint audit analysis model with a historical case library to form a dynamically evolving knowledge base; generating a final audit document based on the updated knowledge base.
2. The method for extracting enterprise audit information based on artificial intelligence according to claim 1, characterized in that: The multimodal conversion processing includes the following steps. Performing grayscale processing on the PDF file of the contract text and the JPG file of the scanned financial statement, applying the Otsu algorithm to calculate the adaptive threshold, generating a black-and-white binary image, and performing noise elimination; Performing optical character recognition on the denoised binary image, segmenting by text line area, and outputting text in JSON format with character coordinates; Performing spectral subtraction noise reduction and voice endpoint detection on the WAV file of the meeting minutes recording, segmenting valid voice segments, and transcribing them into text; Performing sentence splitting and entity tagging on the CSV file of the business system log to generate a corrected text after dependency syntax analysis.
3. The method for extracting enterprise audit information based on artificial intelligence according to claim 1, characterized in that: The construction of the structured audit knowledge graph includes the following steps. Based on the text modal data in the standard data set, extracting audit doubt entities through entity recognition to generate an intermediate data set; Inputting the audit doubt entities in the intermediate data set into a sequence annotation model driven by a bidirectional long short-term memory network, matching a pre-constructed database of laws and regulations clauses, and generating a mapping data table of clause index codes; Performing integrity verification and conflict detection on the mapping data table, and dynamically supplementing missing entity attribute fields; Based on the supplemented mapping data table, using the BERT-base model to identify organizational entity entities in the text, and constructing an association matrix between audit doubt entities and organizational structures; Combining the association matrix with the clause index data, calculating the clause applicability score through a multi-head attention mechanism, and generating a verified structured audit knowledge graph.
4. The method for extracting enterprise audit information based on artificial intelligence according to claim 1, wherein: The construction of the joint audit analysis model includes the following steps. Extracting text feature vectors, image feature vectors, and time series feature vectors from the audit knowledge graph; Performing cross-modal alignment on the text feature vectors, image feature vectors, and time series feature vectors in a shared embedding space, and dynamically adjusting the feature weights; Fusing table structure features and audit text features through a multi-head attention mechanism to generate a joint representation; Training a fully connected network branch based on the joint representation, and parallelly outputting audit problem classification, risk level prediction, and clause matching results to generate a structured audit report.
5. The method for extracting enterprise audit information based on artificial intelligence according to claim 1, characterized in that: The formation of the dynamically evolving knowledge base includes the following steps. Converting the structured audit report into an XML data packet, calling a pre-trained embedding model to generate semantic vectors, retrieving similar cases in the historical case library, and generating a verification report; Based on the difference annotation results in the verification report, receiving a corrected feedback file, parsing the clause index changes, and generating a difference comparison table; Using the proximal policy optimization algorithm to fine-tune the joint audit analysis model, updating the training data set, and reconstructing the audit knowledge graph; Associate the updated audit suspicion entities with the clause index nodes, automatically inherit the cross-modal localization information, and form a dynamically evolving knowledge base.
6. The method for extracting enterprise audit information based on artificial intelligence according to claim 1, wherein: The generation of the final audit document includes the following steps: Locate the original voucher screenshots according to the audit suspicion entity codes in the dynamically evolving knowledge base and generate a visualized image with a red annotation box; Retrieve the legal clause content through hash indexing and construct the paragraph framework of the clause content; Generate rectification process nodes based on the risk level parameters in the latest version of the audit suspicion entity node set; Input the visualized image, clause content, and rectification process into the natural language generation system, fill in the semantic template and reconstruct the sentences to generate a structured document; Associate the text entities with the multi-modal evidence chain through the coordinate binding algorithm, trigger the interactive display of the original vouchers, data tables, and chronological graphs, and generate the final audit document.
7. The method for extracting enterprise audit information based on artificial intelligence according to claim 2, characterized in that: The further processing of the grayscale financial statement scans includes the following steps: Use a table recognition algorithm based on deep learning to locate the table bounding box, segment the cells through projection analysis, and extract the row numbers, column numbers, and content to generate two-dimensional structured table data; Associate the two-dimensional structured table data with the chronological data of the business system logs to generate chronological data.
8. The method for extracting enterprise audit information based on artificial intelligence according to claim 3, characterized in that: The dynamic completion of the missing entity attribute fields includes the following steps: Call the extractor through coordinate positioning of the audit original data file to complete the amount, time, and subject fields; compare the confidence difference of the clause mapping of the same entity to generate a completion request instruction; Calculate the relationship strength between the audit suspicion entity and the organizational structure based on the co-occurrence frequency and syntactic distance.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the artificial intelligence-based enterprise audit information extraction method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the artificial intelligence-based enterprise audit information extraction method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method, device and equipment for obtaining supervision recognition result in multiple modes and storage medium
CN111428044A
Auditing knowledge mapping method and system based on digital auditing platform
CN118822464A
Insurance clause auditing method and device, computer equipment and storage medium
CN119205365A
Power grid service intelligent auditing method and system based on AI enhancement
CN119599821A
Multi-modal auditing method based on large model
CN120011543A
Cited By
Multi-model collaborative enterprise entity hybrid identification method and device
CN120822521A
Financial account distribution and anomaly detection method and system based on machine learning
CN120894164A
Automatic official document generation method based on dynamic rules
CN120975050A
Bidding field information extraction method and device
CN121074929A
Method and system for identifying abnormal transactions in enterprise financial audit based on granularity weighting and dynamic attention
CN121117800A