Qualitative methods, devices, equipment and storage media for audit issues

By converting audit question text into a word vector set and using a pre-trained model to automatically determine the violation type and qualitative regulatory labels, the problem of low efficiency in the qualitative characterization of traditional audit questions is solved, and more efficient and accurate audit question qualitative characterization is achieved.

CN119671482BActive Publication Date: 2025-09-26SHENZHEN POWER SUPPLY BUREAU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411722944.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-09-26
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Traditional qualitative methods for audit issues are inefficient and lack accuracy. Limited by the experience of auditors, they are unable to quickly and accurately determine the type of violation, qualitative regulations, and handling and penalty regulations.

Method used

By converting the audit question text into a word vector set and using the pre-trained target violation type prediction model and target regulation prediction model, the violation type label and qualitative regulation label, including the processing penalty regulation label, are automatically determined.

Benefits of technology

It improves the efficiency and accuracy of the qualitative analysis of audit issues and reduces the time and human errors of manual qualitative analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119671482B_ABST
    Figure CN119671482B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device, and storage medium for qualitative analysis of audit issues. The method includes: obtaining an audit issue text of an audited unit on an audited project; converting the audit issue text into a word vector set using a target word vector model; inputting the word vector set into a target violation type prediction model to obtain a first violation type label corresponding to the audit issue text; inputting the word vector set and the first violation type label into a target regulation prediction model, determining a first qualitative regulation label corresponding to the audit issue text through a qualitative regulation prediction task, and determining a first processing and penalty regulation label corresponding to the audit issue text and the first qualitative regulation label through a processing and penalty regulation prediction task. The present application is conducive to improving the efficiency of qualitative analysis of audit issues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audit data processing technology, and in particular to a qualitative method, apparatus, device and storage medium for audit issues. Background Art

[0002] During the audit of the audited projects of the audited units, auditors need to audit the financial data, contract data, engineering data and other project data of the audited projects to discover audit issues in the project data, and then conduct qualitative analysis of the audit issues to determine the violation types, qualitative regulations and handling and punishment regulations of the audit issues, so as to form audit opinions and audit decisions.

[0003] Traditional qualitative audit problem analysis typically involves manual qualitative analysis based on the auditor's experience. However, due to the wide range of audit issues and the large number of regulations, auditors' limited experience often requires significant time to analyze and objectively identify the violation type, the qualitative regulations, and the disciplinary and penalty regulations. This results in a low efficiency in the qualitative analysis of audit issues. Summary of the Invention

[0004] In order to solve the above-mentioned problems existing in the prior art, the embodiments of the present application provide a qualitative method, apparatus, device and storage medium for audit issues, which converts the audit issue text of the audited unit on the audited project into a word vector set, and inputs the word vector set into a target violation type prediction model to obtain a first violation type label. Then, the word vector set and the first violation type label are input into a target regulation prediction model to obtain a first qualitative regulation label and a first processing and punishment regulation label, thereby qualitatively analyzing the audit issue text through the pre-trained target violation type prediction model and the target regulation prediction model, thereby improving the efficiency of qualitative analysis of audit issues.

[0005] In a first aspect, an embodiment of the present application provides a qualitative method for auditing problems, including:

[0006] Acquire the audit question text of the audited unit on the audited project; the audit question text is a text describing the problems found in the project data of the audited project during the audit process;

[0007] Converting the audit question text into a word vector set through a target word vector model; the word vector set includes multiple feature word vectors;

[0008] Inputting the word vector set into a target violation type prediction model to obtain a first violation type label corresponding to the audit question text; the first violation type label is used to indicate the violation type corresponding to the audit question text;

[0009] The word vector set and the first violation type label are input into the target regulation prediction model, and the first qualitative regulation label corresponding to the audit question text is determined through the qualitative regulation prediction task, and the first processing and penalty regulation label corresponding to the audit question text and the first qualitative regulation label is determined through the processing and penalty regulation prediction task; the qualitative regulation prediction task and the processing and penalty regulation prediction task are both subtasks of the target regulation prediction model; the first qualitative regulation label is used to indicate the qualitative regulation article corresponding to the audit question text, and the first processing and penalty regulation label is used to indicate the processing and penalty regulation article corresponding to the audit question text and the qualitative regulation article.

[0010] In a second aspect, an embodiment of the present application provides a device for qualitatively determining audit questions, including:

[0011] An acquisition module is used to acquire the audit problem text of the audited unit on the audited project; the audit problem text is a text describing the problems found in the project data of the audited project during the audit process;

[0012] A processing module, configured to convert the audit question text into a word vector set through a target word vector model; the word vector set includes a plurality of feature word vectors;

[0013] Inputting the word vector set into a target violation type prediction model to obtain a first violation type label corresponding to the audit question text; the first violation type label is used to indicate the violation type corresponding to the audit question text;

[0014] The word vector set and the first violation type label are input into the target regulation prediction model, and the first qualitative regulation label corresponding to the audit question text is determined through the qualitative regulation prediction task, and the first processing and penalty regulation label corresponding to the audit question text and the first qualitative regulation label is determined through the processing and penalty regulation prediction task; the qualitative regulation prediction task and the processing and penalty regulation prediction task are both subtasks of the target regulation prediction model; the first qualitative regulation label is used to indicate the qualitative regulation article corresponding to the audit question text, and the first processing and penalty regulation label is used to indicate the processing and penalty regulation article corresponding to the audit question text and the qualitative regulation article.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory, wherein the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device performs the method described in the first aspect.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in the first aspect.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program product is operable to enable a computer to execute the method described in the first aspect.

[0018] The implementation of the embodiments of the present application has the following beneficial effects:

[0019] In an embodiment of the present application, first, the audit question text of the audited unit on the audited project is obtained, and the audit question text is converted into a word vector set through a target word vector model. Then, the word vector set is input into the target violation type prediction model to obtain the first violation type label corresponding to the audit question text. Finally, the word vector set and the first violation type label are input into the target regulation prediction model, and the first qualitative regulation label corresponding to the audit question text is determined through the qualitative regulation prediction task of the target regulation prediction model, and the first processing and punishment regulation label corresponding to the audit question text and the first qualitative regulation label is determined through the processing and punishment regulation prediction task of the target regulation prediction model. Thus, through the pre-trained target violation type prediction model, the first violation type label corresponding to the audit question text can be determined based on the word vector set corresponding to the audit question text to determine the violation type corresponding to the audit question text. Through the pre-trained target regulation prediction model, the first qualitative regulation label and the first processing and punishment regulation label corresponding to the audit question text can be determined based on the word vector set and the first violation type label to determine the qualitative regulation clause and the processing and punishment regulation clause corresponding to the audit question text. Compared with manual qualitative analysis of audit issues, the efficiency of qualitative analysis of audit issues can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 A schematic diagram of an application scenario of a qualitative method for auditing issues provided in an embodiment of the present application;

[0022] Figure 2 A flowchart of a qualitative method for auditing questions provided in an embodiment of the present application;

[0023] Figure 3 A schematic diagram of a target violation type prediction model provided in an embodiment of the present application;

[0024] Figure 4 A schematic diagram of a target regulation prediction model provided in an embodiment of the present application;

[0025] Figure 5 A schematic diagram of a qualitative device for audit questions provided in an embodiment of the present application;

[0026] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0028] The terms "first," "second," "third," and "fourth," etc., in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, rather than to describe a specific order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0029] References herein to "embodiments" mean that a particular feature, result, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0030] First, see Figure 1 , Figure 1 This is a schematic diagram of an application scenario of a qualitative method for auditing problems provided in an embodiment of the present application. Figure 1As shown, during the audited project, the auditing entity must conduct an audit investigation of the project data according to relevant regulations to determine whether any audit issues have occurred during the construction of the audited project. The audited entity can be a power grid company. After determining the audit issues of the audited entity, the auditing entity must characterize the audit issues and determine the corresponding violation type, the qualitative regulatory provisions, and the corresponding disciplinary and penalty regulations.

[0031] The project data of the audited projects of the audited unit is stored in a database. The database can be multiple computers or servers used for data storage, reading and writing functions. The computer or server includes a memory. For example, the memory can be a random access memory (RAM), a read-only memory (ROM), a hard disk drive (HDD), a solid state drive (SSD), or other devices used to store and retrieve data in a computer. The database can also store historical audit data of the audited unit, which is not limited in this application.

[0032] Among them, the processing device of the audit unit can read the project data of the audited project from the database. The processing device can be a server with data reading and writing, high-performance computing and data caching functions, including: application server, virtual private server (VPS), rack server, etc. The processing device can also be a computer processor chip, including: central processing unit (CPU), digital signal processor (DSP), embedded microcontroller unit (MCU), etc., which is not limited in this application. Exemplarily, the processing device can be connected to a user interaction device to display the project data of the audited project on the user interaction device. The auditors of the audit unit can audit the project data based on the project data displayed on the user interaction device, discover audit issues in the project data of the audited project, and thus determine the audit issue text corresponding to the project data of the audited project, and input the audit issue text through the user interaction device. The processing device can obtain the audit issue text of the audited unit on the audited project from the user interaction device to characterize the audit issue text.

[0033] It's important to note that traditional qualitative methods for audit issues typically rely on manual qualitative analysis based on the experience of the audit firm's auditors. When there are numerous audit issues, covering a wide range of areas, and involving numerous regulations, auditors need to spend a significant amount of time conducting qualitative analysis to obtain objective and accurate qualitative results. These qualitative results include the type of violation, the qualitative regulatory provisions, and the disciplinary and penalty provisions. Therefore, traditional qualitative methods for audit issues suffer from low qualitative efficiency. Furthermore, manual qualitative analysis of audit issues is influenced by the auditor's experience, resulting in low qualitative accuracy.

[0034] To this end, in the qualitative method of audit questions provided in this application, the processing device obtains the audit question text of the audited unit on the audited project;

[0035] The processing device converts the audit question text into a word vector set through a target word vector model;

[0036] The processing device inputs the word vector set into the target violation type prediction model to obtain a first violation type label corresponding to the audit question text;

[0037] The processing device inputs the word vector set and the first violation type label into the target regulation prediction model, determines the first qualitative regulation label corresponding to the audit question text through the qualitative regulation prediction task, and determines the first processing penalty regulation label corresponding to the audit question text and the first qualitative regulation label through the processing penalty regulation prediction task.

[0038] It can be seen that, when applied to the above scenario, the processing device can obtain the audit problem text of the audited unit on the audited project and convert the audit problem text into a word vector set. Then, the word vector set is input into the target violation type prediction model to obtain the first violation type label corresponding to the audit problem text. Finally, the word vector set and the first violation type label are input into the target regulation prediction model to obtain the first qualitative regulation label and the first handling and penalty regulation label corresponding to the audit problem text, so as to determine the qualitative regulation provisions and handling and penalty regulation provisions corresponding to the audit problem text. Compared with manual characterization of audit problems, the qualitative characterization of audit problems by pre-trained target violation type prediction model and target regulation prediction model can improve the efficiency and accuracy of the qualitative characterization of audit problems.

[0039] See Figure 2 , Figure 2 This is a flowchart of a qualitative method for auditing problems provided in an embodiment of the present application. The method is applied to the processing device in the above scenario, and the method includes but is not limited to the following steps:

[0040] 201: Obtain the audit question text of the audited entity on the audited project.

[0041] In an embodiment of the present application, the audited unit may be the construction unit of the audited project, and the audited projects may include: cable tunnel engineering projects, high-altitude power transmission line construction projects, power station construction projects, etc. The audit question text is a text that describes the problems found in the project data of the audited project during the audit process. Project data may include: financial data, contract data, engineering construction data, etc. For example, the audit question text may be "The contracts of some engineering projects of the audited unit in the audited project lack the seal of the approval unit." The format of the audit question text is a JSON format that can be input into the target word vector model.

[0042] Optionally, the processing device may directly obtain the audit question text input by the user from the user interaction device. The processing device may also determine the audit question text based on the project data of the audited project, which is not limited in this application.

[0043] 202: Convert the audit question text into a word vector set through the target word vector model.

[0044] In an embodiment of the present application, a word vector set includes multiple feature word vectors. The semantics of each feature word vector can be determined by the feature word vector value and direction corresponding to the feature word vector. The target word vector model is pre-trained based on the reference word vector model.

[0045] For example, before converting the audit question text into a word vector set using the target word vector model, the following steps may also be included:

[0046] Get the reference word vector model;

[0047] Obtain multiple sets of training data;

[0048] Preprocessing the audit question text corresponding to each set of training data in the multiple sets of training data to obtain multiple sets of target training data;

[0049] The reference word vector model is trained using multiple sets of target training data to obtain the target word vector model.

[0050] In the embodiment of the present application, each set of training data includes the audit question text corresponding to the training data and the word vector set corresponding to the audit question text. The preprocessing operation includes word segmentation operation, stop word removal operation and format conversion operation.

[0051] The reference word vector model can be a Word2vec model, which is a deep learning model used to convert text into feature word vectors. It can map words in the text to a high-dimensional vector space, so that semantically similar words are close to each other in the vector space. The Word2vec model includes the Continuous Bag of Words (CBOW) model and the Skip Gram model. CBOW can predict the current word based on the words in the context, while Skip Gram can predict the words in the context based on the current word.

[0052] Specifically, the processing device first obtains a reference word vector model and then obtains multiple sets of training data from a database. Optionally, the processing device can also crawl multiple sets of training data from the audit qualitative retrieval system using the Fiddler tool. The audit question qualitative retrieval system is a system used to collect, process, and retrieve qualitative data during the audit process. The audit question qualitative retrieval system stores a large amount of audit question text and the word vector sets corresponding to the audit question text, which can be used to train the reference word vector model.

[0053] Then, the processing device will perform a word segmentation operation on the audit question text corresponding to each set of training data in the multiple sets of training data to obtain multiple text words, and remove stop words from the multiple text words to obtain multiple content words. Finally, the multiple content words are converted into a JSON format that can be input into the reference word vector model to obtain multiple target words, thereby obtaining multiple sets of target training data. By using the multiple target words corresponding to each set of target training data in the multiple sets of target training data as the input of the reference word vector model and the word vector corresponding to each set of target training data as the output of the reference word vector model, the reference word vector model is trained to obtain a target word vector model. The number of iterations of model training can be pre-set.

[0054] Therefore, by training the reference word vector model with multiple sets of preprocessed target training data, a target word vector model can be obtained, so that the audit question text can be converted into a word vector set through the target word vector model, and the audit question text can be qualitatively characterized through the word vector set, thereby improving the accuracy and efficiency of the qualitative characterization of audit questions.

[0055] Based on this, the processing device can obtain multiple words by performing preprocessing operations on the audit question text, and then input the multiple words into the target word vector model to obtain a word vector set.

[0056] 203: Input the word vector set into the target violation type prediction model to obtain the first violation type label corresponding to the audit question text.

[0057] In the embodiment of the present application, the first violation type tag is used to indicate the violation type corresponding to the audit question text. The violation types may include: false financial data, contract violations, project management violations, etc.

[0058] The target violation type prediction model is based on a BiLSTM-TextCNN-Attention structure. Exemplarily, the target violation type prediction model includes a Bidirectional Long Short Term Memory (BiLSTM) model, a Convolutional Neural Network for Text (TextCNN) model, an Attention layer, and a classifier. The word vector set is input into the target violation type prediction model to obtain the first violation type label corresponding to the audit question text, which may include:

[0059] The feature vectors in the word vector set are extracted through the BiLSTM model to obtain the global feature set;

[0060] The TextCNN model is used to extract features from the feature word vectors in the word vector set and the global feature word vectors in the global feature set to obtain a local feature set.

[0061] The attention layer is used to determine the attention weight corresponding to each feature word vector in the global feature set and the local feature set. Based on the attention weight corresponding to each feature word vector in the global feature set and the local feature set, the feature word vectors in the global feature set and the local feature set are weighted and summed to obtain the target feature word vector.

[0062] The target feature word vector is classified by the classifier to obtain the first violation type label.

[0063] In an embodiment of the present application, the global feature set includes multiple global feature word vectors, and the local feature set includes at least one local feature word vector.

[0064] Specifically, if Figure 3As shown, the processing device first extracts features from multiple feature word vectors in the word vector set through the BiLSTM model, wherein the feature extraction includes two parts, forward propagation and backward propagation, the forward propagation is processed by the forward LSTM, and the backward propagation is processed by the backward LSTM. In the forward LSTM, the features of each feature word vector are extracted from front to back according to the order in which the multiple feature word vectors are arranged in the audit question text. In the backward LSTM, the features of each feature word vector are extracted from back to front according to the order in which the multiple feature word vectors are arranged in the audit question text. After both the forward LSTM and the backward LSTM have extracted the features of all feature word vectors in the word vector set, the BiLSTM model can output a global feature set, which includes multiple global feature word vectors.

[0065] Then, if Figure 3 As shown, the processing device uses the TextCNN model to jointly extract features from multiple feature word vectors in the word vector set and multiple global feature word vectors in the global feature set to obtain a local feature set, which includes at least one local feature word vector. Among them, the convolution layer in the TextCNN model can extract local features by sliding the convolution kernel on multiple feature word vectors and multiple global feature word vectors, and the size of the convolution kernel can be pre-set according to needs. Then, the local features extracted by the convolution layer are downsampled through the pooling layer to reduce the dimension of the local features. Finally, the features output by the pooling layer are integrated through the fully connected layer to obtain a local feature set.

[0066] Furthermore, if Figure 3As shown, multiple global feature word vectors in the global feature set and multiple local feature sets in the local feature set can be weighted and summed through the Attention layer to obtain a target feature word vector. Specifically, the processing device determines the attention weight corresponding to each feature word vector in the global feature set and the local feature set through the Attention layer. For example, the attention weight corresponding to each feature word vector can be determined by the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm, that is, the number of occurrences of the vocabulary corresponding to each feature word vector in the global feature set and the local feature set in the audit question text is determined, and the total number of occurrences of the vocabulary corresponding to all feature word vectors in the word vector set in the audit question text is determined. Then, based on the number of occurrences of each feature word vector in the global feature set and the local feature set and the total number of occurrences of the vocabulary corresponding to all feature word vectors in the word vector set in the audit question text, the attention weight of each feature word vector in the global feature set and the local feature set is determined. Then, based on the attention weight corresponding to each feature word vector in the global feature set and the local feature set, all feature word vectors in the global feature set and the local feature set are weighted and summed to obtain the target feature word vector.

[0067] Finally, if Figure 3 As shown, the processing device can classify the target feature word vector through a classifier to obtain a first violation type label. Optionally, the classifier can be a softmax classifier that can predict the probability that the target feature word vector is a different violation type label.

[0068] Optionally, the target violation type prediction model can be trained using training data crawled from the audit question qualitative retrieval system. This training data includes word vectors corresponding to the audit question text and violation type labels corresponding to the audit question text. This training can improve the prediction performance of the target violation type prediction model and obtain violation type labels corresponding to the audit question text.

[0069] As can be seen, after inputting the word vector set into the target violation type prediction model, the BiLSTM model extracts features from the feature word vectors in the word vector set to obtain a global feature set. Then, the TextCNN model extracts features from the feature word vectors in the word vector set and the global feature word vectors in the global feature set to obtain a local feature set. Furthermore, the Attention layer performs a weighted summation of each feature word vector in the global and local feature sets to obtain the target feature word vector. Finally, the classifier classifies the target feature word vector to obtain the first violation type label. Therefore, the pre-trained target violation type prediction model can directly determine the first violation type label corresponding to the audit question text based on the word vector set corresponding to the audit question text, thereby obtaining the violation type corresponding to the audit question text. This improves the efficiency of qualitative analysis of audit question texts compared to manual qualitative analysis of audit question texts. Furthermore, training the target violation type prediction model can improve the accuracy of qualitative analysis of audit questions.

[0070] 204: Input the word vector set and the first violation type label into the target regulation prediction model, determine the first qualitative regulation label corresponding to the audit question text through the qualitative regulation prediction task, and determine the first processing penalty regulation label corresponding to the audit question text and the first qualitative regulation label through the processing penalty regulation prediction task.

[0071] In this embodiment of the present application, both the qualitative regulation prediction task and the penalty regulation prediction task are subtasks of the target regulation prediction model. The first qualitative regulation tag is used to indicate the qualitative regulation clause corresponding to the audit question text, and the first penalty regulation tag is used to indicate the penalty regulation clause corresponding to the audit question text and the qualitative regulation clause.

[0072] It should be noted that the target regulation prediction model is a multi-task learning model. In the present embodiment, the target regulation prediction model includes two subtasks: a qualitative regulation prediction task and a penalty regulation prediction task. The target regulation prediction model can effectively leverage the correlation between the two subtasks to improve the accuracy of the prediction of qualitative regulation labels and penalty regulation labels.

[0073] Exemplarily, inputting the word vector set and the first violation type label into the target regulation prediction model, determining the first qualitative regulation label corresponding to the audit question text through the qualitative regulation prediction task, and determining the first processing penalty regulation label corresponding to the audit question text and the first qualitative regulation label through the processing penalty regulation prediction task may include:

[0074] According to the first violation type label, feature extraction is performed on multiple feature word vectors in the word vector set to obtain multiple common feature word vectors;

[0075] Extracting features from multiple feature word vectors through the qualitative law prediction task to obtain multiple first sub-feature word vectors corresponding to the qualitative law prediction task; using the multiple common feature word vectors and the multiple first sub-feature word vectors as task feature word vectors corresponding to the qualitative law prediction task to obtain multiple first task feature word vectors;

[0076] Determining semantic distances between the plurality of first task feature word vectors and a feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a plurality of first semantic distances;

[0077] determining a first qualitative regulatory label from a plurality of preset qualitative regulatory labels based on the plurality of first semantic distances;

[0078] Extract features from multiple feature word vectors by processing the penalty law prediction task to obtain multiple second sub-feature word vectors corresponding to the penalty law prediction task;

[0079] Determine the third sub-feature word vector corresponding to the first qualitative regulatory label;

[0080] The plurality of common feature word vectors, the third sub-feature word vectors, and the plurality of second sub-feature word vectors are collectively used as task feature word vectors corresponding to the penalty regulations prediction task, to obtain a plurality of second task feature word vectors;

[0081] Determining the semantic distances between the plurality of second task feature word vectors and the feature word vector set of each preset processing penalty regulation label in the plurality of preset processing penalty regulation labels to obtain a plurality of second semantic distances;

[0082] According to the plurality of second semantic distances, a first processing penalty regulation label is determined from a plurality of preset processing penalty regulation labels.

[0083] In the embodiments of this application, Figure 4 As shown, the target regulation prediction model may include: a sharing layer, a qualitative regulation prediction task hidden layer, and a processing penalty regulation prediction task hidden layer.

[0084] Specifically, if Figure 4 As shown, the processing device inputs multiple feature word vectors in the word vector set and the first violation type label into the shared layer of the target regulation prediction model. The shared layer can obtain the target feature word vector based on the first violation type label. Then, feature extraction is performed on the target feature word vector and the multiple feature word vectors to extract common features from the target feature word vector and the multiple feature word vectors to obtain multiple common feature word vectors. Optionally, the shared layer may include convolutional neural networks (CNN) for performing feature extraction on the target feature word vector and the multiple feature word vectors.

[0085] After extracting multiple common feature word vectors, the multiple common feature word vectors are input into the hidden layer of the qualitative regulations prediction task and the hidden layer of the processing and punishment regulations prediction task. Among them, the hidden layer of the qualitative regulations prediction task and the hidden layer of the processing and punishment regulations prediction task also input the word vector set corresponding to the audit question text. Specifically, first, feature extraction is performed on multiple feature word vectors through the qualitative regulations prediction task hidden layer to obtain multiple first sub-feature word vectors corresponding to the qualitative regulations prediction task. It can be understood that the number of layers and the hidden layer function of the qualitative regulations prediction task hidden layer can be pre-set according to the actual test results, so that the qualitative regulations prediction task hidden layer can extract multiple first sub-feature word vectors corresponding to the qualitative regulations prediction task. Then, the multiple common feature word vectors and the multiple first sub-feature word vectors are used together as the task feature word vectors corresponding to the qualitative regulations prediction task to obtain multiple first task feature word vectors.

[0086] Furthermore, the processing device determines, through the qualitative regulation prediction task hidden layer, the semantic distances between the plurality of first task feature word vectors and the feature word vector set of each of the plurality of pre-set qualitative regulation labels, thereby obtaining a plurality of first semantic distances. The plurality of pre-set qualitative regulation labels are qualitative regulation labels corresponding to all qualitative regulatory provisions in the audit question qualitative retrieval system.

[0087] Exemplarily, determining the semantic distances between the plurality of first task feature word vectors and the feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain the plurality of first semantic distances may include:

[0088] Determine the intersection of the plurality of first task feature word vectors and the feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a plurality of feature word vector intersections;

[0089] Determine the intersection of the plurality of first task feature word vectors and the feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a union of the plurality of feature word vectors;

[0090] Determine multiple similarity coefficients according to the intersection of multiple feature word vectors and the union of multiple feature word vectors;

[0091] A plurality of first semantic distances are determined according to the plurality of similarity coefficients.

[0092] In an embodiment of the present application, the first semantic distance is used to represent the semantic relevance between the semantics corresponding to multiple first task feature word vectors and the preset regulatory labels corresponding to the first semantic distance. The closer the first semantic distance is, the closer the semantics corresponding to multiple first task feature word vectors are to the semantics of the preset regulatory labels corresponding to the first semantic distance.

[0093] Specifically, the processing device first determines the intersection of multiple first task feature word vectors and the feature word vector set of each preset qualitative regulatory label in multiple preset qualitative regulatory labels to obtain multiple feature word vector intersections. Determine the number of feature word vectors in each feature word vector intersection in the multiple feature word vector intersections to obtain multiple first quantities. Then, determine the union of multiple first task feature word vectors and the feature word vector set of each preset qualitative regulatory label in multiple preset qualitative regulatory labels to obtain multiple feature word vector unions. Determine the number of feature word vectors in each feature word vector union in the multiple feature word vector unions to obtain multiple second quantities.

[0094] Furthermore, the processing device uses the ratio of multiple first quantities to multiple second quantities as a similarity coefficient to obtain multiple similarity coefficients, and multiple preset qualitative regulatory labels correspond one to one with the multiple similarity coefficients. Optionally, the similarity coefficient can be a Jaccard coefficient. The value range of the similarity coefficient is [0, 1]. When the similarity coefficient is 1, it means that the multiple first task feature word vectors completely overlap with the feature word vector set of the preset qualitative regulatory label corresponding to the similarity coefficient. When the similarity coefficient is 0, it means that the multiple first task feature word vectors have no intersection with the feature word vector set of the preset qualitative regulatory label corresponding to the similarity coefficient. Finally, based on the multiple similarity coefficients, multiple first semantic distances can be determined, wherein the first semantic distance is equal to (1-similarity coefficient). The value range of the first semantic distance is [0, 1]. When the first semantic distance is 0, it means that the semantics corresponding to the multiple first task feature word vectors are consistent with the semantics of the preset qualitative regulatory label corresponding to the similarity coefficient. When the first semantic distance is 1, it means that the semantics corresponding to the multiple first task feature word vectors are completely different from the semantics of the preset qualitative regulatory label corresponding to the similarity coefficient. The closer the first semantic distance is, the closer the semantics corresponding to the multiple first task feature word vectors are to the semantics of the preset qualitative regulatory label corresponding to the first semantic distance.

[0095] Thus, by determining the intersection of multiple first task feature word vectors and the feature word vector set of each preset qualitative regulatory label in multiple preset qualitative regulatory labels, multiple feature word vector intersections are obtained. Then, the union of multiple first task feature word vectors and the feature word vector set of each preset qualitative regulatory label in multiple preset qualitative regulatory labels is determined to obtain multiple feature word vector unions. Furthermore, based on the multiple feature word vector intersections and the multiple feature word vector unions, multiple similarity coefficients can be determined. Finally, based on the multiple similarity coefficients, multiple first semantic distances can be determined, thereby determining the correlation between the semantics corresponding to the multiple first task feature word vectors and the semantics of each preset qualitative regulatory label, thereby improving the accuracy of qualitative regulatory label prediction.

[0096] Furthermore, the processing device may use the preset regulatory label with the shortest first semantic distance among the plurality of preset regulatory labels as the first qualitative regulatory label. Figure 4 As shown, the first qualitative regulatory label can also serve as the input of the hidden layer of the processing and penalty regulation prediction task. The processing and penalty regulation prediction task hidden layer will extract features from multiple feature word vectors in the word vector set to obtain multiple second sub-feature word vectors corresponding to the processing and penalty regulation prediction task. It is understandable that the number of layers and hidden layer functions of the processing and penalty regulation prediction task hidden layer can be pre-set based on actual test results so that the processing and penalty regulation prediction task hidden layer can extract multiple second sub-feature word vectors corresponding to the processing and penalty regulation prediction task.

[0097] Furthermore, the processing device extracts features from the first qualitative regulatory label through the hidden layer of the penalty regulation prediction task to obtain a third sub-feature word vector corresponding to the first qualitative regulatory label. The multiple common feature word vectors, the third sub-feature word vector, and the multiple second sub-feature word vectors are then used together as task feature word vectors corresponding to the penalty regulation prediction task to obtain multiple second task feature word vectors.

[0098] Next, the semantic distance between the plurality of second task feature word vectors and the feature word vector set of each preset processing and punishment regulation label in the plurality of preset processing and punishment regulation labels is determined to obtain a plurality of second semantic distances. The second semantic distance is used to represent the semantic relevance between the semantics corresponding to the plurality of second task feature word vectors and the preset processing and punishment regulation label corresponding to the second semantic distance. The closer the second semantic distance is, the closer the semantics corresponding to the plurality of second task feature word vectors are to the semantics of the preset processing and punishment regulation label corresponding to the second semantic distance. Among them, the method for determining the semantic distance between the plurality of second task feature word vectors and the feature word vector set of each preset processing and punishment regulation label in the plurality of preset processing and punishment regulation labels is similar to the method for determining the semantic distance between the plurality of first task feature word vectors and the feature word vector set of each preset processing and punishment regulation label in the plurality of preset qualitative regulation labels in the above-mentioned embodiment, and will not be repeated here.

[0099] Finally, the processing device will use the preset processing and penalty regulation label with the shortest second semantic distance among multiple preset processing and penalty regulation labels as the first processing and penalty regulation label, and thus determine the processing and penalty regulation clause corresponding to the audit question text based on the first processing and penalty regulation label.

[0100] It can be seen that based on the first violation type label, feature extraction can be performed on the feature word vectors in the word vector set to obtain multiple common feature word vectors. Then, by performing feature extraction on multiple feature word vectors through the qualitative regulation prediction task, multiple first sub-feature word vectors corresponding to the qualitative regulation prediction task can be obtained, so as to obtain multiple first task feature word vectors corresponding to the qualitative regulation prediction task based on the multiple common feature word vectors and the multiple first sub-feature word vectors. By determining the semantic distance between the multiple first task feature word vectors and the feature word vector set of each preset qualitative regulation label in the multiple preset qualitative regulation labels, the first qualitative regulation label can be determined from the multiple preset qualitative regulation labels, thereby obtaining the qualitative regulatory provisions corresponding to the audit problem text, thereby improving the efficiency and accuracy of the qualitative characterization of audit problems. Furthermore, by performing feature extraction on multiple feature word vectors through the processing and penalty regulations prediction task, multiple second sub-feature word vectors corresponding to the processing and penalty regulations prediction task are obtained, and a third sub-feature word vector corresponding to the first qualitative regulation label is determined. The multiple common feature word vectors, the third sub-feature word vectors, and the multiple second sub-feature word vectors are collectively used as the task feature word vectors corresponding to the processing and penalty regulations prediction task, so that the processing and penalty regulations prediction task can more accurately predict the processing and penalty regulations label. By determining the semantic distance between the multiple second task feature word vectors and the feature word vector set of each preset processing and penalty regulations label in the multiple preset processing and penalty regulations labels, the first processing and penalty regulations label can be determined from the multiple preset processing and penalty regulations labels, thereby obtaining the processing and penalty regulations clause corresponding to the audit problem text, thereby improving the efficiency and accuracy of the qualitative assessment of audit problems.

[0101] In an optional embodiment, the method may further include:

[0102] Obtain target project data corresponding to the audit question text from the project data;

[0103] Obtain the qualitative regulatory provisions corresponding to the first qualitative regulatory label, and obtain a verification relationship table based on the qualitative regulatory provisions;

[0104] According to the verification relationship table, the target project data is verified to obtain the verification results;

[0105] When the verification result is a verification match, the penalty regulation article corresponding to the first penalty regulation tag is obtained; and multiple penalty levels are determined according to the penalty regulation article;

[0106] According to the target project data, a penalty level is selected from multiple penalty levels as the target penalty level;

[0107] Determine the handling and punishment decision for the audited unit based on the target penalty level.

[0108] The verification relationship table represents the mapping relationship between project data corresponding to qualitative regulatory provisions. The verification relationship table can be pre-defined based on the content of the qualitative regulatory provisions. Verification results include a match or mismatch. It is understood that after determining the first qualitative regulatory tag and the first processing and penalty regulatory tag corresponding to the audit question text, the processing device will verify the first qualitative regulatory tag and the first processing and penalty regulatory tag to ensure a more accurate characterization of the audit question.

[0109] Specifically, the processing device retrieves target project data corresponding to the audit question text from the project data of the audited project. For example, if the audit question text is "The audited project was not completed on the contractual schedule," the processing device retrieves the contractual schedule and the actual completion schedule from the project data and uses these as the target project data. It then retrieves the qualitative regulatory text corresponding to the first qualitative regulatory tag and, based on the qualitative regulatory text, obtains a verification relationship table. For example, if the qualitative regulatory text mentions a project delay, the verification relationship table contains the contractual schedule and the theoretical completion schedule range. The processing device then verifies the target project data based on the verification relationship table and obtains a verification result. For example, it searches the verification relationship table for a contractual schedule that matches the audited project and determines the theoretical completion schedule range for that contractual schedule. If the actual completion schedule of the audited project falls within the theoretical completion schedule range, the project has not been delayed, and the verification result is a verification mismatch. If the actual completion period of the audited project is outside the theoretical completion period, it means that the construction period is delayed, and the verification result is a verification match.

[0110] If the verification result is a match, it indicates that the first qualitative regulatory tag matches the audit question text. The processing device then retrieves the penalty regulation text corresponding to the first penalty regulation tag. Based on the penalty regulation text, multiple penalty levels are determined. For example, if the penalty regulation text indicates a fine range of x1-x2 for the audited entity, the processing device can classify the penalty levels based on this range, with each penalty level corresponding to a specific amount from x1-x2.

[0111] Then, based on the target project data, a penalty level is selected from multiple penalty levels as the target penalty level. For example, if the project delay is 1-10 days, the first penalty level among the multiple penalty levels is used as the target penalty level; if the project delay is 11-30 days, the second penalty level among the multiple penalty levels is used as the target penalty level. In this way, the target penalty level for the audited unit can be determined, and the punishment decision for the audited unit can be determined based on the target penalty level.

[0112] As can be seen, by obtaining the target project data corresponding to the audit question text, the qualitative regulatory provisions corresponding to the first qualitative regulatory label are obtained, and a verification relationship table is obtained based on the qualitative regulatory provisions. Then, the target project data is verified through the verification relationship table to obtain a verification result, which verifies the first qualitative regulatory label and improves the accuracy of the audit question characterization. When the verification result is a verification match, the processing and penalty regulatory provisions corresponding to the first processing and penalty regulatory label are obtained, and multiple penalty levels are determined based on the processing and penalty regulatory provisions. Furthermore, based on the target project data, one penalty level is selected from the multiple penalty levels as the target penalty level, and the processing and penalty decision for the audited unit is determined based on the target penalty level, thereby improving the accuracy of the audit question characterization.

[0113] In an optional embodiment, the method may further include:

[0114] When the verification result is a mismatch, determine the preset relationship corresponding to the audit question text;

[0115] If the target project data satisfies the preset relationship, the target audit question text is determined based on the target project data, and the target audit question text is used as the audit question text to obtain a second violation type label, a second qualitative regulation label, and a second handling and penalty regulation label; based on the second handling and penalty regulation label, a handling and penalty decision for the audited unit is determined;

[0116] If the target project data does not meet the preset relationship, the audit question text is input into the document material library to obtain the target violation type, target qualitative regulatory provisions, and target handling and penalty regulatory provisions;

[0117] If the first violation type label and the label corresponding to the target violation type are inconsistent, determining a model adjustment parameter based on the first violation type label and the label corresponding to the target violation type; adjusting the parameters of the target violation type prediction model based on the model adjustment parameter to obtain an adjusted target violation type prediction model;

[0118] Input the word vector set into the adjusted target violation type prediction model to obtain the third violation type label;

[0119] Input the word vector set and the third violation type label into the target regulation prediction model to obtain the third qualitative regulation label and the third processing penalty regulation label;

[0120] When the third processing and penalty regulation label is consistent with the label corresponding to the target processing and penalty regulation article, the processing and penalty decision on the audited unit is determined based on the third processing and penalty regulation label.

[0121] The document material library is used to retrieve the violation types, qualitative regulatory provisions and handling and punishment regulatory provisions corresponding to the audit question text. Optionally, the document material library can be the audit question qualitative retrieval system in the above embodiment.

[0122] It is understandable that when the verification result is a mismatch, it indicates that the audit question text does not match the first qualitative regulatory label. In this case, one possibility is that the audit question text is incorrect, and another possibility is that the first violation type label predicted by the target violation type prediction model is inaccurate. Therefore, the processing device determines the preset relationship corresponding to the audit question text. For example, if the audit question text is "The project expenditures in the final accounts report of the audited project do not match the reported expenditures," the preset relationship is the mapping relationship between the project expenditures in the final accounts report and the reported expenditures.

[0123] If the target project data satisfies the preset relationship, it indicates that the audit question text is incorrect. The processing device can determine the target audit question text based on the target project data, use the target audit question text as the audit question text, and execute the method described in any of the above embodiments to obtain a second violation type label, a second qualitative regulatory label, and a second disciplinary and penalty regulatory label. The disciplinary and penalty regulatory provision corresponding to the second disciplinary and penalty regulatory label is then determined, and a disciplinary and penalty decision for the audited entity is determined based on the disciplinary and penalty regulatory provision corresponding to the second disciplinary and penalty regulatory label.

[0124] If the target project data does not meet the preset relationship, it indicates that the model prediction may have errors. The processing device will input the audit question text into the document material library to search the document material library for the target violation type, target qualitative regulatory provisions, and target handling and penalty regulatory provisions corresponding to the audit question text.

[0125] If the first violation type label and the label corresponding to the target violation type are inconsistent, it means that the first violation type label predicted by the target violation type prediction model is wrong, and the processing device will determine the model loss of the target violation type prediction model based on the first violation type label and the label corresponding to the target violation type. For example, the model loss can be a mean square error. According to the model loss of the target violation type prediction model, the model adjustment parameters of the target violation type prediction model can be determined. Among them, the mapping relationship between the model loss and the model adjustment parameters can be pre-set according to the actual test results. According to the model adjustment parameters of the target violation type prediction model, the parameters of the target violation type prediction model are adjusted to obtain the adjusted target violation type prediction model. Then, the word vector set is input into the adjusted target violation type prediction model to obtain a third violation type label, wherein the third violation type label is consistent with the target violation type label, and the word vector set and the third violation type label are input into the target regulation prediction model to obtain a third qualitative regulation label and a third processing penalty regulation label.

[0126] When the third processing and penalty regulation label is consistent with the label corresponding to the target processing and penalty regulation article, the processing and penalty regulation article corresponding to the third processing and penalty regulation label is determined, and the processing and penalty decision on the audited unit is determined based on the processing and penalty regulation article corresponding to the third processing and penalty regulation label.

[0127] If the third handling and penalty regulation label is inconsistent with the label corresponding to the target handling and penalty regulation article, the first loss of the target regulation prediction model is determined based on the third handling and penalty regulation label and the target penalty regulation label, and the second loss of the target regulation prediction model is determined based on the third qualitative regulation label and the target qualitative regulation label. The sum of the first loss and the second loss is taken as the total loss of the target regulation prediction model. The model parameters of the target regulation prediction model are adjusted based on the total loss of the target regulation prediction model until the qualitative regulation label predicted by the target regulation prediction model is consistent with the target qualitative regulation label, and the handling and penalty regulation label predicted by the target regulation prediction model is consistent with the target handling and penalty regulation label. The handling and penalty decision for the audited entity is determined based on the target handling and penalty regulation label.

[0128] Therefore, when the audit question text is incorrect, the audit question text can be adjusted; when the predicted results of the target violation type prediction model or the target regulation prediction model are incorrect, the model parameters of the target violation type prediction model or the target regulation prediction model can be adjusted to improve the accuracy of the qualitative characterization of the audit question.

[0129] In summary, in an embodiment of the present application, first, the audit question text of the audited unit on the audited project is obtained, and the audit question text is converted into a word vector set through a target word vector model. Then, the word vector set is input into the target violation type prediction model to obtain a first violation type label corresponding to the audit question text. Finally, the word vector set and the first violation type label are input into the target regulation prediction model, and the first qualitative regulation label corresponding to the audit question text is determined through the qualitative regulation prediction task of the target regulation prediction model, and the first processing and punishment regulation label corresponding to the audit question text and the first qualitative regulation label is determined through the processing and punishment regulation prediction task of the target regulation prediction model. Thus, through the pre-trained target violation type prediction model, the first violation type label corresponding to the audit question text can be determined based on the word vector set corresponding to the audit question text to determine the violation type corresponding to the audit question text. Through the pre-trained target regulation prediction model, the first qualitative regulation label and the first processing and punishment regulation label corresponding to the audit question text can be determined based on the word vector set and the first violation type label to determine the qualitative regulation clause and the processing and punishment regulation clause corresponding to the audit question text. Compared with manual qualitative analysis of audit issues, it can improve the efficiency of qualitative analysis of audit issues.

[0130] See Figure 5 , Figure 5 Schematic diagram of a device for qualitative analysis of audit questions provided in an embodiment of the present application. The device for qualitative analysis of audit questions 500 can be a processing device according to any of the above embodiments. The device for qualitative analysis of audit questions 500 includes an acquisition module 501 and a processing module 502.

[0131] The acquisition module 501 is used to acquire the audit problem text of the audited unit on the audited project; the audit problem text is a text describing the problems found in the project data of the audited project during the audit process;

[0132] Processing module 502, configured to convert the audit question text into a word vector set using a target word vector model; the word vector set includes a plurality of feature word vectors;

[0133] Input the word vector set into the target violation type prediction model to obtain the first violation type label corresponding to the audit question text; the first violation type label is used to indicate the violation type corresponding to the audit question text;

[0134] The word vector set and the first violation type label are input into the target regulation prediction model, and the first qualitative regulation label corresponding to the audit problem text is determined through the qualitative regulation prediction task, and the first processing penalty regulation label corresponding to the audit problem text and the first qualitative regulation label is determined through the processing penalty regulation prediction task; the qualitative regulation prediction task and the processing penalty regulation prediction task are both subtasks of the target regulation prediction model; the first qualitative regulation label is used to indicate the qualitative regulation article corresponding to the audit problem text, and the first processing penalty regulation label is used to indicate the processing penalty regulation article corresponding to the audit problem text and the qualitative regulation article.

[0135] In one possible embodiment, the target violation type prediction model includes: a BiLSTM model, a TextCNN model, an Attention layer, and a classifier. In inputting the word vector set into the target violation type prediction model to obtain a first violation type label corresponding to the audit question text, the processing module 502 is specifically configured to:

[0136] The feature vectors in the word vector set are extracted through the BiLSTM model to obtain a global feature set; the global feature set includes multiple global feature word vectors;

[0137] The TextCNN model is used to extract features from the feature word vectors in the word vector set and the global feature word vectors in the global feature set to obtain a local feature set; the local feature set includes at least one local feature word vector;

[0138] The attention layer is used to determine the attention weight corresponding to each feature word vector in the global feature set and the local feature set. Based on the attention weight corresponding to each feature word vector in the global feature set and the local feature set, the feature word vectors in the global feature set and the local feature set are weighted and summed to obtain the target feature word vector.

[0139] The target feature word vector is classified by the classifier to obtain the first violation type label.

[0140] In one possible embodiment, in inputting the word vector set and the first violation type label into the target regulation prediction model, determining the first qualitative regulation label corresponding to the audit question text through the qualitative regulation prediction task, and determining the first processing penalty regulation label corresponding to the audit question text and the first qualitative regulation label through the processing penalty regulation prediction task, the processing module 502 is specifically configured to:

[0141] According to the first violation type label, feature extraction is performed on multiple feature word vectors in the word vector set to obtain multiple common feature word vectors;

[0142] Extracting features from multiple feature word vectors through the qualitative law prediction task to obtain multiple first sub-feature word vectors corresponding to the qualitative law prediction task; using the multiple common feature word vectors and the multiple first sub-feature word vectors as task feature word vectors corresponding to the qualitative law prediction task to obtain multiple first task feature word vectors;

[0143] Determining semantic distances between the plurality of first task feature word vectors and a feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a plurality of first semantic distances;

[0144] determining a first qualitative regulatory label from a plurality of preset qualitative regulatory labels based on the plurality of first semantic distances;

[0145] Extract features from multiple feature word vectors by processing the penalty law prediction task to obtain multiple second sub-feature word vectors corresponding to the penalty law prediction task;

[0146] Determine the third sub-feature word vector corresponding to the first qualitative regulatory label;

[0147] The plurality of common feature word vectors, the third sub-feature word vectors, and the plurality of second sub-feature word vectors are collectively used as task feature word vectors corresponding to the penalty regulations prediction task, to obtain a plurality of second task feature word vectors;

[0148] Determining the semantic distances between the plurality of second task feature word vectors and the feature word vector set of each preset processing penalty regulation label in the plurality of preset processing penalty regulation labels to obtain a plurality of second semantic distances;

[0149] According to the plurality of second semantic distances, a first processing penalty regulation label is determined from a plurality of preset processing penalty regulation labels.

[0150] In one possible embodiment, in determining the semantic distances between the plurality of first task feature word vectors and the feature word vector set of each of the plurality of predefined regulatory labels to obtain the plurality of first semantic distances, the processing module 502 is specifically configured to:

[0151] Determine the intersection of the plurality of first task feature word vectors and the feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a plurality of feature word vector intersections;

[0152] Determine a union of the plurality of first task feature word vectors and a feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a plurality of feature word vector unions;

[0153] Determine multiple similarity coefficients according to the intersection of multiple feature word vectors and the union of multiple feature word vectors;

[0154] A plurality of first semantic distances are determined according to the plurality of similarity coefficients.

[0155] In a possible embodiment, before converting the audit question text into a word vector set using the target word vector model, the processing module 502 is further configured to:

[0156] Get the reference word vector model;

[0157] Obtain multiple sets of training data, each set of training data including the audit question text corresponding to the training data and the word vector set corresponding to the audit question text;

[0158] Performing preprocessing operations on the audit question text corresponding to each set of training data in the multiple sets of training data to obtain multiple sets of target training data; the preprocessing operations include word segmentation, stop word removal, and format conversion;

[0159] The reference word vector model is trained using multiple sets of target training data to obtain the target word vector model.

[0160] In a possible embodiment, the processing module 502 is further configured to:

[0161] Obtain target project data corresponding to the audit question text from the project data;

[0162] Obtain the qualitative regulatory article corresponding to the first qualitative regulatory tag, and obtain a verification relationship table based on the qualitative regulatory article; the verification relationship table is used to represent the mapping relationship between the project data corresponding to the qualitative regulatory article;

[0163] Verify the target project data according to the verification relationship table to obtain the verification results; the verification results include verification match or verification mismatch;

[0164] When the verification result is a verification match, the penalty regulation article corresponding to the first penalty regulation tag is obtained; and multiple penalty levels are determined according to the penalty regulation article;

[0165] According to the target project data, a penalty level is selected from multiple penalty levels as the target penalty level;

[0166] Determine the handling and punishment decision for the audited unit based on the target penalty level.

[0167] In a possible embodiment, the processing module 502 is further configured to:

[0168] When the verification result is a mismatch, determine the preset relationship corresponding to the audit question text;

[0169] If the target project data satisfies the preset relationship, the target audit question text is determined based on the target project data, and the target audit question text is used as the audit question text to obtain a second violation type label, a second qualitative regulation label, and a second handling and penalty regulation label; based on the second handling and penalty regulation label, a handling and penalty decision for the audited unit is determined;

[0170] If the target project data does not meet the preset relationship, the audit question text is input into the document material library to obtain the target violation type, target qualitative regulatory provisions, and target handling and penalty regulatory provisions; the document material library is used to retrieve the violation type, qualitative regulatory provisions, and handling and penalty regulatory provisions corresponding to the audit question text;

[0171] If the first violation type label and the label corresponding to the target violation type are inconsistent, determining a model adjustment parameter based on the first violation type label and the label corresponding to the target violation type; adjusting the parameters of the target violation type prediction model based on the model adjustment parameter to obtain an adjusted target violation type prediction model;

[0172] Input the word vector set into the adjusted target violation type prediction model to obtain the third violation type label;

[0173] Input the word vector set and the third violation type label into the target regulation prediction model to obtain the third qualitative regulation label and the third processing penalty regulation label;

[0174] When the third processing and penalty regulation label is consistent with the label corresponding to the target processing and penalty regulation article, the processing and penalty decision on the audited unit is determined based on the third processing and penalty regulation label.

[0175] See Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 6 As shown, electronic device 600 includes a transceiver 601, a processor 602, and a memory 603. These are connected via a bus 604. Memory 603 is used to store computer programs and data and transmit the data stored in memory 603 to processor 602. Electronic device 600 can be the qualitative device 500 for audit questions. Electronic device 600 can also be the processing device of any of the above embodiments.

[0176] The processor 602 is configured to read the computer program in the memory 603 and perform the following operations:

[0177] Obtain the audit question text of the audited unit on the audited project;

[0178] The audit question text is converted into a word vector set through the target word vector model; the word vector set includes multiple feature word vectors;

[0179] Input the word vector set into the target violation type prediction model to obtain the first violation type label corresponding to the audit question text;

[0180] The word vector set and the first violation type label are input into the target regulation prediction model, and the first qualitative regulation label corresponding to the audit problem text is determined through the qualitative regulation prediction task. The first processing penalty regulation label corresponding to the audit problem text and the first qualitative regulation label is determined through the processing penalty regulation prediction task.

[0181] The above mainly introduces the solution of the embodiment of the present application from the perspective of the execution process of the method side. It is understandable that, in order to realize the above functions, the electronic device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiment provided herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0182] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement part or all of the steps of any one of the methods described in the above method embodiments.

[0183] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any one of the methods described in the above method embodiments.

[0184] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required for this application.

[0185] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0186] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0187] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0188] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of software program modules.

[0189] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0190] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0191] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for those skilled in the art, according to the idea of ​​the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A qualitative approach to audit questions, characterized by: include: Obtain the audit question text of the audited unit on the audited project; The audit question text is a text describing the problems found in the project data of the audited project during the audit process; Converting the audit question text into a word vector set through a target word vector model; the word vector set includes multiple feature word vectors; Inputting the word vector set into a target violation type prediction model to obtain a first violation type label corresponding to the audit question text; The first violation type label is used to indicate the violation type corresponding to the audit question text; Input the word vector set and the first violation type label into the target regulation prediction model, determine the first qualitative regulation label corresponding to the audit question text through the qualitative regulation prediction task, and determine the first processing penalty regulation label corresponding to the audit question text and the first qualitative regulation label through the processing penalty regulation prediction task; the qualitative regulation prediction task and the processing penalty regulation prediction task are both subtasks of the target regulation prediction model; the first qualitative regulation label is used to indicate the qualitative regulation article corresponding to the audit question text, and the first processing penalty regulation label is used to indicate the processing penalty regulation article corresponding to the audit question text and the qualitative regulation article; The step of inputting the word vector set and the first violation type label into a target regulation prediction model, determining a first qualitative regulation label corresponding to the audit question text through a qualitative regulation prediction task, and determining a first processing penalty regulation label corresponding to the audit question text and the first qualitative regulation label through a processing penalty regulation prediction task includes: performing feature extraction on a plurality of feature word vectors in the word vector set according to the first violation type label to obtain a plurality of common feature word vectors; Performing feature extraction on the multiple feature word vectors through the qualitative law prediction task to obtain multiple first sub-feature word vectors corresponding to the qualitative law prediction task; using the multiple common feature word vectors and the multiple first sub-feature word vectors together as task feature word vectors corresponding to the qualitative law prediction task to obtain multiple first task feature word vectors; Determining semantic distances between the plurality of first task feature word vectors and a feature word vector set of each of a plurality of preset qualitative regulatory labels to obtain a plurality of first semantic distances; determining the first qualitative regulatory label from the plurality of preset qualitative regulatory labels based on the plurality of first semantic distances; Performing feature extraction on the plurality of feature word vectors through the processing penalty law prediction task to obtain a plurality of second sub-feature word vectors corresponding to the processing penalty law prediction task; Determining a third sub-feature word vector corresponding to the first qualitative regulatory label; The plurality of common feature word vectors, the third sub-feature word vectors, and the plurality of second sub-feature word vectors are collectively used as task feature word vectors corresponding to the processing penalty regulations prediction task, to obtain a plurality of second task feature word vectors; Determining the semantic distances between the plurality of second task feature word vectors and the feature word vector set of each preset processing penalty regulation label in the plurality of preset processing penalty regulation labels to obtain a plurality of second semantic distances; The first processing penalty regulation label is determined from the plurality of preset processing penalty regulation labels according to the plurality of second semantic distances.

2. The method according to claim 1, characterized in that The target violation type prediction model includes: a BiLSTM model, a TextCNN model, an Attention layer, and a classifier; inputting the word vector set into the target violation type prediction model to obtain a first violation type label corresponding to the audit question text includes: Extracting features from the feature word vectors in the word vector set using the BiLSTM model to obtain a global feature set; the global feature set includes multiple global feature word vectors; Performing feature extraction on the feature word vectors in the word vector set and the global feature word vectors in the global feature set using the TextCNN model to obtain a local feature set; the local feature set includes at least one local feature word vector; Determining the attention weight corresponding to each feature word vector in the global feature set and the local feature set through the Attention layer; performing a weighted summation of the feature word vectors in the global feature set and the local feature set according to the attention weight corresponding to each feature word vector in the global feature set and the local feature set to obtain a target feature word vector; The target feature word vector is classified by the classifier to obtain the first violation type label.

3. The method according to claim 1, characterized in that Determining the semantic distances between the plurality of first task feature word vectors and the feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a plurality of first semantic distances includes: Determine the intersection of the plurality of first task feature word vectors and the feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a plurality of feature word vector intersections; Determining a union of the plurality of first task feature word vectors and a feature word vector set of each of the plurality of preset qualitative regulatory labels to obtain a plurality of feature word vector unions; Determining a plurality of similarity coefficients according to the intersection of the plurality of feature word vectors and the union of the plurality of feature word vectors; The plurality of first semantic distances are determined according to the plurality of similarity coefficients.

4. The method according to any one of claims 1 to 3, characterized in that Before converting the audit question text into a word vector set through the target word vector model, the method further includes: Get the reference word vector model; Obtain multiple sets of training data, each set of training data including the audit question text corresponding to the training data and the word vector set corresponding to the audit question text; Performing a preprocessing operation on the audit question text corresponding to each set of training data in the multiple sets of training data to obtain multiple sets of target training data; the preprocessing operation includes a word segmentation operation, a stop word removal operation, and a format conversion operation; The reference word vector model is trained using the multiple sets of target training data to obtain the target word vector model.

5. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Acquire target project data corresponding to the audit question text from the project data; Obtaining the qualitative regulatory article corresponding to the first qualitative regulatory tag, and obtaining a verification relationship table based on the qualitative regulatory article; the verification relationship table is used to represent the mapping relationship between the project data corresponding to the qualitative regulatory article; Verify the target project data according to the verification relationship table to obtain a verification result; the verification result includes a verification match or a verification mismatch; When the verification result is the verification match, obtaining the processing and punishment regulation article corresponding to the first processing and punishment regulation article; and determining multiple punishment levels according to the processing and punishment regulation article; selecting a penalty level from the plurality of penalty levels as a target penalty level according to the target item data; Determine the handling and punishment decision for the audited unit based on the target penalty level.

6. The method according to claim 5, characterized in that The method further comprises: When the verification result is the verification mismatch, determining a preset relationship corresponding to the audit question text; If the target project data satisfies the preset relationship, a target audit question text is determined based on the target project data, the target audit question text is used as the audit question text, and a second violation type label, a second qualitative regulation label, and a second handling and penalty regulation label are obtained; and a handling and penalty decision for the audited unit is determined based on the second handling and penalty regulation label; If the target project data does not satisfy the preset relationship, the audit question text is input into the document material library to obtain the target violation type, target qualitative regulatory provisions, and target handling and punishment regulatory provisions; the document material library is used to retrieve the violation type, qualitative regulatory provisions, and handling and punishment regulatory provisions corresponding to the audit question text; If the first violation type label and the label corresponding to the target violation type are inconsistent, determining a model adjustment parameter based on the first violation type label and the label corresponding to the target violation type; adjusting the parameters of the target violation type prediction model based on the model adjustment parameter to obtain an adjusted target violation type prediction model; Inputting the word vector set into the adjusted target violation type prediction model to obtain a third violation type label; Inputting the word vector set and the third violation type label into the target regulation prediction model to obtain a third qualitative regulation label and a third processing penalty regulation label; When the third processing and punishment regulation label is consistent with the label corresponding to the target processing and punishment regulation article, the processing and punishment decision for the audited unit is determined according to the third processing and punishment regulation label.

7. A qualitative device for audit questions, characterized in that include: The acquisition module is used to obtain the audit question text of the audited unit on the audited project; The audit question text is a text describing the problems found in the project data of the audited project during the audit process; A processing module, configured to convert the audit question text into a word vector set through a target word vector model; the word vector set includes a plurality of feature word vectors; Inputting the word vector set into a target violation type prediction model to obtain a first violation type label corresponding to the audit question text; The first violation type label is used to indicate the violation type corresponding to the audit question text; Input the word vector set and the first violation type label into the target regulation prediction model, determine the first qualitative regulation label corresponding to the audit question text through the qualitative regulation prediction task, and determine the first processing penalty regulation label corresponding to the audit question text and the first qualitative regulation label through the processing penalty regulation prediction task; the qualitative regulation prediction task and the processing penalty regulation prediction task are both subtasks of the target regulation prediction model; the first qualitative regulation label is used to indicate the qualitative regulation article corresponding to the audit question text, and the first processing penalty regulation label is used to indicate the processing penalty regulation article corresponding to the audit question text and the qualitative regulation article; The step of inputting the word vector set and the first violation type label into a target regulation prediction model, determining a first qualitative regulation label corresponding to the audit question text through a qualitative regulation prediction task, and determining a first processing penalty regulation label corresponding to the audit question text and the first qualitative regulation label through a processing penalty regulation prediction task includes: performing feature extraction on a plurality of feature word vectors in the word vector set according to the first violation type label to obtain a plurality of common feature word vectors; Performing feature extraction on the multiple feature word vectors through the qualitative law prediction task to obtain multiple first sub-feature word vectors corresponding to the qualitative law prediction task; using the multiple common feature word vectors and the multiple first sub-feature word vectors together as task feature word vectors corresponding to the qualitative law prediction task to obtain multiple first task feature word vectors; Determining semantic distances between the plurality of first task feature word vectors and a feature word vector set of each of a plurality of preset qualitative regulatory labels to obtain a plurality of first semantic distances; determining the first qualitative regulatory label from the plurality of preset qualitative regulatory labels based on the plurality of first semantic distances; Performing feature extraction on the plurality of feature word vectors through the processing penalty law prediction task to obtain a plurality of second sub-feature word vectors corresponding to the processing penalty law prediction task; Determining a third sub-feature word vector corresponding to the first qualitative regulatory label; The plurality of common feature word vectors, the third sub-feature word vectors, and the plurality of second sub-feature word vectors are collectively used as task feature word vectors corresponding to the processing penalty regulations prediction task, to obtain a plurality of second task feature word vectors; Determining the semantic distances between the plurality of second task feature word vectors and the feature word vector set of each preset processing penalty regulation label in the plurality of preset processing penalty regulation labels to obtain a plurality of second semantic distances; The first processing penalty regulation label is determined from the plurality of preset processing penalty regulation labels according to the plurality of second semantic distances.

8. An electronic device, characterized in that: include: A processor and a memory, the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Auditing problem legal basis analysis method and system based on large model

    CN118445430A

  • Distributed rules based hierarchical auditing system for management of health care services

    US20220198413A1