Deep learning based computer document intelligent compliance detection system
By using deep learning technology, combined with convolutional neural networks and Transformer models, the efficiency and accuracy problems of traditional document compliance detection methods have been solved. This approach enables the fusion of multimodal features of documents and the dynamic parsing of complex rules, thereby improving the efficiency and accuracy of document compliance detection.
Patent Information
- Application Number
- CN202511101919.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-07
Smart Images

Figure CN120596657B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a computer document intelligent compliance detection system based on deep learning. BACKGROUND
[0002] With the deepening of digital transformation, computer documents as the core carrier of information transmission and business flow, their compliance is directly related to the risk control and business standardization in key fields such as finance, law and medicine. For example, the contract documents of the financial industry need to comply with the compliance provisions of the regulatory agencies, legal documents need to strictly follow the provisions of the logic, and medical records need to meet the requirements of privacy protection and data specifications. However, the diversification of document types, the complexity of content, and the dynamic updating of rules make it difficult for traditional document compliance detection methods.
[0003] Some traditional document compliance detection mainly relies on manual review and automatic verification based on fixed rules, which has limitations. Manual review is limited by low efficiency, subjective bias and misjudgment risk caused by fatigue, which cannot adapt to the real-time detection needs of large-scale documents. Some technologies cannot analyze the spatial structure characteristics, context semantic association and complex rule nested logic of documents, affecting the detection accuracy of unstructured documents and multi-modal fusion documents. SUMMARY
[0004] The technical problem to be solved by the present application is to provide a computer document intelligent compliance detection system based on deep learning, which effectively improves the efficiency of document compliance detection.
[0005] To solve the above technical problems, the technical scheme of the present application is as follows:
[0006] In a first aspect, the computer document intelligent compliance detection system based on deep learning comprises:
[0007] A feature fusion module is used to input the document to be detected, extract the spatial structure feature vector of the document image through a convolutional neural network, obtain the text sequence through an optical character recognition engine, and generate the context semantic vector through a Transformer large model encoder. The document logic level is parsed to generate a structure descriptor. The three types of feature vectors are executed through a feature space mapping matching algorithm, and the mapped features are fused through a cross-modal attention mechanism to generate a joint feature tensor.
[0008] A feature calculation module is used to load the latest rule set based on the formal rule base and the joint feature tensor, calculate the logical satisfaction degree between the entities and the rule predicates in the feature tensor through a semantic alignment module, and perform core verification operations through a differentiable logic layer to obtain an abnormal set and a corresponding decision path node sequence.
[0009] a structure analysis module, configured to group the abnormal entities into a node index to construct a relation adjacency matrix, analyze a topology of the adjacency matrix by using a graph neural network to detect abnormalities to obtain a structure deviation degree, and parse a decision path node sequence by using a pre-trained language model to generate a semantic conflict feature;
[0010] an optimization parameter module, configured to generate an augmented rule set according to the risk feature vector, update a formal rule library, update a feature encoding component parameter by gradient back propagation, optimize a differentiable logic layer judgment threshold and a weight of the analysis module, and obtain an updated feature encoding component, a rule library version and a risk quantization parameter.
[0011] Further, a joint feature tensor is generated, including:
[0012] The document image to be detected is input into a pre-trained convolutional neural network, multi-scale spatial features are extracted by alternating processing of convolutional layers and pooling layers, a spatial structure feature vector is obtained through a full connection layer, a document text sequence is extracted by using an optical character recognition engine, after sub-word segmentation and word vector conversion, the document text sequence is input into a Transformer encoder for multi-layer self-attention calculation and feature stabilization processing, and finally a context semantic vector is generated; a document logical topology is organized based on document layout analysis, and a structure descriptor matrix is obtained through calculation;
[0013] The spatial structure feature vector, the context semantic vector and the structure descriptor matrix are respectively subjected to linear transformation, affine transformation and vector projection to generate image feature vectors, text semantic vectors and structure description vectors after projection of the same dimension;
[0014] The structure description vector is taken as a query core, an association weight between the structure description vector and image and text features is calculated by using a cross-modal attention mechanism, visual and semantic information is weighted and aggregated, and finally a joint feature tensor is generated.
[0015] Further, based on the latest rule set loaded by the formal rule library and the joint feature tensor, a semantic alignment module is used to calculate a logical satisfaction degree between entities and rule predicates in the feature tensor; a differentiable logic layer is used to perform a core verification operation to obtain an abnormality set and a corresponding decision path node sequence, including:
[0016] Based on the latest rule set loaded by the formal rule library and the joint feature tensor, a rule is parsed into a structured predicate unit and an entity segment is located, three-level progressive alignment operations of accurate name matching, semantic similarity matching and context reasoning matching are sequentially performed, and an entity-rule predicate binding pair set is generated.
[0017] According to the entity-rule predicate binding pair in the set, the satisfaction degree calculation operation based on the differentiable comparison is performed respectively according to the types of the numerical rules, the relational rules and the composite rules, and a logical satisfaction degree set is generated;
[0018] Based on the logical satisfaction degree set, the candidate abnormal entity is marked by a dynamic decision threshold, the abnormal list of the violation evidence and the decision path node sequence chain is generated by backtracking, and an abnormal set is obtained.
[0019] Further, a risk feature vector is generated, including:
[0020] Based on the structured abnormal set, the original document logical relationship adjacency matrix is inherited with the abnormal entity as a graph node, the cross-entity association edge is added through the rule violation evidence, and the decision path node sequence chain is integrated, and a topologically enhanced abnormal association matrix is generated;
[0021] Based on the topologically enhanced abnormal association matrix, the global graph representation vector is generated by performing multi-layer graph convolution aggregation and graph attention calculation through the pre-trained graph neural network, the cosine similarity with the compliance document benchmark graph vector is calculated to generate a structure deviation scalar value; based on the decision path node sequence chain, the semantic segmentation, vectorization and self-attention conflict weight calculation are performed on the path segment converted into natural language description through the pre-trained language model, the highest conflict node type identifier, the rule predicate logical contradiction frequency and the language model conflict probability value are extracted, and a multi-dimensional semantic conflict feature vector is generated;
[0022] Based on the structure deviation scalar value and the multi-dimensional semantic conflict feature vector, the risk feature vector is generated through numerical normalization processing, feature splicing and double-layer fully connected layer fusion.
[0023] Further, an augmented rule set is generated according to the risk feature vector, and the formal rule library is updated; the feature encoding component parameters are updated through gradient back propagation; the differentiable logic layer decision threshold and the weight of the analysis module are optimized; the updated feature encoding component, rule library version and risk quantization parameter are obtained, including:
[0024] Based on the risk feature vector, an augmented rule draft set containing a confidence score is generated, if the sample verification false positive rate decreases and the recall rate improves, an effective augmented rule set is generated, and rule redundancy processing is performed according to the confidence, and finally a versioned and updated formal rule library is obtained;
[0025] Based on the risk feature vector, the gradient is calculated and the chain back propagation is processed, the parameters of the fusion layer, the graph network, the cross-modal module and the encoder in the feature encoding component chain are updated, and the optimized feature encoding component is output.
[0026] Based on the feature encoding component and the satisfaction set, the dynamic decision threshold boundary is adjusted through the satisfaction probability density distribution, and the offset is fine-tuned combined with the false positive rate index, the graph neural network edge weight coefficient is recalibrated based on the update component, and the semantic conflict feature dimension weight is redistributed, to obtain the updated decision threshold and analysis module weight;
[0027] The versioned rule base, the optimized feature encoding component, and the decision threshold and analysis module weight are fused, the rule base change summary is packaged, the component parameter snapshot is stored, and the risk quantization parameter set is recorded.
[0028] Further, based on the latest rule set loaded by the formal rule base and the joint feature tensor, the rules are parsed into structured predicate units and the entity fragments are located, and the three-level progressive alignment operations of accurate name matching, semantic similarity matching and context reasoning matching are sequentially performed, to generate an entity-rule predicate binding pair set, including:
[0029] Load the latest rule set in the formal rule base, parse each rule into a structured predicate unit, generate a rule subject name set, and perform entity recognition on the joint feature tensor to extract key entity fragments to generate an entity name set;
[0030] If the entity name and the rule subject name are completely consistent, the entity and the corresponding predicate unit are directly bound; a primary binding pair is generated; and unmatched entities and rule subject names generate a candidate set;
[0031] Based on the candidate set, a semantic binding pair is generated through the cosine similarity of the semantic vector, and a set to be inferred is generated; based on the set to be inferred, according to the context semantics and structural descriptors of the joint feature tensor, an inference binding pair is generated through the differentiable reasoning module; the binding pairs of the three levels of matching are integrated, the predicates are selected according to the priority of accurate> semantic> inference, and the entity-rule predicate binding pair set is generated.
[0032] Further, based on the feature encoding component and the satisfaction set, the dynamic decision threshold boundary is adjusted through the satisfaction probability density distribution, and the offset is fine-tuned combined with the false positive rate index, the graph neural network edge weight coefficient is recalibrated based on the update component, and the semantic conflict feature dimension weight is redistributed, to obtain the updated decision threshold and analysis module weight, including:
[0033] Based on the logical satisfaction set, the satisfaction values are arranged in ascending order to generate an ordered satisfaction sequence, the satisfaction value corresponding to the preset quantile position is calculated to generate an initial boundary value, the false positive rate in the sample verification stage is loaded, and the corrected boundary value is obtained after calculation;
[0034] Based on the feature encoding component and the revised boundary value, a new node feature vector is obtained, and an edge weight coefficient is generated through edge weight calculation and sparsification processing; based on the multi-dimensional semantic conflict feature vector and the edge weight coefficient, a final dimension weight vector is obtained through conflict feature contribution degree statistics, weight base generation and weight attenuation;
[0035] Based on the revised boundary value, the edge weight coefficient and the final dimension weight vector, the dynamic decision threshold of the differentiable logic layer, the adjacency weight matrix of the graph convolution layer and the weight tensor of the semantic conflict feature fusion layer are updated respectively to complete the parameter integration of the analysis module.
[0036] Further, based on the risk feature vector, an augmented rule draft set containing a confidence score is generated, and if the sample verification false positive rate decreases and the recall rate increases, an effective augmented rule set is generated, and the rules are sorted according to the confidence and executed to remove redundancy, and finally a versioned updated formal rule library is obtained, including:
[0037] In the risk feature vector, abnormal entity type identifiers, rule violation feature identifiers and context semantic fingerprints are extracted, parameterized rule expressions are generated by matching structured rule data; and based on the risk vector and the historical rule violation frequency weighted, an augmented rule draft set containing a confidence score is generated;
[0038] The draft set is verified by calculating the change amount of the recall rate and the false positive rate to obtain an effective rule subset, and adjacent rules are detected after confidence sorting to generate a deduplication augmented rule set;
[0039] According to the deduplication augmented rule set, a change summary containing a new rule ID, a confidence and an abnormal type identifier is generated through version number iteration to obtain a versioned updated rule library.
[0040] In a second aspect, a computing device includes:
[0041] One or more processors;
[0042] A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the system.
[0043] In a third aspect, a computer readable storage medium stores a program, which is executed by a processor to implement the system.
[0044] The above-mentioned scheme of the present application at least has the following beneficial effects:
[0045] The spatial structure feature of the document image is extracted through the convolutional neural network, the text context semantic vector is generated through the Transformer model, the structure descriptor is generated by analyzing the logic level, and the fusion of the three types of features is realized by combining the cross-modal attention mechanism, so that the problem that the traditional method cannot capture the complete feature of the multi-modal document is solved, and the joint feature tensor can comprehensively reflect the structure, semantic and spatial information of the document.
[0046] Based on the latest rule set of the formal rule base, the logical satisfaction degree of the entity and the rule predicate is quantified through the semantic alignment module, and the core verification is performed by using the differentiable logic layer, the limitation that the traditional fixed rule matching can only process shallow semantics is improved, the nested logic and dynamic updating of the rule can be effectively analyzed, and the abnormal set and the decision path node sequence are output, the logical rigor and the result explainability of the compliance detection are improved. BRIEF DESCRIPTION OF DRAWINGS
[0047] Figure 1 It is a schematic diagram of a computer document intelligent compliance detection system based on deep learning provided by an embodiment of the present application.
[0048] Figure 2 It is a flowchart of generating an entity-rule predicate binding pair set. DETAILED DESCRIPTION
[0049] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art.
[0050] As Figure 1 shown, an embodiment of the present application proposes a computer document intelligent compliance detection system based on deep learning, which comprises:
[0051] The feature fusion module is used for inputting a document to be detected, extracting a spatial structure feature vector of the document image through a convolutional neural network, obtaining a text sequence through an optical character recognition engine, generating a context semantic vector through a Transformer large model encoder, and analyzing a document logic level to generate a structure descriptor; the three types of feature vectors are subjected to a feature space mapping matching algorithm, the mapped features are fused through a cross-modal attention mechanism, and a joint feature tensor is generated;
[0052] The feature calculation module is configured to calculate logical satisfaction degrees between entities in the feature tensor and rule predicates based on a latest rule set loaded by the formal rule base and the joint feature tensor, and through the semantic alignment module; and perform core verification operations by using a differentiable logic layer to obtain an exception set and a corresponding decision path node sequence;
[0053] The structure analysis module is configured to construct a relation adjacency matrix by taking the exception entity set as a node index; analyze the adjacency matrix by using a graph neural network to detect an exception to obtain a structure deviation degree, and parse the decision path node sequence by using a pre-trained language model to generate a semantic conflict feature; and fuse the structure deviation degree and the semantic conflict feature to generate a risk feature vector;
[0054] The optimization parameter module is configured to generate an augmented rule set according to the risk feature vector, update the formal rule base, update parameters of the feature encoding component by gradient back propagation, optimize a decision threshold of the differentiable logic layer and weights of the analysis module, and obtain an updated feature encoding component, a rule base version and risk quantization parameters.
[0055] In the embodiment of the application, the convolutional neural network is used to extract spatial structure features of a document image, the Transformer model is used to generate a text context semantic vector, the logical level is parsed to generate a structure descriptor, and a cross-modal attention mechanism is combined to realize fusion of the three types of features, thereby solving the problem of incomplete feature capture of a multi-modal document in a traditional method, and enabling the joint feature tensor to comprehensively reflect the structure, semantics and spatial information of the document.
[0056] Based on a latest rule set of the formal rule base, the semantic alignment module is used to quantify logical satisfaction degrees between entities and rule predicates, and the differentiable logic layer is used to perform core verification, thereby improving the limitation of a traditional fixed rule matching that can only process shallow semantics, effectively parsing nested logic and dynamic updating of rules, simultaneously outputting an exception set and a decision path node sequence, and improving logical rigor and result interpretability of compliance detection.
[0057] In a preferred embodiment of the application, the joint feature tensor is generated, including:
[0058] The document image to be detected is input into a pre-trained convolutional neural network, multi-scale spatial features are extracted by alternating processing of convolutional layers and pooling layers, and a spatial structure feature vector is obtained through a full connection layer; a document text sequence is extracted by using an optical character recognition engine, and after sub-word segmentation and word vector conversion, the document text sequence is input into a Transformer encoder for multi-layer self-attention calculation and feature stabilization processing, and finally a context semantic vector is generated; a document logical topology is organized based on document layout analysis, and a structure descriptor matrix is obtained through calculation;
[0059] The spatial structure feature vector, the context semantic vector and the structure descriptor matrix are respectively subjected to linear transformation, affine transformation and vector projection, to generate image feature vectors, text semantic vectors and structure description vectors after projection in the same dimension;
[0060] The structure description vector is taken as a query core, the association weight of the structure description vector with the image and the text feature is calculated through a cross-modal attention mechanism, and a joint feature tensor is finally generated after weighted aggregation of visual and semantic information.
[0061] In the embodiment of the application, the model is constructed: an improved ResNet-50 architecture is adopted as a basic network, including 5 convolution stages (each stage is composed of 1-3 residual blocks), each residual block includes 2 3x3 convolution layers (step 1), a batch normalization layer and a ReLU activation function, and down-sampling is realized between stages through a 1x1 convolution layer with a step of 2; a global average pooling layer and one fully connected layer (output dimension 512) are arranged at the end of the network.
[0062] The training data set: a public data set (such as RVL-CDIP) containing 100,000+ document images and industry-specific document images (such as prosecution document scans and case document scans) are adopted, covering different page layouts (single column / multi-column, mixed text and graphics), resolutions (300-600 dpi) and noise types (blur, tilt).
[0063] The document image is unified to 224x224 pixels, and after being converted into a grayscale image, standardization (mean 0.5, standard deviation 0.25) is performed.
[0064] The optimizer adopts Adam (initial learning rate 1e-4, decay to 1 / 10 of the previous value every 5 epochs), the loss function is cross-entropy loss (pretrained for document layout classification task), the training round is 30 rounds, and the batch size is 32; based on the pre-trained model (pre-trained on ImageNet), the last 3 convolution stages and the fully connected layer are fine-tuned using industry document data, to improve the feature capture ability for specific field document structures.
[0065] Implementation process: after the document image to be detected is preprocessed and input into the model, feature extraction and down-sampling are alternately performed through the 5 convolution stages (the output feature map sizes are 112x112, 56x56, 28x28, 14x14 and 7x7 in sequence), the feature matrix of 7x7x2048 is compressed through the global average pooling layer, and finally the 512-dimensional spatial structure feature vector is output through the fully connected layer.
[0066] Context semantic vector generation (based on the Transformer large model encoder)
[0067] Model Construction: Adopting a 6-layer Transformer encoder architecture, each layer contains 1 multi-head self-attention sublayer (8 attention heads, each head dimension 64) and 1 feedforward neural network sublayer (hidden layer dimension 2048, output dimension 512), and setting residual connection and layer normalization (ε = 0.00001), using public text corpus (such as Wikipedia, legal / financial field regulation text) and historical compliance document text (about 500,000 pieces), to build a semantic training set containing entity relationship and rule clauses.
[0068] Extracting document text sequence through Tesseract OCR engine (supporting multiple languages, character recognition accuracy ≥98%), subword segmentation through Byte Pair Encoding (BPE) (vocabulary size 32000), converting to 128-dimensional word vector (using pre-trained GloVe word vector initialization).
[0069] Optimizer using AdamW (learning rate 0.00005, weight decay 0.01), loss function is mask language model loss (random mask 15% subwords), training rounds 20, batch size 64.
[0070] Using industry compliance rule text to fine-tune the encoder, strengthening the coding ability of specific semantics such as "prohibition clauses" and "obligation clauses".
[0071] Implementation process: After subword segmentation and word vector conversion of the text sequence extracted by OCR, input the Transformer encoder, each layer calculates the context association between subwords through multi-head self-attention (attention weight range 0-1), after nonlinear transformation by feedforward network, finally output 512-dimensional context semantic vector (sequence length fixed at 512, fill in the gap if not enough, truncate if too long).
[0072] Layout analysis: Using a layout analysis module based on LayoutLM, identifying elements such as titles, paragraphs, tables, and signatures in the document (accuracy ≥95%), and building a logical topology structure through element coordinates and hierarchical relationships (such as "chapter-section-article").
[0073] Matrix construction: Taking elements as nodes, node attributes include type (title / paragraph, etc., represented by one-hot encoding), position (normalized coordinates, range 0-1), length (number of characters, normalized to 0-1); taking logical associations between elements (such as "contains" and "parallel") as edges, edge weights are association strength (based on position distance and semantic similarity calculation, range 0-1), generating an N×N structure descriptor matrix (N is the number of document elements, usually 50-200).
[0074] Feature space mapping: spatial structure feature vector (512 dimensions) is mapped to the same dimension image feature vector through linear transformation (weight matrix is randomly initialized, dimension 512x512).
[0075] Contextual semantic vector (512 dimensions): mapped to the same dimension text semantic vector through affine transformation (including linear transformation + bias term, bias range [-0.1, 0.1]).
[0076] Structure descriptor matrix (N x N) is mapped to a 512-dimensional structure description vector through vector projection (concatenated by rows and compressed by linear transformation).
[0077] The structure description vector is used as a query (Q), and the image feature vector and the text semantic vector are respectively the key (K) and the value (V). The attention weight is calculated (through the scaled dot-product attention mechanism, the weight range is 0-1).
[0078] The spatial structure feature vector as input comes from the processing result of the convolutional neural network, with a total of 512 dimensions. The numerical value of each dimension is standardized to control the range between -1 and 1. This step is to avoid the difference between different dimensions affecting the subsequent calculation.
[0079] The weight matrix used for linear transformation is a 512x512 square matrix, and its initial value is generated by the Xavier initialization method. According to the dimensions of the input and output (both are 512), the appropriate value range is calculated. Specifically, the initial value of each element will fall between about -0.075 and 0.075. The purpose is to keep the overall numerical distribution stable after the input features are transformed, and there will be no large or small fluctuations.
[0080] During calculation, the 512-dimensional input vector will be matched with the weight matrix: each element in the input vector will be associated with all elements in the corresponding column of the weight matrix to obtain the numerical value of each dimension in the output vector. Since the range of input elements is -1 to 1 and the range of weight matrix elements is -0.075 to 0.075, after calculation, the numerical value of each element in the output image feature vector (still 512 dimensions) is approximately between -38.4 and 38.4.
[0081] The contextual semantic vector as input comes from the processing result of the Transformer large model encoder, with a total of 512 dimensions. Each dimension of the numerical value is processed by layer normalization to stabilize the range between -1 and 1.
[0082] The affine transformation includes two parts: linear transformation and bias adjustment; the weight matrix used in linear transformation is the same as the linear transformation of the spatial structure feature vector, which is also a 512x512 square matrix, initialized by Xavier, with elements ranging from -0.075 to 0.075; the bias term is a 512-dimensional vector, and the initial value of each element is generated by uniform distribution, limited between -0.1 and 0.1.
[0083] During calculation, linear transformation is performed on the input context semantic vector (consistent with the linear transformation of the spatial structure feature), obtaining an intermediate vector; each element of the intermediate vector is added to the element at the corresponding position in the bias term, and finally a 512-dimensional text semantic vector is generated; since the element range of the intermediate vector after linear transformation is about -38.4 to 38.4, after superimposing the bias term, the range of each element of the text semantic vector is about -38.5 to 38.5.
[0084] The structure descriptor matrix as input is an N x N square matrix (N represents the number of elements in the document, usually between 50 and 200), and the elements in the matrix contain two types of information: one is node attribute (such as the type, position, and length of elements such as title and paragraph in the document), where the type is represented by 0 or 1, and the position and length are normalized to range from 0 to 1; the second is the edge weight (representing the association strength between elements), which also ranges from 0 to 1; therefore, all elements in the entire matrix are stable within the range of 0 to 1.
[0085] Vector projection is divided into two steps: the first step is "row concatenation vectorization", which expands the N x N matrix row by row and concatenates it into a one-dimensional vector with a length of N squared (for example, when N = 100, a vector with a length of 10000 is generated); the second step is "linear transformation compression", which uses a weight matrix with a dimension of "N squared x 512" to compress the long vector into 512 dimensions.
[0086] The initial value of the weight matrix is also generated by Xavier initialization, and the specific range is related to the size of N: when N = 100, the range of the weight matrix elements is about -0.0237 to 0.0237; when N = 200, the range is about -0.0122 to 0.0122 (the larger N is, the smaller the weight range is to avoid excessive compression values); during calculation, the concatenated long vector is matched with this weight matrix to finally generate a 512-dimensional structure description vector; the range of each element is related to N: when N = 100, it is about -237 to 237; when N = 200, it is about -488 to 488.
[0087] The image features and the text semantics are aggregated based on weights (the sum of the weights is 1), and are spliced with the structure description vector after layer normalization to generate a 512*3=1536-dimensional joint feature tensor.
[0088] Through the improved ResNet and Transformer models, combined with industry data fine-tuning, the fine-grained capture of document spatial structure and context semantics is realized, the structure descriptor matrix accurately describes the logical level of the document, the dimensions are unified through feature space mapping, and the cross-modal attention mechanism based on the structure features enhances the relevance of visual, semantic and structure features, solves the problem of the "semantic gap" of multi-modal features, and makes the joint feature tensor more comprehensively reflect the comprehensive properties of the document.
[0089] In a preferred embodiment of the application, based on the latest rule set loaded based on the formal rule base and the joint feature tensor, the logical satisfaction degree between the entities and the rule predicates in the feature tensor is calculated through the semantic alignment module; the core verification operation is performed by using the differentiable logic layer to obtain an abnormal set and a corresponding decision path node sequence, including:
[0090] Based on the latest rule set loaded based on the formal rule base and the joint feature tensor, the rule is parsed into a structured predicate unit and the entity segment is located, and a three-level progressive alignment operation of accurate name matching, semantic similarity matching and context reasoning matching is sequentially performed to generate an entity-rule predicate binding pair set;
[0091] According to each binding pair in the entity-rule predicate binding pair set, the satisfaction degree calculation operation based on differentiable comparison is performed according to the types of numerical rules, relational rules and composite rules to generate a logical satisfaction degree set;
[0092] Based on the logical satisfaction degree set, the candidate abnormal entities are marked by a dynamic judgment threshold, the abnormal list of the rule violation evidence and the decision path node sequence chain is generated by backtracking, and the abnormal set is obtained.
[0093] A CFG-based rule parser is used to parse the latest rule set (such as "the contract amount shall not exceed 5 million yuan" and "the signature of party A shall be consistent with the recorded name") in the formal rule base into a structured predicate unit, each unit contains three elements of "subject (such as contract amount) + predicate (such as ≤) + object (such as 5 million yuan)", and is stored as a {subject, predicate, object} triple (the parsing accuracy is ≥98%).
[0094] Entity Positioning: Extract entity segments (e.g., "contract amount = 600 million" "party A signature = XX company") from the joint feature tensor using a pre-trained entity recognition model (based on the BERT-base architecture). The model training data is industry documents annotated with "amount, subject, signature" entities (100,000+ samples). Training parameters: learning rate 0.00002, batch size 16, training rounds 15, entity recognition F1 value ≥ 0.92.
[0095] String exact match between entity segments and rule subjects (e.g., "contract amount" and "contract amount" match completely), and if the match is successful, a binding pair is generated. The matching success rate is about 60%-70% (for standardized entities).
[0096] For entity-predicate pairs that do not match exactly, use the Sentence-BERT model to calculate semantic similarity (the model is based on pre-trained all-MiniLM-L6-v2 fine-tuning, and the training data is industry entity-predicate synonym pairs (50,000+), fine-tuning learning rate 0.00001, batch size 32, training rounds 10). The similarity threshold is set to 0.7 (range 0-1), and if it exceeds the threshold, a binding pair is generated, and the matching success rate is supplemented to 85%-90%.
[0097] Context reasoning matching: For cases that do not match in the first two stages, use a GPT-2-based context reasoning model (fine-tuning dataset contains context samples containing entity-predicate associations (30,000+), learning rate 0.00005, batch size 16, training rounds 8) to infer the implicit association between entities and predicates (e.g., "total price of the agreement" and "contract amount" have an implicit equivalent relationship). If the reasoning confidence is ≥ 0.85 (range 0-1), a binding pair is generated, and the overall binding success rate is ≥ 95%.
[0098] The rules in the binding pair set are divided into three categories by the rule parser:
[0099] Numerical rules: involving quantitative comparison (e.g., "amount ≤ 500 million" "validity period ≥ 3 years"), accounting for about 40%-50%; relationship rules: involving entity relationship constraints (e.g., "party A = payer" "signature date ≤ contract effective date"), accounting for about 30%-40%; composite rules: containing multiple conditions nested (e.g., "amount > 100 million and requires approval by general manager"), accounting for about 10%-20%.
[0100] Differentiable logic layer construction and training:
[0101] Model structure: A 3-layer MLP (Multi-Layer Perceptron) is used as the core of the differentiable logic layer. The input is a concatenated vector of entity features and predicate features (dimension 1024). The hidden layer dimensions are 512 and 256. The output layer is a single neuron (output range 0-1, representing satisfaction: 1 for complete satisfaction, 0 for complete dissatisfaction).
[0102] Training data: Entity-rule pairs labeled with "satisfy / don't satisfy" labels (100,000+), including 50,000+ numerical, 30,000+ relational, and 20,000+ compound rules.
[0103] Training parameters: The optimizer is Adam (learning rate 0.0003), the loss function is binary cross-entropy, the training rounds are 20, the batch size is 64, and the model's satisfaction prediction accuracy on the validation set is ≥0.93.
[0104] Binary cross-entropy as the loss function of the differentiable logic layer, is used to quantify the difference between the model's predicted logical satisfaction and the actual label. The specific calculation process is as follows:
[0105] Label definition: For each entity-rule predicate binding pair in the training data, a binary label is manually labeled (1 represents "satisfy the rule", 0 represents "don't satisfy the rule"), forming a real label set.
[0106] Prediction value range: The output of the differentiable logic layer is a numerical value (range 0-1), representing the model's predicted probability of "satisfying the rule" for the binding pair (i.e., logical satisfaction).
[0107] For a single binding pair, if the real label is 1 (satisfy the rule), the loss component of this sample is calculated by "-log(predicted satisfaction)" (the closer the predicted satisfaction is to 1, the closer the loss is to 0); if the real label is 0 (don't satisfy the rule), the loss component is calculated by "-log(1-predicted satisfaction)" (the closer the predicted satisfaction is to 0, the closer the loss is to 0); the average of all loss components for the batch training data (batch size is 64 in this step) is taken to get the binary cross-entropy loss value of this batch (all numerical values in the range of 0 and above, the smaller the value, the more consistent the predicted result is with the real label).
[0108] Training application: During the model training process, the weight parameters of the differentiable logic layer are continuously adjusted through the backpropagation algorithm, so that the binary cross-entropy loss value gradually decreases (the validation set loss value stabilizes at 0.15 or below at the end of training), thereby improving the model's prediction accuracy of logical satisfaction.
[0109] Classification calculation of logical satisfaction:
[0110] Numerical rule: Calculate the difference between entity value (e.g., 6 million) and rule object (e.g., 5 million) through a differentiable comparator (embedded in the first layer of MLP), and map it to the satisfaction degree (e.g., the satisfaction degree of 6 million to 5 million is 0.3).
[0111] Input the relationship features between entities (e.g., date difference, name similarity) into the intermediate layer of MLP, and output the satisfaction degree of relationship matching (e.g., the satisfaction degree of the signature date being later than the effective date is 0.2).
[0112] Perform differentiable logical operations (e.g., "and" operation takes the minimum value, "or" operation takes the maximum value) on multiple condition satisfaction degrees, and output the comprehensive satisfaction degree (e.g., condition A satisfaction degree 0.9, condition B satisfaction degree 0.3, "and" operation result is 0.3).
[0113] Dynamic decision threshold: The initial threshold is set to 0.6 (range 0-1), and is dynamically adjusted based on historical detection data (e.g., increase to 0.7 when false positive rate is too high, decrease to 0.5 when false negative rate is too high), entities with current logical satisfaction degree lower than the threshold are marked as candidate anomalies.
[0114] For candidate anomaly entities, backtrack the three-level alignment process (record matching level: accurate / semantic / reasoning), rule type (numerical / relationship / complex), and intermediate calculation results of differentiable logic layer (e.g., any condition not met), generate decision path node sequence chain (e.g., "entity 'agreement total price = 6 million' → semantic matching rule 'contract amount ≤ 5 million' → numerical comparison not met → satisfaction degree 0.3 < threshold 0.6").
[0115] Integrate candidate anomaly entities and decision paths to generate an anomaly set containing entity ID, violation evidence (e.g., specific numerical / relationship), and path sequence, with anomaly labeling accuracy ≥ 0.91.
[0116] Three-level progressive alignment combined with accurate matching, semantic similarity, and context reasoning solves the "name heterogeneity" problem between entities and rule predicates, improving the binding success rate. The differentiable logic layer converts discrete rule judgment into continuous numerical calculation, supporting gradient backpropagation (providing a basis for parameter optimization in step 4), while retaining the rigor of rule logic. The decision path node sequence chain records the whole process of anomaly judgment (matching method, rule type, and reason for not meeting), supports classification processing of numerical, relationship, and complex rules, and adapts to complex and variable compliance rules in finance, law, and other fields.
[0117] For example, Figure 2As shown, based on the latest rule set loaded in the formal rule base and the joint feature tensor, through the analysis of the rules for structured predicate units and the positioning of entity fragments, the three-level progressive alignment operations of accurate name matching, semantic similarity matching and context reasoning matching are sequentially performed to generate the entity-rule predicate binding pair set, including:
[0118] Load the latest rule set in the formal rule base, parse each rule into a structured predicate unit, generate a rule subject name set; and perform entity recognition operation on the joint feature tensor to extract key entity fragments and generate an entity name set;
[0119] If the entity name and the rule subject name are completely consistent, the entity and the corresponding predicate unit are directly bound; the primary binding pair is generated; and the unmatched entity and rule subject name generate a candidate set;
[0120] Based on the candidate set, the semantic binding pair is generated by the cosine similarity of the semantic vector, and the inference set is generated; based on the inference set, the reasoning binding pair is generated by the micro-inference module according to the context semantics and structural descriptors of the joint feature tensor; the binding pairs of the three levels of matching are integrated, the predicates are selected according to the priority of accurate> semantic> reasoning, and the entity-rule predicate binding pair set is generated.
[0121] In the embodiment of the present application, the latest rule set in the formal rule base (such as "the contract amount shall not exceed 5 million yuan" and "the signature of party A shall be consistent with the recorded name", and the number of rules is usually 50-200).
[0122] The rule parser based on CFG (context-free grammar) includes preset syntax rules (such as "rule = subject + predicate + object" and "subject = noun phrase").
[0123] Process: parse rules one by one, extract "subject name" (such as "contract amount" and "party A signature"), remove modifiers (such as the numerical part in "shall not exceed 500 million yuan"), generate a rule subject name set (format is a string list, such as ["contract amount", "party A signature",...]), and the parsing accuracy is ≥98%.
[0124] Joint feature tensor (1536 dimensions, integrating image, text and structural features); entity recognition model: named entity recognition (NER) model based on BERT-base architecture, adding CRF (conditional random field) layer to optimize entity boundary prediction.
[0125] Model construction: the input layer receives the text sequence (converted from the text semantic vector of the joint feature tensor), the 12-layer Transformer coding (hidden layer dimension 768), and the output layer is the entity label (such as "amount entity", "subject entity", etc., a total of 10 categories).
[0126] Training process:
[0127] Dataset: Industry document texts labeled with entity types (100,000+ samples, each sample containing 1-5 entities).
[0128] Parameters: Optimizer is Adam (learning rate 2e-5), loss function is CRF negative log-likelihood, training epochs are 15, batch size is 16, and validation set F1 score is ≥0.93.
[0129] Implementation process: Extract text fragments from the joint feature tensor, input them into the model, and output entity names (such as "total price of the agreement" and "signature of Party B"). Remove redundant characters (such as punctuation and spaces) to generate a set of entity names (a list of strings, such as ["total price of the agreement", "signature of Party B",...]). The recognition accuracy is ≥92%.
[0130] The matching logic is as follows: the entity name set and the rule subject name set are compared sentence by sentence to see if the strings of the entity name and the rule subject name are completely identical (case-sensitive, such as "contract amount" matches "contract amount", but does not match "total contract amount").
[0131] Primary binding pair: The format is (entity name, rule body name, match level = "exact"), such as ("contract amount", "contract amount", "exact").
[0132] Candidate set: Unmatched entity names (such as "total price of agreement") and unmatched rule subject names (such as "contract amount") are divided into two sublists, which are usually 30%-50% of the initial set.
[0133] Model: Sentence-BERT, used to convert text into fixed-dimensional semantic vectors (384 dimensions).
[0134] Fine-tuning process:
[0135] Dataset: Manually labeled "entity-rule subject synonym pairs" (50,000+, such as "total agreement price" and "contract amount" are synonyms); learning rate 1e-5, 10 training rounds, batch size 32, vector similarity is optimized by contrastive loss; implementation process: input the entity names and rule subject names in the candidate set into the model respectively to generate a 384-dimensional semantic vector (vector element range [-1, 1]).
[0136] Similarity range: 0-1 (0 means completely unrelated, 1 means completely identical in meaning); Matching threshold: set to 0.7 (optimized with validation set, the accuracy is ≥0.85 at this time).
[0137] Process: Calculate the similarity of each entity name vector in the candidate set with the rule subject name vector, keep the pairs with similarity ≥ 0.7, generate semantic binding pairs (format: (entity name, rule subject name, matching level = "semantic", similarity value), such as ("agreement total price", "contract amount", "semantic", 0.82)), and the scale is 20%-30% of the candidate set.
[0138] Model: Small sample inference model based on GPT-2, input context information and matched entities and rule subjects, output correlation probability.
[0139] Model construction: 12-layer Transformer decoder (hidden layer dimension 768), input is "context + entity + rule subject" text sequence, output layer is binary classification (associated / not associated).
[0140] Training process:
[0141] Entity-rule subject association samples containing context (30,000+, such as "context: 'agreement total price is the total amount agreed in the contract' + entity: 'agreement total price' + rule subject: 'contract amount' → associated").
[0142] Parameters: Optimizer AdamW (learning rate 5e-5, weight decay 0.01), training rounds 8, batch size 16, validation set accuracy ≥ 0.88.
[0143] Inference matching process: entities and rule subjects in the inference set, combined with the context semantics of the joint feature tensor (such as sentence association in the document) and the structure descriptor (such as the position of the entity in the chapter); the model generates the confidence of "whether the entity and rule subject are associated" (range 0-1), and the threshold is set to 0.85 (to ensure accuracy ≥ 0.9); inference binding pairs (format: (entity name, rule subject name, matching level = "inference", confidence), such as ("project total cost", "contract amount", "inference", 0.89)).
[0144] Accurate matching (first level) > semantic matching (second level) > inference matching (third level), that is, the same entity appears in multiple levels of matching, only the high-level binding pair is retained; remove duplicate binding pairs (entity and rule subject are exactly the same), keep the high-confidence items (such as the pair with higher similarity in semantic matching).
[0145] Output: Entity-rule predicate binding pair set, format is a list, containing all binding pairs (such as [("contract amount", "contract amount", "accurate"), ("agreement total price", "contract amount", "semantic"),...]), the overall matching success rate ≥ 95%.
[0146] The three-stage progressive matching covers scenarios from precise to implicit association, solves the matching problem of "name isomorphism but semantic equivalence", and improves the success rate compared with single matching method. Preferentially, precise matching and semantic matching are used to process most scenarios at low cost, and reasoning matching is only enabled for complex cases, which improves the overall accuracy and reduces the calculation time.
[0147] In a preferred embodiment of the present application, the risk feature vector is generated, comprising:
[0148] Based on the structured anomaly set, the original document logical relationship adjacency matrix is inherited as the graph node of the anomaly entity, the cross-entity association edge is added through the rule violation evidence, and the decision path node sequence chain is integrated to generate a topologically enhanced anomaly association matrix;
[0149] Based on the topologically enhanced anomaly association matrix, a pre-trained graph neural network is used to perform multi-layer graph convolution aggregation and graph attention calculation to generate a global graph representation vector, and a cosine similarity between the global graph representation vector and a compliance document benchmark graph vector is calculated to generate a structure deviation scalar value; based on the decision path node sequence chain, a pre-trained language model is used to perform semantic segmentation, vectorization and self-attention conflict weight calculation on the path segment converted into natural language description, to extract the highest conflict node type identifier, rule predicate logical contradiction frequency and language model conflict probability value, and generate a multi-dimensional semantic conflict feature vector;
[0150] Based on the structure deviation scalar value and the multi-dimensional semantic conflict feature vector, a risk feature vector is generated through numerical normalization processing, feature splicing and double-layer full connection layer fusion.
[0151] In an embodiment of the present application, the input basis includes a structured anomaly set (containing anomaly entity ID, type, violation evidence, etc.) and a logical relationship adjacency matrix of the original document (recording the hierarchical and parallel relationships between document elements).
[0152] Matrix construction:
[0153] Node definition: the anomaly entity is used as the graph node, and the node attributes include entity type (such as "contract amount", "seal", etc., identified by an integer, ranging from 1 to 10, corresponding to different entity categories) and violation severity (based on rule satisfaction degree conversion, ranging from 0 to 1, 0 for minor violation and 1 for serious violation).
[0154] The edges between the anomaly entities in the original adjacency matrix (such as the association edge between "contract amount" and "payment clause") are retained, and the edge weight follows the original association strength (ranging from 0 to 1).
[0155] Add new edges between entities with common violation logic (e.g., "excessive amount" and "missing approval" are associated by the "large amount needs approval" rule) with edge weights representing the evidence relevance (calculated based on semantic similarity, ranging from 0.3 to 0.9, with higher values indicating closer association).
[0156] Integrate decision path: Convert the sequential relationship of entities in the decision path node sequence chain into directed edges (e.g., the path "Entity A -> Entity B -> Violation Conclusion" corresponds to a directed edge from A to B), with edge weights being the inverse of the path step number (ranging from 0.1 to 1, with fewer steps resulting in higher weights).
[0157] Finally, form an N x N matrix (N is the number of abnormal entities, usually 10-50), with matrix elements being edge weights (ranging from 0 to 1, with 0 indicating no association and 1 indicating strong association), and the diagonal being the node self-loop weight (fixed at 1, representing the entity's own attributes).
[0158] Graph neural network construction and training:
[0159] Model structure: Use a 6-layer graph attention network (GAT) with 2 attention heads (each with dimension 64) per layer, input node attribute vector (dimension 32), hidden layer dimensions 128, 64, output layer dimension 512 (global graph representation vector), and residual connection and layer normalization (ε = 0.00001) for each layer.
[0160] Training data: A graph structure dataset containing 100,000+ samples, of which 50,000+ are "normal topology graphs" of compliance documents (adjacency matrix and corresponding label "compliant"), and 50,000+ are abnormal document graphs labeled with "structural deviation" (such as hierarchical disorder, missing entity association, etc.).
[0161] Training parameters: Optimizer is Adam (learning rate 0.0002), loss function is contrastive loss (pulling similar graphs closer and pulling dissimilar graphs farther apart), training rounds 30, batch size 32, and the model's topology classification accuracy on the validation set ≥0.92.
[0162] Structural deviation degree calculation process:
[0163] Input the topology-enhanced abnormal association matrix into the pre-trained GAT, and output a 512-dimensional global graph representation vector through multi-layer graph convolution (aggregating neighbor node features) and graph attention calculation (assigning high weights to important associated edges, with attention weights ranging from 0 to 1).
[0164] Load the baseline graph vector of the compliant document (generated by processing a large number of compliant documents through the same GAT), and calculate the cosine similarity between the abnormal graph representation vector and the baseline vector (ranging from -1 to 1, with closer to 1 indicating more similar structure).
[0165] Deviation generation: structural deviation = 1 - cosine similarity (range 0-1, 0 means no deviation, 1 means complete deviation; usually the deviation of abnormal samples ≥0.3).
[0166] Pre-training of language model:
[0167] Model base: 12-layer RoBERTa pre-training model (vocabulary size 50000) is used, and a conflict recognition head (output layer dimension 3, corresponding to three conflict dimensions) is added.
[0168] Training data: 30,000+ decision path samples labeled with "semantic conflict", each sample contains a decision path node sequence chain (such as "entity A → violates rule B → entity C conflict") and corresponding conflict dimension label (node type, contradiction frequency, conflict probability).
[0169] Training parameters: AdamW optimizer (learning rate 0.00005, weight decay 0.01), loss function is joint loss (classification loss + regression loss), training rounds 15, batch size 16, conflict feature extraction accuracy of the model on the validation set ≥0.91.
[0170] Feature extraction process:
[0171] Convert the decision path node sequence chain into a natural language description (such as "contract amount entity (type 3) violates the 'not more than 5 million' rule, and there is a logical contradiction with the approval entity (type 7), conflict frequency 2 times").
[0172] Input RoBERTa model, perform semantic segmentation (split sentences into phrases, accuracy ≥96%), vectorization (generate 768-dimensional phrase vectors), and self-attention conflict weight calculation (assign high weights to conflict keywords, range 0-1).
[0173] Three-dimensional feature extraction:
[0174] Extract the type code of the conflict-involved entity (integer, range 1-10, corresponding to 10 core entities); count the number of occurrences of keywords such as "violate" and "conflict" in the path (non-negative integer, range 0-10, 0 means no contradiction); conflict confidence output by the model (range 0-1, ≥0.5 means there is a significant conflict); combine the three types of features into a 3-dimensional semantic conflict feature vector (such as [3, 2, 0.85]).
[0175] Structural deviation: already in the range of 0-1, no additional processing is needed;
[0176] Semantic conflict features: node type identifier is normalized by min-max (mapped to 0-1), logical contradiction frequency is normalized by logarithm (avoid large values dominate, mapped to 0-1), conflict probability value remains in the range of 0-1.
[0177] When performing min-max normalization on node type identifier, the specific process is as follows: determine the original value range of node type identifier. This kind of identifier is used to distinguish the integer value of abnormal entity category (such as "contract amount" corresponding to 1, "seal" corresponding to 2, "date" corresponding to 3, etc.), and its value range is fixed to 1 to 10 (covering all common entity types) in actual processing, that is, the minimum value is 1 and the maximum value is 10; for each specific node type identifier value (such as the type identifier of a certain abnormal entity is 5), it is converted to the range of 0 to 1 by the following logic: first calculate the difference between the value and the minimum value (such as the difference between 5 and 1 is 4), then calculate the proportion of this difference and the total difference between the maximum value and the minimum value (the difference between 10 and 1 is 9), this proportion is the normalized result (such as the proportion of 4 divided by 9 is about 0.44); in this way, the node type identifier of the original range 1 to 10 will be uniformly mapped to the range of 0 to 1: the entity with identifier 1 corresponds to 0, the entity with identifier 10 corresponds to 1, and the intermediate identifiers (such as 3, 7, etc.) are proportionally distributed between 0 and 1 (such as 3 corresponds to about 0.22, 7 corresponds to about 0.67).
[0178] The structure deviation (scalar, 0-1) is spliced with the normalized 3-dimensional semantic conflict feature vector to form a 4-dimensional intermediate vector.
[0179] Fully connected layer fusion: the first fully connected layer: input 4-dimensional vector, hidden layer dimension 64, (output value range is all values greater than or equal to 0, normalized to 0-1 by layer); the second fully connected layer: input 64-dimensional vector, output 32-dimensional risk feature vector (each element range 0-1).
[0180] By capturing the topological association deviation of abnormal entities through graph neural networks and analyzing semantic conflicts through language models, the limitations of single structure or semantic analysis are broken through, and the collaborative detection of "structural abnormalities-semantic contradictions" is realized. The topological enhancement matrix preserves the association logic of abnormal entities, and the pre-trained model is fine-tuned on industry data, which effectively improves the recognition accuracy of entity types and rule conflicts in the fields of finance and law. Both graph neural networks and language models support incremental training, which can continuously optimize new types of abnormalities (such as new entities and unknown rule conflicts).
[0181] In a preferred embodiment of the present application, the augmented rule set is generated based on the risk feature vector, the formal rule library is updated, the feature encoding component parameters are updated through gradient back propagation, the threshold of the differentiable logic layer and the weight of the analysis module are optimized, and the updated feature encoding component, rule library version and risk quantification parameters are obtained, including:
[0182] Based on the risk feature vector, an augmented rule draft set containing confidence scores is generated, and if the sample verification false positive rate decreases and the recall rate improves, an effective augmented rule set is generated, and rule redundancy processing is performed according to the confidence score, and finally a versioned updated formal rule library is obtained;
[0183] Based on the risk feature vector, the gradient is calculated and the chain back propagation is processed, the parameters of the fusion layer, the graph network, the cross-modal module and the encoder in the feature encoding component chain are updated, and the optimized feature encoding component is output;
[0184] Based on the feature encoding component and the satisfaction set, the dynamic decision threshold boundary is adjusted through the satisfaction probability density distribution, and the offset is fine-tuned combined with the false positive rate index, the edge weight coefficient of the graph neural network is recalibrated based on the updated component, and the semantic conflict feature dimension weight is redistributed, to obtain the updated decision threshold and analysis module weight;
[0185] The versioned rule library, the optimized feature encoding component and the decision threshold and analysis module weight are fused, and the rule library change summary is packaged, the component parameter snapshot is stored, and the risk quantification parameter set is recorded.
[0186] In an embodiment of the present application, based on a 32-dimensional risk feature vector (each element ranges from 0 to 1, and the higher the value, the more significant the risk), the abnormal mode corresponding to the high-value element in the vector (such as "node type 3 + conflict frequency 2 + deviation 0.8" corresponding to "large contract without approval") is extracted, a natural language rule draft (such as "if the contract amount > 5 million and there is no general manager's signature, then it is determined as a violation") is generated combined with the rule template, and a confidence score is calculated for each draft (based on the mean value of the risk feature vector, ranging from 0 to 1, ≥0.7 is considered as high confidence).
[0187] Sample verification and effectiveness screening: apply the draft rule set to the verification sample pool (containing 10,000+ historical documents), and calculate the false positive rate (target ≤5%) and the recall rate (target ≥90%); only keep the drafts with "false positive rate decrease ≥3% and recall rate increase ≥5%", and form an effective augmented rule set.
[0188] Calculate the semantic similarity between the effective rules (based on Sentence-BERT, threshold ≥0.8 is considered as repetition), keep the high-confidence rules (such as keeping the confidence of 0.85 in two similar rules, and discarding the confidence of 0.72), and finally control the number of newly added rules to 10%-20% of the original rule set.
[0189] Add a timestamp (e.g., "V20250721") and change notes (X new, Y deleted) to the updated rule library; store in JSON format, including rule ID, content, confidence, and effective date, supporting backtracking to historical versions (retain the last 10 versions).
[0190] Feature encoding component composition: includes cross-modal fusion layer, graph neural network, cross-modal attention module, Transformer encoder, and convolutional neural network, with a total parameter size of approximately 5 million.
[0191] Gradient backpropagation process:
[0192] The difference between the risk feature vector and the "ideal risk distribution" (the vector of labeled compliance / non-compliance label conversion) is used as the loss value (range 0-1, the larger the value, the greater the deviation).
[0193] The loss signal is sequentially passed back to the double-layer fully connected fusion layer, semantic conflict feature extraction module, graph neural network, cross-modal attention module, Transformer encoder / convolutional neural network, and each layer calculates the parameter gradient (representing the impact of parameters on loss, range -0.1-0.1).
[0194] Adam optimizer (learning rate 0.00001, weight decay 0.001) is used to adjust parameters in the gradient direction, with a single update amplitude not exceeding 5% of the original parameter value (to avoid drastic fluctuations), with a focus on optimizing the weights of the cross-modal attention module (accounting for 40% of the total update) and the convolution kernel parameters of the graph neural network (accounting for 30%).
[0195] After 10 rounds of iterative updates, the feature matching accuracy of the feature encoding component improved by 6-8% (from 89% to over 95% on the validation set), and the consistency of cross-modal fused features (cosine similarity) increased by 10%.
[0196] Dynamic decision threshold adjustment:
[0197] Satisfaction distribution analysis: Calculate the probability density distribution of the historical logic satisfaction set (e.g., compliance sample satisfaction is concentrated in 0.7-1, and abnormal sample satisfaction is concentrated in 0-0.5), and calculate the lower bound of the 95% confidence interval as the initial threshold (usually 0.6-0.7).
[0198] False positive rate fine-tuning: If the current false positive rate (on the validation set) exceeds 5% (preset upper limit), generate an offset (0.05) by multiplying the excess proportion (e.g., false positive rate 6% exceeds 1%) by 0.05 (fixed coefficient), and increase the threshold to 0.65-0.75; if the false positive rate is less than 3%, reduce the threshold by 0.03-0.05 (to improve recall rate).
[0199] Based on the updated feature encoding component, the edge weight (range 0-1) of the anomaly correlation matrix is recalculated, and the high contribution correlation edges (such as "amount over standard - approval missing") in the risk feature are increased by 10%-20% of the weight.
[0200] According to the contribution degree (calculated by variance analysis, range 0-1) of each semantic dimension (node type / contradiction frequency / conflict probability) in the risk feature vector, the dimension weight (sum is 1) is adjusted, such as the conflict probability dimension contribution degree is high, the weight is increased from 0.3 to 0.4.
[0201] Record 10-20 new rules, delete 3-5 duplicate / low confidence rules (confidence <0.6), and generate a structured change log (including rule ID, change type, and reason).
[0202] Component parameter snapshot: store the weight matrix of each layer of the feature encoding component (focus on saving the cross-modal layer and the graph neural network layer), and use quantization storage (precision 16 bits) to reduce the storage space (compressed to 50% of the original size).
[0203] Including dynamic decision threshold (0.6-0.8), graph neural network edge weight range (0-1), and semantic conflict dimension weight (such as [0.3, 0.2, 0.5]), stored in JSON format, and supports real-time calling.
[0204] Through risk feature feedback rule library update and parameter optimization, the system can automatically adjust to new risk patterns (such as new types of violations), without the need for manual retraining, and the adaptation efficiency is improved. Feature encoding component parameter optimization reduces feature extraction bias, dynamic threshold and weight adjustment balances false positive rate and recall rate, and rule library redundancy avoidance avoids repeated judgment.
[0205] The embodiment of the application also provides a computing device, comprising a processor and a memory storing a computer program, wherein the computer program is run by the processor to execute the system as described above. All implementation manners in the above system embodiment are applicable to this embodiment, and the same technical effects can also be achieved.
[0206] The embodiment of the application also provides a computer readable storage medium storing instructions, which, when executed on a computer, cause the computer to execute the system as described above. All implementation manners in the above system embodiment are applicable to this embodiment, and the same technical effects can also be achieved.
[0207] The above is the preferred embodiment of the application. It should be noted that for those skilled in the art, without departing from the principles of the application, a number of improvements and refinements can be made, which should also be considered within the scope of protection of the application.
Claims
1. A deep learning based computer document intelligent compliance detection system, characterized in that, The method comprises the following steps: a feature fusion module is used to input a document to be detected, extract a spatial structure feature vector of a document image through a convolutional neural network, obtain a text sequence by using an optical character recognition engine, generate a context semantic vector by using a Transformer large model encoder, and parse a document logical level to generate a structure descriptor; three types of feature vectors are subjected to a feature space mapping matching algorithm, and the mapped features are fused through a cross-modal attention mechanism to generate a joint feature tensor; a feature calculation module is used to load a latest rule set based on a formal rule library and calculate a logical satisfaction degree between entities and rule predicates in the joint feature tensor through a semantic alignment module; a differentiable logic layer is used to perform a core verification operation to obtain an abnormality set and a corresponding decision path node sequence; a structure analysis module is used to construct a relationship adjacency matrix by taking the abnormal entity set as a node index; a graph neural network is used to analyze the topological structure of the relationship adjacency matrix to detect abnormalities to obtain a structure deviation degree, and a pre-trained language model is used to parse the decision path node sequence to generate a semantic conflict feature; the structure deviation degree and the semantic conflict feature are fused to generate a risk feature vector; an optimization parameter module is used to generate an augmented rule set according to the risk feature vector and update the formal rule library; the parameters of the feature encoding components are updated through gradient backpropagation; the decision threshold of the differentiable logic layer and the weight of the analysis module are optimized; updated feature encoding components, rule library versions and risk quantification parameters are obtained, including: generating an augmented rule draft set containing a confidence score based on the risk feature vector, generating an effective augmented rule set if the sample verification false positive rate decreases and the recall rate increases, performing rule redundancy processing according to the confidence, and finally obtaining a versioned and updated formal rule library; based on the risk feature vector, the gradient is calculated and processed in chain reverse propagation, the parameters of the fusion layer, the graph network, the cross-modal module and the encoder in the feature encoding component chain are updated, and the optimized feature encoding components are output; based on the feature encoding components and the logical satisfaction degree set, the dynamic decision threshold boundary is adjusted through the probability density distribution of the satisfaction degree, and the offset is fine-tuned combined with the false positive rate index; the graph neural network edge weight coefficient is recalibrated based on the updated components, and the semantic conflict feature dimension weight is redistributed, to obtain the updated decision threshold and analysis module weight; the versioned rule library, the optimized feature encoding components and the decision threshold and analysis module weight are fused, and the rule library change summary is encapsulated, the component parameter snapshot is stored, and the risk quantification parameter set is recorded.
2. The deep learning based computer document intelligent compliance detection system of claim 1, wherein, The joint feature tensor is generated, including: a pre-trained convolutional neural network is used to input a document image to be detected, multi-scale spatial features are extracted through the alternating processing of convolutional layers and pooling layers, and a spatial structure feature vector is obtained through a full connection layer; a document text sequence is extracted through an optical character recognition engine, and after sub-word segmentation and word vector conversion, the document text sequence is input into a Transformer encoder for multi-layer self-attention calculation and feature stabilization processing, and finally a context semantic vector is generated; a document logical topology is organized based on document layout analysis, and a structure descriptor matrix is obtained through calculation; The spatial structure feature vector, the context semantic vector, and the structure descriptor matrix are respectively subjected to linear transformation, affine transformation, and vector projection to generate image feature vectors, text semantic vectors, and structure description vectors with the same dimension after projection; Taking the structure description vector as the query core, the cross-modal attention mechanism is used to calculate the association weight of the structure description vector with the image and text features, and the visual and semantic information is weighted and aggregated to finally generate a joint feature tensor.
3. The deep learning based computer document intelligent compliance detection system of claim 2, wherein, Based on the latest rule set loaded from the formal rule base and the joint feature tensor, the semantic alignment module is used to calculate the logical satisfaction degree between the entities and the rule predicates in the joint feature tensor; The core verification operation is performed by using the differentiable logic layer to obtain an abnormality set and a corresponding decision path node sequence, including: Based on the latest rule set loaded from the formal rule base and the joint feature tensor, the rule is parsed into a structured predicate unit and the entity fragment is located, and a three-stage progressive alignment operation of accurate name matching, semantic similarity matching, and context reasoning matching is sequentially performed to generate an entity-rule predicate binding pair set. According to each binding pair in the entity-rule predicate binding pair set, the satisfaction degree calculation operation based on the differentiable comparison is performed according to the types of numerical rules, relational rules, and composite rules to generate a logical satisfaction degree set. Based on the logical satisfaction degree set, the candidate abnormal entities are marked by a dynamic judgment threshold, the abnormal list of the rule violation evidence and the decision path node sequence chain is generated by backtracking, and the abnormality set is obtained.
4. The deep learning based computer document intelligent compliance detection system of claim 3, wherein, The risk feature vector is generated, including: Based on the structured abnormality set, the original document logical relationship adjacency matrix is inherited with the abnormal entity as the graph node, the cross-entity association edge is added through the rule violation evidence, and the decision path node sequence chain is integrated to generate a topologically enhanced abnormal association matrix; Based on the topologically enhanced abnormal association matrix, the global graph representation vector is generated by performing multi-layer graph convolution aggregation and graph attention calculation through the pre-trained graph neural network, the cosine similarity between the global graph representation vector and the compliance document benchmark graph vector is calculated to generate a structure deviation scalar value; based on the decision path node sequence chain, the path fragment converted into natural language description is subjected to semantic segmentation, vectorization, and self-attention conflict weight calculation through the pre-trained language model to extract the highest conflict node type identifier, rule predicate logical contradiction frequency, and language model conflict probability value, and a multi-dimensional semantic conflict feature vector is generated; Based on the structure deviation scalar value and the multi-dimensional semantic conflict feature vector, the risk feature vector is generated by numerical normalization processing, feature splicing, and double-layer full connection layer fusion.
5. The deep learning based computer document intelligent compliance detection system of claim 4, wherein, Based on the latest rule set loaded from the formal rule base and the joint feature tensor, the rule is parsed into a structured predicate unit and the entity fragment is located, and a three-stage progressive alignment operation of accurate name matching, semantic similarity matching, and context reasoning matching is sequentially performed to generate an entity-rule predicate binding pair set, including: The latest rule set in the formal rule base is loaded, each rule is parsed into a structured predicate unit, a rule subject name set is generated; and the joint feature tensor is subjected to entity recognition to extract key entity fragments, and an entity name set is generated; If the entity name and the rule subject name are completely consistent, the entity is directly bound with the corresponding predicate unit; a primary binding pair is generated; unmatched entity names and rule subject names generate a candidate set; Based on the candidate set, a semantic binding pair is generated through the cosine similarity of the semantic vector, and a set to be inferred is generated; based on the set to be inferred, the reasoning binding pair is generated through the differentiable reasoning module according to the joint feature tensor context semantics and the structure descriptor; the binding pairs of the three levels of matching are integrated, the predicates are selected according to the priority of accuracy> semantics> reasoning, and the entity-rule predicate binding pair set is generated.
6. The deep learning based computer document intelligent compliance detection system of claim 5, wherein, Based on the feature encoding component and the logical satisfaction degree set, the dynamic judgment threshold boundary is adjusted through the satisfaction probability density distribution, and the offset is fine-tuned combined with the false positive rate index; the graph neural network edge weight coefficient is recalibrated based on the update component, and the semantic conflict feature dimension weight is redistributed, to obtain the updated judgment threshold and analysis module weight, including: Based on the logical satisfaction degree set, the satisfaction degree values are arranged in ascending order to generate an ordered satisfaction degree sequence, the satisfaction degree value corresponding to the preset quantile position is calculated to generate an initial boundary value, the false positive rate of the sample verification stage is loaded, and the modified boundary value is obtained after calculation; Based on the feature encoding component and the modified boundary value, a new node feature vector is obtained, the edge weight coefficient is generated through edge weight calculation and sparsification processing; based on the multi-dimensional semantic conflict feature vector and the edge weight coefficient, the final dimension weight vector is generated through conflict feature contribution degree statistics, weight base calculation and weight attenuation processing; Based on the modified boundary value, the edge weight coefficient and the final dimension weight vector, the dynamic judgment threshold of the differentiable logic layer, the adjacent weight matrix of the graph convolution layer and the weight tensor of the semantic conflict feature fusion layer are updated respectively to complete the integration of the analysis module parameter set.
7. The deep learning based computer document intelligent compliance detection system of claim 6, wherein, Based on the risk feature vector, an augmented rule draft set containing confidence scores is generated; if the sample verification false positive rate decreases and the recall rate improves, an effective augmented rule set is generated, and rule redundancy processing is performed according to the confidence order, finally obtaining a versioned updated formal rule library, including: In the risk feature vector, abnormal entity type identifiers, rule violation feature identifiers and context semantic fingerprints are extracted to generate parameterized rule expressions; and based on the risk vector and the historical violation frequency weighting, an augmented rule draft set with confidence scores is generated; The augmented rule draft set is verified through the calculation of the change amount of the recall rate and the false positive rate to obtain an effective rule subset, and adjacent rules are detected after confidence sorting to generate a deduplicated augmented rule set; According to the deduplicated augmented rule set, a change summary containing a new rule ID, a confidence score and an abnormal type identifier is generated through version number iteration to obtain a versioned updated rule library.
8. A computing device, comprising: comprising: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the system as claimed in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a program which, when executed by a processor, implements the system as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Text privacy policy compliance detection method and system
CN118626645A
Intelligent case generation method based on judicial knowledge domain graph
CN119443249A