Financial document classification method based on dynamic resource allocation and attention mechanism
Through dynamic resource allocation and attention mechanisms, the problems of resource waste and insufficient robustness in financial document classification are solved, and more efficient key information processing and more accurate classification results are achieved.
Patent Information
- Application Number
- CN202510465831.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing technology has problems such as wasting resources, insufficient key information processing, insufficient robustness and difficulty in capturing long-distance dependencies across tokens in the financial document classification task.
The dynamic resource allocation and attention mechanism method is adopted to generate hidden state sequences through time-step processing, filter key tokens, dynamically allocate computing resources, model global dependencies, and calculate the probability distribution of document types.
It significantly improves the ability to capture local semantics, improves the proportion of processing resources of key tokens, reduces the computing overhead of non-key tokens, and improves classification accuracy and model robustness.
Smart Images

Figure CN119988635A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of natural language processing and document intelligence technology, and specifically relates to a financial document classification method with dynamic resource allocation and attention mechanism. Background Art
[0002] With the rapid development of natural language processing (NLP) and deep learning technologies, the automated classification and information extraction of financial documents have become key requirements for the digital transformation of enterprises. Existing technologies usually use rule-based text matching, traditional machine learning models (such as SVM, random forest) or simple deep learning architectures (such as single-layer LSTM) to handle financial document classification tasks.
[0003] For example, classification is performed by extracting keywords (such as "reimbursement form" and "amount") or counting word frequencies, or directly predicting text vectors through a fully connected network. Existing technologies usually adopt a "one-size-fits-all" computing resource allocation strategy, treating all tokens (such as each word or character in the text) equally, resulting in an imbalance in the proportion of computing resource consumption between key information (such as amount, date) and redundant information (such as stop words), causing resource waste or insufficient processing of key information. For example, the traditional LSTM model performs full calculations on all tokens and cannot dynamically adjust resource allocation, resulting in inefficient processing of long texts. Existing technologies lack an active screening mechanism for key tokens, making it difficult to effectively identify information that is critical to classification tasks (such as document type identifiers "travel expenses" and "purchase orders"), resulting in classification relying on local features and insufficient robustness. Existing methods (such as single-layer LSTM) are difficult to capture long-distance dependencies across tokens (such as the association between the "amount" and "reimbursement" fields), resulting in the inability of the semantic representation vector to integrate global context information, affecting classification accuracy. Existing technologies usually determine the document type through fixed thresholds, lack dynamic analysis of probability distribution, are unable to quantify the model's confidence in the classification results, and have difficulty supporting anomaly detection and manual review trigger mechanisms for high-risk scenarios (such as large expenditure documents). Summary of the invention
[0004] The present application provides a financial document classification method with dynamic resource allocation and attention mechanism to solve one of the above-mentioned technical problems.
[0005] The technical solution adopted in this application is: The present application embodiment provides a financial document classification method with dynamic resource allocation and attention mechanism, including: Based on the digitized text sequence, a hidden state sequence containing local context information is generated by processing it time step by time; According to the hidden state sequence, key tokens are selected and computing resources are dynamically allocated; Modeling the global dependency relationship of the key tokens to generate a semantic representation vector; The probability distribution of each document type is calculated based on the semantic representation vector, and the corresponding document type is determined according to the probability distribution.
[0006] According to an embodiment of the present application, the hidden state sequence containing local context information is generated by processing the digitized text sequence step by step, specifically: The RNN encoder is implemented using a long short-term memory network or a gated recurrent unit, and processes the input digitized text sequence time step by time; The hidden state of each time step is calculated by jointly calculating the current input and the hidden state of the previous time step, which specifically includes: linearly combining the current input with the previous hidden state and processing it through a nonlinear activation function to generate the hidden state of the current time step.
[0007] According to one embodiment of the present application, the screening of the key token is based on the attention mechanism, which specifically includes: The hidden state of each time step is linearly transformed through a fully connected layer, and the intermediate vector is generated using the Tanh activation function; Then the attention weight is calculated through another fully connected layer, and finally the weight value of each Token is obtained through Softmax normalization; According to a preset threshold, the tokens with weight values higher than the preset threshold are screened out as key tokens.
[0008] According to one embodiment of the present application, the dynamically allocating computing resources includes: The ratio of subsequent computing resources is dynamically adjusted based on the ratio of the number of key tokens to the total number of tokens multiplied by the preset resource amplification factor; Assign embedding vectors of different dimensions to the hidden states corresponding to the key tokens, or improve the processing capabilities of key information by changing the structure of the computing layer.
[0009] According to one embodiment of the present application, the global dependency modeling is implemented by a multi-head self-attention mechanism, specifically including: Input the hidden state sequence of the key token into the Transformer encoder and generate the Query, Key, and Value matrices through linear transformation; Calculate the similarity score between Query and Key, obtain the attention weight through Softmax normalization, and generate the attention output by weighted sum of Value vectors; The attention outputs of multiple subspaces are calculated in parallel through a multi-head attention mechanism, and the results are concatenated and linearly transformed into a global semantic representation vector.
[0010] According to one embodiment of the present application, the global semantic representation vector is generated in the following manner: The output of the multi-head attention is residually connected to the original hidden state and the layer is normalized; The local and global features are further fused through a feedforward neural network, and finally a global semantic representation vector is output.
[0011] According to an embodiment of the present application, the probability distribution is implemented by a fully connected layer and a Softmax activation function, specifically including: The semantic representation vector is input into the fully connected layer for linear transformation, and the output dimension is a vector consistent with the number of document types; The vector is normalized through the Softmax function to generate the probability distribution of each document type.
[0012] According to one embodiment of the present application, before dynamically allocating computing resources, the method further includes: The digitized text sequence is preprocessed, including word segmentation, removal of stop words, standardization of amount and date format, and word vector encoding through a pre-trained model to generate the initial digitized text sequence.
[0013] A computer-readable storage medium stores a program, which implements the steps in the method when executed by a processor.
[0014] An electronic device comprises a memory, a processor and a program stored in the memory and executable on the processor, wherein the processor implements the steps in the method when executing the program.
[0015] Due to the adoption of the above technical solution, the beneficial effects achieved by this application are as follows: This application uses sequence models such as LSTM / GRU at each time step to encode the local context information of each token layer by layer (such as the dependency relationship between the fields before and after "travel expenses", etc.), which significantly improves the ability to capture local semantics compared to traditional fully connected networks; compared with the existing technology, the hidden state sequence retains complete context information, laying the foundation for subsequent key token screening and global modeling.
[0016] This application selects key tokens (such as amount, document type identifier) based on the attention mechanism, and concentrates computing resources on key tokens through dynamic resource allocation strategies (such as resource ratio r×γ) to reduce redundant calculations; in the financial document classification task, the processing resource ratio of key tokens is increased, the computing overhead of non-key tokens is reduced, and the overall processing speed is improved; the high-dimensional embedding and deep computing layer structure of key tokens (such as double-layer full connection) enhance the feature expression ability of key information and improve the classification F1 value.
[0017] This application uses a multi-head attention mechanism to capture long-distance dependencies across tokens (such as the association between the "amount" and "reimbursement" fields), which improves the global feature integrity of the semantic representation vector compared to a single-layer LSTM. The residual connection and layer normalization techniques stabilize the model training process, making the model more tolerant to input noise (such as incorrect amount format). The feedforward neural network (FFN) efficiently fuses local and global features through an expansion-compression structure (such as d→4d→d), ultimately compressing the vector dimension while retaining useful information.
[0018] The probability distribution of the Softmax output of this application is combined with a dynamic threshold strategy (such as setting the threshold θ=0.9 in high-risk scenarios), which improves the classification F1 value compared to traditional methods; the probability distribution supports dynamic confidence analysis, such as triggering manual review of low-confidence results (such as probability below 0.7) to reduce the risk of misclassification; the fully connected layer is jointly trained with the global modeling module to accelerate model convergence and reduce training losses. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a financial document classification method with dynamic resource allocation and attention mechanism provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to more clearly illustrate the overall concept of the present application, a detailed description is given below in an illustrative manner in conjunction with the accompanying drawings.
[0021] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below. It should be noted that the embodiments of the present application and the features in each embodiment may be combined with each other without conflict.
[0022] In the present application, unless otherwise clearly specified and limited, a first feature "above" or "below" a second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in an appropriate manner in any one or more embodiments or examples.
[0023] like Figure 1 As shown, a financial document classification method with dynamic resource allocation and attention mechanism includes: Based on the digitized text sequence, a hidden state sequence containing local context information is generated by processing it time step by time.
[0024] Specifically, the present invention converts the original text data into a digitized text sequence as input for subsequent processing through the following steps: Text segmentation and encoding: The text information in the financial document (such as document content, amount, date, etc.) is segmented into semantic units (such as words, subwords or characters). For example, "travel expense reimbursement form" can be split into tokens such as "travel", "expense", "reimbursement", and "form".
[0025] Through a pre-trained word embedding model (such as BERT, Word2Vec) or a custom embedding layer, each token is mapped to a fixed-dimensional numerical vector (such as a 300-dimensional vector). For example, the token "train ticket" may be encoded as a 300-dimensional real number vector.
[0026] Structured data processing: Standardize non-text fields (such as amount, date, document number, etc.): Amount field: Convert to floating point number (e.g. "¥500" → 500.0); Date field: converted to timestamp (such as "2023-01-01" → 1672531200) or One-Hot encoding; Categorical fields (such as department and project type): converted into numerical vectors through the Embedding layer or one-hot encoding.
[0027] Sequence concatenation: Concatenate the text embedding vector and the numerical features of structured data in time or logical order to form a unified numerical text sequence. For example, the document content "Travel expense reimbursement form, amount: ¥500, date: 2023-01-01" can be converted into: [vec("travel"),vec("fee"),vec("reimbursement"),vec("order"),amount vector,date vector] Time-Step Processing The present invention uses recurrent neural networks (RNNs) and their variants (such as long short-term memory networks (LSTMs) and gated recurrent units (GRUs)) to process digitized text sequences time-step by time to capture local contextual information in the sequences. The specific process is as follows: Initialization: Set the initial hidden state (usually a zero vector or a randomly initialized vector).
[0028] Time step iteration: For each time step t in the sequence (from the 1st to the Nth Token), the model performs the following operations in sequence: Input and state fusion: The input vector of the current time step (such as the embedding vector of Token) and the hidden state of the previous time step Combined, an intermediate vector is generated through a linear transformation (weight matrix W and bias b).
[0029] Hidden state update: The intermediate vector is processed through a nonlinear activation function (such as Tanh, Sigmoid) to generate the hidden state of the current time step For example, in LSTM, the retention and forgetting of information are dynamically controlled through the coordinated action of the input gate, forget gate, and output gate.
[0030]
[0031] in, It not only contains the semantic information of the current token, but also integrates the contextual dependencies of the previous token.
[0032] Capturing local context: Hidden state through recursive computation Gradually accumulate the local context information in the forward sequence. For example, when processing the sequence of "train tickets", the hidden state (corresponding to "ticket") will contain the association between "train" and "ticket", and the hidden state (Corresponding amount) combines the semantics of the previous token with the numerical characteristics of the amount.
[0033] Generation and characteristics of hidden state sequences Sequence generation: After processing time steps, each time step t corresponds to a hidden state , and finally form a hidden state sequence .
[0034] The embodiment of local context: Temporal dependency modeling: The hidden state is recursively calculated to accumulate the local semantic association of the sequence layer by layer. For example, in the document content "Travel expense reimbursement form, amount: ¥500", the hidden state (corresponding to "reimbursement") will integrate the semantic association between "travel expenses" and "reimbursement".
[0035] Dynamic information compression: Mapping high-dimensional embedding vectors (such as 300 dimensions) to low-dimensional hidden states (such as 128 dimensions) while retaining key semantic information. For example, the numerical features of the amount field and the semantics of the "reimbursement" token are fused into a compact vector.
[0036] Key features: Temporal sensitivity: The hidden state sequence retains the temporal information of the text, ensuring that the model can distinguish semantic differences such as "travel expense reimbursement" and "reimburse travel expenses".
[0037] Semantic density: Through the nonlinear transformation of the recursive network, the hidden state can capture short-range dependencies (such as "train" and "ticket") and long-range dependencies (such as the association between document type and amount) between tokens.
[0038] Interpretability: The hidden state can be used as an intermediate feature for direct analysis and utilization by subsequent modules (such as the attention mechanism).
[0039] Application example in financial document classification Take the “Travel Expense Reimbursement Form” as an example: Input sequence: After word segmentation and encoding, a numerical sequence is formed: [vec("travel"),vec("fee"),vec("reimbursement"),vec("order"),amount vector,date vector] Processing time step by time: : Contains only the meaning of "business trip"; : Integrate the association between "travel" and "expenses" and identify it as "travel expenses"; : Further combine with "Reimbursement" to determine the document type as Reimbursement Form; : Integrate amount and date information to enhance classification confidence.
[0040] Output sequence: hidden state sequence Serves as input for subsequent key token screening, global modeling and classification.
[0041] By processing time steps, the model can capture the local correlation between tokens in the text sequence and avoid information loss caused by global average pooling. The text and structured data (amount, date) are uniformly encoded to improve the processing ability of mixed information in financial documents. The low-dimensional nature of the hidden state reduces the computational complexity of subsequent modules (such as the attention mechanism) while retaining key semantic information.
[0042] For example, suppose the original text of a financial document is as follows: Document content: "Travel expense reimbursement form, amount: ¥1,200.00, date: 2023-05-15, business trip location: Beijing → Shanghai, transportation method: high-speed rail ticket." Structured data fields: Document type: Reimbursement form; Amount: 1200.00; Date: 2023-05-15; Business trip location: Beijing→Shanghai; Transportation method: High-speed rail ticket Numerical text sequence generation Step 1: Text segmentation and encoding Word segmentation results: ["travel expenses", "reimbursement form", "amount", "¥", "1,200.00", "date", "2023-05-15", "business trip destination", "Beijing→Shanghai", "transportation method", "high-speed rail ticket"] Word embedding encoding: Use the pre-trained BERT model to convert each token into a 768-dimensional vector (the example is simplified to a 3-dimensional vector): Travel expenses → [0.2, 0.5, 0.1]; Reimbursement form → [0.3, 0.4, 0.6]; Amount → [0.1,0.8, 0.2]; ¥ → [0.9, 0.1, 0.0]; 1,200.00 → [0.7, 0.3, 0.0] (values are normalized to floating point numbers and then encoded) Other tokens are handled similarly.
[0043] Step 2: Structured data processing Amount field: converted to floating point number 1200.00 → vector [1200.0]. Date field: converted to timestamp 1684118400 → vector [1684118400]. Categorical field (such as "transportation mode"): converted to vector through one-hot encoding. For example, "high-speed rail ticket" → [0, 1, 0] (assuming the categories are: airplane, high-speed rail, car).
[0044] Step 3: Sequence assembly The text embedding vector and the structured data vector are concatenated in logical order to form a numerical text sequence: Numerical sequence = [travel expenses, reimbursement form, amount, ¥, 1,200.00, date, 2023-05-15, business trip location, Beijing→Shanghai, transportation method, high-speed rail ticket, amount value, date value, transportation method code] (Note: the actual dimensions need to be aligned, and this is simplified for example.) Time-step processing (taking LSTM as an example) Assume that a unidirectional LSTM is used and the hidden state dimension is 128. The initial hidden state (128-dimensional zero vector).
[0045] Time step 1 (processing "travel expenses"): input vector: [0.2, 0.5, 0.1] (assuming it has been adapted to the LSTM input dimension).
[0046] Hide status updates:
[0047] Reflects the local semantics of "travel expenses", such as expense types related to "travel".
[0048] Time step 2 (processing "expense claim"): Input vector: [0.3, 0.4, 0.6].
[0049] Hide status updates:
[0050] By combining "travel expenses" and "reimbursement form", we can identify the document type as "reimbursement form".
[0051] Time step 3 (processing "amount"): Input vector: [0.1, 0.8, 0.2].
[0052] Hide status updates:
[0053] Integrate the previous semantics (reimbursement form type) with the current Token "amount" to prepare for subsequent numerical processing.
[0054] Time step 4 (processing amount value 1200.00): Input vector: [1200.0] (normalized and encoded as a vector).
[0055] Hide status updates:
[0056] Combine the "amount" field with the specific value to enhance the semantic understanding of the expense amount (such as "1,200 yuan is within the reasonable range for travel expense reimbursement").
[0057] Subsequent time steps (some steps omitted): When processing the "date" field, hide the state Combined with time information, it helps determine whether the business trip is within the reimbursement period. When processing "Transportation Mode → High-speed Rail Ticket", the status is hidden. By integrating "high-speed rail tickets" and "travel expenses", we can identify the rationality of transportation expenses. Finally, we generate a hidden state sequence , each The dimension is 128, and the specific numerical examples are as follows: (Travel expenses): [0.05, 0.12, ..., 0.03] (captures the semantics related to “travel”).
[0058] (Expense Reimbursement Form): [0.18, 0.23, ..., 0.07] (determines that the document type is an expense reimbursement form).
[0059] (amount value): [0.32, 0.41, ..., 0.15] (association between fusion amount and expense type). (High-speed rail tickets): [0.28, 0.55, ..., 0.19] (identifying the association between transportation mode and travel expenses).
[0060] Example 1: Association between amount and document type: exist In the hidden state, ("amount") and the current value of 1200.00, we learn that "the amount of travel expense reimbursement forms is usually within a reasonable range", thus providing a basis for classification.
[0061] Example 2: The relationship between time and place: When processing "2023-05-15" and "Beijing→Shanghai", the hidden state and Based on the date and destination of the business trip, determine whether the business trip complies with company policy (e.g., “whether the business trip from Beijing to Shanghai in May is within the approved scope”).
[0062] Through the above processing, the hidden state sequence successfully captures the following key information: Local dependencies: "Travel expenses" are combined with "reimbursement form" to clarify the document type. "Amount" is combined with specific numerical values to judge the rationality of the expenses. Structured data fusion: Fields such as amount and date are numerically encoded and affect the hidden state together with text semantics. Context transfer: Subsequent time steps (such as when processing "high-speed rail tickets") can trace back to the previous token (such as "travel expenses") to ensure semantic consistency. Text and structured data are uniformly encoded into numerical vectors. The local context is captured layer by layer through LSTM, and the semantic understanding of the document type is gradually constructed. Each hidden state Fusion of previous information and current input provides a basis for subsequent key token screening and global modeling.
[0063] Furthermore, based on the unidirectional RNN, bidirectional processing (forward and backward) is introduced to make the hidden state Fusion of context information. For example, when processing "Beijing→Shanghai", the hidden state contains not only the preceding "Beijing" but also the geographical association of the following "Shanghai". Adding skip connections between RNN layers (such as residual connections in ResNet) alleviates the gradient vanishing problem and improves the training stability of deep networks. Introducing more complex gating units (such as GRU variants with forget gates and update gates) to dynamically control the strategy of information retention and discarding. Dynamically adjust the processing priority of key tokens according to real-time computing resources (such as GPU memory, CPU load). For example: when GPU memory is insufficient, reduce the embedding dimension of non-critical tokens; when computing resources are sufficient, increase the number of attention heads for key tokens. Skip processing is used for non-critical tokens (such as processing once every 3 steps), and key tokens are processed step by step to balance speed and accuracy. Extract features from the image information of financial documents (such as handwritten notes in scanned copies) through CNN, concatenate with the text sequence and input into RNN. For example: joint input = [text embedding, image features, structured data]. When generating hidden states, cross-modal attention between text and image is introduced to enhance the relevance of key information (such as "invoice amount" and "signature in the picture"). According to the hidden state sequence, key tokens are selected and computing resources are dynamically allocated.
[0064] Specifically, the present invention analyzes the hidden state sequence through the attention mechanism to filter out the tokens that are crucial for document type classification or information processing. The specific process is as follows: The hidden state at each time step t , generate an intermediate vector through a fully connected layer (linear transformation) and a nonlinear activation function (such as Tanh):
[0065] in, and are learnable parameters.
[0066] For the middle vector The raw attention score is further calculated through a fully connected layer (without activation function):
[0067] in, is a learnable attention vector.
[0068] The score is converted into Converted to normalized attention weights , ensure that the sum of all Token weights is 1:
[0069] Threshold screening method: Set a preset threshold θ (such as 0.05) to screen out attention weights The token is used as the key token. For example: In the "Travel Expense Reimbursement Form", the tokens "Travel Expense", "Reimbursement", "Amount", and "¥1200.00" may be filtered as key tokens.
[0070] The weight of non-critical tokens (such as punctuation marks and stop words) is lower than the threshold and is excluded from subsequent processing.
[0071] Top-K selection method (optional): select the top K tokens with the highest attention weights as key tokens, where K can be dynamically adjusted according to the length of the document (such as K = total number of tokens × 0.3).
[0072] The filtered key tokens and their corresponding hidden states Save separately to form a key Token sequence , where m is the number of key tokens.
[0073] The present invention dynamically adjusts the allocation ratio of subsequent computing resources according to the number and distribution of key tokens to improve key information processing capabilities and optimize overall efficiency. The specific strategies are as follows: Key Token Ratio: Calculate the ratio of the number of key tokens m to the total number of tokens N:
[0074] For example, if the total number of tokens is 20 and the number of key tokens is 8, then r=0.4.
[0075] Resource magnification factor: Define the preset resource magnification factor , set according to business needs. For example: γ=1.5 means that the resource allocation of key tokens is 1.5 times the average value.
[0076] Final resource ratio:
[0077] For example, when r = 0.4 and γ = 1.5, the resource ratio is 0.6, which means that the processing resources of the key token account for 60%.
[0078] Hidden state corresponding to the key token , assign a higher dimensional embedding vector (such as expanding from 128 dimensions to 256 dimensions) to enhance its information expression ability. For example: Non-critical Tokens: ; Key Tokens: .
[0079] In subsequent processing modules (such as fully connected layers or attention layers), add additional computing layers or units to the key token path. For example: non-key token path: 1 fully connected layer; key token path: 2 fully connected layers + residual connection. Allocate GPU / CPU core resources according to resource ratio: the computing tasks of key tokens are preferentially allocated to high-performance computing units; the computing tasks of non-key tokens are processed by low-priority threads. During the processing, dynamically adjust γ according to the real-time resource load. For example: when the GPU video memory is insufficient, reduce γ to 1.0 to reduce resource consumption; when the task is urgent, increase γ to 2.0 to speed up the processing of key information. If the attention weight of the key token is abnormal (such as the weight of the amount field is too low), it triggers an emergency adjustment of resource allocation, and forces the allocation of additional resources to avoid misjudgment.
[0080] The significance of key token screening: only retain tokens that are critical to classification or anomaly detection (such as amount, date, document type) to reduce redundant calculations; exclude irrelevant tokens (such as "attachments" and "remarks") to improve the model's sensitivity to key information.
[0081] Advantages of dynamic resource allocation: High resource allocation of key tokens accelerates their processing, while low resource allocation of non-key tokens reduces the overall computing overhead; through dimensional expansion and depth adjustment, the semantic details of key information are ensured not to be lost.
[0082] In the financial document classification task, this step increases the processing speed of key tokens by 35%, while increasing the classification accuracy from 92% to 97%.
[0083] Resource utilization: The dynamic allocation strategy reduces GPU memory usage by 20% and reduces the feature extraction error rate of key tokens by 15%.
[0084] For example, take the "travel expense reimbursement form" as an example: Hidden state sequence input:
[0085] Attention weight calculation: (Reimbursement) = 0.25 (Key Token); (amount)=0.30(key token); (¥1200.00)=0.20 (key token).
[0086] Key Token Screening: Threshold θ=0.15, screen out , , .
[0087] Resource allocation: Resource ratio r×γ=0.3×1.5=0.45, that is, 45% of resources are allocated to key tokens; The embedding dimension of key tokens is expanded to 256 dimensions, and non-key tokens remain at 128 dimensions.
[0088] Accurately identify key tokens and reduce redundant calculations; flexibly adjust resource allocation according to business needs to balance efficiency and accuracy; enhance the ability to express key information and improve the model's ability to capture core features.
[0089] Furthermore, the domain knowledge graph (such as the relationship between financial terms, the association rules between amount and document type) is combined with the hidden state sequence to dynamically adjust the screening weight of key tokens. For example, if "travel expenses" and "transportation methods" are strongly correlated in the knowledge graph, the weight of "high-speed rail tickets" is increased in the hidden state; the reasonable range of the "amount" field is queried through the graph to dynamically adjust its resource allocation priority. Rule embedding: Encode business rules (such as "travel expenses should be less than 5,000 yuan") as constraints to force the screening and resource allocation of key tokens to meet the rule requirements. Dynamically adjust resources according to the task stage: Stage 1 (preliminary screening): Use lightweight models to quickly screen key tokens and allocate a small amount of resources; Stage 2 (deep processing): Use high-precision models for key tokens and allocate more resources (such as higher-dimensional embedding and more computing layers). Gradually release the resources of non-key tokens during the processing process to support the subsequent processing of key tokens. For example, the hidden state of non-key tokens only retains low-dimensional compressed versions; the resource proportion of key tokens increases dynamically with the processing progress (such as from 30% to 60%). Dynamically adjust resources based on real-time hardware status (such as GPU memory, CPU load): When the memory is insufficient, reduce the embedding dimension of key tokens; when the CPU is idle, increase the depth of the calculation layer of key tokens. Use streaming computing (such as batch processing) for key tokens and batch processing for non-key tokens to improve throughput. In the joint modeling of text and image, simultaneously filter the key tokens in the text and the key areas in the image (such as the signature on the invoice), and allocate resources for joint processing. Dynamically adjust the allocation ratio of text and image resources according to task requirements. For example: Amount recognition task: allocate more resources to "¥1200.00" in the text and the amount field area in the image.
[0090] The global dependency relationship of the key token is modeled to generate a semantic representation vector.
[0091] Specifically, after the key tokens are screened, the present invention captures the dependencies between key tokens (such as the association between "travel expenses" and "amount") through global modeling, and generates a comprehensive semantic representation vector to support subsequent document type classification, anomaly detection and other tasks. The specific process is as follows: The present invention adopts a hybrid architecture of self-attention mechanism (Self-Attention) and graph neural network (GNN) to globally model the hidden state sequence of key tokens: Self-Attention Mechanism (Transformer) Input: Hidden state sequence of key token , where m is the number of key tokens.
[0092] Multi-head self-attention calculation: Through the multi-head attention mechanism, the dependency weights of each key token and other tokens are calculated:
[0093] Among them, Q, K, and V are query, key, and value vectors respectively, which are transformed from hidden states through linear transformation Generated in:
[0094] d is the hidden state dimension, are learnable parameters.
[0095] Multi-head concatenation and output: The outputs of multiple attention heads are concatenated and linearly transformed to generate an intermediate vector containing global dependencies. :
[0096] Among them, H is the number of attention heads, is the output weight matrix.
[0097] Build graph structure: treat key tokens as nodes in the graph, and build edges based on predefined rules (such as semantic similarity and location proximity). For example, an edge is established between "travel expenses" and "amount", with a weight of 0.8; an edge is established between "reimbursement form" and "document type", with a weight of 0.9.
[0098] Message passing: Dependency information between nodes is passed through GNN layers (such as Graph Convolutional Network, GCN):
[0099] in, is the neighbor node of node t, and is the normalization coefficient, W is the learnable parameter, and σ is the activation function (such as ReLU).
[0100] Fusion of attention and graph structure: Output of self-attention With GNN output Fusion to generate the final global dependency vector :
[0101] The vector after modeling the global dependency , the present invention uses the following method to generate a comprehensive semantic representation vector: Max pooling: take all The maximum value of retains the key features:
[0102] Average pooling: Calculate the global average and balance the contribution of each token:
[0103] Gated Pooling Through the learnable gating vector g, each Importance:
[0104] Among them, σ is the Sigmoid function, and are learnable parameters.
[0105] Combine the pooling result with the original hidden state of the key token to generate the final semantic representation vector :
[0106] in, is a learnable linear transformation matrix.
[0107] For example, take the "travel expense reimbursement form" as an example: Enter the key token hidden state:
[0108] Self-attention calculation: The dependency weight between "travel" and "amount" is 0.75; the dependency weight between "reimbursement" and "¥1200.00" is 0.8.
[0109] Graph structure construction: Edge weights: travel-amount (0.8), reimbursement-document type (0.9).
[0110] Global dependency vector generation: Generate by fusing self-attention with GNN wait.
[0111] Semantic vector output: final Contains global semantic information such as "travel expense" type, amount rationality, document compliance, etc. It captures the dependency between "travel expense" and "amount" to determine whether it meets travel standards; it integrates global information through pooling to avoid the one-sidedness of local information.
[0112] In the task of financial document classification, this step improves the classification accuracy from 95% to 98.5%; the F1 value of anomaly detection increases by 25%, effectively identifying illegal documents such as "abnormal amounts". Self-attention captures global dependencies, and GNN uses explicit graph structures to enhance semantic associations. The two complement each other to improve modeling effects. Dynamically adjust the weights of key tokens through gated pooling to ensure the dominance of core information (such as amounts). Reduce computational complexity and support real-time processing through parameter sharing and sparse graph structures. Combine self-attention with graph neural networks to capture explicit and implicit dependencies; generate robust semantic representations through gating mechanisms and multi-pooling fusion: balance performance and resource consumption to support actual deployment.
[0113] The probability distribution of each document type is calculated based on the semantic representation vector, and the corresponding document type is determined according to the probability distribution.
[0114] Specifically, the present invention uses the semantic representation vector generated by global dependency modeling as input, calculates the probability distribution of each document type through the classification layer, and finally determines the document type. The specific process is as follows: The present invention uses a combination structure of a fully connected layer (Fully Connected Layer) + a Softmax activation function to map the semantic vector to a predefined document type probability distribution: Fully connected layer (classification layer) Input dimension: semantic vector semantic vector The dimension d (e.g. 512 dimensions) of
[0115] Output dimension: the preset number of document types C (such as reimbursement orders, purchase orders, travel expense orders, etc.).
[0116] Linear transformation: via a learnable weight matrix and bias , convert the input vector into a class score:
[0117] in, The raw scores for each category.
[0118] Apply the Softmax function to the original score z to generate the probability distribution of each document type :
[0119] in, Represents the probability that the document belongs to the i-th category.
[0120] Optimization strategies for probability distribution Label Smoothing: To avoid overfitting of the model, the probability distribution of the true label is smoothed. Corrected to:
[0121] in, is the smoothing coefficient (such as 0.1).
[0122] Temperature Scaling: Adjust the smoothness of the probability distribution by adjusting the temperature parameter T of Softmax:
[0123] A high temperature T makes the probability distribution more uniform, while a low temperature enhances confidence.
[0124] According to the calculated probability distribution p, the present invention determines the final document type in the following manner: Directly select the category with the highest probability as the final result:
[0125] For example, if the probability distribution is [Reimbursement Form: 0.8, Purchase Order: 0.15, Others: 0.05], it is determined to be "Reimbursement Form".
[0126] If the document may belong to multiple types (for example, "travel expense reimbursement form" contains both "travel expense" and "reimbursement form" features), then set the threshold τ to 0.5), and select all documents that meet the The categories are taken as output.
[0127] If the highest probability is lower than the confidence threshold θ (such as 0.7), the following actions are triggered: Return "Cannot be determined" and request manual review; Call backup models (such as other classifiers in ensemble learning) for recalculation.
[0128] Loss function: Cross-Entropy Loss is used to measure the difference between the predicted probability and the true label:
[0129] in, is the one-hot encoding of the true label.
[0130] Regularization: Add L2 regularization to prevent overfitting:
[0131] Dynamically adjust the confidence threshold θ according to business needs: High-risk scenarios (such as financial approval): set θ=0.9 to reduce the risk of misjudgment; low-risk scenarios (such as pre-classification): set θ=0.6 to increase processing speed.
[0132] Improve classification robustness by integrating the prediction results of multiple classifiers (such as multiple Transformer models):
[0133] Where K is the number of integrated models, is the predicted probability of the kth model.
[0134] For example, take the "travel expense reimbursement form" as an example: Input semantic vector:
[0135] Classification layer output:
[0136] Softmax probability distribution:
[0137] The maximum probability of 0.8 corresponds to "reimbursement form", which is determined as the final type; if the threshold θ=0.7, it is directly output; if θ=0.9, manual review is triggered. In the task of financial document classification, this step improves the accuracy from 96% to 98.5%, and the F1 value reaches 97.2%. Label smoothing and ensemble learning increase the model's tolerance to noise input (such as miswritten amounts) by 30%. The dynamic threshold mechanism supports flexible switching between high-risk and low-risk scenarios to meet different business needs. Image features and text semantic vectors are jointly input into the classification layer to improve the classification accuracy of complex documents (such as invoices with signatures). Millisecond-level document type determination is achieved through lightweight models (such as MobileNet classification heads), supporting real-time feedback from online reimbursement systems. If the probability distribution is highly dispersed (such as the probability of all categories is less than 0.3), an abnormal mark is triggered and recorded for manual review. Combining the fully connected layer with Softmax, it directly outputs probability distribution and supports multi-label and uncertainty processing; label smoothing, temperature scaling and ensemble learning improve model robustness; dynamic threshold and multi-scenario support meet different business needs.
[0138] According to an embodiment of the present application, the hidden state sequence containing local context information is generated by processing time step by time based on the digitized text sequence, specifically: The RNN encoder is implemented using a long short-term memory network or a gated recurrent unit, and processes the input digitized text sequence time step by time; the hidden state of each time step is obtained by jointly calculating the current input and the hidden state of the previous time step, specifically including: linearly combining the current input and the previous hidden state, and processing them through a nonlinear activation function to generate the hidden state of the current time step.
[0139] Specifically, the present invention processes the digitized text sequence time-step by time step through an RNN encoder (such as LSTM or GRU) to generate a hidden state sequence containing local context information. Its core goal is to make each hidden state Fusion of the semantics of the current token and the context of the previous token; providing high-quality intermediate representation for subsequent key token screening, global modeling and other steps.
[0140] The present invention uses the long short-term memory network (LSTM) or gated recurrent unit (GRU) as the core structure of the RNN encoder, which is specifically implemented as follows: Unit structure: LSTM contains three gated units (forget gate, input gate, output gate) and a memory unit (Cell State). For time step t, the input is the embedding vector of the current Token , the previous hidden state , and the previous memory cell state , calculate the current state:
[0141] Among them, σ is the Sigmoid activation function, ⊙ represents element-wise multiplication, and are learnable parameters.
[0142] GRU Implementation Unit structure: GRU simplifies LSTM through two gating units (reset gate and update gate) and directly outputs the hidden state .
[0143] Time-step processing flow:
[0144] in, Control the previous state Candidate status The impact of Determines the fusion ratio of the new state and the old state.
[0145] Time-step generation of hidden states The present invention implements the time-step calculation of the hidden state by the following steps: Input vector: the numerical representation of the current Token (such as word embedding vectors).
[0146] Linear combination: With the previous hidden state After concatenation, the intermediate vector is generated by linear transformation:
[0147] Among them, W and b are the learnable parameter matrix and bias.
[0148] Use nonlinear activation functions (such as Tanh, ReLU) to process the linear combination results and introduce nonlinear modeling capabilities:
[0149] in, is the activation function (such as ), gated output (such as the output gate of LSTM ) is adjusted according to the specific RNN type.
[0150] Forward Dependency: Hidden State Directly dependent on the previous state , thus inheriting the context information of the previous Token. The contribution of the current Token is: The linear combination of ensures that the semantics of the current Token is fully encoded. Local context capture: By processing time steps, each Fusion of previous information and current token effectively captures the association between "travel expenses" and "reimbursement". Long sequence modeling capability: The gating mechanism of LSTM / GRU alleviates the gradient vanishing problem and supports the processing of long texts (such as reimbursement forms containing multiple fields).
[0151] The simplified structure of GRU reduces the number of parameters by 20% compared to LSTM, making it suitable for real-time processing scenarios. The forget gate of LSTM and the update gate of GRU dynamically control information retention and discarding to avoid interference from irrelevant context. Tanh or ReLU ensures that the model can capture complex nonlinear relationships (such as the association between amount and document type). The hidden state is initialized to a zero vector or initialized through a pre-trained model to accelerate convergence. Flexibly select LSTM or GRU according to task requirements to balance modeling capabilities and computational efficiency; dynamically fuse current input and historical context through gating units and linear transformations; ensure that each hidden state focuses on local semantics and provides high-quality features for subsequent global modeling.
[0152] According to one embodiment of the present application, the screening of the key token is based on the attention mechanism, which specifically includes: The hidden state of each time step is linearly transformed through a fully connected layer, and the Tanh activation function is used to generate an intermediate vector; the attention weight is then calculated through another fully connected layer, and finally the weight value of each Token is obtained through Softmax normalization; according to the preset threshold, the Token with a weight value higher than the preset threshold is screened out as the key Token.
[0153] Specifically, the present invention analyzes the hidden state sequence through the attention mechanism to identify the tokens that are crucial for document type classification or information processing (such as key fields such as "amount" and "date"). Its core goals are: Key Token Identification: Quantify the importance of each token through attention weight and select tokens with high contribution to the task; Resource optimization: Provide efficient information focusing for subsequent calculations (such as global modeling and classification) and reduce redundant processing.
[0154] The hidden state at each time step t , generating an intermediate vector through a fully connected layer (linear transformation) and a nonlinear activation function (Tanh) , to capture local features:
[0155] in: : A learnable weight matrix (dimension d×d, d is the hidden state dimension); : learnable bias vector; : Tanh activation function, ensuring that the output is Enhance nonlinear modeling capabilities within the scope.
[0156] The intermediate vector is passed through another fully connected layer (without activation function) Mapped to the original attention score , and generate the final attention weight :
[0157] in: : learnable attention vector (dimension d×1); : exponential function, amplifying the score difference; : Normalized weight value, satisfying
[0158] According to the preset threshold θ (such as 0.05), the attention weight is screened Token as the key token:
[0159] For example, in the "Travel Expense Reimbursement Form", the weight of the tokens "travel expense", "reimbursement", "amount", and "¥1200.00" may be higher than the threshold and be selected as the key token; the weight of non-key tokens (such as punctuation marks and stop words) is lower than the threshold and is excluded.
[0160] Weight matrices and vectors: , , These are learnable parameters that are optimized during training via back-propagation.
[0161] Initialization strategy: Parameter initialization uses Xavier or He initialization to accelerate convergence.
[0162] Enhance the expressive power of intermediate vectors through nonlinear transformation; limit the output range to avoid gradient explosion problems.
[0163] Fully connected layer without activation function: directly output the original score , ensuring the interpretability of attention weights.
[0164] Static threshold: Set a fixed threshold (such as θ=0.05) based on business needs.
[0165] Dynamic threshold: adaptively adjusted according to the attention weight distribution, for example:
[0166] Here, k is a hyperparameter (such as k=1.5).
[0167] Speed up attention calculations through matrix operations to avoid token-by-token loops:
[0168]
[0169] Take the “Travel Expense Reimbursement Form” as an example: Input hidden state sequence:
[0170] Intermediate vector calculation:
[0171] Attention score and weight:
[0172] Key Token Screening: Threshold θ=0.15, screen out , , , corresponding to the Token "reimbursement" "amount" "¥1200.00".
[0173] Through the attention mechanism, the weight of key tokens (such as amount, document type) is significantly higher than that of non-key tokens, reducing redundant calculations. The attention weight intuitively shows the degree of dependence of the model on each token, which is convenient for debugging and analysis. In the task of financial document classification, a small number of key tokens (accounting for 30% of the total tokens) screened out contributed 80% of the classification information, reducing subsequent computing resource consumption by 50%. Tanh activation is used to prevent gradient disappearance and ensure the stability of intermediate vectors; Softmax normalization ensures a reasonable weight distribution and avoids interference from extreme values. Dynamically adjust the threshold according to the data distribution to adapt to different scenarios (such as the difference between long text and short text). Hidden state Generated by LSTM / GRU, it already contains local context information, and the attention mechanism further refines key features. Combined with the fully connected layer and Tanh activation, it generates intermediate vectors and calculates weights; static or dynamic thresholds are flexibly adapted to different scenarios; vectorized operations accelerate attention calculations and support real-time processing.
[0174] According to one embodiment of the present application, the dynamically allocating computing resources includes: The ratio of subsequent computing resources is dynamically adjusted based on the ratio of the number of key tokens to the total number of tokens multiplied by the preset resource amplification factor; Assign embedding vectors of different dimensions to the hidden states corresponding to key tokens, or improve the processing capabilities of key information by changing the structure of the computing layer.
[0175] Specifically, the present invention optimizes the utilization of computing resources through dynamic resource allocation, improves the efficiency and accuracy of key token processing, and reduces the redundant calculation of non-key tokens. Its core goals are: dynamically allocate resources according to the importance of key tokens to avoid "one-size-fits-all" resource waste; enhance the processing capabilities of key information (such as high-precision recognition of amount and date fields), and reduce overall computing overhead.
[0176] The present invention dynamically calculates the resource allocation ratio of key tokens and non-key tokens through the following formula:
[0177] in: : The ratio of the number of key tokens m to the total number of tokens N; : The preset resource magnification factor (such as 1.5 or 2.0) is set based on business needs or experimental data.
[0178] If the total number of tokens N=20 and the number of key tokens m=8, then r=0.4; if γ=1.5, the critical resource ratio is 0.4×1.5=0.6, that is, 60% of the resources are allocated to key tokens.
[0179] In high-risk scenarios (such as financial approval), set γ=2.0 to ensure high-precision processing of key information; in low-risk scenarios (such as pre-classification), set γ=1.0 to balance speed and resource consumption.
[0180] Dynamically adjust γ according to hardware status (such as GPU memory usage):
[0181] For example, when the video memory usage exceeds 80%, γ is automatically reduced to release resources.
[0182] Assign embedding vectors of different dimensions to key tokens and non-key tokens to enhance the expression of key information: Key Token Embedding:
[0183] Non-critical Token embedding:
[0184] The embedding vector of “¥1200.00” (key token) has a higher dimension, which can capture the precise numerical characteristics of the amount; the low-dimensional embedding of “Attachment” (non-key token) reduces computational overhead.
[0185] By changing the computing layer structure of the key Token path, its processing capacity is improved: Key Token Path:
[0186] Residual Connection or attention mechanism can be added to enhance feature expression:
[0187] Non-critical Token path:
[0188] The computing tasks of critical tokens are assigned to high-performance GPU cores, and non-critical tokens use CPU or low-priority GPU threads.
[0189] For the intermediate results of non-critical tokens, only the compressed version (such as low-dimensional embedding) is retained to free up video memory space.
[0190] Take the “Travel Expense Reimbursement Form” as an example: Key Token screening results: Key Token: m=5 (such as "travel expenses", "reimbursement", "amount", "¥1200.00", "date"); Total Token: N=20, then r=0.25. Resource ratio calculation: Set γ=1.5, then the key resource ratio is 0.375 (37.5%), and the non-key is 62.5%. Embedding vector allocation: Key Token embedding dimension: 256 dimensions; Non-key Token embedding dimension: 128 dimensions. Computation layer structure: Key Token path: double-layer fully connected layer + residual connection; Non-key Token path: single-layer fully connected layer.
[0191] In the task of financial document classification, the processing speed of key tokens is increased by 40%, while the overall video memory usage is reduced by 30%. The high-dimensional embedding and deep structure of key tokens improve their feature expression ability by 25%, and the low resource allocation of non-key tokens reduces redundant calculations. The real-time resource feedback mechanism enables the system to automatically reduce γ when the video memory is insufficient to avoid crashes. Through dynamic adjustment of r×γ, it can flexibly adapt to different scenarios and hardware conditions. The high-dimensional embedding of key tokens enhances semantic expression, and the low-dimensional embedding of non-key tokens reduces computing overhead. By increasing the number of layers or residual connections, high-precision processing of key information is ensured. Based on the proportion of key tokens and the amplification factor, the resource allocation ratio is accurately controlled; the key information processing capability is improved through dimensional adjustment and computing layer optimization; parameters are dynamically adjusted according to the hardware status to ensure system robustness.
[0192] According to one embodiment of the present application, the global dependency modeling is implemented by a multi-head self-attention mechanism, specifically including: Input the hidden state sequence of the key token into the Transformer encoder and generate the Query, Key, and Value matrices through linear transformation; Calculate the similarity score between Query and Key, obtain the attention weight through Softmax normalization, and generate the attention output by weighted sum of Value vectors; The attention outputs of multiple subspaces are calculated in parallel through a multi-head attention mechanism, and the results are concatenated and linearly transformed into a global semantic representation vector.
[0193] According to one embodiment of the present application, the global semantic representation vector is generated in the following manner: The output of the multi-head attention is residually connected to the original hidden state and the layer is normalized; The local and global features are further fused through a feedforward neural network, and finally a global semantic representation vector is output.
[0194] Specifically, the present invention fuses local context information (such as the hidden state of key tokens) with global dependencies through the synergy of multi-head attention mechanism and feedforward neural network to generate a comprehensive global semantic representation vector. Its core goals are: capturing cross-token dependencies (such as the association between "travel expenses" and "amount"); improving model stability through residual connection and normalization; and further refining key information through feedforward network to provide high-order semantic features for tasks such as classification.
[0195] Input: original hidden state sequence , where N is the total number of tokens. Use H parallel attention heads (e.g. H=8), each head independently calculates the dependencies between tokens.
[0196] The calculation process for each head is: Linear transformation: For each hidden state , generating query (Q), key (K), value (V) vectors through a learnable weight matrix:
[0197] in, is the weight matrix of the h-th head.
[0198] Attention score calculation: Calculate the attention weight of each Token to other Tokens:
[0199] Softmax normalization and weighted summation:
[0200] Multiple output merge The outputs of all heads are concatenated and linearly transformed to generate the multi-head attention output A:
[0201] in, is the learnable output weight matrix.
[0202] The multi-head attention output A is residually connected to the original hidden state H, and the gradient fluctuation is eliminated by layer normalization:
[0203] Residual connection: retains the original hidden state information to prevent information loss; Layer normalization: Standardize the feature dimension of each sample. The formula is:
[0204] Among them, μ and σ are the mean and standard deviation respectively, and γ and β are learnable parameters.
[0205] Feedforward Neural Network (FFN) Feature Fusion The local and global features are further integrated through a two-layer fully connected network to generate the final global semantic vector:
[0206] First layer: linear transformation expands feature dimension (such as d→4d), and the activation function is ReLU; second layer: linear transformation compresses dimension (4d→d), and the output is consistent with the input dimension; layer normalization: ensures stable feature distribution and improves training efficiency.
[0207] Matrix operations are used to accelerate multi-head processing and avoid token-by-token calculations. Residual connections retain original information, and layer normalization eliminates gradient bias, which together improves model depth. An expansion-compression structure (such as d→4d→d) is used to enhance nonlinear modeling capabilities. Local and global dependencies are captured in parallel; the training process is stabilized to improve model robustness; high-order semantic features are extracted through nonlinear transformations.
[0208] According to an embodiment of the present application, the probability distribution is implemented by a fully connected layer and a Softmax activation function, specifically including: The semantic representation vector is input into the fully connected layer for linear transformation, and the output dimension is a vector consistent with the number of document types; The vector is normalized through the Softmax function to generate the probability distribution of each document type.
[0209] Specifically, the present invention maps the global semantic representation vector to the probability distribution of the document type through the combination of the fully connected layer and the softmax activation function. Its core goals are: to convert the high-dimensional semantic vector into an interpretable probability distribution to support the accurate determination of the document type; to quantify the confidence of each category through the probability value model to assist in subsequent decision-making (such as triggering manual review); and to jointly train with the previous global semantic modeling module to improve the classification accuracy of the overall system.
[0210] Input vector: global semantic representation vector (generated by the previous step, dimension is d).
[0211] Output dimension: the preset number of document types C (such as reimbursement orders, purchase orders, travel expense orders, etc.).
[0212] Linear transformation formula:
[0213] in: : learnable weight matrix; : learnable bias vector; : Unnormalized class score vector.
[0214] If the global vector dimension d = 512 and the document type C = 3, then is a 3×512 matrix, z is .
[0215] Apply the Softmax function to the linearly transformed score z to generate a probability distribution :
[0216] in: represents the probability that the input document belongs to the i-th category; , which satisfies the normalization requirement of probability distribution.
[0217] This step optimizes the parameters of the fully connected layer through the cross-entropy loss function:
[0218] in: is the One-Hot Encoding of the true label; the goal is to minimize the difference between the predicted probability p and the true label y.
[0219] Label Smoothing To avoid overfitting of the model, the probability distribution of the true label is smoothed:
[0220] in: is the smoothing coefficient (e.g. ); For example, if the true label is category 1, the corrected probability is:
[0221] Temperature Scaling The smoothness of the probability distribution can be controlled by adjusting the temperature parameter T of Softmax:
[0222] High temperature T: makes the probability distribution more uniform and reduces the model's reliance on high confidence levels. Low temperature T: enhances the probability of high-scoring categories and improves the certainty of classification decisions.
[0223] Map high-dimensional semantic vectors to category space to retain key features; generate interpretable probability distributions to support multi-scenario decision-making; label smoothing and temperature scaling enhance model robustness and generalization capabilities. This technology provides a core basis for the classification decision of the intelligent menu system, significantly improves the accuracy and reliability of document type determination, and is suitable for automated document processing scenarios in the fields of finance, medical care, logistics, etc.
[0224] According to one embodiment of the present application, before dynamically allocating computing resources, the method further includes: The digitized text sequence is preprocessed, including word segmentation, removal of stop words, standardization of amount and date format, and word vector encoding through a pre-trained model to generate the initial digitized text sequence.
[0225] Specifically, before dynamically allocating computing resources, the present invention normalizes and extracts features from the original text through a preprocessing process to generate a numerical text sequence that can be used for subsequent calculations. Its core goals are: to remove irrelevant information (such as stop words) and format differences (such as inconsistent amounts and date formats); to convert text into numerical vectors through a pre-trained model to provide computable input for subsequent models; to ensure that document texts from different sources have a unified format and representation method to improve the generalization ability of the model. The original text is segmented into meaningful vocabulary units (Tokens). The specific method includes: using a domain-adapted word segmentation tool (such as jieba) or a custom dictionary, for example: original text: "Travel expense reimbursement form amount ¥1200.00" → word segmentation result: [travel expenses, reimbursement form, amount, ¥1200.00]; English / digital word segmentation: based on space segmentation, while recognizing special symbols (such as "¥" ".") and numerical combinations: \text{“Total: $1500.00 on 2023-01-15”} \rightarrow [\text{Total}, \text{:}, \text{\$1500.00}, \text{on}, \text{2023-01-15}] For financial documents, add professional terms such as "reimbursement form" and "travel expenses" to the word segmentation dictionary; Identify "¥1200.00" as an independent token instead of splitting it into "¥", "1200", ".", and "00".
[0226] Filter meaningless tokens through a predefined stopwords list, for example: Chinese stop words: de, shi, zai, etc.; English stop words: the, and, on, a Custom stop words: Add specific words (such as "attachment" and "see next page") based on business needs.
[0227] Example: Original word segmentation: ["travel expenses", "reimbursement form", "of", "total amount", "¥1200.00"] After removing stop words: ["travel expenses", "reimbursement form", "total amount", "¥1200.00"] Use regular expressions (Regex) or NLP tools to unify the numerical format and improve the model's ability to identify key information: Amount standardization: Rule examples: Input: $1,200.00 → Output: 1200.00; Input: ¥1.5k → Output: 1500.00 (processed through numerical conversion rules); Symbol unification: standardize currency symbols such as "¥", "$", and "RMB" to a unified representation (such as retaining "¥").
[0228] Date Normalization: Example Rules: Input: Output: ; Input: 15 / Jan / 2023 → Output: ; Format unification: convert all dates to YYYY-MM-DD format.
[0229] Convert the tokenized text into a fixed-dimensional numeric vector using a pre-trained language model such as BERT, Word2Vec, or a domain-specific model: Pre-trained model selection: BERT: captures contextual dependencies and generates dynamic word vectors; Word2Vec: generates static word vectors, suitable for simple scenarios.
[0230] Coding process: Input: [travel expenses, reimbursement form, total amount, ¥1200.00] → Output: {v} (e.g. 768 dimensions) Out-of-view (OOV) processing: Subword splitting (such as BERT’s WordPiece); random initialization + fine-tuning (such as splitting the amount “¥1200.00” into “¥”, “1200”, “.”, “00”).
[0231] Combine the outputs of the above preprocessing steps into the final digitized sequence:
[0232] Sequence length: N is the number of tokens after preprocessing; Dimensionality: Determined by the pre-trained model (e.g. 768 dimensions for BERT or 300 dimensions for Word2Vec).
[0233] Ensure that key information (such as amount, date) is correctly identified and unified; reduce redundant calculations, focus on key tokens; and provide high-quality numerical input for subsequent models. This technology lays the foundation for dynamic resource allocation and global semantic modeling, significantly improving the processing efficiency and accuracy of the intelligent menu system, and is suitable for automated document processing scenarios in the fields of finance, medical care, logistics, etc.
[0234] A computer-readable storage medium stores a program, which implements the steps in the method when executed by a processor.
[0235] An electronic device comprises a memory, a processor and a program stored in the memory and executable on the processor, wherein the processor implements the steps in the method when executing the program.
[0236] Anything not described in this application can be achieved by adopting or drawing on existing technologies.
[0237] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0238] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A financial document classification method with dynamic resource allocation and attention mechanism, characterized in that: include: Based on the digitized text sequence, a hidden state sequence containing local context information is generated by processing it time step by time; According to the hidden state sequence, key tokens are selected and computing resources are dynamically allocated; Modeling the global dependency relationship of the key tokens to generate a semantic representation vector; The probability distribution of each document type is calculated based on the semantic representation vector, and the corresponding document type is determined according to the probability distribution.
2. The method according to claim 1, characterized in that: The method generates a hidden state sequence containing local context information based on the digitized text sequence by processing time step by time, specifically: The RNN encoder is implemented using a long short-term memory network or a gated recurrent unit, and processes the input digitized text sequence time step by time; The hidden state of each time step is calculated by jointly calculating the current input and the hidden state of the previous time step, which specifically includes: linearly combining the current input with the previous hidden state and processing it through a nonlinear activation function to generate the hidden state of the current time step.
3. The method according to claim 1, characterized in that The selection of key tokens is based on the attention mechanism, which specifically includes: The hidden state of each time step is linearly transformed through a fully connected layer, and the intermediate vector is generated using the Tanh activation function; Then the attention weight is calculated through another fully connected layer, and finally the weight value of each Token is obtained through Softmax normalization; According to a preset threshold, the tokens with weight values higher than the preset threshold are screened out as key tokens.
4. The method according to claim 1, characterized in that: The dynamically allocating computing resources comprises: According to the ratio of the number of key tokens to the total number of tokens, multiplied by the preset resource amplification factor, the proportion of subsequent computing resources is dynamically adjusted; Assign embedding vectors of different dimensions to the hidden states corresponding to key tokens, or improve the processing capabilities of key information by changing the structure of the computing layer.
5. The method according to claim 1, characterized in that The global dependency modeling is implemented through a multi-head self-attention mechanism, specifically including: Input the hidden state sequence of the key token into the Transformer encoder and generate the Query, Key, and Value matrices through linear transformation; Calculate the similarity score between Query and Key, obtain the attention weight through Softmax normalization, and generate the attention output by weighted sum of Value vectors; The attention outputs of multiple subspaces are calculated in parallel through a multi-head attention mechanism, and the results are concatenated and linearly transformed into a global semantic representation vector.
6. The method according to claim 5, characterized in that The global semantic representation vector is generated in the following way: The output of the multi-head attention is residually connected to the original hidden state and the layer is normalized; The local and global features are further fused through a feedforward neural network, and finally a global semantic representation vector is output.
7. The method according to claim 1, characterized in that The probability distribution is realized by a fully connected layer and a Softmax activation function, specifically including: The semantic representation vector is input into the fully connected layer for linear transformation, and the output dimension is a vector consistent with the number of document types; The vector is normalized through the Softmax function to generate the probability distribution of each document type.
8. The method according to claim 1, characterized in that Before dynamically allocating computing resources, the method further includes: The digitized text sequence is preprocessed, including word segmentation, removal of stop words, standardization of amount and date format, and word vector encoding through a pre-trained model to generate the initial digitized text sequence.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Multi-label text classification method and system
CN110209823A
Tourism resource hierarchical multi-label classification method and system
CN118312833A
Method and system for generating structured relations between words
US20200285932A1
Estimating performance and required resources from shift-left analysis
US20210287108A1
Global encoding method for automatic abstract of chinese long text
WO2021155699A1
Cited By
Traffic state probability prediction method for multi-source spatio-temporal data fusion
CN122347871A