Enterprise intelligent document identification and automatic classification filing management processing method and system

By building a hierarchical cognitive attention network and an adaptive feedback optimization mechanism, the problem of insufficient deep semantic information capture in the existing technology is solved, high accuracy and stability of document classification are achieved, and the automation level of enterprise intelligent document management is improved.

CN120579003APending Publication Date: 2025-09-02STATE GRID HEILONGJIANG ELECTRIC POWER COMPANY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510735621.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively capture the deep semantic information of documents, lacks adaptive optimization capabilities, resulting in low classification accuracy and lack of consistency evaluation and optimization feedback mechanisms, which cannot guarantee the stability and accuracy of classification results.

Method used

A hierarchical cognitive attention network is constructed, and feature extraction is performed using capsule networks and dynamic routing algorithms. Combined with memory-enhanced neural networks and metacognitive controllers, it is mapped to non-Euclidean manifold space through Riemann metric tensors, and adaptive feedback optimization is performed using particle swarm optimization and reinforcement learning to generate multi-dimensional classification vectors and comprehensive evaluation vectors.

Benefits of technology

It realizes accurate capture of deep semantic features of the document, improves the processing ability and adaptability of the classification model, enhances the accuracy and consistency of classification results, and reduces the cost of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579003A_ABST
    Figure CN120579003A_ABST
Patent Text Reader

Abstract

The invention provides an enterprise intelligent document identification and automatic classification archiving management processing method and system, and relates to the technical field of intelligent document processing, and the method comprises the steps: constructing a hierarchical cognitive attention network, extracting initial features of a document through a capsule network, and carrying out semantic processing through a memory enhancement neural network; and adaptively adjusting feature extraction parameters based on document complexity to obtain cognitive feature representation. Mapping the cognitive feature representation to a non-Euclidean manifold space, determining an optimal classification boundary by using an improved particle swarm optimization algorithm, and generating a multi-dimensional classification vector; and constructing an adaptive feedback optimization network, fusing the cognitive features and the classification vectors, carrying out iterative optimization based on the consistency evaluation score until a preset threshold value is reached, and outputting a final classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent document processing technology, and in particular to a method and system for enterprise intelligent document recognition and automatic classification and archiving management processing. Background Art

[0002] As enterprises advance their digital transformation, intelligent document management systems are playing an increasingly important role in improving management efficiency and business processing capabilities. The vast volume, diverse formats, and complex content of various document data generated in daily enterprise operations place higher demands on automatic document identification, classification, and archiving. Traditional document classification methods rely primarily on keyword matching and rule templates, often requiring extensive manual intervention when processing unstructured documents, making them difficult to meet the actual needs of enterprises for intelligent document management.

[0003] However, the existing technology still has several shortcomings. Traditional feature extraction methods are difficult to effectively capture the deep semantic information and contextual associations of documents, resulting in low classification accuracy; static classification models cannot adapt to the dynamic changes of document content and business scenarios, and lack adaptive optimization capabilities; a single classification algorithm is prone to ambiguous judgments when processing complex documents, reducing the reliability of classification results; existing methods lack consistency evaluation and optimization feedback mechanisms for classification results, and cannot guarantee the stability and accuracy of classification results.

[0004] In summary, there is an urgent need for an enterprise intelligent document recognition and automatic classification and archiving management method. This method involves building a document feature extraction model with multi-level cognitive understanding capabilities to accurately capture the deep semantic features of documents; designing an adaptive optimization algorithm based on non-Euclidean space to improve the classification model's ability to handle complex documents; and establishing a feedback optimization mechanism based on reinforcement learning to improve the accuracy and consistency of classification results through dynamic adjustment and continuous optimization. This invention can solve the problems in the existing technology. Summary of the Invention

[0005] The embodiment of the present invention provides an enterprise intelligent document recognition and automatic classification and archiving management processing method and system, which can solve the problems in the prior art.

[0006] According to a first aspect of the embodiments of the present invention, Provided is an enterprise intelligent document recognition and automatic classification and archiving management processing method, comprising: A hierarchical cognitive attention network consisting of a perception layer, a comprehension layer, and a decision-making layer is constructed. At the perception layer, a capsule network and a dynamic routing algorithm are used to extract features from the document being processed, construct feature space relationships, and obtain an initial feature map. At the comprehension layer, the long-term memory module and working memory module of the memory-augmented neural network are used to semantically process the initial feature map to obtain a semantic feature vector. At the decision layer, a metacognitive controller calculates document complexity based on the semantic feature vector and adaptively adjusts feature extraction parameters to obtain a cognitive feature representation of the document. The cognitive feature representation is mapped to a non-Euclidean manifold space via the Riemann metric tensor. The projection of the classification scheme on the tangent space is used as the particle position, and the projection of the update direction on the tangent plane is used as the particle velocity to generate an initial particle swarm. The particle search trajectory is recorded, the adaptive inertia weight is calculated, and the search step size is dynamically adjusted. The optimal solution is iteratively updated by combining the learning factor and the random factor, the optimal classification boundary is determined, and a multi-dimensional classification vector is generated. An adaptive feedback optimization network is constructed to perform attention-weighted mapping on cognitive feature representations and multidimensional classification vectors to generate a comprehensive evaluation vector for the document. A multidimensional similarity calculation model is used to obtain the consistency evaluation score between documents. When the evaluation score is less than the preset threshold, reinforcement learning is used to construct a strategy network to generate an optimized strategy combination and provide feedback. The iteration is repeated until the consistency evaluation score reaches the preset threshold, and the classification result of the document is output.

[0007] In an optional embodiment, a hierarchical cognitive attention network is constructed, comprising a perception layer, a comprehension layer, and a decision layer. A capsule network and a dynamic routing algorithm are used in the perception layer to extract features of the document to be processed, construct feature space relationships, and obtain an initial feature map. The comprehension layer performs semantic processing on the initial feature map using the long-term memory module and the working memory module of the memory-enhanced neural network to obtain a semantic feature vector. A metacognitive controller in the decision layer calculates document complexity based on the semantic feature vector and adaptively adjusts feature extraction parameters to obtain a cognitive feature representation of the document, including: A capsule network consisting of a main capsule layer and an auxiliary capsule layer is constructed, wherein the main capsule layer contains multiple parallel convolution capsules and the auxiliary capsule layer contains multiple attention capsules. The document to be processed is input into the capsule network, and the local feature patterns are captured and local feature vectors are extracted through parallel convolution capsules. The dynamic routing coefficients between local feature vectors are calculated through attention capsules. The coupling coefficients between feature vectors are iteratively calculated based on the dynamic routing coefficients. The local feature vectors are adaptively combined to generate an initial feature mapping matrix. A memory-enhanced neural network consisting of a long-term memory module and a working memory module is constructed. The initial feature mapping matrix is ​​input, and the query vector, key vector, and value vector are calculated using a multi-head attention mechanism. The attention weight is calculated based on the query vector and key vector. The working memory module updates the attention weight based on the historical feature state to obtain an updated attention weight. The memory attention mechanism is used to calculate the memory read weight of the feature mapping matrix and the pre-trained semantic knowledge base. The product of the updated attention weight and the value vector is fused with the product of the memory read weight and the semantic knowledge base to obtain a semantic feature vector. A metacognitive controller is constructed to calculate the document complexity index based on the semantic feature vector, including a weighted combination of the bi-norm and information entropy of the semantic feature vector. The convolution kernel size and the number of iterations of the dynamic routing coefficient are dynamically adjusted according to the document complexity index. The semantic feature vector is normalized and updated through a feedforward neural network to obtain a cognitive feature representation.

[0008] In an optional embodiment, the working memory module updates the attention weight based on the historical feature state, and obtaining the updated attention weight includes: The attention weight is input into the working memory module, and the difference of attention weights of adjacent time steps is calculated to obtain the differential state matrix. The forget gate parameters and update gate parameters are obtained through nonlinear transformation. Multiply the forget gate parameter by the historical feature state bit by bit to obtain the filtered historical feature state, multiply the update gate parameter by the attention weight of the previous time step to obtain the feature state to be updated; perform a weighted combination of the filtered historical feature state and the feature state to be updated to obtain the updated attention weight; The document complexity index is calculated based on the semantic feature vector, including the weighted combination of the bi-norm and information entropy of the semantic feature vector: The semantic feature vector is divided into local neighborhoods by K-nearest neighbor clustering, and the variance of the feature vector in each local neighborhood is calculated to obtain the vector distribution density; The cosine similarity of adjacent feature vectors is calculated based on the sliding window to obtain local correlation; The vector distribution density and local correlation are nonlinearly mapped to obtain the adaptive weight; the bi-norm and information entropy of the semantic feature vector are weightedly fused according to the adaptive weight to obtain the document complexity index.

[0009] In an optional embodiment, the cognitive feature representation is mapped to a non-Euclidean manifold space via a Riemannian metric tensor, the projection of the classification scheme on the tangent space is used as the particle position, and the projection of the update direction on the tangent plane is used as the particle velocity. Generating an initial particle swarm includes: Constructing a search neighborhood for cognitive feature representation, calculating the Euclidean distance matrix of sample pairs within the search neighborhood, obtaining a similarity matrix through kernel function transformation, performing a logarithmic transformation on the similarity matrix and obtaining a Hessian matrix to obtain a Riemannian metric tensor; The geodesic distance between sample pairs is calculated based on the Riemannian metric tensor, a local distance-preserving matrix is ​​constructed, and eigenvalue decomposition is performed to obtain the principal eigenvectors and eigenvalues. The local curvature index is calculated based on the distribution of the eigenvalues, and the local curvature of the non-Euclidean manifold space is adjusted through conformal transformation. A reference point is selected within the search neighborhood, and a tangent space centered at the reference point is constructed. The data points in the non-Euclidean manifold space are projected onto the tangent space through exponential mapping, and the geodesic distances before and after the projection are calculated to maintain the geometric structure. Calculating the local distance-preserving matrix condition number to determine the optimal segmentation threshold, dividing the tangent space into subspaces, calculating the singular value distribution of the local distance-preserving matrix, and dynamically adjusting the shape and size of the subspace; Analyze the spectral distribution of the local distance-preserving matrix and calculate the sampling density function of each subspace to determine the number of particle allocations; A tangent plane is constructed in each subspace, and the update direction of the classification scheme is calculated. The updated direction is projected onto the tangent plane through the Riemann metric tensor. Within the preset maximum velocity range, the velocity vector is obtained based on the projection sampling of the tangent plane. The position vector is determined by combining the sampling density function, and the initial particle swarm is formed by pairing.

[0010] In an optional embodiment, recording the particle search trajectory, calculating the adaptive inertia weight and dynamically adjusting the search step size, combining the learning factor and the random factor to iteratively update the optimal solution, determining the optimal classification boundary, and generating a multi-dimensional classification vector include: Record the position sequence of the initial particle swarm within the preset memory length to form a historical search trajectory, calculate the displacement vectors of adjacent moments in the historical search trajectory, and determine the overall displacement. The ratio of the overall displacement to the cumulative sum of displacements at adjacent moments is used as a motion trend indicator. Construct a quantum annealing control matrix, where each element represents the probability of quantum state superposition in the corresponding search dimension. Calculate the wave function collapse probability based on the number of iterations, calculate the mean of the motion trend index of all particles, and couple it with the wave function collapse probability to obtain the quantum control coefficient. Substitute the quantum control coefficient into the exponential decay function, calculate the adaptive inertia weight, multiply it with the particle's current velocity vector, determine the first difference vector with the individual optimal position, and the second difference vector with the global optimal position, and then perform a weighted combination to obtain the velocity update; Based on the speed update, the updated position vector is calculated, the corresponding classification accuracy and category overlap are determined, and the weighted combination is used to obtain the fitness value; When the fitness value increases for a preset number of consecutive times are less than a preset threshold, the position vector corresponding to the maximum fitness value is determined as the optimal classification boundary, and a multi-dimensional classification vector of the document to be processed is generated through linear transformation.

[0011] In an optional embodiment, an adaptive feedback optimization network is constructed to perform attention-weighted mapping on cognitive feature representations and multidimensional classification vectors to generate a comprehensive evaluation vector for the documents. A multidimensional similarity calculation model is used to obtain consistency evaluation scores between documents, including: It receives the document's cognitive feature representation and multidimensional classification vector and constructs a deep feature fusion layer containing local attention units and global attention units. The local attention unit calculates the correlation between the cognitive feature representation and the multidimensional classification vector using a multi-scale sliding window to obtain a feature component weight matrix. The global attention unit calculates cross-dimensional semantic dependencies based on a hierarchical semantic tree to obtain a feature combination pattern. The feature component weight matrix and feature combination pattern are input into a residual memory module with forget gate. The historical feature combination patterns are prioritized according to temporal correlation and selectively stored to guide the allocation of feature component weight matrices. A separable convolutional network with channel attention is used to perform adaptive feature recalibration to generate a comprehensive evaluation vector for the document. A feature projection layer is constructed to map the comprehensive evaluation vector to multiple semantic subspaces through orthogonal transformation. An adaptive kernel function network is constructed to calculate local similarity. Regularization constraints based on mutual information minimization are imposed to maintain the complementarity of semantic subspaces. Dynamic weighting based on confidence is used to integrate the local similarity scores of semantic subspaces to obtain a consistency evaluation score.

[0012] In an optional embodiment, when the evaluation score is less than a preset threshold, reinforcement learning is used to build a policy network, generate an optimized policy combination and provide feedback, and iterate repeatedly until the consistency evaluation score reaches the preset threshold. The output document classification results include: When the consistency evaluation score is less than the convergence threshold, a two-layer policy network with a memory enhancement mechanism is constructed. The outer network sets up an experience replay buffer pool to store historical optimization trajectories and select macro optimization directions based on the similarity of historical optimization trajectories. The inner network adopts a dual policy gradient architecture consisting of a value function network and a policy network. The value function is used to evaluate the current state value and calculate the advantage function. The policy network generates a parameter adjustment plan based on the advantage function and the comprehensive evaluation vector, generating an optimization strategy combination. The optimization strategy combination is optimized and adjusted through a bidirectional feedback path with a feedback strength adaptive adjustment mechanism, acting on the hierarchical cognitive attention network and the particle swarm optimization process respectively, and the updated parameters are input into the next round of iteration; An adaptive early stopping mechanism is set based on the consistency evaluation score. When the improvement in the consistency evaluation score over a preset number of consecutive iterations is less than the dynamically adjusted convergence threshold, a parameter rollback operation is triggered. The convergence threshold is adaptively adjusted based on the historical fluctuations of the consistency evaluation score. Repeat until the consistency evaluation score is greater than or equal to the convergence threshold, and output the current classification result as the final document classification.

[0013] In an optional embodiment, the optimization strategy combination is optimized and adjusted through a bidirectional feedback path with a feedback strength adaptive adjustment mechanism, acting on the hierarchical cognitive attention network and the particle swarm optimization process respectively, and the updated parameters are input into the next round of iteration, including: The first feedback path of the bidirectional feedback path acts on the hierarchical cognitive attention network, adjusts the feature component weight matrix based on the gradient sensitivity analysis, and obtains an updated feature component weight matrix; updates the feature extraction parameters based on the importance score of the comprehensive evaluation vector, and obtains the updated feature extraction parameters; optimizes the fusion ratio of the feature combination pattern based on the information gain of the consistency evaluation score, and obtains the updated feature combination pattern; The second feedback path of the bidirectional feedback path acts on the particle swarm optimization process, dynamically adjusts the search step size based on the diversity of the optimization strategy combination, and obtains an updated search step size parameter; adaptively updates the inertia weight based on the convergence speed of the consistency evaluation score, and obtains an updated inertia weight parameter; reconstructs the local optimal position based on the local landscape analysis of the comprehensive evaluation vector, and obtains an updated local optimal position; The updated feature component weight matrix, updated feature extraction parameters, updated feature combination mode, updated search step parameters, updated inertia weight parameters and updated local optimal position are input into the next round of iterative optimization.

[0014] According to a second aspect of the embodiments of the present invention, Provided is an enterprise intelligent document recognition and automatic classification and archiving management and processing system, including: The first unit is used to construct a hierarchical cognitive attention network consisting of a perception layer, a comprehension layer, and a decision layer. In the perception layer, a capsule network and a dynamic routing algorithm are used to extract features from the document being processed, construct feature space relationships, and obtain an initial feature map. In the comprehension layer, the long-term memory module and working memory module of the memory-augmented neural network are used to semantically process the initial feature map to obtain a semantic feature vector. In the decision layer, a metacognitive controller calculates document complexity based on the semantic feature vector and adaptively adjusts feature extraction parameters to obtain a cognitive feature representation of the document. The second unit is used to map the cognitive feature representation to a non-Euclidean manifold space via the Riemann metric tensor, generate an initial particle swarm using the projection of the classification scheme on the tangent space as the particle position and the projection of the update direction on the tangent plane as the particle velocity; record the particle search trajectory, calculate the adaptive inertia weight and dynamically adjust the search step size; iteratively update the optimal solution by combining the learning factor and the random factor, determine the optimal classification boundary, and generate a multidimensional classification vector; The third unit is used to build an adaptive feedback optimization network, perform attention-weighted mapping on cognitive feature representations and multidimensional classification vectors, generate a comprehensive evaluation vector for the document, and use a multidimensional similarity calculation model to obtain the consistency evaluation score between documents; when the evaluation score is less than the preset threshold, reinforcement learning is used to build a policy network, generate an optimized strategy combination and feedback, repeat the iteration until the consistency evaluation score reaches the preset threshold, and output the classification result of the document.

[0015] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0016] In an embodiment of the present invention, a capsule network and a dynamic routing algorithm are used for feature extraction to construct feature space relationships, which can better capture the local and overall features of a document. Compared with traditional feature extraction methods, it can more accurately express the semantic information of a document and improve the effectiveness of feature extraction. At the same time, a metacognitive controller is used to adaptively adjust feature extraction parameters according to document complexity, which can further improve the accuracy of feature extraction. Document features are mapped to a non-Euclidean manifold space based on the Riemannian metric, and the optimal classification boundary is determined using an improved particle swarm optimization algorithm. This can more accurately characterize the differences between documents and achieve accurate document classification. At the same time, multidimensional classification vectors are automatically generated, reducing the cost of manual intervention and achieving automated document classification. An adaptive feedback optimization network is constructed, which continuously optimizes the classification results through reinforcement learning methods and iteratively optimizes based on consistency evaluation scores, which can effectively improve the robustness of the classification results. At the same time, the network can adaptively adjust parameters according to different types of documents, enhancing the adaptability of document classification and enabling it to better handle various types of documents. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the process flow of the enterprise intelligent document recognition and automatic classification and archiving management processing method according to an embodiment of the present invention; Figure 2 This is a comparison chart of the accuracy of document complexity calculation methods; Figure 3 This is the effect diagram of adaptive search step size optimization. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0019] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0020] Figure 1 FIG. 1 is a flow chart of an enterprise intelligent document recognition and automatic classification and archiving management processing method according to an embodiment of the present invention. Figure 1 As shown, the method includes: A hierarchical cognitive attention network consisting of a perception layer, a comprehension layer, and a decision-making layer is constructed. At the perception layer, a capsule network and a dynamic routing algorithm are used to extract features from the document being processed, construct feature space relationships, and obtain an initial feature map. At the comprehension layer, the long-term memory module and working memory module of the memory-augmented neural network are used to semantically process the initial feature map to obtain a semantic feature vector. At the decision layer, a metacognitive controller calculates document complexity based on the semantic feature vector and adaptively adjusts feature extraction parameters to obtain a cognitive feature representation of the document. The cognitive feature representation is mapped to a non-Euclidean manifold space via the Riemann metric tensor. The projection of the classification scheme on the tangent space is used as the particle position, and the projection of the update direction on the tangent plane is used as the particle velocity to generate an initial particle swarm. The particle search trajectory is recorded, the adaptive inertia weight is calculated, and the search step size is dynamically adjusted. The optimal solution is iteratively updated by combining the learning factor and the random factor, the optimal classification boundary is determined, and a multi-dimensional classification vector is generated. An adaptive feedback optimization network is constructed to perform attention-weighted mapping on cognitive feature representations and multidimensional classification vectors to generate a comprehensive evaluation vector for the document. A multidimensional similarity calculation model is used to obtain the consistency evaluation score between documents. When the evaluation score is less than the preset threshold, reinforcement learning is used to construct a strategy network to generate an optimized strategy combination and provide feedback. The iteration is repeated until the consistency evaluation score reaches the preset threshold, and the classification result of the document is output.

[0021] In an optional embodiment, a hierarchical cognitive attention network is constructed, comprising a perception layer, a comprehension layer, and a decision layer. A capsule network and a dynamic routing algorithm are used in the perception layer to extract features from the document to be processed, construct feature space relationships, and obtain an initial feature map. The comprehension layer performs semantic processing on the initial feature map using the long-term memory module and the working memory module of the memory-augmented neural network to obtain a semantic feature vector. A metacognitive controller in the decision layer calculates document complexity based on the semantic feature vector and adaptively adjusts feature extraction parameters to obtain a cognitive feature representation of the document, including: A capsule network consisting of a main capsule layer and an auxiliary capsule layer is constructed, wherein the main capsule layer contains multiple parallel convolution capsules and the auxiliary capsule layer contains multiple attention capsules. The document to be processed is input into the capsule network, and the local feature patterns are captured and local feature vectors are extracted through parallel convolution capsules. The dynamic routing coefficients between local feature vectors are calculated through attention capsules. The coupling coefficients between feature vectors are iteratively calculated based on the dynamic routing coefficients. The local feature vectors are adaptively combined to generate an initial feature mapping matrix. A memory-enhanced neural network consisting of a long-term memory module and a working memory module is constructed. The initial feature mapping matrix is ​​input, and the query vector, key vector, and value vector are calculated using a multi-head attention mechanism. The attention weight is calculated based on the query vector and key vector. The working memory module updates the attention weight based on the historical feature state to obtain an updated attention weight. The memory attention mechanism is used to calculate the memory read weight of the feature mapping matrix and the pre-trained semantic knowledge base. The product of the updated attention weight and the value vector is fused with the product of the memory read weight and the semantic knowledge base to obtain a semantic feature vector. A metacognitive controller is constructed to calculate the document complexity index based on the semantic feature vector, including a weighted combination of the bi-norm and information entropy of the semantic feature vector. The convolution kernel size and the number of iterations of the dynamic routing coefficient are dynamically adjusted according to the document complexity index. The semantic feature vector is normalized and updated through a feedforward neural network to obtain a cognitive feature representation.

[0022] In one specific embodiment, a capsule network is constructed in the perception layer, comprising a primary capsule layer and a secondary capsule layer. The primary capsule layer consists of multiple parallel convolutional capsules, each responsible for capturing different local feature patterns in the document. For example, assuming a document consists of a series of words, a convolutional capsule can use a convolution kernel of size 3 to extract local features consisting of three consecutive words. The secondary capsule layer consists of multiple attention capsules, which are used to calculate the relationships between local feature vectors. Assuming the primary capsule layer outputs 10 local feature vectors, each attention capsule calculates a dynamic routing coefficient between these 10 vectors. The dynamic routing algorithm is an iterative process that continuously updates the coupling coefficient to ultimately determine which local feature vectors should be combined to form a higher-level feature representation. For example, if the terms "artificial intelligence" and "machine learning" appear in a document, the coupling coefficients between their corresponding local feature vectors will be relatively high, and they will ultimately be combined. The dynamic routing algorithm adaptively combines the local feature vectors to generate the document's initial feature map matrix, which reflects the spatial relationships between different local features in the document. Assume that the dimension of the initial feature map matrix is ​​10×256, which means it consists of 10 local feature vectors, each with a dimension of 256.

[0023] Next, a memory-augmented neural network is constructed in the understanding layer. This network consists of a long-term memory module and a working memory module. The initial feature map matrix obtained from the perception layer is input into this network. The network first calculates the query vector, key vector, and value vector using a multi-head attention mechanism. The multi-head attention mechanism can focus on different parts of the input in parallel. Then, attention weights are calculated based on the query vector and key vector, reflecting the importance of different features. The working memory module updates the attention weights based on the historical feature states to capture the document's context. For example, if a keyword appears repeatedly in a document, the working memory module increases the attention weight of the feature corresponding to that keyword. Simultaneously, the memory attention mechanism calculates the memory read weights for the initial feature map matrix and a pre-trained semantic knowledge base. The semantic knowledge base can be a knowledge graph containing a large number of words and concepts. The memory read weights reflect the knowledge base content that is relevant to the document content. Finally, the updated attention weights are multiplied by the value vector and then fused with the memory read weights and the semantic knowledge base to generate a semantic feature vector. This vector incorporates the document's local features, contextual information, and background knowledge, providing a more comprehensive representation of the document's semantics. Assume that the dimension of the semantic feature vector is 512.

[0024] A metacognitive controller is constructed at the decision layer. This controller calculates a document complexity index based on the semantic feature vector output by the understanding layer. The document complexity index can be a weighted combination of the bi-norm and information entropy of the semantic feature vector. For example, a larger bi-norm indicates a larger document contains more information, while a higher information entropy indicates a more dispersed document. Based on the document complexity index, the convolution kernel size and the number of iterations of the dynamic routing coefficients in the perception layer are dynamically adjusted. For example, if the document complexity is high, the convolution kernel size can be increased to capture longer-term contextual information, and the number of dynamic routing iterations can be increased to more finely combine local features. Finally, the semantic feature vector is normalized and updated using a feedforward neural network to obtain the final document cognitive feature representation. This feature representation integrates both semantic and complexity information of the document and can be used for various downstream tasks, such as document classification and information retrieval. Assume that the dimensionality of the final document cognitive feature representation is also 512.

[0025] In this embodiment, by combining capsule networks, memory-enhanced neural networks and metacognitive control mechanisms, the method can more effectively capture the deep semantic information and complexity of documents, thereby improving the accuracy of document feature representation; the dynamic routing algorithm and memory-enhanced mechanism can effectively handle noise and ambiguity in documents, thereby enhancing the robustness of the model; the metacognitive controller can adaptively adjust feature extraction parameters according to the complexity of the document, thereby achieving more accurate document understanding and analysis.

[0026] In an optional embodiment, the working memory module updates the attention weight based on the historical feature state, and the updated attention weight includes: The attention weight is input into the working memory module, and the difference of attention weights of adjacent time steps is calculated to obtain the differential state matrix. The forget gate parameters and update gate parameters are obtained through nonlinear transformation. Multiply the forget gate parameter by the historical feature state bit by bit to obtain the filtered historical feature state, multiply the update gate parameter by the attention weight of the previous time step to obtain the feature state to be updated; perform a weighted combination of the filtered historical feature state and the feature state to be updated to obtain the updated attention weight; The document complexity index is calculated based on the semantic feature vector, including the weighted combination of the bi-norm and information entropy of the semantic feature vector: The semantic feature vector is divided into local neighborhoods by K-nearest neighbor clustering, and the variance of the feature vector in each local neighborhood is calculated to obtain the vector distribution density; The cosine similarity of adjacent feature vectors is calculated based on the sliding window to obtain local correlation; The vector distribution density and local correlation are nonlinearly mapped to obtain the adaptive weight; the bi-norm and information entropy of the semantic feature vector are weightedly fused according to the adaptive weight to obtain the document complexity index.

[0027] In one specific embodiment, a semantic feature vector of a text is obtained. For example, a pre-trained language model (such as BERT or RoBERT) can be used to encode the text, obtaining a vector representation for each word or subword. A pooling operation (such as average pooling or max pooling) is then performed to obtain a semantic feature vector for the entire text. Assume that the semantic feature vector of a text is [0.2, 0.5, 0.1, 0.8, 0.4].

[0028] Next, the attention weights are fed into the working memory module for state update. The initial attention weights can be set to a uniform distribution. For example, for a vector containing five features, the initial attention weights can be set to [0.2, 0.2, 0.2, 0.2]. Assume that the attention weights at the previous time step are [0.1, 0.3, 0.2, 0.3, 0.1]. Calculate the difference between the attention weights at the current time step and the previous time step to obtain the differential state matrix. For example, subtract each element to obtain [0.1, -0.1, 0.0, -0.1, 0.1]. Then, perform a nonlinear transformation on the differential state matrix, such as using a sigmoid function, to obtain the forget gate parameters and update gate parameters. Assume that the forget gate parameters are [0.55, 0.45, 0.5, 0.45, 0.55] and the update gate parameters are [0.55, 0.45, 0.5, 0.45, 0.55]. Multiply the forget gate parameter by the historical feature state (i.e., the attention weight at the previous time step) bitwise to obtain the filtered historical feature state, for example, [0.055, 0.135, 0.1, 0.135, 0.055]. Multiply the update gate parameter by the attention weight at the current time step bitwise to obtain the feature state to be updated, for example, [0.11, 0.09, 0.1, 0.09, 0.11]. Finally, perform a weighted combination of the filtered historical feature state and the feature state to be updated, for example, adding them together to obtain the updated attention weights [0.165, 0.225, 0.2, 0.225, 0.165].

[0029] The document complexity index is calculated based on the semantic feature vector. The semantic feature vector is divided into local neighborhoods using the K-nearest neighbor clustering method. Assuming the K value is 3, the feature vector and its three most similar vectors form a local neighborhood. The variance of the feature vectors in each local neighborhood is calculated to obtain the vector distribution density. For example, the variance of a local neighborhood is 0.02. Next, the cosine similarity of adjacent feature vectors is calculated based on a sliding window to obtain the local correlation. For example, the cosine similarity of two adjacent feature vectors is 0.8. The vector distribution density and local correlation are nonlinearly mapped, for example, using an exponential function, to obtain an adaptive weight. Assume that the adaptive weight is 0.6. Finally, the bi-norm and information entropy of the semantic feature vector are weighted and fused based on the adaptive weight to obtain the document complexity index. For example, if the bi-norm of the semantic feature vector is 1.1 and the information entropy is 2.2, the document complexity index is 0.6×1.1+(1-0.6)×2.2=1.54.

[0030] The updated attention weights are weighted and fused with the semantic feature vector to obtain the final text representation. For example, the updated attention weights [0.165, 0.225, 0.2, 0.225, 0.165] are bitwise multiplied and summed with the semantic feature vector [0.2, 0.5, 0.1, 0.8, 0.4] to obtain a final text representation value of 0.438.

[0031] In this embodiment, by dynamically adjusting the attention weight through the working memory module, the key information of the text can be captured more accurately, thereby improving the accuracy of text representation; the introduction of the document complexity index can enable the model to better adapt to different types of text and enhance the robustness of the model; more accurate and robust text representation can effectively improve the performance of downstream tasks (such as text classification, sentiment analysis, etc.).

[0032] like Figure 2As shown in the figure, the accuracy comparison of different document complexity calculation methods on various types of documents is demonstrated. This technical solution has shown significant advantages in all test document fields. In the news article category, this technical solution achieved the highest accuracy of 95.8%, which is 13.4 percentage points higher than the 82.4% of the traditional TF-IDF method. Especially in the legal text, a recognized complex document type, this technical solution still maintained an accuracy of 89.5%, while the single entropy measurement method was only 80.6%, demonstrating the powerful ability of this technical solution in processing highly specialized and complex structural documents. This technical solution divides the local neighborhood of the semantic feature vector through K-nearest neighbor clustering, and combines the vector distribution density with local correlation for adaptive weight calculation, so that it achieves an accuracy of 94.2% on educational material documents, which is much higher than the 84.1% of the word vector averaging method. Judging from the average value, this technical solution achieved an accuracy of 92.35%, which is 8.1 percentage points higher than the second-place single entropy measurement method and 16.13 percentage points higher than the weakest traditional TF-IDF method, fully demonstrating the versatility and reliability of this technical solution in document complexity assessment in different fields.

[0033] In an optional embodiment, the cognitive feature representation is mapped to a non-Euclidean manifold space via a Riemannian metric tensor, the projection of the classification scheme on the tangent space is used as the particle position, and the projection of the update direction on the tangent plane is used as the particle velocity. Generating an initial particle swarm includes: Constructing a search neighborhood for cognitive feature representation, calculating the Euclidean distance matrix of sample pairs within the search neighborhood, obtaining a similarity matrix through kernel function transformation, performing a logarithmic transformation on the similarity matrix and obtaining a Hessian matrix to obtain a Riemannian metric tensor; The geodesic distance between sample pairs is calculated based on the Riemannian metric tensor, a local distance-preserving matrix is ​​constructed, and eigenvalue decomposition is performed to obtain the principal eigenvectors and eigenvalues. The local curvature index is calculated based on the distribution of the eigenvalues, and the local curvature of the non-Euclidean manifold space is adjusted through conformal transformation. A reference point is selected within the search neighborhood, and a tangent space centered at the reference point is constructed. The data points in the non-Euclidean manifold space are projected onto the tangent space through exponential mapping, and the geodesic distances before and after the projection are calculated to maintain the geometric structure. Calculating the local distance-preserving matrix condition number to determine the optimal segmentation threshold, dividing the tangent space into subspaces, calculating the singular value distribution of the local distance-preserving matrix, and dynamically adjusting the shape and size of the subspace; Analyze the spectral distribution of the local distance-preserving matrix and calculate the sampling density function of each subspace to determine the number of particle allocations; A tangent plane is constructed in each subspace, and the update direction of the classification scheme is calculated. The updated direction is projected onto the tangent plane through the Riemann metric tensor. Within the preset maximum velocity range, the velocity vector is obtained based on the projection sampling of the tangent plane. The position vector is determined by combining the sampling density function, and the initial particle swarm is formed by pairing.

[0034] The Hessian matrix is ​​a matrix that describes the behavior of a function near a point, capturing the function's second-order derivatives. It reflects the changing curvature of a function and can be used to analyze local geometric properties and determine the nature of extreme points during optimization.

[0035] The geodesic distance is the distance between two points on a surface or manifold, measured along the shortest path along the surface or manifold. It is a distance measurement based on the intrinsic geometric properties of the surface. Unlike straight-line distance in Euclidean space, the geodesic distance is applicable to non-flat geometries.

[0036] The conformal transformation specifically refers to a geometric transformation that preserves angles but may change distance ratios. Through this transformation, local shape characteristics can remain unchanged, while the overall curvature and ratio may be dynamically adjusted to meet specific geometric requirements.

[0037] The tangent space is a local linear space around a point on a manifold, used to approximate the local properties of the manifold near that point. It contains all the tangent directions at that point and provides the basis for linear algebra operations on the manifold.

[0038] The Riemann metric tensor is a mathematical tool used to define distances and angles on a manifold. It gives the manifold its geometric structure by providing the inner product of the tangent space at each point, helping to describe properties such as geodesics, volume, and curvature.

[0039] The tangent plane is a specific representation of the tangent space, which is used to approximate the local flat structure of the manifold near a certain point. Operations performed on the tangent plane are usually performed to simplify calculations on complex manifolds, such as converting nonlinear problems into linear problems.

[0040] In one embodiment, a search neighborhood for cognitive feature representation is constructed. For example, assuming there are 100 samples, each with 5 features, we can define a search neighborhood where the neighbors of each sample are the 10 samples with the closest Euclidean distance to it.

[0041] Calculate the Euclidean distance matrix of sample pairs in the search neighborhood. Taking the above example, for each sample, calculate the Euclidean distance between it and its 10 neighbors, and obtain a 100×10 distance value.

[0042] Applying a kernel function to the Euclidean distance matrix yields a similarity matrix. For example, a Gaussian kernel function can be used to convert Euclidean distances to similarities. The parameters of the Gaussian kernel function can be adjusted based on the data distribution; for example, the parameters can be set to the inverse of the mean distance. For example, if the Euclidean distances between a sample and its neighbors are 1, 2, 3, ..., 10, respectively, after applying the Gaussian kernel function, the resulting similarity values ​​might be 0.9, 0.8, 0.7, ..., 0.0, respectively. On this basis, the similarity matrix is ​​logarithmically transformed and the Hessian matrix is ​​calculated to obtain the Riemann metric tensor. The logarithmic transformation can enhance the nonlinear characteristics of the similarity matrix. The Hessian matrix describes the second-order derivative of the similarity matrix and can be used to measure the rate of change of data in different directions.

[0043] The geodesic distance between pairs of samples is calculated based on the Riemannian metric tensor, and a local distance-preserving matrix is ​​constructed. The geodesic distance is the shortest distance between two points measured along the surface of a manifold. For example, the geodesic distance between two points on the Earth's surface is the length of a great circle arc.

[0044] Perform eigenvalue decomposition on the local distance-preserving matrix to obtain the principal eigenvectors and eigenvalues. For example, after eigenvalue decomposition of a 3×3 matrix, three eigenvalues ​​and corresponding three eigenvectors can be obtained.

[0045] The local curvature index is calculated based on the distribution of eigenvalues. Eigenvalues ​​reflect the degree of variation in the data across different directions. For example, if the eigenvalues ​​are relatively close, the data varies similarly across all directions, resulting in low curvature. Large differences in the eigenvalues ​​indicate that the data varies more significantly in certain directions, resulting in high curvature. Assuming the eigenvalues ​​are 1, 2, and 3, the local curvature index can be defined as the variance of the eigenvalues.

[0046] The local curvature of a non-Euclidean manifold space is dynamically adjusted through conformal transformations based on the local curvature index. Conformal transformations can preserve angles but change the area or length of a local region. For example, if the local curvature index is large, a conformal transformation can reduce the area of ​​that region, thereby reducing the local curvature.

[0047] A reference point is selected within the search neighborhood and a tangent space centered at the reference point is constructed. The tangent space is a linear space tangent to the manifold at the reference point.

[0048] Data points in a non-Euclidean manifold space are projected onto the tangent space through the exponential mapping. The exponential mapping can map vectors in the tangent space onto the manifold.

[0049] Computing the geodesic distance before and after the projection ensures that the geometry remains unchanged. For example, if the geodesic distance between two points before projection is 1, then the Euclidean distance between the two points after projection should also be 1.

[0050] Calculate the condition number of the local distance-preserving matrix and determine the optimal segmentation threshold. The condition number is the ratio of the largest eigenvalue to the smallest eigenvalue of a matrix and can be used to measure the stability of the matrix.

[0051] The tangent space is divided into subspaces based on the optimal segmentation threshold. For example, the tangent space can be divided into multiple subspaces based on the feature vector.

[0052] Calculate the singular value distribution of the local distance-preserving matrix and dynamically adjust the shape and size of the subspace based on the singular value distribution. The singular values ​​can reflect the degree of expansion and contraction of the matrix in different directions.

[0053] The spectral distribution characteristics of the local distance-preserving matrix are analyzed, and the sampling density function of each subspace is calculated based on the spectral distribution characteristics. The spectral distribution characteristics refer to the distribution of eigenvalues ​​or singular values.

[0054] The number of particles allocated to each subspace is determined based on the sampling density function. For example, if the sampling density function value of a subspace is larger, more particles can be allocated to the subspace.

[0055] Construct a tangent plane in each subspace. A tangent plane is a subspace of the tangent space.

[0056] Calculate the update direction of the classification scheme and project the update direction onto the tangent plane through the Riemann metric tensor. For example, if the update direction is (1,1) and the Riemann metric tensor is the identity matrix, then the projected direction is still (1,1).

[0057] The velocity vector is obtained by sampling based on the tangent plane projection within a preset maximum velocity range. For example, if the maximum velocity is 1, the velocity vector can be randomly sampled within the unit circle.

[0058] Importance sampling is performed in conjunction with a sampling density function to determine the position vector. Importance sampling can weight samples according to the sampling density function, thereby improving sampling efficiency.

[0059] Pair the position vector and velocity vector to form an initial particle group. For example, a particle can be represented as (position vector, velocity vector).

[0060] In this embodiment, by mapping the data to a non-Euclidean manifold space, the nonlinear structure of the data can be better captured, thereby improving the classification accuracy; by dynamically adjusting the local curvature and the shape and size of the subspace, it can adapt to different types of data and improve the robustness of the model; through importance sampling and particle swarm optimization algorithms, the convergence speed of the model can be accelerated.

[0061] In an optional embodiment, recording the particle search trajectory, calculating the adaptive inertia weight and dynamically adjusting the search step size, iteratively updating the optimal solution by combining the learning factor and the random factor, determining the optimal classification boundary, and generating a multidimensional classification vector include: Record the position sequence of the initial particle swarm within the preset memory length to form a historical search trajectory, calculate the displacement vectors of adjacent moments in the historical search trajectory, and determine the overall displacement. The ratio of the overall displacement to the cumulative sum of displacements at adjacent moments is used as a motion trend indicator. Construct a quantum annealing control matrix, where each element represents the probability of quantum state superposition in the corresponding search dimension. Calculate the wave function collapse probability based on the number of iterations, calculate the mean of the motion trend index of all particles, and couple it with the wave function collapse probability to obtain the quantum control coefficient. Substitute the quantum control coefficient into the exponential decay function, calculate the adaptive inertia weight, multiply it with the particle's current velocity vector, determine the first difference vector with the individual optimal position, and the second difference vector with the global optimal position, and then perform a weighted combination to obtain the velocity update; Based on the speed update, the updated position vector is calculated, the corresponding classification accuracy and category overlap are determined, and the weighted combination is used to obtain the fitness value; When the fitness value increases for a preset number of consecutive times are less than a preset threshold, the position vector corresponding to the maximum fitness value is determined as the optimal classification boundary, and a multi-dimensional classification vector of the document to be processed is generated through linear transformation.

[0062] The complementary value specifically refers to the relationship between the degree of overlap between categories during the classification process and their complementarity. Specifically, the category overlap measures the degree of overlap between features in different categories. The higher the overlap, the worse the classification effect. The complementary value, on the other hand, represents the opposite degree of overlap, that is, the degree of separation between categories. The higher the complementary value, the better the distinction between categories, and the higher the classification accuracy. In the algorithm, the complementary value is used as an optimization metric, combined with the classification accuracy, to comprehensively evaluate the classification effect. It emphasizes the importance of reducing overlap between categories, thereby improving overall classification performance.

[0063] In one specific embodiment, a particle swarm is initialized, with each particle representing a candidate solution for a document classification boundary. Each particle has a position vector, representing a point in multidimensional space and a potential classification boundary. A preset memory length is set for each particle to record its historical search trajectory. For example, if the memory length is set to 5, the position information of each particle for the past five moments is recorded.

[0064] Calculate the adaptive inertia weight for each particle. Calculate the displacement vector of each particle within the memory length. For example, if the memory length is 5, calculate the position difference between each two adjacent moments in the past 5 moments, obtaining 4 displacement vectors. Then, calculate the cumulative sum of these 4 displacement vectors to represent the overall displacement of the particle. Simultaneously, calculate the modulus of each displacement vector and add these moduli together to represent the cumulative sum of the displacements at adjacent moments. The ratio of the overall displacement to the cumulative sum of the displacements at adjacent moments is used as the motion trend indicator for the particle. Calculate the average motion trend indicator for all particles.

[0065] Construct a quantum annealing control matrix with the same dimensions as the document's feature dimensions. Each element in the matrix represents the probability of a quantum state superposition in the corresponding dimension, and the initial value can be set to a uniform distribution. Calculate the probability of wave function collapse based on the current number of iterations. For example, you can use an increasing function so that the probability of wave function collapse gradually increases with the number of iterations. Multiply the mean of the motion trend indicators of all particles by the probability of wave function collapse to obtain the quantum control coefficient. Substitute the quantum control coefficient into the exponential decay function to calculate the adaptive inertia weight. For example, use a negative exponential function with the quantum control coefficient as the exponent so that the inertia weight decreases as the quantum control coefficient increases.

[0066] Update the position and velocity of each particle. Multiply the calculated adaptive inertia weight by the particle's current velocity vector. Simultaneously, calculate the difference vector between the particle's current position and its individual historical optimal position, as well as the difference vector between the particle's current position and its global optimal position. Weight these two difference vectors using a preset learning factor and a random factor, and add them to the inertia weight term to obtain the particle's updated velocity. Add the velocity update to the particle's current position vector to obtain the updated position vector.

[0067] Based on the updated position vector, the fitness value of each particle is calculated. The updated position vector is used as the classification boundary to calculate the accuracy of the document classification. The overlap between different categories is also calculated, for example, using the distance between category centers. The classification accuracy and the complementary value of category overlap (for example, 1-overlap) are weighted together to obtain the particle's fitness value.

[0068] Determine whether the termination condition has been met. If the fitness value increases for a preset number of consecutive times (e.g., 10 times) are less than a preset threshold (e.g., 0.001), the iteration is terminated and the position vector with the maximum fitness value is determined as the optimal classification boundary. Otherwise, the iteration continues to update the particle position and velocity.

[0069] The optimal classification boundary is linearly transformed and processed through a normalization function (e.g., a sigmoid function) to generate a multidimensional classification vector for the document to be processed. For example, the hyperplane parameters represented by the optimal classification boundary can be used as the initial values ​​of the multidimensional classification vector, which is then mapped to a specific range through linear transformation and normalization.

[0070] For example, assume there are two categories, a document feature dimension of 2, a particle swarm size of 3, and a memory length of 2. The initial positions are (1, 1), (2, 2), and (3, 3), respectively, with initial velocities of (0, 0). After the first iteration, the positions are updated to (1.5, 1.5), (2.5, 2.5), and (3.5, 3.5). Assuming the first particle has the highest fitness value, and after 10 consecutive iterations, the fitness value increases by less than 0.001, the boundary defined by (1.5, 1.5) is linearly transformed and normalized to generate a multidimensional classification vector for the document.

[0071] In this embodiment, the adaptive inertia weight and quantum annealing mechanism can more effectively search for the optimal classification boundary, thereby improving the accuracy of document classification; by optimizing the classification boundary, the overlap between different categories can be effectively reduced, making the classification results clearer; by combining the advantages of particle swarm optimization and quantum annealing, the robustness and global search capability of the algorithm are improved, and it is not easy to fall into the local optimal solution.

[0072] In an optional embodiment, an adaptive feedback optimization network is constructed to perform attention-weighted mapping on cognitive feature representations and multidimensional classification vectors to generate a comprehensive evaluation vector for the document. A multidimensional similarity calculation model is used to obtain consistency evaluation scores between documents, including: It receives the document's cognitive feature representation and multidimensional classification vector and constructs a deep feature fusion layer containing local attention units and global attention units. The local attention unit calculates the correlation between the cognitive feature representation and the multidimensional classification vector using a multi-scale sliding window to obtain a feature component weight matrix. The global attention unit calculates cross-dimensional semantic dependencies based on a hierarchical semantic tree to obtain a feature combination pattern. The feature component weight matrix and feature combination pattern are input into a residual memory module with forget gate. The historical feature combination patterns are prioritized according to temporal correlation and selectively stored to guide the allocation of feature component weight matrices. A separable convolutional network with channel attention is used to perform adaptive feature recalibration to generate a comprehensive evaluation vector for the document. A feature projection layer is constructed to map the comprehensive evaluation vector to multiple semantic subspaces through orthogonal transformation. An adaptive kernel function network is constructed to calculate local similarity. Regularization constraints based on mutual information minimization are imposed to maintain the complementarity of semantic subspaces. Dynamic weighting based on confidence is used to integrate the local similarity scores of semantic subspaces to obtain a consistency evaluation score.

[0073] In a specific embodiment, a cognitive feature representation and a multidimensional classification vector of a document to be evaluated are received. The cognitive feature representation can be a sequence of word vectors of the document, such as the vector representation of each word obtained using a pre-trained BERT model. The multidimensional classification vector can be the classification probability distribution of a document under a predefined category system, such as the probability that a document belongs to categories such as "sports", "entertainment", and "technology". Assume that the cognitive feature representation of document A is a series of 768-dimensional vectors with a length of 512, indicating that there are 512 words in the document, and each word is represented by a 768-dimensional vector. The multidimensional classification vector of document A is a 10-dimensional vector, which respectively represents the probability that the document belongs to 10 different categories.

[0074] A deep feature fusion layer is constructed, consisting of local and global attention units. The local attention unit uses a multi-scale sliding window mechanism to calculate the correlation between cognitive feature representations and multidimensional classification vectors. For example, using sliding windows of sizes 3, 5, and 7, the correlation between the word vectors within the window and the multidimensional classification vector is calculated, resulting in multiple feature component weight matrices. The global attention unit constructs a hierarchical semantic tree at the document level. For example, it decomposes documents into different levels, such as paragraphs, sentences, and phrases, and calculates cross-dimensional semantic dependencies to obtain feature combination patterns. For example, a feature combination pattern indicates a high correlation between the category "sports" and the phrase "basketball." The feature component weight matrix and feature combination patterns are input into a residual memory module with a forget gate mechanism. This module prioritizes historical feature combination patterns based on temporal correlation and selectively stores them to guide the allocation of feature component weight matrices. For example, if the category "sports" and the phrase "basketball" frequently co-occur in previous documents, a higher weight will be assigned to similar combinations in the current document. A separable convolutional network with a channel-wise attention mechanism then adaptively recalibrates the feature component weight matrices and feature combination patterns, integrating them to generate a comprehensive evaluation vector for the document. For example, the convolutional network can learn which feature combination patterns are more important and adjust the corresponding feature component weights, ultimately generating a fixed-length vector, such as 512 dimensions, as the comprehensive evaluation vector for the document.

[0075] A feature projection layer is constructed to map the comprehensive evaluation vector into multiple semantic subspaces through an orthogonal transformation. For example, a 512-dimensional comprehensive evaluation vector is mapped into four 128-dimensional semantic subspaces. An adaptive kernel function network is constructed in each semantic subspace to calculate local similarity. Regularization constraints based on mutual information minimization are imposed on the semantic subspaces to maintain their complementarity. For example, by minimizing the mutual information between different semantic subspaces, different semantic information is captured. A confidence-based dynamic weighting is used to combine the local similarity scores of multiple semantic subspaces to obtain the final consistency evaluation score. For example, if the confidence level of the local similarity score of a semantic subspace is higher, it is given a higher weight. Assuming that the local similarity scores of documents A and B in four semantic subspaces are 0.8, 0.6, 0.9, and 0.7, respectively, and the confidence levels are 0.9, 0.7, 0.8, and 0.6, respectively, the final consistency evaluation score can be calculated as a weighted average.

[0076] In this embodiment, through the modeling of deep feature fusion layers and multi-semantic subspaces, the semantic information of the document can be captured more comprehensively, thereby improving the accuracy of consistency assessment; the forgetting gating mechanism and adaptive feature recalibration strategy enable the model to better adapt to different types of documents and topics, thereby enhancing the robustness of the model; the application of technologies such as separable convolutional networks and orthogonal transformations reduces computational complexity and improves computational efficiency while ensuring model performance.

[0077] In an optional embodiment, when the evaluation score is less than a preset threshold, reinforcement learning is used to build a policy network, generate an optimized policy combination and provide feedback, and iterate repeatedly until the consistency evaluation score reaches the preset threshold. The output document classification results include: When the consistency evaluation score is less than the convergence threshold, a two-layer policy network with a memory enhancement mechanism is constructed. The outer network sets up an experience replay buffer pool to store historical optimization trajectories and select macro optimization directions based on the similarity of historical optimization trajectories. The inner network adopts a dual policy gradient architecture consisting of a value function network and a policy network. The value function is used to evaluate the current state value and calculate the advantage function. The policy network generates a parameter adjustment plan based on the advantage function and the comprehensive evaluation vector, generating an optimization strategy combination. The optimization strategy combination is optimized and adjusted through a bidirectional feedback path with a feedback strength adaptive adjustment mechanism, acting on the hierarchical cognitive attention network and the particle swarm optimization process respectively, and the updated parameters are input into the next round of iteration; An adaptive early stopping mechanism is set based on the consistency evaluation score. When the improvement in the consistency evaluation score over a preset number of consecutive iterations is less than the dynamically adjusted convergence threshold, a parameter rollback operation is triggered. The convergence threshold is adaptively adjusted based on the historical fluctuations of the consistency evaluation score. Repeat until the consistency evaluation score is greater than or equal to the convergence threshold, and output the current classification result as the final document classification.

[0078] In a specific embodiment, when the consistency evaluation score of the document classification is lower than the preset threshold, it is necessary to construct a two-layer strategy network with a memory enhancement mechanism. The two-layer strategy network includes an outer network and an inner network. The outer network sets an experience replay buffer pool to store historical optimization trajectories. The buffer pool capacity of the outer network is set to 500 records, each record contains the document feature vector before the operation, the optimization operation performed, and the consistency evaluation score after the operation. When an optimization decision is needed, the system retrieves historical records with a similarity of more than 0.75 to the current document features from the buffer pool, calculates the average benefit of each optimization direction, and selects the macro optimization direction with the highest historical benefit. For example, for a financial report document containing complex tables, the system retrieves the optimization history of similar documents and shows that adjusting the text feature extraction weight can achieve a higher consistency improvement than adjusting the structural feature extraction weight. Based on this, the system tends to optimize in terms of text feature extraction.

[0079] The inner network adopts a dual policy gradient architecture, consisting of a value function network and a policy network. The value function network consists of a three-layer fully connected neural network. The input layer has a node count equal to the document feature dimension (typically 256 dimensions), the hidden layer has 128 nodes, and the output layer has a single node, representing the estimated value of the current state. The policy network is a four-layer neural network. The input layer receives the document feature vector and the consistency evaluation vector, and the hidden layer has 192 nodes and 96 nodes, respectively. The output layer dimension matches the number of adjustable parameters. In practice, adjustable parameters include attention weights (64 parameters), feature extraction thresholds (8 parameters), and classification decision boundaries (number of categories × 2 parameters). The value function network evaluates the current state and calculates an advantage function, which is calculated by subtracting the baseline value from the current state value. This function represents the advantage of the current state relative to the average state. Based on the advantage function and the comprehensive evaluation vector, the policy network generates parameter adjustment plans, with the adjustment magnitude proportional to the advantage function value. For example, when processing a corporate rules and regulations document, if the system detects that the advantage function value of the text keyword extraction parameter is 0.42 (higher than the average 0.15), it will generate an adjustment plan to significantly increase the text keyword weight.

[0080] After the optimization strategy combination is generated, it is optimized and adjusted through a bidirectional feedback pathway with an adaptive feedback strength adjustment mechanism. The bidirectional feedback pathway connects the hierarchical cognitive attention network and the particle swarm optimization process. The feedback strength is determined by an initial baseline value of 0.3 and an adaptive adjustment coefficient. The adaptive adjustment coefficient is calculated based on the change in consistency assessment scores over the past five iterations. When the rate of change is less than 0.05, the feedback strength is increased by 1.5 times, and when the rate of change is greater than 0.15, the feedback strength is reduced by 0.8 times. In the hierarchical cognitive attention network, feedback adjustments are applied to the attention weights of each layer. The adjustment amplitude is proportional to the corresponding parameter value output by the policy network. The adjusted weights are renormalized to ensure that they sum to 1. For example, for a document containing minutes of a multi-department meeting, the policy network may output an adjustment parameter indicating an increase in the attention weight of the organizational structure. The feedback pathway adjusts the weight of this layer from 0.25 to 0.33, while proportionally reducing the weights of other layers. In the particle swarm optimization process, feedback adjustments are applied to the particle's velocity vector and position constraints, adjusting the weight of the global optimal position and the particle's exploration range. For example, for the optimization of the feature extraction threshold, if the policy network determines that the current threshold is set too high, the feedback path will adjust the global optimal position of the particle swarm toward a lower threshold and increase the exploration range of this dimension.

[0081] An adaptive early stopping mechanism is set based on the consistency assessment score. The trigger condition is that the increase in the consistency assessment score in the preset number of iterations is less than the dynamically adjusted convergence threshold. The preset number of rounds is initially set to 5, and the dynamic convergence threshold is initially set to 0.02, which is adaptively adjusted according to the historical fluctuations of the assessment score. The system maintains a historical assessment score window of length 10 and calculates the standard deviation of the scores within the window. When the standard deviation is greater than 0.08, the convergence threshold is increased to 1.2 times the original value, and the maximum does not exceed 0.05; when the standard deviation is less than 0.03, the convergence threshold is reduced to 0.8 times the original value, and the minimum is not less than 0.005. When the early stopping mechanism is triggered, the system performs a parameter rollback operation and rolls back to the parameter configuration with the highest historical assessment score. For example, when processing a batch of corporate training material documents, if the consistency assessment scores from the 18th to the 22nd iterations are 0.823, 0.827, 0.829, 0.830, and 0.831, respectively, and the growth rate is continuously lower than the current convergence threshold of 0.01, the system triggers the early stopping mechanism and falls back to the parameter configuration of the 22nd round.

[0082] The optimization process is repeated until the consistency score is greater than or equal to the convergence threshold, at which point the current classification result is output as the final document classification. For example, in a typical enterprise document classification task, the initial consistency score for a mixed document containing financial data, contract terms, and product descriptions was 0.65, below the preset threshold of 0.85. After 28 rounds of iterative optimization, adjusting parameters such as text feature weights from 0.4 to 0.55, structural feature weights from 0.3 to 0.25, and semantic relevance thresholds from 0.6 to 0.75, the system achieved a final consistency score of 0.87, indicating that the document belonged to both the "Financial Statement" and "Product Contract" categories.

[0083] In the prior art, enterprise document classification methods mainly use rule-based methods or simple machine learning models, such as support vector machines or decision trees. These methods have problems of low consistency and poor adaptability when dealing with complex and multi-faceted enterprise documents. Traditional methods often use a single gradient descent or genetic algorithm in the parameter optimization process, which lacks memory capacity and strategic adjustment mechanism, resulting in easy trapping of local optimal solutions in the face of complex document types. This embodiment solves the problem of consistency evaluation in enterprise document classification, introduces a two-layer policy network combined with a memory enhancement mechanism, and realizes intelligent optimization of classification parameters through the experience replay of the outer network and the dual policy gradient architecture of the inner network. At the same time, a two-way feedback path with adaptive feedback intensity is designed, so that the optimization process can make targeted adjustments according to the characteristics of different documents.

[0084] In an optional embodiment, the optimization strategy combination is optimized and adjusted through a bidirectional feedback path with a feedback strength adaptive adjustment mechanism, acting on the hierarchical cognitive attention network and the particle swarm optimization process respectively, and the updated parameters are input into the next round of iteration, including: The first feedback path of the bidirectional feedback path acts on the hierarchical cognitive attention network, adjusts the feature component weight matrix based on the gradient sensitivity analysis, and obtains an updated feature component weight matrix; updates the feature extraction parameters based on the importance score of the comprehensive evaluation vector, and obtains the updated feature extraction parameters; optimizes the fusion ratio of the feature combination pattern based on the information gain of the consistency evaluation score, and obtains the updated feature combination pattern; The second feedback path of the bidirectional feedback path acts on the particle swarm optimization process, dynamically adjusts the search step size based on the diversity of the optimization strategy combination, and obtains an updated search step size parameter; adaptively updates the inertia weight based on the convergence speed of the consistency evaluation score, and obtains an updated inertia weight parameter; reconstructs the local optimal position based on the local landscape analysis of the comprehensive evaluation vector, and obtains an updated local optimal position; The updated feature component weight matrix, updated feature extraction parameters, updated feature combination mode, updated search step parameters, updated inertia weight parameters and updated local optimal position are input into the next round of iterative optimization.

[0085] In one specific embodiment, the bidirectional feedback loop comprises two feedback pathways, one for the hierarchical cognitive attention network and the other for the particle swarm optimization process. The first feedback pathway primarily adjusts the hierarchical cognitive attention network. This pathway first adjusts the feature component weight matrix based on gradient sensitivity analysis, calculating the impact of each feature component on the consistency assessment score. Specifically, each element in the current feature component weight matrix is ​​perturbed slightly by 0.01 of the current weight value, and the change in consistency assessment score before and after the perturbation is recorded. The score change caused by the perturbation is normalized to form a sensitivity matrix. The weight adjustment is proportional to the sensitivity, calculated as the current weight plus the sensitivity multiplied by a learning rate factor (typically set between 0.05 and 0.15). For example, for a corporate financial report document, the system might detect a sensitivity of 0.42 for the "balance sheet" feature, significantly higher than the 0.08 sensitivity of the "chart" feature. Based on this, the weight of the "balance sheet" feature is increased from 0.25 to 0.31, while the weight of the "chart" feature is only slightly adjusted from 0.15 to 0.16. The adjusted weight matrix needs to be normalized to ensure that the sum of the weights of each dimension is 1.

[0086] The first feedback path also updates the feature extraction parameters based on the importance score of the comprehensive evaluation vector. The comprehensive evaluation vector contains multidimensional evaluation indicators such as classification accuracy, structural recognition, and content completeness, and each indicator component corresponds to a specific feature extraction parameter. The correlation strength between each evaluation indicator and the feature extraction parameter is calculated. The correlation strength is calculated using the Pearson correlation coefficient, and the value range is -1 to 1. The adjustment direction of the feature extraction parameter is determined by the positive or negative correlation strength, and the adjustment amplitude is determined by the absolute value of the correlation strength and the difference between the current consistency assessment score. For example, for the content completeness assessment component and the text segmentation threshold parameter, if the correlation strength is 0.65 (positive correlation), the current text segmentation threshold is 0.35, and the consistency assessment score is 0.72 (lower than the target threshold of 0.85), the text segmentation threshold is adjusted to 0.41, an increase of approximately 17%. In this way, the feature extraction parameters are adaptively adjusted for different types of corporate documents, such as enhancing the clause segmentation accuracy of contract documents and enhancing the topic clustering accuracy of meeting minutes.

[0087] The first feedback path also optimizes the fusion ratio of feature combination modes based on the information gain of the consistency assessment score. Feature combination modes include serial mode, parallel mode, hierarchical mode and other combination methods, and each mode has different applicability to different types of documents. The system constructs a decision tree structure, records the changes in consistency assessment scores under different fusion ratios, and calculates the contribution of each adjustment to information gain. The information gain calculation is based on the change in entropy of the consistency assessment score distribution before and after the adjustment, and selects the fusion ratio adjustment scheme that can maximize the information gain. For example, when processing a batch of corporate rules and regulations documents, it may be found that increasing the fusion ratio of the hierarchical mode from 0.4 to 0.65 and reducing the ratio of the serial mode from 0.35 to 0.2 can increase the consistency assessment score from 0.78 to 0.84, and the information gain reaches 0.32, which is higher than the information gain of other adjustment schemes.

[0088] The second feedback path of the bidirectional feedback loop acts on the particle swarm optimization process. This path dynamically adjusts the search step size based on the diversity of the optimization strategy combinations. The degree of dispersion of each particle position in the current particle swarm is calculated, and the standard deviation is used to measure the level of diversity. When the diversity index falls below the preset threshold of 0.15, indicating that the particle swarm may be trapped in a local optimum, the search step size is increased, typically by 50% to 100%, to promote global exploration. When the diversity index rises above the preset threshold of 0.4, indicating that the particle distribution is too dispersed, the system reduces the search step size by 20% to 40% to strengthen local exploration. The upper and lower limits of the search step size are set to 3 times and 0.3 times the initial step size, respectively. For example, when optimizing the classification parameters of a multi-departmental, complex enterprise document, if the current particle position diversity index is 0.12, which is below the threshold of 0.15, the search step size is increased from 0.08 to 0.14, prompting the particles to explore a wider parameter space and avoid being trapped in a local optimum.

[0089] The second feedback path also adaptively updates the inertia weight parameter based on the convergence rate of the consistency evaluation score. A historical consistency evaluation score window of length 5 is maintained, and the rate of change of the evaluation score is calculated. When the rate of change of the evaluation score for three consecutive rounds is less than 0.02, it indicates slow convergence. The inertia weight is reduced, typically by 15% to 25%, to increase the particle's attraction to the global and local optimal positions. When the rate of change of the evaluation score is greater than 0.08, it indicates that the convergence rate is too fast and may miss a better solution. The inertia weight is increased by 10% to 20%, enhancing the particle's ability to maintain the original search direction due to inertia. The upper and lower limits of the inertia weight are set to 0.9 and 0.2, respectively. For example, when processing an enterprise technical document classification task, if the evaluation scores for four consecutive rounds are 0.76, 0.77, 0.775, and 0.778, respectively, and the rate of change remains below 0.01, the system will reduce the inertia weight from 0.5 to 0.4, accelerating the particle's convergence to known good areas.

[0090] The second feedback pathway also reconstructs local optima based on a local landscape analysis of the integrated evaluation vector. The parameter space is divided into multiple subregions, and the distribution of consistency evaluation scores is fitted within each region to construct a parameter-score response surface. Potential local optima are identified by analyzing the gradient and curvature of the response surface. For each particle, its current local optimal position is weighted and synthesized with the potential optimal position obtained from the analysis. The weight of the synthesis is related to the particle's historical performance and the evaluation score of the potential optimal position. For example, when processing enterprise personnel file document classification, a potential local optimal point (0.45, 0.35, 0.2) is identified in the feature weight parameter space with a predicted evaluation score of 0.86. The particle's current local optimal position is (0.4, 0.3, 0.3) with an evaluation score of 0.82. A new local optimal position (0.435, 0.335, 0.23) is synthesized based on a 70% weight for the new position and a 30% weight for the original position.

[0091] The updated feature component weight matrix, feature extraction parameters, feature combination mode, search step size parameter, inertia weight parameter, and local optimal position are input into the next round of iterative optimization. For example, in a complete iteration of enterprise document classification optimization, the first feedback path adjusted the text feature weight from 0.45 to 0.52, the document structure feature weight from 0.3 to 0.25, the semantic relevance threshold from 0.6 to 0.68, and the feature level fusion ratio from 0.4 to 0.55. The second feedback path increased the search step size from 0.1 to 0.15, reduced the inertia weight from 0.6 to 0.5, and adjusted the local optimal position from the original parameter combination to a new, more optimal parameter combination. These updated parameters acted together in the next iteration, increasing the consistency assessment score from the current 0.75 to 0.79.

[0092] In the prior art, enterprise document classification systems generally adopt a one-way feedback mechanism, which lacks the ability to make two-way collaborative adjustments to classification parameters and optimization processes. In traditional methods, feedback adjustments are mainly focused on feature weights or classification decision thresholds, ignoring the optimization of feature extraction parameters and combination patterns, and parameter adjustments often adopt a fixed step size strategy, which cannot be adaptively adjusted according to the dynamic characteristics of the optimization process. This embodiment solves the problem of lack of coordination and adaptability in the parameter optimization process in enterprise document classification. Through the first feedback path, the feature component weight matrix, feature extraction parameters and feature combination patterns are simultaneously optimized, and the second feedback path dynamically adjusts the search step size, inertia weight and local optimal position, thereby achieving two-way collaborative optimization of the classification algorithm and the optimization algorithm. This improvement significantly improves the optimization efficiency and classification accuracy of enterprise document classification, improves the average classification consistency evaluation score of complex enterprise documents, increases the optimization convergence speed, and significantly enhances the adaptability to diversified enterprise documents.

[0093] like Figure 3 The figure shows the dynamic changes in the search step size parameter during the optimization process and its impact on search efficiency. The chart clearly demonstrates the superiority of the second feedback path in this technical solution, which dynamically adjusts the search step size based on the diversity of optimization strategy combinations. In the early stages of optimization (iterations 1-20), this technical solution uses a larger search step size (starting from an initial value of 0.42 and reaching a peak of 0.53 at the 10th iteration), ensuring extensive exploration of the parameter space and rapidly improving the search efficiency from 0.28 to 0.67. As the iterations progress toward the middle stages (iterations 21-60), the search step size gradually decreases (from 0.48 to 0.23), enabling the algorithm to more carefully explore promising areas and further improving the search efficiency to 0.83. In the later stages of optimization (iterations 61-100), the search step size stabilizes within a smaller range (around 0.20, ultimately decreasing to 0.17), ensuring stable convergence and ultimately achieving a search efficiency of 0.92. In contrast, the fixed step size method kept the search step size constant (0.30) throughout the entire process, resulting in poor search efficiency in the early stages (only 0.15), and in the later stages, it was difficult to fine-tune due to the large step size (the final search efficiency was only 0.75); and although the simple decaying step size method was improved, its preset decay rate could not be adjusted according to the actual optimization progress, resulting in an unsmooth search efficiency curve and a final efficiency of 0.83. These data fully demonstrate the effectiveness of the dynamic adjustment of the search step size mechanism in this technical solution. By coordinating with the optimization process, it achieves a balance between global exploration and local fine search, significantly improving the efficiency of parameter optimization. Figure 2 FIG. 1 is a structural diagram of an enterprise intelligent document recognition and automatic classification and archiving management processing system according to an embodiment of the present invention. Figure 2 As shown, the system includes: The first unit is used to construct a hierarchical cognitive attention network including a perception layer, an understanding layer, and a decision layer; a capsule network and a dynamic routing algorithm in the perception layer are used to extract features of the document to be processed, construct feature space relationships, and obtain an initial feature map; a memory-enhanced neural network in the understanding layer is used to semantically process the initial feature map using a long-term memory module and a working memory module to obtain a semantic feature vector; a metacognitive controller in the decision layer calculates document complexity based on the semantic feature vector, and adaptively adjusts feature extraction parameters according to the document complexity to obtain a cognitive feature representation of the document; The second unit is used to map the cognitive feature representation to a non-Euclidean manifold space through the Riemann metric tensor, use the projection of the classification scheme in the tangent space as the particle position, and use the projection of the update direction of the classification scheme on the tangent plane as the particle velocity to generate an initial particle swarm; record the historical search trajectory of each particle in the multidimensional search space, calculate the adaptive inertia weight, and dynamically adjust the search step size of the particle according to the adaptive inertia weight; combine the preset learning factor and random factor, iteratively update the local optimal solution and the global optimal solution, determine the optimal classification boundary, and generate a multidimensional classification vector for the document to be processed based on the optimal classification boundary; The third unit is used to construct an adaptive feedback optimization network, perform attention-weighted mapping on the cognitive feature representation and the multidimensional classification vector through the deep feature fusion layer, generate a comprehensive evaluation vector for the document, and adopt a multidimensional similarity calculation model based on the comprehensive evaluation vector to obtain the consistency evaluation score between the documents; when the consistency evaluation score is less than a preset threshold, a reinforcement learning method is used to construct a strategy network, and an optimization strategy combination is generated based on historical optimization records. The optimization strategy combination is fed back to the hierarchical cognitive attention network and the particle swarm optimization process, and the iterative execution is repeated until the consistency evaluation score is greater than or equal to the preset threshold, and the classification result output of the document is completed.

[0094] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0095] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. Enterprise intelligent document recognition and automatic classification and archiving management processing method, characterized by: include: Construct a hierarchical cognitive attention network consisting of perception, understanding, and decision-making layers; In the perception layer, capsule networks and dynamic routing algorithms are used to extract features from the document being processed, construct feature space relationships, and obtain an initial feature map. In the understanding layer, the long-term memory module and working memory module of the memory-augmented neural network are used to perform semantic processing on the initial feature map to obtain a semantic feature vector. In the decision layer, the metacognitive controller calculates the document complexity based on the semantic feature vector and adaptively adjusts the feature extraction parameters to obtain a cognitive feature representation of the document. The cognitive feature representation is mapped to a non-Euclidean manifold space via the Riemann metric tensor. The projection of the classification scheme on the tangent space is used as the particle position, and the projection of the update direction on the tangent plane is used as the particle velocity to generate an initial particle swarm. The particle search trajectory is recorded, the adaptive inertia weight is calculated, and the search step size is dynamically adjusted. The optimal solution is iteratively updated by combining the learning factor and the random factor, the optimal classification boundary is determined, and a multi-dimensional classification vector is generated. An adaptive feedback optimization network is constructed to perform attention-weighted mapping on cognitive feature representations and multidimensional classification vectors to generate a comprehensive evaluation vector for the document. A multidimensional similarity calculation model is used to obtain the consistency evaluation score between documents. When the evaluation score is less than the preset threshold, reinforcement learning is used to construct a strategy network to generate an optimized strategy combination and provide feedback. The iteration is repeated until the consistency evaluation score reaches the preset threshold, and the classification result of the document is output.

2. The method according to claim 1, characterized in that A hierarchical cognitive attention network consisting of a perception layer, an understanding layer, and a decision layer is constructed. In the perception layer, a capsule network and a dynamic routing algorithm are used to extract features from the document being processed, construct feature space relationships, and obtain an initial feature map. In the understanding layer, the long-term memory module and working memory module of the memory-enhanced neural network are used to perform semantic processing on the initial feature map to obtain a semantic feature vector. In the decision layer, a metacognitive controller calculates document complexity based on the semantic feature vector and adaptively adjusts feature extraction parameters to obtain the document's cognitive feature representation, including: A capsule network consisting of a main capsule layer and an auxiliary capsule layer is constructed, wherein the main capsule layer contains multiple parallel convolution capsules and the auxiliary capsule layer contains multiple attention capsules. The document to be processed is input into the capsule network, and the local feature patterns are captured and local feature vectors are extracted through parallel convolution capsules. The dynamic routing coefficients between local feature vectors are calculated through attention capsules. The coupling coefficients between feature vectors are iteratively calculated based on the dynamic routing coefficients. The local feature vectors are adaptively combined to generate an initial feature mapping matrix. A memory-enhanced neural network consisting of a long-term memory module and a working memory module is constructed. The initial feature mapping matrix is ​​input, and the query vector, key vector, and value vector are calculated using a multi-head attention mechanism. The attention weight is calculated based on the query vector and key vector. The working memory module updates the attention weight based on the historical feature state to obtain an updated attention weight. The memory attention mechanism is used to calculate the memory read weight of the feature mapping matrix and the pre-trained semantic knowledge base. The product of the updated attention weight and the value vector is fused with the product of the memory read weight and the semantic knowledge base to obtain a semantic feature vector. A metacognitive controller is constructed to calculate the document complexity index based on the semantic feature vector, including a weighted combination of the bi-norm and information entropy of the semantic feature vector. The convolution kernel size and the number of iterations of the dynamic routing coefficient are dynamically adjusted according to the document complexity index. The semantic feature vector is normalized and updated through a feedforward neural network to obtain a cognitive feature representation.

3. The method according to claim 2, characterized in that The working memory module updates the attention weight based on the historical feature state, and the updated attention weight includes: The attention weight is input into the working memory module, and the difference of attention weights of adjacent time steps is calculated to obtain the differential state matrix. The forget gate parameters and update gate parameters are obtained through nonlinear transformation. Multiply the forget gate parameter by the historical feature state bit by bit to obtain the filtered historical feature state, multiply the update gate parameter by the attention weight of the previous time step to obtain the feature state to be updated; perform a weighted combination of the filtered historical feature state and the feature state to be updated to obtain the updated attention weight; The document complexity index is calculated based on the semantic feature vector, including the weighted combination of the bi-norm and information entropy of the semantic feature vector: The semantic feature vector is divided into local neighborhoods by K-nearest neighbor clustering, and the variance of the feature vector in each local neighborhood is calculated to obtain the vector distribution density; The cosine similarity of adjacent feature vectors is calculated based on the sliding window to obtain local correlation; The vector distribution density and local correlation are nonlinearly mapped to obtain the adaptive weight; the bi-norm and information entropy of the semantic feature vector are weightedly fused according to the adaptive weight to obtain the document complexity index.

4. The method according to claim 1, wherein The cognitive feature representation is mapped to the non-Euclidean manifold space via the Riemann metric tensor. The projection of the classification scheme on the tangent space is used as the particle position, and the projection of the update direction on the tangent plane is used as the particle velocity. The initial particle swarm is generated by: Constructing a search neighborhood for cognitive feature representation, calculating the Euclidean distance matrix of sample pairs within the search neighborhood, obtaining a similarity matrix through kernel function transformation, performing a logarithmic transformation on the similarity matrix and obtaining a Hessian matrix to obtain a Riemannian metric tensor; The geodesic distance between sample pairs is calculated based on the Riemannian metric tensor, a local distance-preserving matrix is ​​constructed, and eigenvalue decomposition is performed to obtain the principal eigenvectors and eigenvalues. The local curvature index is calculated based on the distribution of the eigenvalues, and the local curvature of the non-Euclidean manifold space is adjusted through conformal transformation. A reference point is selected within the search neighborhood, and a tangent space centered at the reference point is constructed. The data points in the non-Euclidean manifold space are projected onto the tangent space through exponential mapping, and the geodesic distances before and after the projection are calculated to maintain the geometric structure. Calculating the local distance-preserving matrix condition number to determine the optimal segmentation threshold, dividing the tangent space into subspaces, calculating the singular value distribution of the local distance-preserving matrix, and dynamically adjusting the shape and size of the subspace; Analyze the spectral distribution of the local distance-preserving matrix and calculate the sampling density function of each subspace to determine the number of particle allocations; A tangent plane is constructed in each subspace, and the update direction of the classification scheme is calculated. The updated direction is projected onto the tangent plane through the Riemann metric tensor. Within the preset maximum velocity range, the velocity vector is obtained based on the projection sampling of the tangent plane. The position vector is determined by combining the sampling density function, and the initial particle swarm is formed by pairing.

5. The method according to claim 1, wherein Record the particle search trajectory, calculate the adaptive inertia weight and dynamically adjust the search step size, combine the learning factor and random factor to iteratively update the optimal solution, determine the optimal classification boundary, and generate a multi-dimensional classification vector including: Record the position sequence of the initial particle swarm within the preset memory length to form a historical search trajectory, calculate the displacement vectors of adjacent moments in the historical search trajectory, and determine the overall displacement. The ratio of the overall displacement to the cumulative sum of displacements at adjacent moments is used as a motion trend indicator. Construct a quantum annealing control matrix, where each element represents the probability of quantum state superposition in the corresponding search dimension. Calculate the wave function collapse probability based on the number of iterations, calculate the mean of the motion trend index of all particles, and couple it with the wave function collapse probability to obtain the quantum control coefficient. Substitute the quantum control coefficient into the exponential decay function, calculate the adaptive inertia weight, multiply it with the particle's current velocity vector, determine the first difference vector with the individual optimal position, and the second difference vector with the global optimal position, and then perform a weighted combination to obtain the velocity update; Based on the speed update, the updated position vector is calculated, the corresponding classification accuracy and category overlap are determined, and the weighted combination is used to obtain the fitness value; When the fitness value increases for a preset number of consecutive times are less than a preset threshold, the position vector corresponding to the maximum fitness value is determined as the optimal classification boundary, and a multi-dimensional classification vector of the document to be processed is generated through linear transformation.

6. The method according to claim 1, wherein An adaptive feedback optimization network is constructed to perform attention-weighted mapping on cognitive feature representation and multidimensional classification vectors to generate a comprehensive evaluation vector for the document. A multidimensional similarity calculation model is used to obtain consistency evaluation scores between documents, including: It receives the document's cognitive feature representation and multidimensional classification vector and constructs a deep feature fusion layer containing local attention units and global attention units. The local attention unit calculates the correlation between the cognitive feature representation and the multidimensional classification vector using a multi-scale sliding window to obtain a feature component weight matrix. The global attention unit calculates cross-dimensional semantic dependencies based on a hierarchical semantic tree to obtain a feature combination pattern. The feature component weight matrix and feature combination pattern are input into a residual memory module with forget gate. The historical feature combination patterns are prioritized according to temporal correlation and selectively stored to guide the allocation of feature component weight matrices. A separable convolutional network with channel attention is used to perform adaptive feature recalibration to generate a comprehensive evaluation vector for the document. A feature projection layer is constructed to map the comprehensive evaluation vector to multiple semantic subspaces through orthogonal transformation. An adaptive kernel function network is constructed to calculate local similarity. Regularization constraints based on mutual information minimization are imposed to maintain the complementarity of semantic subspaces. Dynamic weighting based on confidence is used to integrate the local similarity scores of semantic subspaces to obtain a consistency evaluation score.

7. The method according to claim 6, characterized in that When the evaluation score is less than the preset threshold, reinforcement learning is used to build a policy network, generate an optimized policy combination and provide feedback. The iteration is repeated until the consistency evaluation score reaches the preset threshold. The classification results of the output documents include: When the consistency evaluation score is less than the convergence threshold, a two-layer policy network with a memory enhancement mechanism is constructed. The outer network sets up an experience replay buffer pool to store historical optimization trajectories and select macro optimization directions based on the similarity of historical optimization trajectories. The inner network adopts a dual policy gradient architecture consisting of a value function network and a policy network. The value function is used to evaluate the current state value and calculate the advantage function. The policy network generates a parameter adjustment plan based on the advantage function and the comprehensive evaluation vector, generating an optimization strategy combination. The optimization strategy combination is optimized and adjusted through a bidirectional feedback path with a feedback strength adaptive adjustment mechanism, acting on the hierarchical cognitive attention network and the particle swarm optimization process respectively, and the updated parameters are input into the next round of iteration; An adaptive early stopping mechanism is set based on the consistency evaluation score. When the improvement in the consistency evaluation score over a preset number of consecutive iterations is less than the dynamically adjusted convergence threshold, a parameter rollback operation is triggered. The convergence threshold is adaptively adjusted based on the historical fluctuations of the consistency evaluation score. Repeat until the consistency evaluation score is greater than or equal to the convergence threshold, and output the current classification result as the final document classification.

8. The method according to claim 7, characterized in that The optimization strategy combination is optimized and adjusted through a bidirectional feedback path with a feedback strength adaptive adjustment mechanism, acting on the hierarchical cognitive attention network and the particle swarm optimization process respectively. The updated parameters are input into the next round of iteration, including: The first feedback path of the bidirectional feedback path acts on the hierarchical cognitive attention network, adjusts the feature component weight matrix based on the gradient sensitivity analysis, and obtains an updated feature component weight matrix; updates the feature extraction parameters based on the importance score of the comprehensive evaluation vector, and obtains the updated feature extraction parameters; optimizes the fusion ratio of the feature combination pattern based on the information gain of the consistency evaluation score, and obtains the updated feature combination pattern; The second feedback path of the bidirectional feedback path acts on the particle swarm optimization process, dynamically adjusts the search step size based on the diversity of the optimization strategy combination, and obtains an updated search step size parameter; adaptively updates the inertia weight based on the convergence speed of the consistency evaluation score, and obtains an updated inertia weight parameter; reconstructs the local optimal position based on the local landscape analysis of the comprehensive evaluation vector, and obtains an updated local optimal position; The updated feature component weight matrix, updated feature extraction parameters, updated feature combination mode, updated search step parameters, updated inertia weight parameters and updated local optimal position are input into the next round of iterative optimization.

9. An enterprise intelligent document recognition and automatic classification and filing management processing system, used to implement the method according to any one of claims 1 to 8, characterized in that: include: The first unit is used to build a hierarchical cognitive attention network consisting of perception layer, understanding layer and decision layer; In the perception layer, capsule networks and dynamic routing algorithms are used to extract features from the document being processed, construct feature space relationships, and obtain an initial feature map. In the understanding layer, the long-term memory module and working memory module of the memory-augmented neural network are used to perform semantic processing on the initial feature map to obtain a semantic feature vector. In the decision layer, the metacognitive controller calculates the document complexity based on the semantic feature vector and adaptively adjusts the feature extraction parameters to obtain a cognitive feature representation of the document. The second unit is used to map the cognitive feature representation to a non-Euclidean manifold space via the Riemann metric tensor, generate an initial particle swarm using the projection of the classification scheme on the tangent space as the particle position and the projection of the update direction on the tangent plane as the particle velocity; record the particle search trajectory, calculate the adaptive inertia weight and dynamically adjust the search step size; iteratively update the optimal solution by combining the learning factor and the random factor, determine the optimal classification boundary, and generate a multidimensional classification vector; The third unit is used to build an adaptive feedback optimization network, perform attention-weighted mapping on cognitive feature representations and multidimensional classification vectors, generate a comprehensive evaluation vector for the document, and use a multidimensional similarity calculation model to obtain the consistency evaluation score between documents; when the evaluation score is less than the preset threshold, reinforcement learning is used to build a policy network, generate an optimized strategy combination and feedback, repeat the iteration until the consistency evaluation score reaches the preset threshold, and output the classification result of the document.

10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Classified management method, system and equipment for multi-source electric power appeal data and medium

    CN121093054A

  • Archive data automatic classification method and system based on machine learning

    CN121211129A

  • Text scene deduction evaluation method and system

    CN121257730A

  • A text scenario deduction evaluation method and system

    CN121257730B