Government affair file intelligent classification and retrieval system

Through deep learning and knowledge graph technology, combined with dynamic authority management, the efficiency, accuracy and security of the government document management system are solved, and efficient and accurate classification and retrieval of government document are achieved, meeting multi-level security needs and providing comprehensive knowledge support.

CN120407796AActive Publication Date: 2025-08-01NO 15 INST OF CHINA ELECTRONICS TECH GRP

Patent Information

Application Number
CN202510635147.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-01
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing government document management system has shortcomings in efficiency, accuracy, security and compatibility. Especially when facing complex and diverse government document formats, it is difficult to ensure the accuracy and security of classification and retrieval, and the authority management is not refined enough to meet multi-level security needs.

Method used

Deep learning technology is adopted, combining the feature extraction modules of convolutional neural networks and recurrent neural networks, and a knowledge graph module is built for semantic understanding, an urgency assessment module and permission management module are designed to realize dynamic permission adjustment and encryption collaboration, and support multimodal information fusion and cross-domain knowledge fusion, and dynamic update mechanisms.

Benefits of technology

It significantly improves the classification and retrieval efficiency of government affairs documents, improves the accuracy rate to more than 95%, enhances security, refined authority management, and improves compatibility, meets the complex needs of government affairs office, provides comprehensive knowledge services, and improves the intelligence level of government affairs office.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407796A_ABST
    Figure CN120407796A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent government affair file classification and retrieval system, and belongs to the technical field of file management and information. Aiming at the problems that a traditional government affair file management mode is low in efficiency and high in error rate, and an existing general file management system is insufficient in security management and control in the aspect of government affair file management, low in retrieval accuracy and the like, intelligent classification is performed by adopting a CNN and RNN combined deep learning architecture, and accurate retrieval is realized through a word vector model and an attention mechanism. The system constructs a knowledge graph in the government affair field to provide intelligent decision support, designs a multi-level authority management system to guarantee file security, has a real-time data updating and self-learning mechanism, and can adapt to rapid updating of government affair knowledge. Compared with the prior art, the system has the remarkable advantages in the aspects of classification and retrieval efficiency, accuracy, safety, compatibility, intelligent expansibility and the like, intelligence, high efficiency and standardization of government affair file management can be achieved, and the complex requirements of government affair office are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of document management and information technology, and in particular relates to an intelligent classification and retrieval system for government documents. Background Art

[0002] Throughout the history of government document management, traditional manual methods have long dominated. Staff rely on experience to manually categorize documents using simple catalogs, and searches rely on keywords in file names to find documents in shared folders or filing cabinets. This approach is not only inefficient, but according to research, a 50-person department spends 800-1200 hours per month on document classification, with each search taking an average of 30-40 minutes. Furthermore, the subjectivity of manual judgment leads to a classification error rate of 12%-18%.

[0003] With the development of information technology, universal document management systems have emerged, such as SharePoint and WPS Enterprise Cloud Documents. They have basic storage, category browsing, and keyword search functions, and some also have simple permission management. However, there are obvious flaws in the management of government documents. In terms of security control, the permission settings are not detailed enough. More than 70% of the government departments that use them have reported that they cannot meet the multi-level permission requirements of documents with different levels of confidentiality. At the same time, there is insufficient understanding of the professional terms in government documents, and the retrieval accuracy rate is only 60%-70%. The system involved in the patent with publication number CN119068494A also has similar problems in terms of government targeting.

[0004] While academic research has seen numerous cutting-edge explorations in text classification and information retrieval, including deep learning models based on CNNs and RNNs and improved vector space retrieval algorithms, these technologies face numerous challenges in practical application. For one thing, the models lack self-learning and adaptive capabilities in the face of the rapid evolution of government knowledge. The deep learning model introduced by the research team saw its classification accuracy drop from 85% to 60% within a month after the release of the new policy. The patented technology, published with publication number CN118861213B, also failed to effectively address this issue. Furthermore, compatibility and accuracy are difficult to ensure when processing complex and diverse government document formats, with an average of 20%-30% of files experiencing parsing errors or content loss. Summary of the Invention

[0005] The present invention solves the difficulties in government document processing by deeply integrating advanced technologies such as deep learning, semantic understanding, and index optimization algorithms; it realizes an organic overall solution to meet the complex needs of document management in government office scenarios, thereby establishing an intelligent classification and retrieval system for government documents.

[0006] The present invention discloses an intelligent classification and retrieval system for government affairs documents. The intelligent classification and retrieval system for government affairs documents includes a feature extraction module, a knowledge graph module, a retrieval module, an urgency assessment module, and a permission management module;

[0007] The feature extraction module adopts a deep learning architecture that combines a convolutional neural network model and a recurrent neural network model. The convolutional neural network model sets multiple convolutional layers, and the size of the convolutional kernel is dynamically adjusted according to the file features to extract the local features of the file. The recurrent neural network model uses LSTM units and integrates the self-attention mechanism to capture the logical relationships of the text sequence;

[0008] The knowledge graph module constructs a knowledge graph for the government affairs field, integrates cross-domain knowledge such as government affairs, economy, and society, and has a dynamic update mechanism to monitor government affairs data sources in real time and update entities and relationships;

[0009] The retrieval module converts files and retrieval terms into vectors by means of a word vector model, calculates the semantic similarity through the cosine similarity formula, and adaptively adjusts the attention weights according to the retrieval term popularity, the user's historical retrieval behavior, and the file timeliness in combination with the attention mechanism to achieve accurate retrieval;

[0010] The urgency assessment module clarifies the timeliness of government affairs documents based on a fuzzy rule base and a dynamic rule weight adjustment mechanism to reflect the urgency of government affairs documents;

[0011] The permission management module combines the organizational structure of government departments and the confidentiality level of files, designs an access control algorithm based on roles and confidentiality levels, introduces a risk assessment mechanism at the same time, dynamically adjusts permissions according to file access risks, and coordinates permission management with file encryption. Users with different permissions obtain different decryption keys.

[0012] Furthermore, the intelligent classification and retrieval system for government affairs documents includes a preprocessing module. The preprocessing module preprocesses the uploaded government affairs documents, including removing duplicate and invalid documents, classifying and storing them according to departments, years, etc.; and converting them into a unified text format and removing the headers, footers, and advertising information in the documents.

[0013] Furthermore, the convolutional neural network model is based on a dynamic convolutional kernel selection method, and the specific steps are as follows:

[0014] Extract the number of words L, the number of paragraphs P, the title level H, and the nested list depth D of the file, construct a multi-dimensional feature vector X = [L, P, H, D], calculate the complexity score, and divide the file into high-complexity files, medium-complexity files, and low-complexity files based on the complexity score;

[0015] Semantic features of high-complexity files, medium-complexity files, and low-complexity files are extracted using convolutional kernels of different scales, and multi-scale features are fused through weighted splicing;

[0016] In the training stage, the convolutional neural network uses a labeled dataset in the government affairs field, and the loss function is represented by cross-entropy with a class balance factor:

[0017]

[0018] Among them, To alleviate the class imbalance problem, c is the class serial number, and N c is the number of samples in class c; The optimizer uses AdamW, and the learning rate is η = 0.01.

[0019] Furthermore, the complexity score is expressed as:

[0020] C = w T ·X norm + b

[0021] Among them, w = [0.35, 0.25, 0.20, 0.20] T is the weight vector, X norm is the feature after normalizing the number of words L, the number of paragraphs P, the title level H, and the nested list depth D, and b is the bias term;

[0022] Among them, when C ≥ 1.2, it is a high-complexity file; when 0.8 ≤ C < 1.2, it is a medium-complexity file; when C < 0.8, it is a low-complexity file.

[0023] Furthermore, the feature extraction processes for high-complexity files, medium-complexity files, and low-complexity files are as follows:

[0024] For high-complexity files, a 7×7 convolutional kernel is used to extract long-range semantic features, and dilated convolution is superimposed to expand the receptive field, with a dilation rate r = 2, and the high-complexity feature map F large ;

[0025] For medium-complexity files, a 5×5 convolutional kernel is used to extract semantic features, and a gating mechanism is combined to filter out noise features. The gating weight g i,j is:

[0026] g i,j = σ(W g ·F i,j + b g )

[0027] Among them, σ is the Sigmoid function, and W g is the weight matrix in the gating mechanism, which is used to weight the feature F i,j , and F i,jis the feature vector at the (i, j) position in the feature map obtained after the convolution operation, which contains the local feature information of the file, b g is the bias term in the gating mechanism, used to adjust the overall level of the gating output; the complexity feature map in the output:

[0028] F mid = g i,j ⊙F raw

[0029] F raw is the original feature map without being processed by the gating mechanism when using a 5x5 convolution kernel to extract the semantic features of medium-complexity files; for low-complexity files, a 3×3 convolution kernel is used to extract semantic features, and an attention module is inserted to strengthen the key regions, and the spatial attention weight A spatial is:

[0030] A spatial = Sigmoid(W s ·AvgPool(F sm ) + W c ·MaxPool(F sm ))

[0031] F small = A spatial ⊙F raw

[0032] Sigmoid is the normalization function, W s and W c are the weight matrices, AvgPool is the average pooling operation, F sm is the original feature map; F small is the low-complexity feature map weighted by attention, and MaxPool is the max pooling operation;

[0033] Finally, the multi-scale features are fused by weighted splicing:

[0034] F fusion = α·F large + β·F mid + γ·F small

[0035] where, F fusion represents the fused feature map;

[0036] The weight coefficients α, β, γ are dynamically calculated by the complexity score C, and the calculation formula is:

[0037]

[0038] k1k2k3 are the parameters used to dynamically calculate the weight coefficients, which are constants determined during the model training or design process.

[0039] Furthermore, in the training phase, an annotation dataset in the government affairs field is adopted. The annotation dataset in the government affairs field includes 10 categories and 500,000 documents;

[0040] Specifically, the annotation dataset in the government affairs field contains ten core annotation categories, covering document type classification, policy field division, institutional entity recognition, policy term extraction, service process step annotation, people's livelihood demand classification, laws and regulations clause positioning, time node annotation, responsible entity association, and urgency level classification.

[0041] Furthermore, the urgency evaluation module constructs a fuzzy logic model through a mixed membership function;

[0042] The mixed membership function includes three variables: time sensitivity, impact scope, and associated event level;

[0043] For time sensitivity, a right - half triangle function is used to define the membership degree of time sensitivity:

[0044]

[0045] μ High (T s ) represents the membership degree when the time sensitivity is "high", and is used to measure the matching degree between the urgency of a government affairs document in the time dimension and the "high" urgency level. Its value range is between 0 and 1, and the larger the value, the closer the document is to the "high" urgency level standard in terms of time; T s is the time sensitivity variable, used to describe the deadline or time - related attributes of a government affairs document, and is the input variable for calculating the membership degree;

[0046] This function grows linearly when T s ≥2, and the maximum time sensitivity is at T s =3, reflecting that the closer the deadline, the higher the urgency;

[0047] For the impact scope, a trapezoidal function is used to define the membership degree of the impact scope:

[0048]

[0049] μ Regional (S i ) represents the membership degree when the impact scope is "regional", and is used to measure the matching degree between the impact scope of a government affairs document and the "regional" scope. The value range is between 0 and 1. When the value is 1, it means that the impact scope of the document fully conforms to the definition of "regional", and the closer to 0, the lower the matching degree with the "regional" scope. S iTo represent the influence scope variable, which is used to describe the size of the influence scope of government affairs documents and is the input variable for calculating this membership degree;

[0050] This function maintains a membership degree of 1 within the interval S i ∈[0.5, 2.5], and linearly decays on both sides, which is suitable for describing the smooth process of the influence scope transitioning from within the department to the city-wide policy;

[0051] For the associated time level, a Gaussian function is used to define the membership degree of the associated event level:

[0052]

[0053] μ Important (E l ) represents the membership degree when the associated event level is "important", which is used to measure the matching degree between the level of the event associated with the government affairs document and the "important" level. The larger its value, the more the associated event conforms to the "important" standard, and its value range is between 0 and 1. E l is the associated event level variable, which is used to describe the importance level of the event associated with the government affairs document and is the input variable for calculating this membership degree.

[0054] Furthermore, a fuzzy rule base containing 27 rules is constructed based on time sensitivity, influence scope, and associated event level;

[0055] Among them, the generation process of the rule base starts from the in-depth analysis and fuzzy set definition of the input variables. Among them, time sensitivity, influence scope, and associated event level are respectively divided into three fuzzy sets according to the value range: time sensitivity is divided into "low", "medium", and "high", influence scope is divided into "local", "regional", and "global", and associated event level is divided into "routine", "important", and "urgent".

[0056] Furthermore, on the basis of the fuzzy set definition, the construction of the rule base follows the IF-THEN logical structure, covering all possible combinations of input variables:

[0057] Among them, low, local, and routine correspond to the urgency level of ordinary documents;

[0058] Low, local, and important correspond to the urgency level of ordinary documents;

[0059] Low, local, and urgent correspond to the urgency level of urgent documents;

[0060] Low, regional, and routine correspond to the urgency level of ordinary documents;

[0061] Low, regional, and important correspond to the urgency level of urgent documents;

[0062] Low, regional, and urgent correspond to the urgency level of urgent documents;

[0063] Low, global, routine corresponds to the urgency level of ordinary;

[0064] Low, global, important corresponds to the urgency level of urgent;

[0065] Low, global, urgent corresponds to the urgency level of extremely urgent;

[0066] Medium, local, routine corresponds to the urgency level of ordinary;

[0067] Medium, local, important corresponds to the urgency level of ordinary;

[0068] Medium, local, urgent corresponds to the urgency level of urgent;

[0069] Medium, regional, routine corresponds to the urgency level of ordinary;

[0070] Medium, regional, important corresponds to the urgency level of urgent;

[0071] Medium, regional, urgent corresponds to the urgency level of extremely urgent;

[0072] Medium, global, routine corresponds to the urgency level of urgent;

[0073] Medium, global, important corresponds to the urgency level of extremely urgent;

[0074] Medium, global, urgent corresponds to the urgency level of extremely urgent;

[0075] High, local, routine corresponds to the urgency level of urgent;

[0076] High, local, important corresponds to the urgency level of urgent;

[0077] High, local, urgent corresponds to the urgency level of extremely urgent;

[0078] High, regional, routine corresponds to the urgency level of urgent;

[0079] High, regional, important corresponds to the urgency level of extremely urgent;

[0080] High, regional, urgent corresponds to the urgency level of extremely urgent;

[0081] High, global, routine corresponds to the urgency level of extremely urgent;

[0082] High, global, important corresponds to the urgency level of extremely urgent;

[0083] High, global, urgent corresponds to the urgency level of extremely urgent.

[0084] Furthermore, the dynamic rule weight adjustment mechanism includes:

[0085] The initial weight of each rule They are all set to 1, indicating that all rules contribute equally to the output in the initial state; by collecting the real-time feedback of users on the permission allocation results and combining with the matching degree r of the emergency level predicted by the model predicted , the rule weights are dynamically adjusted:

[0086]

[0087] represents the weight of the i-th rule at time t; represents the weight of the i-th rule at time t + 1; η represents the learning rate; r actual represents the actual matching degree of the emergency level;

[0088] The dynamic weight adjustment also works in coordination with the risk assessment model, and the coordination is achieved through the weight scaling factor λ R λ R The calculation formula of is:

[0089]

[0090] R represents the real-time risk value, R crit represents the risk threshold;

[0091] To prevent the model from deviating from the initial design due to long-term weight drift, the system sets a weight decay mechanism:

[0092]

[0093] Among them, the decay coefficient γ = 0.001, ensuring that the weights gradually return to the initial values without continuous feedback, maintaining the stability of the system.

[0094] The beneficial effects achieved by the present invention are:

[0095] In terms of efficiency and accuracy, the present system demonstrates outstanding advantages. Based on the intelligent classification architecture of CNN and RNN, it subverts the low-efficiency mode of traditional manual classification, can complete the classification of a large number of government documents in an instant, and greatly shortens the document classification cycle. The retrieval link is equally excellent. The retrieval driven by semantic understanding combined with the attention mechanism abandons the inefficient way of traditional file name keyword retrieval, shortens the original retrieval time of 30 - 40 minutes per time to seconds or minutes, and the classification accuracy is increased to more than 95%, and the retrieval accuracy is increased from 60% - 70% of the general system to more than 90%, accurately meeting the needs of government affairs office for rapid and accurate file acquisition.

[0096] From the perspectives of security and compatibility, the system effectively solves the problems of the existing technologies. The multi-level permission management system conducts refined permission allocation for files with different security levels and user roles. Compared with the extensive permission settings of general file management systems, it greatly strengthens the file security management and prevents the leakage of government affairs file information. At the same time, by applying the OCR technology and format conversion processing flow, it can accurately identify and process various complex-format government affairs files such as PDF, Word, Excel, pictures, and scanned documents, solves the format compatibility problem, and ensures the integrity and accuracy of file processing.

[0097] At the level of intelligence and scalability, the system brings a new experience to government affairs office work. The intelligent decision-making supported by the knowledge graph enables the system to associate relevant information when classifying and retrieving files. For example, when retrieving "environmental protection policies in a certain region", it can be associated with law enforcement cases and enterprise environmental protection measures, etc., providing comprehensive knowledge services. The real-time data update and self-learning mechanism enable the system to keep up with the pace of knowledge update in the government affairs field. In addition, the scalable design using the microservices architecture facilitates the flexible expansion of functions according to business needs. The rich reserved interfaces are convenient for integration with other government affairs systems, breaking the situation of isolated operation of existing systems, improving the overall collaborative efficiency and intelligent level of government affairs office work, providing convenience for staff to work anytime and anywhere, and comprehensively optimizing the user experience. Brief Description of the Drawings

[0098] Figure 1 It is a classification flowchart of an intelligent classification and retrieval system for government affairs files;

[0099] Figure 2 It is a retrieval flowchart of an intelligent classification and retrieval system for government affairs files. Detailed Embodiments

[0100] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description progresses. However, these embodiments are exemplary only and do not constitute any limitation to the scope of the present invention. Those skilled in the art should understand that without departing from the spirit and scope of the present invention, the details and forms of the technical solutions of the present invention can be modified or replaced, but such modifications and replacements all fall within the protection scope of the present invention.

[0101] The present invention aims to construct an intelligent classification and retrieval system for government affairs files to solve many problems in the management of government affairs files in the existing technologies. The system is deployed in the local area network of government affairs office work. The server adopts a high-performance multi-core processor, a large-capacity memory, and high-speed storage devices to meet the storage and rapid processing requirements of a large number of government affairs files. The clients include desktop computers and laptop computers used by staff, which conduct data interaction with the server through the local area network.

[0102] Specifically, the intelligent classification and retrieval system for government documents includes a preprocessing module, a feature extraction module, a knowledge graph module, a retrieval module, an urgency assessment module, and a permission management module.

[0103] The preprocessing module conducts data collection and collation: It collects government documents generated by various departments of a certain city government in the past five years, including policy regulations, administrative approval documents, meeting minutes, etc., totaling approximately 500,000 copies. The file formats cover PDF, Word, Excel, pictures, and scanned copies. These files are preliminarily collated to remove duplicate and invalid files and are classified and stored according to departments, years, etc.

[0104] The feature extraction module builds an intelligent classification model: It uses the collated file data to train an intelligent classification model based on the convolutional neural network model (CNN) and the recurrent neural network model (RNN). In the CNN part, multiple convolutional layers are set, and the convolutional kernel sizes are 3x3 and 5x5 respectively. Through convolutional operations, local features in the files, such as policy terms and key clauses, are extracted. In the RNN part, LSTM units are adopted to capture long-distance dependencies in the text and understand the logical order between policy clauses. After multiple rounds of training, the classification accuracy of the model reaches 96%.

[0105] In the traditional convolutional neural network model (CNN), the convolutional kernel size is usually fixed, making it difficult for the model to flexibly adapt to the feature extraction requirements of different types of government documents. The present invention innovatively designs a dynamic convolutional kernel selection method. In the file preprocessing stage, the system will automatically analyze features such as the number of words and structural complexity of the file. For policy and regulation documents with a large number of words and complex structures, the system will dynamically select larger convolutional kernels, such as 7x7 or 9x9, which can more comprehensively capture long sequence information and complex semantic structures in the file; while for notice documents with short content and simple structures, the system will select smaller convolutional kernels, such as 2x2 or 3x3, to accurately locate key information and reduce unnecessary computational effort. Through this dynamic adjustment mechanism, the model can more efficiently extract features of different files, significantly improving the accuracy and efficiency of classification.

[0106] Specifically, the dynamic convolutional kernel selection method (Dynamic Kernel Selection Algorithm, DKSA) solves the problem of semantic granularity differences in government document classification through a multi-scale feature fusion mechanism. The traditional CNN model uses a fixed-size convolutional kernel (such as 5×\alpha), which is difficult to capture both the long paragraph dependencies of policies and regulations and the local keywords of notices and announcements, resulting in a significant decrease in classification accuracy in complex file scenarios. The core innovation of the improved dynamic convolutional kernel selection method lies in introducing a file structure complexity quantification index and an adaptive kernel size mapping function. The specific implementation steps are as follows:

[0107] 1. Structural Complexity Quantification Model: In the preprocessing stage, the number of words L, the number of paragraphs P, the title hierarchy H, and the nested list depth D of the file are extracted to construct a multi-dimensional feature vector X = [L, P, H, D]. After dimensionality reduction by principal component analysis (PCA), the complexity score C is calculated as follows:

[0108] C = w T ·X norm +b

[0109] where w = [0.35, 0.25, 0.20, 0.20] T is the weight vector, and X norm is the feature after normalization of the number of words L, the number of paragraphs P, the title hierarchy H, and the nested list depth D (e.g., ), and b = 0.1 is the bias term. According to the C value, the files are divided into three categories: high complexity (C ≥ 1.2), medium complexity (0.8 ≤ C < 1.2), and low complexity (C < 0.8).

[0110] 2. Dynamic Convolution Kernel Mapping and Multi-scale Feature Fusion: For files with different complexities, DKSA dynamically selects the convolution kernel size and fuses multi-scale features. In this embodiment, the files are divided into three categories, and the processing process is as follows:

[0111] High-complexity files: Use a 7×7 convolution kernel to extract long-range semantic features, and stack dilated convolution to expand the receptive field with a dilation rate r = 2, and output the feature map F large .

[0112] Medium-complexity files: Use a 5×5 convolution kernel, combined with a gating mechanism to filter out noise features. The gating weight g i,j = σ(W g ·F i,j +b g ), where σ is the Sigmoid function, and the output F mid = g ⊙ F raw .

[0113] Low-complexity files: Use a 3×3 convolution kernel and insert an attention module (CBAM) to strengthen the key area. The calculation method of the spatial attention weight is:

[0114] A spatial = Sigmoid(W s ·AvgPool(F small ) + W c ·MaxPool(F small ))

[0115] Finally, multi-scale features are fused through weighted splicing, and the splicing method is:

[0116] F fusion =α·F large +β·F mid +γ·F small

[0117] The weight coefficients α, β, and γ are dynamically calculated by the complexity score C, and the calculation formula is:

[0118]

[0119]

[0120] 3. Model training and deployment. The training phase uses a government affairs annotation dataset (10 categories, 500,000 documents). The government affairs annotation dataset includes 10 categories and 500,000 documents. Specifically, the government affairs annotation dataset contains ten core annotation categories, covering document type classification, policy field division, institutional entity identification, policy terminology extraction, service process step annotation, people's livelihood demand classification, legal and regulatory clause positioning, time node annotation, responsible party association, and urgency classification.

[0121] The loss function is the cross entropy with a category balance factor, and the definition formula is:

[0122]

[0123] in, Alleviate the class imbalance problem (N c is the number of samples of category c). The optimizer uses AdamW, and the learning rate is η = 0.01.

[0124] The retrieval module uses word embedding models (such as Word2Vec) to convert text in files into vector representations and build a file vector library. Simultaneously, it trains an attention mechanism model to accurately focus on key information during the retrieval process. By learning from a large number of user search records, the retrieval model has achieved an accuracy rate of 92%.

[0125] Specifically, optimize the retrieval model into a personalized retrieval model: Traditional retrieval models usually adopt a unified retrieval strategy and cannot meet the personalized needs of different users. Based on user roles and historical retrieval behaviors, the present invention constructs a personalized retrieval model. The system deeply analyzes the user's historical retrieval records, extracts the user's retrieval preferences, such as the fields of interest, commonly used retrieval terms, etc., and converts these preferences into a user preference vector. When the user conducts a retrieval, the system not only converts the retrieval terms into vectors but also combines the user preference vector with them to jointly calculate the similarity with the document vectors. For example, for a user responsible for environmental protection work, the system will preferentially display environmental protection-related documents in the retrieval results and perform more accurate sorting of relevant documents according to the user's historical preferences. This personalized retrieval method can better meet the diverse needs of different users and significantly improve the retrieval experience and efficiency.

[0126] The knowledge graph module constructs a knowledge graph: Entities and relationships are extracted from the collected government affairs documents and government public data to construct a government affairs knowledge graph. For example, entity information such as policy names, issuing departments, and applicable scopes, as well as relationships such as issuance and applicability, are extracted from policy and regulation documents. The knowledge graph covers approximately 100,000 entities and 500,000 relationships.

[0127] Fuse multi-modal information: Existing technologies often rely solely on text information when classifying documents and ignore other important information dimensions in the documents. The present invention proposes a method for fusing multi-modal information. In addition to processing the document text, it also makes full use of the document's metadata and image information (if the document contains images). For the metadata of the document, such as the issuing department and the issuing date, the system encodes it and converts it into a vector form that can be understood by the model. For the image information in the document, the system first uses image processing technologies, such as edge detection and feature point extraction, to convert the image into a feature vector. Then, these metadata vectors, image feature vectors, and text feature vectors are fused, enabling the model to comprehensively consider information from multiple aspects when making classification decisions, thereby more accurately determining the document category, especially suitable for processing complex government affairs documents containing multiple information forms.

[0128] Semantic Expansion Combining with Knowledge Graph: Existing retrieval technologies have limitations in semantic understanding and are difficult to fully explore the potential semantics of retrieval terms. The present invention utilizes the knowledge graph in the government affairs field to expand the semantics of retrieval terms. When a user inputs a retrieval term, such as "new energy subsidy policy", the system will automatically associate related concepts, such as "new energy vehicle types" and "basis for formulating subsidy standards", based on the knowledge graph, and use these related concepts as expanded retrieval terms. At the same time, a new semantic fusion algorithm is introduced, which can comprehensively consider the similarity between the original retrieval term, the expanded retrieval terms and the document vector, as well as their association strength in the knowledge graph, and re-rank the retrieval results. In this way, the retrieval results are more comprehensive and accurate, and can provide more valuable information for users.

[0129] Dynamic Entity and Relationship Update: Due to the fast pace of knowledge update in the government affairs field, traditional knowledge graph update methods often cannot keep up with the changes in a timely manner. The present invention establishes a dynamic entity and relationship update mechanism. The system monitors government affairs data sources in real time, such as government official websites and policy release platforms. Once new information such as new policies and regulations, department adjustments, etc. are found, it will immediately start the entity and relationship extraction program. Using natural language processing technologies and information extraction algorithms, new entities and relationships are extracted from the newly released information, and the structure and content of the knowledge graph are updated in a timely manner. For example, when a new environmental protection policy is released, the system can quickly identify new entities such as new environmental protection supervision measures and new relationships such as the responsibility relationship between the policy and relevant enterprises, so as to ensure that the knowledge graph always maintains the latest state and provides accurate knowledge support for the classification and retrieval of the system.

[0130] Cross-Domain Knowledge Fusion: In order to broaden the application scope of the government affairs document management system, the present invention proposes a method of cross-domain knowledge fusion. In addition to integrating knowledge within the government affairs field, government affairs knowledge is also fused with knowledge in fields such as economy and society. For example, in terms of environmental protection policies, the system not only focuses on the content of the policy itself, but also associates economic data of enterprises, such as the market scale of the environmental protection industry and the investment of enterprises in environmental protection projects, as well as social and people's livelihood impacts, such as the improvement degree of the policy on the living environment of residents and the impact on employment. By constructing a cross-domain knowledge fusion model and exploring the deep relationships between knowledge in different fields, the system can provide more comprehensive and in-depth knowledge support for government affairs decision-making and meet the complex decision-making needs in government affairs work.

[0131] The urgency assessment module is used to clarify the timeliness of government affairs documents: In the design and implementation of the fuzzy logic model of this patent, through the refined processing of multi-dimensional input variables, the innovative application of the hybrid membership function, the construction of a comprehensively covered fuzzy rule base, and the introduction of a dynamic rule weight adjustment mechanism, the accuracy and adaptability of the urgency assessment of government affairs documents are significantly improved.

[0132] The fuzzification process of the multi-dimensional input variables and the design of the hybrid membership function include: For the time sensitivity (T s ), the influence scope (S i ), and the associated event level (E l ) of the three input variables, a hybrid membership function is designed;

[0133] Specifically, the hybrid membership function includes using a right semi-triangle function for the time sensitivity to define the membership degree of the "high" urgency level:

[0134]

[0135] This function linearly increases when T s ≥2, and the vertex is located at T s = 3 (the maximum time sensitivity), intuitively reflecting the actual scenario of "the closer the deadline, the higher the urgency".

[0136] Specifically, the hybrid membership function includes using a trapezoidal function for the influence scope to define the membership degree of the "area" range:

[0137]

[0138] This function maintains a membership degree of 1 in the interval Si ∈ [0.5, 2.5], and linearly decays on both sides, which is suitable for describing the smooth process of the influence scope transitioning from within the department to the city-wide policy.

[0139] Specifically, the hybrid membership function includes using a Gaussian function for the associated time level to define the membership degree of the "important" event:

[0140]

[0141] The symmetry and smoothness of the Gaussian function can effectively handle the uncertainty of the associated event level (such as the subjectivity in the evaluation of the emergency event level). The setting of the standard deviation σ = 0.5 makes the membership degree relatively high in the interval E l = 2 ± 0.5, which conforms to the fuzzy boundary characteristics of the "important" event.

[0142] Regarding the parameter settings, the purposes of the settings in this embodiment include but are not limited to: The vertex position (such as T s = 2) is determined based on the statistical analysis of the historical government document processing data to ensure that the time window of more than 80% of the extremely urgent documents is covered; the standard deviation (such as σ = 0.5) is optimized through cross-validation to balance the sensitivity and noise resistance of the model; the function type selection (triangle, trapezoid, Gaussian) is based on the variable characteristics: the time sensitivity requires a clear threshold, the influence scope requires a wide coverage, and the associated event requires a smooth transition.

[0143] Furthermore, aiming at the problem that traditional permission management systems rely only on simple rules and cannot handle complex scenarios with multi-factor coupling, this patent constructs a fuzzy rule base containing 27 rules that cover all combinations of input variables, and the specific implementation is as follows:

[0144] The goal of the comprehensively covering fuzzy rule base is to construct a comprehensive and detailed rule system by covering all possible combinations of input variables, so as to accurately map the emergency level determination logic in complex business scenarios. The generation process of the rule base starts with the in-depth analysis of input variables and the definition of fuzzy sets, where the time sensitivity (T s ), scope of influence (S i ), and associated event level (E l ) are respectively divided into three fuzzy sets: the time sensitivity is divided into "low", "medium", and "high", the scope of influence is divided into "local", "regional", and "global", and the associated event level is divided into "routine", "important", and "urgent". Each fuzzy set is mathematically expressed through specific types of membership functions (triangle, trapezoid, and Gaussian functions) to reflect the differences in the characteristics of different variables.

[0145] Based on the definition of fuzzy sets, the construction of the rule base follows the "IF-THEN" logical structure, covering all possible combinations of input variables, a total of 3×3×3 = 27 rules. Each rule maps the combination of fuzzy sets of input variables to the fuzzy set of the output variable (emergency level E); among them, the emergency level E includes three types: extremely urgent, urgent, and normal, and the specific rules corresponding to the emergency level are shown in Table 1.

[0146] Table 1 is the emergency level judgment logic

[0147]

[0148]

[0149] Regarding the dynamic rule weight adjustment mechanism, this patent proposes a feedback-driven rule weight optimization mechanism to solve the problem that traditional static rule bases are difficult to adapt to changes in business scenarios. The initial weight of each rule is set to 1, indicating that all rules contribute equally to the output in the initial state. During the operation of the system, by collecting the real-time feedback of users on the permission allocation results (such as a satisfaction score from 1 to 5), combined with the matching degree r predicted (calculated from the defuzzified score E and the actual marked emergency level), the rule weights are dynamically adjusted.

[0150]

[0151] $w_{i}(t)$ represents the weight of the $i$-th rule at time $t$. It reflects the influence degree of this rule on the final output result at the current moment. Initially, the weights of all rules are usually set to 1, and as the system runs, they are continuously adjusted according to the feedback. $w_{i}(t + 1)$ represents the weight of the $i$-th rule at time $t + 1$. It is the new weight value of the $i$-th rule after one weight adjustment. $\eta$ represents the learning rate, which is used to control the adjustment amplitude of the rule weights. It is a preset constant, and the reference value is 0.01. $r$ actual $m$ represents the actual matching degree of the urgency level. It is a value determined according to the matching situation between the actually marked urgency level and the urgency level predicted by the model, used to measure the degree of conformity between the model prediction and the actual situation, and further guide the adjustment of the rule weights.

[0152] Among them, the learning rate $\eta = 0.01$ controls the adjustment amplitude to prevent the weights from fluctuating violently. For example, if a certain rule frequently leads to excessive permission allocation (low user score) in the EPA scenario, its weight will gradually decrease, thus reducing its impact on the final output; conversely, in the rules for handling emergencies, if user feedback shows that permission elevation helps with rapid response, the weight will increase accordingly. This mechanism significantly improves the system's scenario adaptability by adaptively learning the business characteristics of different departments (e.g., the EPA pays more attention to the associated event level, and the finance department attaches more importance to time sensitivity).

[0153] Furthermore, the dynamic weight adjustment also works in coordination with the risk assessment model. When the real-time risk value $RR$ (calculated from the access frequency, IP anomaly index, etc.) is relatively high, the system will temporarily increase the weights of the rules related to security. For example, the weight of the rule "IF $E$ l = Emergency THEN Demote Weight" is increased to give priority to blocking potential risks. This coordination is achieved through the weight scaling factor $\lambda$ R $\lambda$ R The calculation formula of $\lambda$ is:

[0154]

[0155] $R$ represents the real-time risk value, $R$ crit represents the risk threshold;

[0156] Among them, $R$ crit = 1.5 is the risk threshold. For example, when $R = 1.8$, $\lambda$ R = 2.2, and the weights of the relevant security rules are amplified to 2.2 times the original value, thus strengthening the permission control in high-risk scenarios. At the same time, to prevent the model from deviating from the initial design due to long-term weight deviation, the system sets a weight decay mechanism, and the formula is:

[0157]

[0158] Among them, the attenuation coefficient γ = 0.001, which ensures that the weight gradually returns to the initial value without continuous feedback and maintains the stability of the system.

[0159] Through the design of the rule base to generate collaborative optimization with dynamic weight adjustment, the fuzzy logic model of this patent has achieved technological breakthroughs in many aspects in the evaluation of the urgency of government documents. Compared with traditional methods, 27 full-coverage rules combined with the dynamic weight mechanism have expanded the permission allocation granularity from level 3 to level 6, and the classification accuracy has been greatly improved. At the same time, the learning rate η and attenuation coefficient γ of weight adjustment are optimized through grid search and cross-validation to ensure a balance between the fast adaptability and long-term stability of the system. For example, in the simulation test, after the system receives 100 user feedbacks, the weight adjustment improves the emergency handling efficiency in the scenario of the Environmental Protection Bureau by 50%, and the processing delay of time-sensitive documents in the financial department is shortened to within 1 second. This technical advantage is not only reflected in the theoretical innovation of the mathematical model, but also verified through actual business scenarios, demonstrating its wide application potential in government intelligent management.

[0160] Dynamic permission adjustment of the permission management module based on risk assessment: Set up a permission management system according to the organizational structure of government departments and the confidentiality level of documents. Users are divided into roles such as ordinary clerks, section leaders, and department heads, and corresponding file access permissions are assigned to each role. For example, ordinary clerks can only access public and some internal documents, section leaders can access documents below the confidential level, and department heads have the highest permissions.

[0161] Traditional permission management usually adopts a static permission allocation method and cannot respond in a timely manner to potential risks during the file access process. The present invention introduces a dynamic permission adjustment mechanism based on risk assessment. The system will monitor information such as the access frequency of files and the access behavior patterns of users in real time, and quantitatively evaluate the file access risk through a risk assessment algorithm. When abnormal access behaviors are detected, such as frequent access to a large number of high-confidentiality files in a short period of time, access from an abnormal IP address, etc., the system will immediately start the dynamic permission adjustment program to temporarily restrict or freeze the access permissions of relevant users until the risk is lifted. This dynamic permission management method can more effectively guarantee the security of government documents and prevent data leakage and illegal access.

[0162] Collaborative management of encryption and permissions: To further enhance the security of government documents, the present invention collaboratively designs permission management and file encryption. The system uses advanced encryption algorithms to encrypt and store government documents, and files of different confidentiality levels use encryption keys of different strengths. In terms of permission management, only users with corresponding permissions can obtain specific decryption keys. For example, ordinary clerks can only obtain decryption keys for public files and some internal files, and the validity period and number of uses of the keys will also be limited according to the file confidentiality level and access risk. Through this method of collaborative management of encryption and permissions, even if the file data is illegally obtained, the file content cannot be viewed without the correct decryption key, thereby greatly improving the security of the file.

[0163] 3. System Operation and Operation

[0164] like Figure 1 As shown, the file classification process:

[0165] File upload: Staff upload newly generated government documents to the system through the client. For example, a staff member of the Municipal Environmental Protection Bureau wrote a policy document on river pollution control and uploaded it in Word format.

[0166] Preprocessing: The system automatically preprocesses the uploaded files, including format conversion (converting Word files into a unified text format) and denoising (removing irrelevant information such as headers, footers, advertisements, etc. in the file).

[0167] Feature extraction: CNN is used to extract features from the preprocessed file. The convolution kernel slides over the file text, extracting local features such as "river pollution control" and "environmental protection measures." These features are then passed to the RNN.

[0168] Classification decision: The RNN uses the features extracted by the CNN and its own understanding of the text sequence to determine the file's category. In this example, the model determines that the file belongs to the "Environmental Protection Policies and Regulations" category and stores it in the corresponding database folder.

[0169] like Figure 2 As shown, the file retrieval process:

[0170] Search term input: The staff enters the search term in the client, such as "relevant policies for pollution control of a certain river".

[0171] Semantic Understanding: The system converts search terms into vectors using a word embedding model and calculates semantic similarity with vectors in the document embedding library. Simultaneously, an attention mechanism assigns weights to different parts of the document based on the focus of the search term, highlighting key information related to the search term.

[0172] Result Screening and Sorting: The system screens out relevant documents based on semantic similarity and attention weights and sorts them in descending order of relevance. For example, in addition to direct pollution control policy documents, it is also associated with environmental monitoring reports of the river, relevant law enforcement cases, and other documents.

[0173] Result Display: The search results are displayed in a list form on the client side, and each document shows the file name, file abstract, and relevance score. Staff can click on the document to view it.

[0174] Permission Management and Security Assurance: When staff attempt to access a document, the system verifies it according to the permission management system. For example, when an ordinary clerk attempts to access a confidential-level financial budget document, the system will prompt insufficient permissions and deny access. Only the department head with the corresponding permissions can successfully open the document.

[0175] Format Compatibility Processing: When staff upload a document in scanned format, the system first uses OCR technology to convert the text in the image into editable text. After steps such as character segmentation, feature extraction, and classification recognition, the recognition accuracy reaches 98%. Subsequently, the system converts it into a unified format for subsequent processing to ensure the integrity and accuracy of the document content.

[0176] IV. System Maintenance and Optimization

[0177] Data Update: Regularly collect new government affairs documents and perform incremental training on the classification model and retrieval model to adapt to new policies, regulations, and business requirements. For example, update the data once a month, and the accuracy of the model remains stable during long-term operation.

[0178] User Feedback Processing: Collect user feedback on the search results, such as whether they are accurate and whether they meet the needs. According to the feedback information, adjust the parameters of the retrieval algorithm to optimize the search results. For example, if users feedback that the relevance of a certain type of search result is low, by analyzing the feedback data, adjust the weight distribution of the attention mechanism to improve the retrieval accuracy.

[0179] System Performance Monitoring: Real-time monitor performance metrics such as the CPU, memory, and disk I / O of the server to ensure the stable operation of the system. When performance bottlenecks are found, optimize in a timely manner, such as increasing the server memory and optimizing the database query statements.

[0180] Improve the accuracy of classification and retrieval, use intelligent algorithms to accurately understand the semantics of government affairs documents, reduce classification errors caused by the subjectivity of manual judgment, reduce the classification error rate to an extremely low level, and increase the retrieval accuracy to over 90%.

[0181] Aiming at the problem of insufficient security control in the general file management system, the present invention will build a refined permission management system to meet the multi-level security requirements of government affairs documents with different confidentiality levels. In the face of the rapid update of government affairs knowledge and complex file formats, the system has strong self-learning and adaptive capabilities, as well as excellent format compatibility, ensuring that it can adapt to new policies and regulations in a timely manner, accurately process government affairs documents of various formats, and achieve the intelligentization, high efficiency and standardization of government affairs document management.

[0182] The above are only the specific steps of the present invention and do not constitute any limitation to the protection scope of the present invention; all technical solutions formed by equivalent transformation or equivalent replacement fall within the scope of the present invention's rights protection; the parts not elaborated in detail in the present invention belong to the well-known technologies of those skilled in the art.

Claims

1. An intelligent classification and retrieval system for government affairs documents, characterized in that, The intelligent classification and retrieval system for government affairs documents includes a feature extraction module, a knowledge graph module, a retrieval module, an urgency assessment module, and a permission management module; The feature extraction module adopts a deep learning architecture that combines a convolutional neural network model and a recurrent neural network model. The convolutional neural network model sets multiple convolutional layers, and the size of the convolutional kernel is dynamically adjusted according to document features to extract local features of the document. The recurrent neural network model uses LSTM units and integrates a self-attention mechanism to capture the logical relationships of text sequences; The knowledge graph module constructs a knowledge graph for the government affairs domain, integrates cross-domain knowledge such as government affairs, economy, and society, and has a dynamic update mechanism to monitor government affairs data sources in real time and update entities and relationships; The retrieval module converts documents and retrieval terms into vectors with the help of a word vector model, calculates semantic similarity through the cosine similarity formula, and adaptively adjusts the attention weights according to the retrieval term popularity, user historical retrieval behavior, and document timeliness in combination with the attention mechanism to achieve accurate retrieval; The urgency assessment module clarifies the timeliness of government affairs documents based on a fuzzy rule base and a dynamic rule weight adjustment mechanism to reflect the urgency of government affairs documents; The permission management module combines the organizational structure of government departments and the document confidentiality level, designs an access control algorithm based on roles and confidentiality levels, introduces a risk assessment mechanism at the same time, dynamically adjusts permissions according to document access risks, and coordinates permission management with document encryption. Users with different permissions obtain different decryption keys.

2. The intelligent classification and retrieval system for government affairs documents according to claim 1, wherein The intelligent classification and retrieval system for government affairs documents includes a preprocessing module. The preprocessing module preprocesses the uploaded government affairs documents, including removing duplicate and invalid documents, and classifying and storing them according to departments, years, etc.; And convert them into a unified text format and remove the header, footer, and advertisement information in the documents.

3. The intelligent classification and retrieval system for government affairs documents according to claim 1, wherein The convolutional neural network model is based on a dynamic convolutional kernel selection method, and the specific steps are as follows: Extract the number of words L, the number of paragraphs P, the title level H, and the nested list depth D of the document, construct a multi-dimensional feature vector X = [L, P, H, D], calculate the complexity score, and divide the document into high-complexity documents, medium-complexity documents, and low-complexity documents based on the complexity score; Use convolutional kernels of different scales to extract semantic features of high-complexity documents, medium-complexity documents, and low-complexity documents respectively, and fuse the multi-scale features through weighted splicing; The convolutional neural network uses a government affairs domain labeled data set in the training stage, and the loss function is represented by cross-entropy with a class balance factor: Among them, to alleviate the problem of class imbalance, c is the class serial number, and N c is the number of samples in class c; the optimizer uses AdamW, and the learning rate is η = 0.

01.

4. The intelligent classification and retrieval system for government affairs documents according to claim 3, wherein The complexity score is expressed as: C = w T ·X norm + b Among them, w = [0.35, 0.25, 0.20, 0.20] T is the weight vector, X norm is the feature after normalization of the number of words L, the number of paragraphs P, the title level H, and the nested list depth D, and b is the bias term; Among them, when C ≥ 1.2, it is a high-complexity document; when 0.8 ≤ C < 1.2, it is a medium-complexity document; when C < 0.8, it is a low-complexity document.

5. The intelligent classification and retrieval system for government affairs documents according to claim 3, wherein The feature extraction processes of high-complexity documents, medium-complexity documents, and low-complexity documents are respectively: For high-complexity documents, a 7×7 convolutional kernel is used to extract long-range semantic features, and dilated convolution is stacked to expand the receptive field with a dilation rate r = 2, and the high-complexity feature map F is output large ; For medium-complexity files, a 5×5 convolutional kernel is used to extract semantic features, and a gating mechanism is combined to filter out noise features. The gating weight g i,j is as follows: g i,j = σ(W g ·F i,j + b g ) Among them, σ is the Sigmoid function, and W g is the weight matrix in the gating mechanism, which is used to weight the feature F i,j F i,j is the feature vector at the (i, j) position in the feature map obtained after the convolution operation, which contains the local feature information of the file, and b g is the bias term in the gating mechanism, which is used to adjust the overall level of the gating output; the complexity feature map in the output: F mid = g i,j ⊙F raw F raw is the original feature map without being processed by the gating mechanism when extracting semantic features of medium-complexity files using a 5x5 convolutional kernel; for low-complexity files, a 3×3 convolutional kernel is used to extract semantic features, and an attention module is inserted to strengthen key regions, and the spatial attention weight A spatial is as follows: A spatial = Sigmoid(W s · AvgPool(F sm ) + W c · MaxPool(F sm )) F small = A spatial ⊙F raw Sigmoid is a normalization function, W s and W c are weight matrices, AvgPool is an average pooling operation, F sm is the original feature map; F small is the low-complexity feature map weighted by attention, and MaxPool is a max pooling operation; Finally, the multi-scale features are fused through weighted splicing: F fusion = α·F large + β·F mid + γ·F small Among them, F fusion represents the fused feature map; The weight coefficients α, β, γ are dynamically calculated from the complexity score C, and the calculation formula is: k1, k2, k3 are parameters used to dynamically calculate the weight coefficients, and are constants determined during model training or design.

6. The intelligent classification and retrieval system for government affairs documents according to claim 3, wherein, The training phase uses a government affairs annotated dataset, which includes 10 categories and 500,000 documents. Specifically, the government affairs annotation dataset contains ten core annotation categories, covering document type classification, policy field division, institutional entity identification, policy term extraction, service process step annotation, people's livelihood demand classification, legal and regulatory clause positioning, time node annotation, responsible party association, and urgency classification.

7. The intelligent classification and retrieval system for government affairs documents according to claim 1, wherein The emergency assessment module constructs a fuzzy logic model through a mixed membership function; The hybrid membership function includes three variables for time sensitivity, impact range and associated event level; The right semi-triangular function is used to define the membership of time sensitivity: μ High (T s ) is the membership degree when the time sensitivity is "high", which is used to measure the matching degree between the urgency of government affairs documents in the time dimension and the "high" urgency level. Its value range is between 0 and 1. The larger the value, the closer the document is to the standard of "high" urgency level in terms of time; T s is the time sensitivity variable, which is used to describe the deadline or time-related attributes of government affairs documents and is the input variable for calculating the membership degree; This function grows linearly when T s ≥ 2, and the maximum time sensitivity is at T s = 3, indicating that the closer the deadline is, the higher the urgency; The degree of membership of the influence range is defined using the trapezoidal function for the influence range: μ Regional (S i ) is the membership degree when the influence range is "region", used to measure the matching degree between the influence range of government affairs documents and the "region" range. The value range is between 0 and 1. When the value is 1, it means that the influence range of the document completely conforms to the definition of "region", and the closer to 0, the lower the matching degree with the "region" range, S i is the influence range variable, used to describe the size of the influence range of government affairs documents, and is the input variable for calculating this membership degree; This function remains with a membership degree of 1 within the interval S i ∈[0.5, 2.5], linearly decays on both sides, and is applicable to describe the smooth process of the influence range transitioning from within the department to the city-wide policy; The membership of the correlation event level is defined using a Gaussian function for the correlation time level: μ Important (E l ) is the membership degree when the associated event level is "important", which is used to measure the matching degree between the level of the event associated with the government document and the "important" level. The larger its value, the more the associated event conforms to the "important" standard, and its value range is between 0 and 1. E l is the variable representing the level of the associated event, which is used to describe the importance level of the event associated with the government document and is the input variable for calculating this membership degree.

8. The intelligent classification and retrieval system for government affairs documents according to claim 7, wherein A fuzzy rule base containing 27 rules was constructed based on time sensitivity, impact scope and level of associated events; The rule base generation process begins with an in-depth analysis of the input variables and the definition of fuzzy sets. Time sensitivity, impact range, and associated event level are divided into three fuzzy sets based on their value ranges: time sensitivity is divided into "low," "medium," and "high," impact range is divided into "local," "regional," and "global," and associated event level is divided into "routine," "important," and "urgent." 9. The intelligent classification and retrieval system for government affairs documents according to claim 8, wherein Based on the definition of fuzzy sets, the construction of the rule base follows the IF-THEN logical structure, covering all possible combinations of input variables: Among them, low, local, and regular correspond to the urgency level of flat items; Low, local, and important correspond to the urgency level of flat items; Low, Local, and Urgent correspond to the urgency level of expedited; Low, regional, and regular correspond to the urgency level of flat items; Low, Regional, and Important correspond to the urgency level of expedited; Low, Regional, and Emergency correspond to the urgency level of expedited; Low, global, and regular correspond to the urgency level of flat items; Low, Global, and Important correspond to the urgency level of expedited; Low, Global, and Emergency correspond to the level of urgency as Extremely Urgent; Medium, local, and regular correspond to the urgency level of flat items; Medium, partial and important correspond to the urgency level of flat items; Medium, local, and emergency correspond to the highest urgency level; Medium, regional, and regular correspond to the urgency level of flat items; Medium, regional, and important correspondence urgency levels are expedited; The urgency level for medium, regional and emergency responses is extremely urgent; The urgency level of medium, global and regular responses is expedited; The urgency level for medium, global and important correspondence is urgent; The urgency level for medium, global and emergency responses is extremely urgent; High, local, and routine correspond to the urgency level of expedited; High, local, and important correspond to the urgency level of expedited; High, local, and emergency correspond to the level of urgency as extremely urgent; High, regional, and regular correspondence urgency levels are expedited; High, regional, and important correspond to the urgency level of urgent; High, regional, and emergency response levels are urgent; High, global, and regular correspondence urgency levels are urgent; High, global, and important correspond to the urgent level; High, Global, and Emergency correspond to the level of urgency for emergency.

10. The intelligent classification and retrieval system for government affairs documents according to claim 1, characterized in that, The dynamic rule weight adjustment mechanism includes: Initial weight of each rule All are set to 1, indicating that all rules contribute equally to the output in the initial state; by collecting real-time feedback from users on the permission allocation results and combining with the matching degree r of the urgency predicted by the model predicted , dynamically adjust the rule weights: represents the weight of the $i$-th rule at time $t$; represents the weight of the $i$-th rule at time $t + 1$; $\eta$ represents the learning rate; $r$ actual represents the matching degree of the actual urgency; The dynamic weight adjustment also works in concert with the risk assessment model, and is achieved collaboratively through the weight scaling factor λ R λ R The calculation formula of which is as follows: R represents the real-time risk value, R crit represents the risk threshold; To prevent long-term weight drift from causing the model to deviate from the initial design, the system sets a weight decay mechanism: Among them, the attenuation coefficient γ = 0.001, which ensures that the weights gradually return to the initial values without continuous feedback and maintains the stability of the system.

Citation Information

Patent Citations

  • Knowledge graph fine-grained access control method and system based on user behavior dynamic attributes

    CN116628226A

  • Archive system and method based on artificial intelligence

    CN119066201A

  • File classification management method

    CN119474371A

  • Government affair information intelligent retrieval and generation system applying RAG technology

    CN119807261A

  • Government affair information management system based on data matching

    CN119831812A

Cited By

  • Cross-department government affair data security sharing and analysis method based on deep learning

    CN120632954A

  • Intelligent file classification and retrieval method and system based on artificial intelligence

    CN121166924A