An intelligent sorting and archiving method and system for electronic documents based on large AI models
Through the intelligent sorting and archiving method of electronic files based on AI large model, the problem of inefficiency of traditional methods is solved, and the rapid classification and archiving of massive electronic files is achieved, and the classification accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510436221.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-09
AI Technical Summary
Traditional electronic file archiving methods are inefficient, manual processing is time-consuming and labor-intensive, and existing intelligent models are difficult to accurately capture the deep semantics of files, resulting in inaccurate classification, large computing resources consumption, and slow processing speed.
Using an AI big model-based method, we use the method to preprocess by receiving electronic text information, identify titles and paragraphs, refine paragraph text, perform semantic matching and weight calculations, and use the classification path recognition model to automatically determine archive categories and storage paths, reduce computing resource consumption, and improve classification accuracy and efficiency.
It realizes the rapid classification and archiving of massive electronic files, significantly improves work efficiency, accurately judges the file attribute category, and reduces the consumption of computing resources.
Smart Images

Figure CN119988614B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document filing, and specifically to an intelligent sorting and filing method and system for electronic documents based on an AI large model. Background Art
[0002] With the rapid development of information technology, electronic documents have become the main means of information storage and transmission in modern society. Their quantity has increased explosively. Traditional file classification and management methods mainly rely on manual annotation and rule-making based on experience, which can still cope when the number of documents is small and the types are relatively single, but are overwhelmed when faced with a large number and diverse electronic documents. The traditional process of classifying and filing electronic documents and determining the retention period usually involves manually reading the document content, annotating according to a preset classification system, and determining the retention years based on factors such as document type and importance. This method has the problem of low efficiency. Manually processing documents is time-consuming and laborious, especially when the number of documents is huge, the processing speed becomes a bottleneck and it is difficult to meet the needs of efficient management. Currently, some intelligent models are used to automatically classify electronic documents. However, the content of electronic documents is complex and changeable, containing rich semantic information. Existing methods often rely on keyword matching or simple rule judgment, and it is difficult to accurately capture the deep semantics of the documents, resulting in inaccurate classification. In addition, when the document content involves multiple topics or classification tags, existing methods are difficult to accurately judge their belonging categories and often require manual intervention for the final decision. Moreover, existing methods perform word segmentation and text feature extraction on the full text of the document, consuming a large amount of computing resources and having a slow processing speed. Therefore, an intelligent sorting and filing method and system for electronic documents based on an AI large model are needed to solve the above problems. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide an intelligent sorting and filing method and system for electronic documents based on an AI large model to solve the problems in the above background art.
[0004] The present invention is implemented as follows. An intelligent sorting and filing method for electronic documents based on an AI large model, the method includes the following steps:
[0005] Receive electronic text information, perform preprocessing, identify the title and paragraphs of the electronic text information to obtain a title information and several paragraph texts;
[0006] Input all the paragraph texts into the AI large model for refined description to obtain refined statements for each paragraph text;
[0007] Perform semantic matching on the title information and refined statements with classification tags to determine the classification tags corresponding to the title information and each paragraph text;
[0008] Calculate the weight value of each classification label according to the title information weight and the paragraph text weight, and obtain the classification label arrangement information. There are N labels arranged in order in the classification label arrangement information;
[0009] Input the classification label arrangement information into the classification path recognition model to obtain the filing category and storage path. The classification path recognition model is trained based on the classification labels arranged in order;
[0010] Determine the storage life of the electronic text information according to the weight value of each classification label.
[0011] As a further solution of the present invention: The step of inputting all the paragraph texts into the AI large model for refined description to obtain the refined sentences of each paragraph text specifically includes:
[0012] Determine the number of tokens of each paragraph text. When the number of tokens is greater than the threshold M, segment the paragraph text. After segmentation, determine the number of tokens of each paragraph text again;
[0013] Determine the number of tokens of the output sentence according to the number of tokens of the paragraph text, determine the output requirement, and input the paragraph text + output requirement into the AI large model to obtain the refined sentence.
[0014] As a further solution of the present invention: The step of semantically matching the title information and the refined sentences with the classification labels to determine the classification labels corresponding to the title information and each paragraph text specifically includes:
[0015] Retrieve the label introduction information corresponding to each classification label, and perform word segmentation on the title information, the refined sentences, and the label introduction information;
[0016] Use the bag-of-words model, TF-IDF, or word embedding to extract the text features of the text content after word segmentation;
[0017] Based on the text features and the BERT model, perform semantic similarity calculation to determine the classification labels corresponding to the title information and each paragraph text.
[0018] As a further solution of the present invention: The step of calculating the weight value of each classification label according to the title information weight and the paragraph text weight to obtain the classification label arrangement information specifically includes:
[0019] Determine the weight value of the classification label corresponding to the title information according to the title information weight;
[0020] Allocate the total paragraph weight according to the number of characters of each paragraph text to determine the weight value of the classification label corresponding to each paragraph text;
[0021] Merge the same classification labels, and obtain the classification label arrangement information based on the top N classification labels in terms of weight values.
[0022] As a further solution of the present invention: the step of determining the storage life of the electronic text information according to the weight value of each classification label specifically includes:
[0023] Retrieve the corresponding storage time of each classification label, with a total of n classification labels;
[0024] Determine the storage life H of the electronic text information according to the weight values of all the classification labels, , where represents the storage time of the i-th classification label, represents the weight value of the i-th classification label.
[0025] As a further solution of the present invention: the training steps of the classification path recognition model are as follows:
[0026] Collect classification label groups, each classification label group contains N sorted classification labels, and each classification label group is marked with an archiving category and a storage path;
[0027] Perform one-hot encoding on the classification label groups, and divide them into a training set, a validation set and a test set;
[0028] Determine that the input layer is a fully connected layer, add one or more convolutional layers to extract features in the classification label groups, add a pooling layer after the convolutional layer, determine the output layer, and use the softmax activation function and the number of output nodes equal to the number of archiving category and storage path categories;
[0029] Perform model training, determine that the loss function is the cross-entropy loss function, and select Adam or RMSprop as the optimizer;
[0030] Use accuracy, recall and F1 score metrics to evaluate the performance of the model on the test set, and adjust the model structure or hyperparameters.
[0031] Another object of the present invention is to provide an intelligent sorting and archiving system for electronic files based on an AI large model, and the system includes:
[0032] A title paragraph recognition module, which is used to receive electronic text information, perform preprocessing, perform title recognition and paragraph recognition on the electronic text information, and obtain a title information and several paragraph texts;
[0033] A refined statement determination module, which is used to input all the paragraph texts into the AI large model for refined description, and obtain the refined statements of each paragraph text;
[0034] A classification label determination module, configured to perform semantic matching on the title information and the refined statements with classification labels to determine the classification labels corresponding to the title information and each paragraph text;
[0035] A classification label arrangement module, configured to calculate the weight value of each classification label according to the title information weight and the paragraph text weight to obtain classification label arrangement information, where there are N labels arranged in order in the classification label arrangement information;
[0036] An archiving path determination module, configured to input the classification label arrangement information into a classification path recognition model to obtain an archiving category and a storage path, where the classification path recognition model is trained based on the classification labels arranged in order;
[0037] A storage period calculation module, configured to determine the storage period of the electronic text information according to the weight value of each classification label.
[0038] As a further solution of the present invention: The refined statement determination module includes:
[0039] A paragraph token number determination unit, configured to determine the token number of each paragraph text. When the token number is greater than a threshold M, the paragraph text is segmented. After segmentation, the token number of each paragraph text is determined again;
[0040] An output requirement determination unit, configured to determine the token number of the output statement according to the paragraph text token number, determine the output requirement, and input the paragraph text + output requirement into an AI large model to obtain a refined statement.
[0041] As a further solution of the present invention: The classification label determination module includes:
[0042] A text word segmentation processing unit, configured to retrieve the label introduction information corresponding to each classification label, and perform word segmentation processing on the title information, the refined statement, and the label introduction information;
[0043] A text feature extraction unit, configured to extract the text features of the text content after word segmentation processing using a bag-of-words model, TF-IDF, or word embedding;
[0044] A classification label determination unit, configured to calculate the semantic similarity based on the text features and the BERT model to determine the classification labels corresponding to the title information and each paragraph text.
[0045] As a further solution of the present invention: The classification label arrangement module includes:
[0046] A first weight value calculation unit, configured to determine the weight value of the classification label corresponding to the title information according to the title information weight;
[0047] The second weight value calculation unit is used to allocate the total paragraph weight according to the number of characters in each paragraph text, and determine the weight value of the classification label corresponding to each paragraph text;
[0048] The arrangement information determination unit is used to merge the same classification labels, and obtain the classification label arrangement information based on the classification labels with the top N weight values.
[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0050] By obtaining the refined sentences of each paragraph text, the present invention only needs to analyze and process the refined sentences subsequently, without processing all the contents of the paragraph text, greatly reducing the consumption of computing resources. The present invention calculates the weight value of each classification label according to the title information weight and the paragraph text weight, obtains the classification label arrangement information, and inputs the classification label arrangement information into the classification path recognition model, then the filing category and storage path will be automatically output. In this way, when a file involves multiple classification labels, its belonging category can also be accurately judged, improving the rationality of classification. The storage life of the electronic text information will also be determined according to the weight value of each classification label, and through the automatic processing ability of the AI large model, the rapid classification and filing of a large number of electronic files are realized, significantly improving the work efficiency. Brief Description of the Drawings
[0051] Figure 1 It is a flowchart of an intelligent sorting and filing method for electronic files based on an AI large model.
[0052] Figure 2 It is a flowchart of determining refined sentences in an intelligent sorting and filing method for electronic files based on an AI large model.
[0053] Figure 3 It is a flowchart of determining classification labels in an intelligent sorting and filing method for electronic files based on an AI large model.
[0054] Figure 4 It is a flowchart of obtaining classification label arrangement information in an intelligent sorting and filing method for electronic files based on an AI large model.
[0055] Figure 5 It is a flowchart of determining the storage life in an intelligent sorting and filing method for electronic files based on an AI large model.
[0056] Figure 6 It is a flowchart of training a classification path recognition model in an intelligent sorting and filing method for electronic files based on an AI large model.
[0057] Figure 7 It is a schematic structural diagram of an intelligent sorting and filing system for electronic files based on an AI large model. Detailed implementation manners
[0058] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention.
[0059] The following describes in detail the specific implementation of the present invention with reference to specific embodiments.
[0060] As Figure 1 shown, an embodiment of the present invention provides an intelligent sorting and archiving method for electronic files based on an AI large model. The method includes the following steps:
[0061] S100: Receive electronic text information, perform preprocessing, identify the title and paragraphs of the electronic text information, and obtain a title information and several paragraph texts;
[0062] S200: Input all the paragraph texts into the AI large model for refined description to obtain refined statements for each paragraph text;
[0063] S300: Perform semantic matching on the title information and the refined statements with classification labels to determine the classification labels corresponding to the title information and each paragraph text;
[0064] S400: Calculate the weight value of each classification label according to the title information weight and the paragraph text weight to obtain classification label arrangement information, and there are N labels arranged in order in the classification label arrangement information;
[0065] S500: Input the classification label arrangement information into a classification path recognition model to obtain an archiving category and a storage path, and the classification path recognition model is trained based on the classification labels arranged in order;
[0066] S600: Determine the storage life of the electronic text information according to the weight value of each classification label.
[0067] In the embodiments of the present invention, after receiving the electronic text information uploaded by the user, preprocessing is first performed, and the preprocessing includes format unification and data cleaning. Then, title recognition and paragraph recognition are performed on the electronic text information to obtain a title information and several paragraph texts; then all the paragraph texts are input into the AI large model for refined description to obtain the refined sentences of each paragraph text. The AI large model can select ChatGPT, Deepseek, Tongyi Qianwen, etc. Subsequently, only the refined sentences need to be analyzed and processed, without processing all the content of the paragraph texts, significantly reducing the consumption of computing resources. Then, semantic matching is performed between the title information and the refined sentences and the classification labels to determine the classification labels corresponding to the title information and each paragraph text. Specifically, which classification labels need to be set in advance. Then, the weight value of each classification label is calculated according to the title information weight and the paragraph text weight. For example, the title information weight is 30%, and the sum of the weights of all paragraph texts is 70%, obtaining the classification label arrangement information. There are N labels arranged in order in the classification label arrangement information. For example, N is five. Then, the classification label arrangement information is input into the classification path recognition model, and the filing category and storage path will be automatically output. The classification path recognition model is pre-trained based on the classification labels arranged in order. The content and arrangement order of the N labels will affect the final output result. In this way, when a file involves multiple classification labels, its belonging category can also be accurately judged, improving the rationality of classification. Finally, the storage period of the electronic text information is also determined according to the weight value of each classification label, significantly improving work efficiency.
[0068] As Figure 2 shown, as a preferred embodiment of the present invention, the step of inputting all the paragraph texts into the AI large model for refined description to obtain the refined sentences of each paragraph text specifically includes:
[0069] S201, determine the token number of each paragraph text. When the token number is greater than the threshold M, segment the paragraph text. After segmentation, determine the token number of each paragraph text again;
[0070] S202, determine the token number of the output sentence according to the paragraph text token number, determine the output requirement, and input the paragraph text + output requirement into the AI large model to obtain the refined sentence.
[0071] In the embodiments of the present invention, before inputting the paragraph text into the AI large model, it is necessary to determine the number of tokens of each paragraph text. When the number of tokens is greater than the threshold M, for example, M = 4000, the paragraph text is segmented. When segmenting, the number of segments is determined according to the number of tokens divided by M. After segmentation, the number of tokens of each paragraph text is determined again. Then, the number of tokens of the output statement is determined according to the number of tokens of the paragraph text. For example, the number of tokens of the output statement = the number of tokens of the paragraph text / 4, and the number of tokens of the output statement has a minimum limit, for example, 50. Then, the output requirement is determined, and the paragraph text + the output requirement is input into the AI large model to obtain the refined statement. For example, the number of tokens of the output statement corresponding to a certain paragraph text is 337, and 337 / 4 is equal to 84.25. Then the output requirement is: refine the content of the paragraph text and output an answer content of about 84 tokens.
[0072] As Figure 3 shown, as a preferred embodiment of the present invention, the step of semantically matching the title information and the refined statement with the classification label to determine the classification label corresponding to the title information and each paragraph text specifically includes:
[0073] S301, retrieve the label introduction information corresponding to each classification label, and perform word segmentation on the title information, the refined statement, and the label introduction information;
[0074] S302, use the bag-of-words model, TF-IDF, or word embedding to extract the text features of the text content after word segmentation;
[0075] S303, perform semantic similarity calculation based on the text features and the BERT model to determine the classification label corresponding to the title information and each paragraph text.
[0076] In the embodiments of the present invention, many classification labels will be formulated, and each classification label corresponds to label introduction information. First, the label introduction information corresponding to the title information and the classification label of each paragraph text will be retrieved, and then word segmentation will be performed on the title information, the refined statement, and the label introduction information. Here, the WordPiece word segmentation method is used. Then, the bag-of-words model, TF-IDF, or word embedding is used to extract the text features of the text content after word segmentation. These features should be able to reflect information such as the theme, keywords, and semantics of the text. Then, semantic similarity calculation is performed based on the text features and the BERT model to determine the classification label corresponding to the title information and each paragraph text.
[0077] As Figure 4 shown, as a preferred embodiment of the present invention, the step of calculating the weight value of each classification label according to the title information weight and the paragraph text weight to obtain the classification label arrangement information specifically includes:
[0078] S401. Determine the weight value of the classification label corresponding to the title information according to the title information weight;
[0079] S402. Allocate the total paragraph weight according to the number of characters in each paragraph text, and determine the weight value of the classification label corresponding to each paragraph text;
[0080] S403. Merge the same classification labels, and obtain the classification label arrangement information based on the top N classification labels by weight value.
[0081] In the embodiment of the present invention, first, the weight value of the classification label corresponding to the title information is determined according to the title information weight. For example, it is 30% here. Then, the total paragraph weight is allocated according to the number of characters in each paragraph text. The total paragraph weight is 70%. After the total weight allocation is completed, the weight value of the classification label corresponding to each paragraph text can be determined. It is also necessary to merge the same classification labels. Finally, the top N classification labels by weight value are determined to obtain the classification label arrangement information. N is a fixed value. For example, the classification label arrangement information is: label 26, label 41, label 12, label 87, and label 63; when the number of classification labels is less than N, all are arranged.
[0082] As Figure 5 shown, as a preferred embodiment of the present invention, the step of determining the storage life of the electronic text information according to the weight value of each classification label specifically includes:
[0083] S601. Retrieve the storage time corresponding to each classification label, with a total of n classification labels;
[0084] S602. Determine the storage life H of the electronic text information according to the weight values of all the classification labels.
[0085] In the embodiment of the present invention, the storage time is set in advance for each classification label. In order to determine the storage life of the electronic text information, the storage life H of the electronic text information is determined according to the weight values of all the classification labels of the electronic text information. , where represents the storage time of the i-th classification label, represents the weight value of the i-th classification label, .
[0086] As Figure 6 shown, as a preferred embodiment of the present invention, the training steps of the classification path recognition model are as follows:
[0087] S011. Collect classification label groups, where each classification label group contains N sequentially arranged classification labels, and each classification label group is marked with an archiving category and a storage path;
[0088] S012. Perform one-hot encoding on the classification label groups, and divide them into a training set, a validation set, and a test set;
[0089] S013. Determine that the input layer is a fully connected layer, add one or more convolutional layers to extract features from the classification label groups, add a pooling layer after the convolutional layer, determine the output layer, and use the softmax activation function and the number of output nodes equal to the number of archiving categories and storage path categories;
[0090] S014. Conduct model training, determine that the loss function is the cross-entropy loss function, and select Adam or RMSprop as the optimizer;
[0091] S015. Evaluate the performance of the model on the test set using accuracy, recall, and F1-score metrics, and adjust the model structure or hyperparameters.
[0092] In the embodiments of the present invention, it is necessary to pre-train a classification path recognition model. First, a large number of classification label groups need to be collected. Each classification label group contains N or fewer sequentially arranged classification labels, and each classification label group is marked with an archiving category and a storage path. Then, each classification label is converted into a numerical form, one-hot encoding is used, and the data set is divided into a training set, a validation set, and a test set. Then, a CNN model is constructed: determine that the input layer is a fully connected layer, add one or more convolutional layers to extract features from the classification label groups. The size and number of convolutional kernels can be adjusted according to data characteristics and task requirements. Activation functions such as ReLU are used to increase non-linearity. A pooling layer is added after the convolutional layer to reduce the dimension of the feature map, reduce the amount of calculation, and enhance the robustness of the model; determine the output layer, and use the softmax activation function and the number of output nodes equal to the number of archiving categories and storage path categories. Then, model training is carried out: determine that the loss function is the cross-entropy loss function, and select Adam or RMSprop as the optimizer; use the training set data to train the model, adjust the model parameters through backpropagation. At the end of each epoch, use the validation set to evaluate the model performance, and adjust hyperparameters such as the learning rate and batch size as needed. Finally, evaluate the performance of the model on the test set using accuracy, recall, and F1-score metrics, and adjust the model structure (such as the size and number of convolutional kernels, the number of nodes in the fully connected layer, etc.) or hyperparameters (such as the learning rate, batch size, number of training epochs, etc.) according to the evaluation results.
[0093] As Figure 7 shown, the embodiments of the present invention also provide an intelligent electronic file sorting and archiving system based on an AI large model. The system includes:
[0094] The title paragraph recognition module 100 is used to receive electronic text information, perform preprocessing, recognize the title and paragraphs of the electronic text information, and obtain a title information and several paragraph texts;
[0095] The refined statement determination module 200 is used to input all the paragraph texts into the AI large model for refined description to obtain the refined statements of each paragraph text;
[0096] The classification label determination module 300 is used to perform semantic matching on the title information and the refined statements with the classification labels to determine the classification labels corresponding to the title information and each paragraph text;
[0097] The classification label arrangement module 400 is used to calculate the weight value of each classification label according to the title information weight and the paragraph text weight to obtain the classification label arrangement information, and there are N labels arranged in order in the classification label arrangement information;
[0098] The filing path determination module 500 is used to input the classification label arrangement information into the classification path recognition model to obtain the filing category and the storage path, and the classification path recognition model is trained based on the classification labels arranged in order;
[0099] The storage life calculation module 600 is used to determine the storage life of the electronic text information according to the weight value of each classification label.
[0100] As a preferred embodiment of the present invention, the refined statement determination module 200 includes:
[0101] The paragraph token number determination unit is used to determine the token number of each paragraph text. When the token number is greater than the threshold M, the paragraph text is segmented. After segmentation, the token number of each paragraph text is determined again;
[0102] The output requirement determination unit is used to determine the token number of the output statement according to the paragraph text token number, determine the output requirement, and input the paragraph text + output requirement into the AI large model to obtain the refined statement.
[0103] As a preferred embodiment of the present invention, the classification label determination module 300 includes:
[0104] The text word segmentation processing unit is used to retrieve the label introduction information corresponding to each classification label and perform word segmentation processing on the title information, the refined statements, and the label introduction information;
[0105] The text feature extraction unit is used to extract the text features of the text content after word segmentation processing using the bag-of-words model, TF-IDF, or word embedding;
[0106] A classification label determination unit, configured to calculate semantic similarity based on text features and a BERT model, and determine classification labels corresponding to the title information and each paragraph of text.
[0107] As a preferred embodiment of the present invention, the classification label arrangement module 400 includes:
[0108] A first weight value calculation unit, configured to determine the weight value of the classification label corresponding to the title information according to the title information weight;
[0109] A second weight value calculation unit, configured to allocate the total paragraph weight according to the number of characters in each paragraph of text, and determine the weight value of the classification label corresponding to each paragraph of text;
[0110] An arrangement information determination unit, configured to merge the same classification labels, and obtain classification label arrangement information based on the top N classification labels by weight value.
[0111] The above only describes the preferred embodiments of the present invention in detail, and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0112] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0113] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0114] After considering the specification and the disclosure of the embodiments, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. An intelligent sorting and archiving method for electronic documents based on an AI large model, characterized in that The method includes the following steps: Receive electronic text information, perform preprocessing, identify the title and paragraphs of the electronic text information, and obtain a title information and several paragraph texts; Input all the paragraph texts into the AI large model for refined description to obtain refined statements for each paragraph text; Perform semantic matching on the title information and refined statements with classification labels to determine the classification labels corresponding to the title information and each paragraph text; Calculate the weight value of each classification label according to the title information weight and paragraph text weight to obtain classification label arrangement information, in which there are N labels arranged in order; Input the classification label arrangement information into the classification path recognition model to obtain the filing category and storage path, and the classification path recognition model is trained based on the classification labels arranged in order; Determine the storage life of the electronic text information according to the weight values of each classification label, specifically including: retrieving the corresponding storage time of each classification label, with a total of n classification labels; determining the storage life H of the electronic text information according to the weight values of all the classification labels, , where, represents the storage time of the i-th classification label, represents the weight value of the i-th classification label.
2. The intelligent sorting and archiving method of electronic documents based on the AI large model according to claim 1, wherein, The step of inputting all the paragraph texts into the AI large model for refined description to obtain refined statements for each paragraph text specifically includes: Determine the number of tokens of each paragraph text. When the number of tokens is greater than the threshold M, segment the paragraph text. After segmentation, determine the number of tokens of each paragraph text again; Determine the number of tokens of the output statement according to the number of tokens of the paragraph text, determine the output requirement, and input the paragraph text + output requirement into the AI large model to obtain the refined statement.
3. The intelligent sorting and archiving method of electronic documents based on the AI large model according to claim 1, characterized in that, The step of performing semantic matching on the title information and refined statements with classification labels to determine the classification labels corresponding to the title information and each paragraph text specifically includes: Retrieve the label introduction information corresponding to each classification label, and perform word segmentation on the title information, refined statements, and label introduction information; Use the bag-of-words model, TF-IDF or word embedding to extract the text features of the text content after word segmentation; Perform semantic similarity calculation based on the text features and the BERT model to determine the classification labels corresponding to the title information and each paragraph text.
4. The intelligent sorting and archiving method for electronic documents based on the AI large model according to claim 1, characterized in that The step of calculating the weight value of each classification label according to the title information weight and paragraph text weight to obtain classification label arrangement information specifically includes: Determine the weight value of the classification label corresponding to the title information according to the title information weight; Allocate the total paragraph weight according to the character number of each paragraph text to determine the weight value of the classification label corresponding to each paragraph text; Merge the same classification labels, and obtain the classification label arrangement information based on the classification labels with the top N weight values.
5. The intelligent sorting and archiving method of electronic documents based on the AI large model according to claim 1, characterized in that, The training steps of the classification path recognition model are as follows: Collect classification label groups, each classification label group contains N classification labels arranged in order, and each classification label group is marked with a filing category and a storage path; Perform one-hot encoding on the classification label groups, and divide them into a training set, a validation set, and a test set; Determine that the input layer is a fully connected layer, add one or more convolutional layers to extract the features in the classification label groups, add a pooling layer after the convolutional layer, determine the output layer, and use the softmax activation function and the number of output nodes equal to the number of filing category and storage path categories; Perform model training, determine that the loss function is the cross-entropy loss function, and select Adam or RMSprop as the optimizer; Evaluate the performance of the model on the test set using accuracy, recall, and F1-score metrics, and adjust the model structure or hyperparameters.
6. An intelligent electronic document sorting and archiving system based on a large AI model, characterized in that, The system includes: A title paragraph recognition module, which is used to receive electronic text information, perform preprocessing, identify the title and paragraphs of the electronic text information, and obtain a title information and several paragraph texts; A refined statement determination module, which is used to input all the paragraph texts into the AI large model for refined description to obtain the refined statements of each paragraph text; A classification label determination module, which is used to semantically match the title information and the refined statements with the classification labels to determine the classification labels corresponding to the title information and each paragraph text; A classification label arrangement module, which is used to calculate the weight values of each classification label according to the title information weight and the paragraph text weight to obtain classification label arrangement information, and there are N labels arranged in order in the classification label arrangement information; An archiving path determination module, which is used to input the classification label arrangement information into the classification path recognition model to obtain the archiving category and the storage path, and the classification path recognition model is trained based on the classification labels arranged in order; A storage life calculation module, which is used to determine the storage life of the electronic text information according to the weight values of each classification label, specifically including: retrieving the corresponding storage time of each classification label, with a total of n classification labels; determining the storage life H of the electronic text information according to the weight values of all the classification labels, , where, represents the storage time of the i-th classification label, represents the weight value of the i-th classification label.
7. The intelligent electronic document sorting and archiving system based on the AI large model according to claim 6, characterized in that The refined statement determination module includes: A paragraph token number determination unit, which is used to determine the token number of each paragraph text. When the token number is greater than the threshold M, the paragraph text is segmented. After segmentation, the token number of each paragraph text is determined again; An output requirement determination unit, which is used to determine the token number of the output statement according to the paragraph text token number, determine the output requirement, and input the paragraph text + output requirement into the AI large model to obtain the refined statement.
8. The intelligent sorting and archiving system for electronic files based on the AI large model according to claim 6, wherein, The classification label determination module includes: A text tokenization processing unit, which is used to retrieve the label introduction information corresponding to each classification label and perform tokenization processing on the title information, the refined statements, and the label introduction information; A text feature extraction unit, which is used to extract the text features of the text content after tokenization processing using the bag-of-words model, TF-IDF, or word embedding; A classification label determination unit, which is used to calculate the semantic similarity based on the text features and the BERT model to determine the classification labels corresponding to the title information and each paragraph text.
9. The intelligent electronic document sorting and archiving system based on the AI large model according to claim 6, wherein The classification label arrangement module includes: A first weight value calculation unit, which is used to determine the weight value of the classification label corresponding to the title information according to the title information weight; A second weight value calculation unit, which is used to allocate the total paragraph weight according to the character number of each paragraph text to determine the weight value of the classification label corresponding to each paragraph text; An arrangement information determination unit, which is used to merge the same classification labels and obtain the classification label arrangement information based on the top N classification labels with weight values.
Citation Information
Patent Citations
Remote resource arrangement service method and system based on information technology
CN116775972A
Data classification tag identification method and device, storage medium and electronic equipment
CN117609506A
Archiving management method and system for massive small files
CN118964301A
Text clustering method and device, computer equipment and storage medium
CN119128150A