Document classification method and system
Through the method of decoding and extracting document meta information and keywords, documents are classified in multiple dimensions, which solves the problems of low efficiency and low accuracy of document classification in the prior art, and achieves efficient and accurate classification of diversified and complex documents.
Patent Information
- Application Number
- CN202510299498.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is inefficient and inaccurate in document classification, especially for diverse and complex document data.
By detecting the encoding method of the document and decoding process, the meta information and keywords of the document are extracted, and a multi-dimensional classification method is used to classify the document in combination with meta information and keywords.
Accurate classification of diversified and complex documents is achieved and classification efficiency is improved.
Smart Images

Figure CN120216701A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of flaw detection and testing, and particularly relates to a document classification method and system. Background Art
[0002] With the rapid development of information technology, a large number of documents are continuously generated and stored, and effectively managing and utilizing these documents has become an important challenge.
[0003] In related technologies, document classification usually relies on manual rules and statistical methods, which have certain limitations for diversified and complex document data, with low efficiency and low accuracy. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a document classification method, which classifies a target document in multiple dimensions according to the meta-information and target keywords of the target document, and can still accurately classify diversified and complex documents with high efficiency.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A document classification method includes the following steps: detecting the encoding method of a target document using the chardet library, and performing decoding processing on the target document according to the encoding method by adopting a corresponding decoding strategy to obtain a decoded document; respectively extracting document content and meta-information from the decoded document; extracting target keywords from the document content; and classifying the target document in multiple dimensions according to the meta-information and the target keywords.
[0007] In an embodiment of the present invention, the meta-information includes document creation time, document modification time, document author information, document type information, and document access permission information.
[0008] In an embodiment of the present invention, extracting the target keywords from the document content includes the following steps: performing word segmentation processing on the document content to obtain a vocabulary list; obtaining the vocabulary frequency, the first similarity with the topic name, and the position score of each vocabulary in the vocabulary list; calculating the key degree value of each vocabulary according to the vocabulary frequency, the first similarity, and the position score; and determining whether the corresponding vocabulary is the target keyword according to the key degree value.
[0009] In an embodiment of the present invention, the key degree value is calculated by the following formula:
[0010]
[0011] Wherein, G is the key degree value, f is the word frequency, k1 is the first weight, s is the first similarity, k2 is the second weight, d is the position score, e is the natural constant, k3 is the third weight, and α1, α2, β, and α3 are all constants greater than 0.
[0012] In an embodiment of the present invention, classifying the target document according to the target keyword includes: obtaining category keywords related to the field to which the target document belongs; vectorizing each of the target keywords and each of the category keywords; initializing the category scores corresponding to each of the category keywords to zero; respectively calculating the second similarity between each target keyword vector and the category keyword vector, and accumulating the calculated second similarity to the category score of the corresponding category keyword; and determining the category of the target document according to the category keyword with the highest category score.
[0013] A document classification system includes: a decoding module, the decoding device is used to detect the encoding method of the target document by using the chardet library, and perform decoding processing on the target document according to the encoding method by adopting a corresponding decoding strategy to obtain a decoded document; a first extraction module, the first extraction module is used to respectively extract document content and meta-information from the decoded document; a second extraction module, the second extraction module is used to extract target keywords from the document content; a classification module, the classification module is used to perform multi-dimensional classification on the target document according to the meta-information and the target keywords.
[0014] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the above-mentioned document classification method is implemented.
[0015] A non-transitory computer-readable storage medium stores a computer program thereon, wherein when the program is executed by a processor, the above-mentioned document classification method is implemented.
[0016] Advantages of the present invention:
[0017] The present invention performs multi-dimensional classification on the target document according to the meta-information and target keywords of the target document, and can still perform accurate classification on diverse and complex documents with high efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of the document classification method according to an embodiment of the present invention;
[0019] Figure 2 It is a block diagram of the document classification system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] Figure 1 It is a flowchart of the document classification method for the embodiments of the present invention.
[0022] As Figure 1 shown, the document classification method for the embodiments of the present invention may include the following steps:
[0023] S1. Use the chardet library to detect the encoding method of the target document, and perform decoding processing on the target document according to the encoding method by adopting a corresponding decoding strategy to obtain a decoded document.
[0024] S2. Extract the document content and meta-information from the decoded document respectively.
[0025] Among them, the meta-information includes the document creation time, document modification time, document author information, document type information, and document access permission information. Specifically, the meta-information can be obtained by reading the attribute information of the decoded document.
[0026] S3. Extract target keywords from the document content.
[0027] In an embodiment of the present invention, extracting target keywords from the document content may include the following steps:
[0028] S31. Perform word segmentation processing on the document content to obtain a vocabulary list.
[0029] S32. Obtain the vocabulary frequency, the first similarity with the theme name, and the position score of each vocabulary in the vocabulary list.
[0030] In an embodiment of the present invention, the vocabulary frequency of each vocabulary in the vocabulary list can be calculated by the following formula:
[0031]
[0032] Among them, f is the vocabulary frequency, N is the total number of vocabulary in the target document, and n is the number of times the current vocabulary appears in the target document.
[0033] In an embodiment of the present invention, the first similarity can be obtained by calculating the cosine similarity between each vocabulary in the vocabulary list and the theme name.
[0034] In one embodiment of the present invention, the positions where a vocabulary appears in a document may include the topic name position, the first-level heading position, the second-level heading position, the third-level heading position, and other positions. Different position scores may be set according to the importance of the content in different positions. For example, the position score corresponding to the topic name position is d1, the position score corresponding to the first-level heading position is d2, the position score corresponding to the second-level heading position is d3, the position score corresponding to the third-level heading position is d4, and the position score corresponding to other positions is d5, where d1 > d2 > d3 > d4 > d5. It should be noted that in other embodiments of the present invention, according to the different document structures, the corresponding position scores can be adaptively adjusted according to the importance of each position.
[0035] It can be understood that when the same vocabulary appears in multiple positions, the position score corresponding to the vocabulary can be calculated by calculating the average value of the position scores of the vocabulary in each position.
[0036] S33, Calculate the key degree value of each vocabulary according to the vocabulary frequency, the first similarity, and the position score.
[0037] In one embodiment of the present invention, the key degree value can be calculated by the following formula:
[0038]
[0039] Where G is the key degree value, f is the vocabulary frequency, k1 is the first weight, s is the first similarity, k2 is the second weight, d is the position score, e is the natural constant, k3 is the third weight, and α1, α2, β, and α3 are all constants greater than 0.
[0040] S34, Confirm whether the corresponding vocabulary is the target keyword according to the key degree value.
[0041] In one embodiment of the present invention, the vocabulary with a key degree value greater than or equal to the first threshold can be confirmed as the target keyword, where the first threshold can be calibrated according to the actual situation.
[0042] Thus, by calculating the key degree value of each vocabulary in the target document, the target keywords in the target document are extracted, thereby greatly improving the accuracy of keyword extraction in the target document, and further improving the accuracy of target document classification.
[0043] S4, Classify the target document in multiple dimensions according to the meta information and the target keywords.
[0044] In one embodiment of the present invention, classifying the target document according to the target keywords may include the following steps:
[0045] S41, Obtain the category keywords related to the field to which the target document belongs.
[0046] Among them, the category keywords of the documents in each field can be pre - counted to construct a corresponding database. The corresponding category keywords can be extracted from the database according to the field to which the target document belongs.
[0047] S42, vectorize each target keyword and each category keyword.
[0048] S43, initialize the category scores corresponding to each category keyword to zero.
[0049] S44, calculate the second similarity between each target keyword vector and the category keyword vector respectively, and accumulate the calculated second similarity to the category score of the corresponding category keyword.
[0050] S45, confirm the category of the target document according to the category keyword with the highest category score.
[0051] In an embodiment of the present invention, classifying the target document according to the meta - information may include: classification based on time: according to the creation time or modification time of the target document, the target document is divided into categories such as recent documents, historical documents, and long - unupdated documents, etc., to reflect the timeliness and usage frequency of the target document; classification based on author information: according to the author information of the target document, the target document is divided into categories such as personally created documents, team collaborative documents, and documents written by a specific organization, etc., to clarify the ownership relationship of the target document; classification based on document type: the document type information may include attribute information such as document format, purpose, and language. According to attributes such as document format, purpose, and language, the target document is divided into categories such as contracts, reports, emails, research papers, etc., for easy retrieval and management; classification based on access rights: according to the access rights of the target document, the target document is divided into categories such as public documents, internal documents, and confidential documents, etc., to ensure the security and compliance of the documents.
[0052] Thus, by classifying the target document in multiple dimensions according to the meta - information and target keywords of the target document, accurate classification can still be carried out for diverse and complex documents, and the efficiency is relatively high.
[0053] In summary, for the document classification method according to the embodiment of the present invention, the chardet library is used to detect the encoding method of the target document, and the target document is decoded according to the encoding method by adopting a corresponding decoding strategy to obtain the decoded document, and the document content and meta - information are respectively extracted from the decoded document, and the target keywords are extracted from the document content, and the target document is classified in multiple dimensions according to the meta - information and target keywords. Thus, by classifying the target document in multiple dimensions according to the meta - information and target keywords of the target document, accurate classification can still be carried out for diverse and complex documents, and the efficiency is relatively high.
[0054] Corresponding to the document classification method of the above embodiments, the present invention also proposes a document classification system.
[0055] As Figure 2 shown, the document classification system of the embodiments of the present invention may include: a decoding module 100, a first extraction module 200, a second extraction module 300, and a classification module 400.
[0056] Among them, the decoding device 100 is used to detect the encoding method of the target document by using the chardet library, and perform decoding processing on the target document according to the encoding method by adopting a corresponding decoding strategy to obtain a decoded document; the first extraction module 200 is used to extract document content and meta-information from the decoded document respectively; the second extraction module 300 is used to extract target keywords from the document content; the classification module 400 is used to perform multi-dimensional classification on the target document according to the meta-information and the target keywords.
[0057] In an embodiment of the present invention, the second extraction module 300 is specifically used for: performing word segmentation processing on the document content to obtain a vocabulary list; obtaining the vocabulary frequency, the first similarity with the topic name, and the position score of each vocabulary in the vocabulary list; calculating the key degree value of each vocabulary according to the vocabulary frequency, the first similarity, and the position score; and confirming whether the corresponding vocabulary is a target keyword according to the key degree value.
[0058] In an embodiment of the present invention, the second extraction module 300 is specifically used for: calculating the key degree value through the following formula:
[0059]
[0060] where G is the key degree value, f is the vocabulary frequency, k1 is the first weight, s is the first similarity, k2 is the second weight, d is the position score, e is the natural constant, k3 is the third weight, and α1, α2, β, and α3 are all constants greater than 0.
[0061] In an embodiment of the present invention, the classification module 400 is specifically used for: obtaining category keywords related to the field to which the target document belongs; vectorizing each target keyword and each category keyword; initializing the category score corresponding to each category keyword to zero; calculating the second similarity between each target keyword vector and the category keyword vector respectively, and accumulating the calculated second similarity to the category score of the corresponding category keyword; and confirming the category of the target document according to the category keyword with the highest category score.
[0062] It should be noted that for the details not disclosed in the document classification system of the embodiments of the present invention, please refer to the details disclosed in the above document classification method, and details will not be described in detail here.
[0063] According to the document classification system of the embodiments of the present invention, a decoding device uses the chardet library to detect the encoding method of a target document, and adopts a corresponding decoding strategy according to the encoding method to perform decoding processing on the target document to obtain a decoded document. Then, a first extraction module extracts document content and meta-information from the decoded document respectively, and a second extraction module extracts target keywords from the document content. Further, a classification module performs multi-dimensional classification on the target document according to the meta-information and the target keywords. Thus, multi-dimensional classification of the target document is performed according to the meta-information and the target keywords of the target document, and accurate classification can still be carried out for diverse and complex documents with relatively high efficiency.
[0064] Corresponding to the above embodiments, the present invention further provides a computer device.
[0065] The computer device according to the embodiments of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the document classification method of the above embodiments is implemented.
[0066] According to the computer device of the embodiments of the present invention, multi-dimensional classification of the target document is performed according to the meta-information and the target keywords of the target document, and accurate classification can still be carried out for diverse and complex documents with relatively high efficiency.
[0067] Corresponding to the above embodiments, the present invention further provides a non-transitory computer-readable storage medium.
[0068] The non-transitory computer-readable storage medium according to the embodiments of the present invention stores a computer program, and when the program is executed by a processor, the above-mentioned document classification method is implemented.
[0069] According to the non-transitory computer-readable storage medium of the embodiments of the present invention, multi-dimensional classification of the target document is performed according to the meta-information and the target keywords of the target document, and accurate classification can still be carried out for diverse and complex documents with relatively high efficiency.
[0070] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The meaning of "plurality" is two or more, unless otherwise specifically defined.
[0071] In the present invention, unless otherwise clearly defined or limited, terms such as "installed", "connected", "coupled", "fixed", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0072] In the present invention, unless otherwise clearly defined or limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0073] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0074] In addition, in each embodiment of the present invention, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0075] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A document classification method, characterized in that: The following steps are involved: The chardet library is used to detect the encoding method of the target document, and a corresponding decoding strategy is adopted according to the encoding method to decode the target document to obtain a decoded document; Extracting document content and meta information from the decoded document respectively; Extract target keywords from the document content; The target document is classified in multiple dimensions according to the meta information and the target keywords.
2. The document classification method according to claim 1, characterized in that: The meta information includes document creation time, document modification time, document author information, document type information, and document access permission information.
3. The document classification method according to claim 1, characterized in that: Extracting the target keyword from the document content comprises the following steps: Performing word segmentation processing on the document content to obtain a vocabulary list; Obtaining the vocabulary frequency of each word in the vocabulary list, the first similarity with the subject name, and the position score of the word occurrence; Calculate the criticality value of each word according to the word frequency, the first similarity and the position score; It is determined whether the corresponding word is the target keyword according to the criticality value.
4. The document classification method according to claim 3, characterized in that: The criticality value is calculated by the following formula: Among them, G is the criticality value, f is the vocabulary frequency, k1 is the first weight, s is the first similarity, k2 is the second weight, d is the position score, e is a natural constant, k3 is the third weight, α1, α2, β and α3 are all constants greater than 0.
5. The document classification method according to claim 1, characterized in that: Classifying the target document according to the target keyword includes: Obtaining category keywords related to the field to which the target document belongs; Vectorizing each of the target keywords and each of the category keywords; Initializing the category score corresponding to each of the category keywords to zero; Calculating the second similarity between each target keyword vector and the category keyword vector respectively, and adding the calculated second similarity to the category score of the corresponding category keyword; The category of the target document is determined according to the category keyword with the highest category score.
6. A document classification system, characterized in that: include: A decoding module, wherein the decoding device is used to detect the encoding method of the target document using the chardet library, and adopt a corresponding decoding strategy according to the encoding method to decode the target document to obtain a decoded document; A first extraction module, the first extraction module is used to extract document content and meta information from the decoded document respectively; A second extraction module, the second extraction module is used to extract target keywords from the document content; A classification module is used to perform multi-dimensional classification on the target document according to the meta information and the target keywords.
7. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the document classification method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the document classification method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Electronic archive information intelligent management system
CN117556112A
Automatic filing method and device of electronic file, electronic equipment and storage medium
CN118689850A