Document classification method and device, equipment and medium
By automatically processing documents and using machine learning models to build document classifiers, the problems of inefficiency of traditional document classification methods and inaccurate classification results are solved, and efficient and accurate document classification and large-scale document management are achieved.
Patent Information
- Application Number
- CN202510243351.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional document classification methods rely on manual operations and preset rules, resulting in inefficiency, inaccurate classification results, and difficult to meet the needs of large-scale document management.
By obtaining pre-generated classification corpus and domain vocabulary, a document classifier is built based on machine learning models, and documents are processed automatically, which reduces the influence of subjective factors and improves classification accuracy and efficiency.
It significantly improves classification efficiency, reduces manual operation time, improves the objectivity and consistency of classification results, and enhances the management capabilities of large-scale documents.
Smart Images

Figure CN120144761A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and particularly to a document classification method, apparatus, device, and medium. Background Art
[0002] With the rapid development of information technology, enterprises generate a large number of documents every day. These documents are rich in content but lack structure, posing challenges to enterprise information management and decision-making. As an information organization and management technology, document classification can help enterprises quickly identify and retrieve important information, improving work efficiency.
[0003] Traditional document classification methods mainly rely on manual operations and classification based on preset rules. This manual classification method often faces the problem of low efficiency when dealing with a large amount of document materials. Due to the limitations of manual operations, the classification process is easily affected by subjective factors, resulting in inaccurate classification results. In addition, with the continuous growth of the number of documents, this method relying on manpower for classification is difficult to meet the needs of large-scale document management, and errors are extremely likely to occur during the classification process, thus affecting the quality and efficiency of document management. Summary of the Invention
[0004] One or more embodiments of this specification provide a document classification method, apparatus, device, and medium for solving the technical problems raised in the background art.
[0005] One or more embodiments of this specification adopt the following technical solutions:
[0006] A document classification method provided by one or more embodiments of this specification, the method includes:
[0007] Obtain a pre-generated classification corpus, the classification corpus includes multiple documents of a specified enterprise, and the multiple documents include a pre-set training set of documents and a test set of documents;
[0008] Obtain the domain vocabulary of the specified enterprise, the domain vocabulary includes preset vocabulary;
[0009] Based on the relevance between the preset vocabulary in the domain vocabulary and each document, obtain the keywords of each document;
[0010] Convert the training set documents and the test set documents in the multiple documents into first feature vectors of each training set document composed of the relevance of the keywords of the training set documents, and second feature vectors of each test set document composed of the relevance of the keywords of the test set documents;
[0011] Train a document classifier based on the first feature vectors of the training set documents, where the document classifier is constructed by a machine learning model;
[0012] Test and evaluate the document classifier based on the second feature vector;
[0013] If the document classifier passes the test and evaluation, deploy the document classifier so as to classify new documents in the classification corpus through the document classifier.
[0014] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:
[0015] Improve classification efficiency: By automatically processing documents, this method significantly improves classification efficiency, reduces manual operation time, and is especially suitable for processing a large number of document materials.
[0016] Reduce subjective errors: Since this method classifies based on the relevance between the preset vocabulary in the domain vocabulary table and keywords, it reduces the influence of subjective factors in the manual classification process and improves the objectivity and consistency of the classification results.
[0017] Enhance classification accuracy: The machine learning model trained by the feature vectors of the training set documents can capture the features of the documents more precisely, thereby improving the accuracy of classification.
[0018] Adapt to large-scale document management: This method can process large-scale document data and meets the needs of large-scale document management through automated and model-driven classification.
[0019] Furthermore, obtaining the keywords of the documents based on the relevance between the preset vocabulary in the domain vocabulary table and each document includes:
[0020] Use the preset vocabulary in the domain vocabulary table as search terms, search in each of the documents respectively, and determine the relevance of the preset vocabulary in each of the documents;
[0021] Obtain the keywords of each document based on the relevance of the preset vocabulary in each document.
[0022] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:
[0023] Improve the accuracy of keyword extraction: By using the preset vocabulary in the domain vocabulary table as search terms, it can ensure that the extracted keywords are closely related to a specific domain, thereby improving the accuracy of the keywords.
[0024] Enhanced domain adaptability: The domain vocabulary can reflect the professional terms and common words in a specific domain, which makes keyword extraction more in line with the domain characteristics and applicable to document classification in different domains.
[0025] Further, using the preset words in the domain vocabulary as search terms, and searching in each of the documents respectively to determine the relevance of the preset words in each of the documents, includes:
[0026] Using each preset word in the domain vocabulary as a search term, and searching in each of the documents respectively to obtain the ranking of each preset word in each document;
[0027] Based on the ranking of each preset word in each document, determining the relevance of each preset word in each document, where the ranking of each preset word in each document is positively correlated with the relevance of each preset word in each document.
[0028] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects:
[0029] Enhanced keyword pertinence: By using the preset words in the domain vocabulary as search terms, it can ensure that the extracted keywords are highly relevant to the document content, thereby enhancing the pertinence of the keywords.
[0030] Improved accuracy of keyword extraction: By analyzing the ranking of each preset word in the document, it is possible to more accurately identify the key information in the document and improve the accuracy of keyword extraction.
[0031] Optimized document classification performance: Using the ranking-based keywords as features can optimize the performance of document classification because they can better represent the core content of the document.
[0032] Further, obtaining the keywords of each document based on the relevance of the preset words in each document, includes:
[0033] Sorting the relevance of each preset word in each document to obtain the relevance sorting result of each preset word in each document;
[0034] Based on the relevance sorting result, determining the preset number of keywords for each document.
[0035] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects:
[0036] Improved quality of keyword selection: By sorting the relevance of each preset word in each document, it is possible to preferentially select the words most relevant to the document content as keywords, thereby improving the quality of keyword selection.
[0037] Enhance the ability to recognize document themes: Selecting relatively relevant preset vocabulary as keywords helps to more accurately recognize and express the themes of documents, enhancing the document theme recognition ability.
[0038] Simplify the understanding of document content: Using the keywords determined by relevance ranking can simplify the understanding of document content, enabling both readers and systems to quickly grasp the core information of the document.
[0039] Furthermore, obtaining the domain vocabulary of the specified enterprise includes:
[0040] Obtain the preset document materials of the specified enterprise, where the preset document materials include one or more of the existing knowledge bases, relevant standard documents, business technical term documents, and business specification documents of the specified enterprise;
[0041] Perform word segmentation on the preset document materials based on word segmentation technology to obtain the domain vocabulary of the specified enterprise.
[0042] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects:
[0043] Improve the accurate capture of domain knowledge: By extracting vocabulary from the existing knowledge bases, relevant standard documents, business technical term documents, and business specification documents of the enterprise, it is possible to more accurately capture the enterprise-specific domain knowledge.
[0044] Enhance the pertinence and professionalism of the vocabulary: Constructing the vocabulary based on the specific document materials of the enterprise can ensure the professionalism and pertinence of the vocabulary, which is applicable to the specific business and processes of the enterprise.
[0045] Improve the accuracy of text processing: The establishment of the domain vocabulary helps to improve the accuracy of text processing tools (such as search engines, natural language processing systems, etc.), because these tools can use this vocabulary to better understand enterprise documents.
[0046] Furthermore, the machine learning model includes one or more of the Naive Bayes algorithm, decision tree algorithm, KNN nearest neighbor algorithm, center vector algorithm, and support vector machine algorithm.
[0047] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects:
[0048] Diversity: The model combination provides multiple algorithm options, and the most suitable algorithm can be selected according to different data sets and problems, thereby improving the generalization ability and adaptability of the model.
[0049] Robustness: Different algorithms have different sensitivities to data noise and outliers. Combining multiple algorithms can enhance the model's robustness to data noise and outliers.
[0050] Accuracy: By combining the advantages of different algorithms, the prediction accuracy of the model can be improved. Each algorithm has its specific advantages, and combining them can complement each other's deficiencies.
[0051] Further, using the preset vocabulary in the domain vocabulary list as search terms, and searching in each of the documents respectively to determine the relevance of the preset vocabulary in each of the documents, includes:
[0052] Based on a pre-set enterprise-level search engine, using the preset vocabulary in the domain vocabulary list as search terms, and searching in each of the documents respectively to determine the relevance of the preset vocabulary in each of the documents, wherein the enterprise-level search engine uses the ElasticSearch full-text retrieval tool.
[0053] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:
[0054] Improve search efficiency and accuracy: Using a full-text retrieval tool such as ElasticSearch can quickly locate relevant content in the document, thereby improving search efficiency.
[0055] Optimize the performance of the enterprise-level search engine: ElasticSearch is a high-performance search library that can handle a large amount of data and provide fast search response times, which is very valuable for enterprise-level applications.
[0056] Enhance the precision of the search: The preset vocabulary in the domain vocabulary list is constructed based on the knowledge of the enterprise's specific domain, so using it as a search term can more precisely match the relevant content in the document.
[0057] A document classification device provided by one or more embodiments of this specification includes:
[0058] A corpus acquisition unit, which acquires a pre-generated classification corpus, and the classification corpus includes multiple documents of a specified enterprise, and the multiple documents include pre-set training set documents and test set documents;
[0059] A vocabulary list acquisition unit, which acquires the domain vocabulary list of the specified enterprise, and the domain vocabulary list includes preset vocabulary;
[0060] A keyword determination unit, which obtains the keywords of each of the documents based on the relevance of the preset vocabulary in the domain vocabulary list to each of the documents;
[0061] A vector conversion unit that respectively converts the training set documents and the test set documents among the multiple documents into a first feature vector of each training set document composed of the relevance degrees of the keywords of the training set documents, and a second feature vector of each test set document composed of the relevance degrees of the keywords of the test set documents;
[0062] A classifier training unit that trains a document classifier based on the first feature vectors of the respective training set documents, and the document classifier is constructed through a machine learning model;
[0063] A classifier testing unit that tests and evaluates the document classifier based on the second feature vectors;
[0064] A classifier application unit that, if the document classifier passes the test and evaluation, deploys the document classifier so as to classify new documents in the classification corpus through the document classifier.
[0065] A document classification device provided by one or more embodiments of this specification includes:
[0066] At least one processor; and,
[0067] A memory communicatively connected to the at least one processor; wherein,
[0068] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:
[0069] Obtain a pre-generated classification corpus, where the classification corpus includes multiple documents of a specified enterprise, and the multiple documents include a pre-set training set documents and test set documents;
[0070] Obtain the domain vocabulary of the specified enterprise, where the domain vocabulary includes pre-set vocabulary;
[0071] Based on the relevance degrees of the pre-set vocabulary in the domain vocabulary and each document, obtain the keywords of each document;
[0072] Respectively convert the training set documents and the test set documents among the multiple documents into a first feature vector of each training set document composed of the relevance degrees of the keywords of the training set documents, and a second feature vector of each test set document composed of the relevance degrees of the keywords of the test set documents;
[0073] Based on the first feature vectors of the respective training set documents, train a document classifier, and the document classifier is constructed through a machine learning model;
[0074] Testing and evaluating the document classifier based on the second feature vector;
[0075] If the document classifier passes the testing and evaluation, deploy the document classifier to classify new documents in the classified corpus through the document classifier.
[0076] A non - volatile computer storage medium provided by one or more embodiments of this specification stores computer - executable instructions, and when the computer - executable instructions are executed by a computer, they can implement:
[0077] Obtain a pre - generated classified corpus, where the classified corpus includes multiple documents of a specified enterprise, and the multiple documents include pre - set training - set documents and test - set documents;
[0078] Obtain the domain vocabulary of the specified enterprise, where the domain vocabulary includes preset vocabulary;
[0079] Based on the relevance between the preset vocabulary in the domain vocabulary and each document, obtain the keywords of each document;
[0080] Convert the training - set documents and the test - set documents in the multiple documents into the first feature vectors of each training - set document composed of the relevance of the keywords of the training - set documents and the second feature vectors of each test - set document composed of the relevance of the keywords of the test - set documents respectively;
[0081] Train a document classifier based on the first feature vectors of each training - set document, and the document classifier is constructed through a machine - learning model;
[0082] Testing and evaluating the document classifier based on the second feature vector;
[0083] If the document classifier passes the testing and evaluation, deploy the document classifier to classify new documents in the classified corpus through the document classifier.
[0084] The above - mentioned at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:
[0085] Improve classification efficiency: By automatically processing documents, this method significantly improves classification efficiency, reduces manual operation time, and is especially suitable for processing a large number of document materials.
[0086] Reduce subjective errors: Since this method classifies based on a preset domain vocabulary and the relevance of keywords, it reduces the influence of subjective factors in the manual classification process and improves the objectivity and consistency of classification results.
[0087] Enhanced classification accuracy: The machine learning model trained with the feature vectors of the training set documents can capture the features of the documents more precisely, thus improving the classification accuracy.
[0088] Adapt to large-scale document management: This method can handle large-scale document data and meet the requirements of large-scale document management through automated and model-driven classification. Description of the Drawings
[0089] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0090] Figure 1 It is a schematic flowchart of a document classification method provided by one or more embodiments of this specification;
[0091] Figure 2 It is a schematic structural diagram of a document classification device provided by one or more embodiments of this specification;
[0092] Figure 3 It is a schematic structural diagram of a document classification device provided by one or more embodiments of this specification. Detailed Embodiments
[0093] The embodiments of this specification provide a document classification method, device, device and medium.
[0094] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0095] Figure 1 It is a schematic flowchart of a document classification method provided by one or more embodiments of this specification, and this process can be executed by a document classification system. Some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.
[0096] The method process steps of the embodiments of this specification are as follows:
[0097] S101. Obtain a pre-generated classification corpus, where the classification corpus includes multiple documents of a specified enterprise, and the multiple documents include pre-set training set documents and test set documents.
[0098] In the embodiments of this specification, the classification corpus can be a collection of documents of the target classification system and corresponding categories of the specified enterprise's electronic documents.
[0099] It should be noted that regarding the above S101, the following specific implementation methods can be adopted:
[0100] Define the classification system: Clearly define the target classification system for the enterprise's electronic documents. This usually includes determining the categories of documents (such as financial reports, market analysis, technical documents, etc.) and sub-categories under each category.
[0101] Collect documents: Collect the existing electronic documents of the enterprise to ensure that these documents cover all categories and sub-categories in the classification system.
[0102] Document preprocessing: Preprocess the collected documents, including unifying the format, removing irrelevant content, marking document attributes, etc., to facilitate subsequent processing.
[0103] Establish a classification corpus: According to the classification system, allocate the preprocessed documents to the corresponding categories and sub-categories to establish a classification corpus. Use a database or a file system to store these documents to ensure that each document can be quickly retrieved according to its category.
[0104] S102. Obtain the domain vocabulary of the specified enterprise, where the domain vocabulary includes preset vocabulary.
[0105] In the embodiments of this specification, the preset document materials of the specified enterprise can be obtained first. The preset document materials include one or more of the existing knowledge base of the specified enterprise, relevant standard documents, business technical term documents, and business specification documents. Based on the word segmentation technology, perform word segmentation processing on the preset document materials to obtain the domain vocabulary of the specified enterprise.
[0106] It should be noted that regarding the above S101, the following specific implementation methods can be adopted:
[0107] Determine document materials: Determine the types of documents that need to be included in the preset document materials. These may include:
[0108] Existing knowledge base: Contains the knowledge and experience accumulated within the enterprise.
[0109] Relevant standard documents: Such as industry standards, national standards, etc.
[0110] Business technical term documents: Involve professional terms in the specific business field of the enterprise.
[0111] Business specification documents: including the operation processes, management specifications, etc. of the enterprise.
[0112] Collect documents: Collect the above types of documents from the enterprise's internal systems or storage media.
[0113] Document preprocessing: Preprocess the collected documents, including:
[0114] Format conversion: Convert the documents into a unified format, such as TXT or PDF.
[0115] Data cleaning: Remove irrelevant content, such as headers, footers, watermarks, etc.
[0116] Text standardization: Unify the text format, such as unifying punctuation marks, number representations, etc.
[0117] Word segmentation technology selection: Select a suitable word segmentation technology, such as rule-based word segmentation, statistics-based word segmentation, or deep learning models. Ensure that the word segmentation technology can recognize and correctly process the professional terms in the enterprise domain.
[0118] Word segmentation processing: Use the selected word segmentation technology to perform word segmentation on the preprocessed documents. During the word segmentation process, the following strategies can be adopted:
[0119] Retain professional terms: Ensure that the terms unique to the enterprise are not split.
[0120] Remove stop words: Remove common meaningless words, such as "de", "shi", "zai", etc.
[0121] Part-of-speech tagging: Add part-of-speech tags to the word segmentation results, which helps subsequent lexical analysis.
[0122] Generate a domain vocabulary: Extract high-frequency words and professional terms in specific domains from the word segmentation results. Remove duplicates and sort the words to generate the final domain vocabulary.
[0123] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects:
[0124] Improve the accurate capture of domain knowledge: By extracting words from the enterprise's existing knowledge base, relevant standard documents, business technical term documents, and business specification documents, the enterprise-specific domain knowledge can be captured more accurately.
[0125] Enhance the pertinence and professionalism of the vocabulary: Building a vocabulary based on the specific document materials of the enterprise can ensure the professionalism and pertinence of the vocabulary, which is applicable to the specific business and processes of the enterprise.
[0126] Improving the accuracy of text processing: The establishment of a domain vocabulary helps improve the accuracy of text processing tools (such as search engines, natural language processing systems, etc.) because these tools can use this vocabulary to better understand enterprise documents.
[0127] S103, obtaining keywords of each document based on the relevance of preset vocabulary in the domain vocabulary to each document.
[0128] In the embodiments of this specification, the preset vocabulary in the domain vocabulary can be used as search terms to search in each document respectively to determine the relevance of the preset vocabulary in each document; based on the relevance of the preset vocabulary in each document, the keywords of each document are obtained.
[0129] It should be noted that for the above content, the following specific implementation solutions can be adopted:
[0130] Word segmentation processing: Using word segmentation technology to segment the preset vocabulary and document content for matching.
[0131] Matching algorithm: Developing or using existing algorithms to determine the frequency and position of the preset vocabulary in the document, as well as the contextual meaning of the vocabulary.
[0132] Relevance calculation: Calculating a relevance score for the occurrence of each preset vocabulary in each document. This can be achieved through the following methods:
[0133] TF-IDF (Term Frequency - Inverse Document Frequency): Measuring the importance of a vocabulary in a document.
[0134] BM25: A ranking function for information retrieval, used to evaluate the relevance to a query.
[0135] Word embedding similarity: If both the document and the vocabulary use word embedding technology, the similarity between word embeddings can be calculated.
[0136] Screening keywords: Based on the relevance scores of the preset vocabulary, screening out the vocabulary with the highest relevance to the preset vocabulary in each document.
[0137] Result collation: Creating a keyword list for each document to reflect the core content of the document. Storing the keyword list in a database and establishing an index for fast retrieval.
[0138] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects:
[0139] Improving the accuracy of keyword extraction: By using the preset vocabulary in the domain vocabulary as search terms, it can be ensured that the extracted keywords are closely related to a specific domain, thereby improving the accuracy of keywords.
[0140] Enhanced domain adaptability: The domain vocabulary can reflect the professional terms and common words in a specific domain, which makes keyword extraction more in line with the domain characteristics and applicable to document classification in different domains.
[0141] Furthermore, when using the preset vocabulary in the domain vocabulary as search terms to search in each of the documents and determine the relevance of the preset vocabulary in each of the documents, each preset vocabulary in the domain vocabulary can be used as a search term to search in each of the documents respectively to obtain the ranking of each preset vocabulary in each document; based on the ranking of each preset vocabulary in each document, determine the relevance of each preset vocabulary in each document, and the ranking of each preset vocabulary in each document is positively correlated with the relevance of each preset vocabulary in each document.
[0142] It should be noted that regarding the above content, the following specific implementation solutions can be adopted:
[0143] Word segmentation and search: Perform word segmentation on the preset vocabulary and document content for accurate matching. Implement a search mechanism that uses each preset vocabulary as a search term to search in all documents.
[0144] Ranking determination: Assign a ranking to the position where each preset vocabulary appears in each document. The ranking can range from 1 to n, where n is the total number of times the word appears in the document. For each document, calculate the ranking of each preset vocabulary, considering the order and frequency of the word appearance.
[0145] Relevance calculation: Define a formula to calculate the relevance of each preset vocabulary in each document. Since the ranking is positively correlated with the relevance, the following formula can be used:
[0146] Relevance = ranking / maximum ranking; where the maximum ranking is the maximum number of times the word appears in all documents. To make the relevance between 0 and 1, the above formula can be normalized.
[0147] Result sorting: Create a matrix to record the relevance of each preset vocabulary in each document. Sort the keywords of each document according to the relevance and screen out the most relevant keywords.
[0148] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects:
[0149] Enhanced keyword pertinence: By using the preset vocabulary in the domain vocabulary as search terms, it can ensure that the extracted keywords are highly relevant to the document content, thereby enhancing the pertinence of the keywords.
[0150] Improve the accuracy of keyword extraction: By analyzing the ranking of each preset term in the document, key information in the document can be identified more accurately, improving the accuracy of keyword extraction.
[0151] Optimize document classification performance: Using the ranking-based keywords as features can optimize the performance of document classification because they can better represent the core content of the document.
[0152] Furthermore, when obtaining the keywords of each document based on the relevance of the preset terms in each document, the relevance of each preset term in each document can be sorted to obtain the relevance ranking result of each preset term in each document; based on the relevance ranking result, a preset number of keywords for each document are determined.
[0153] It should be noted that for the above content, the following specific implementation solutions can be adopted:
[0154] Sorting: Sort the relevance of each preset term. This can be achieved in the following ways:
[0155] For each document, sort the relevance values of all preset terms in descending order. Built-in sorting functions or libraries (such as the sorted() function in Python) can be used to perform the sorting.
[0156] Determine the number of keywords: Determine the number of keywords for each document according to the relevance ranking result of each document. The following are some methods for determining the number of keywords:
[0157] Fixed number method: Set a fixed number of keywords for each document, such as the top 5 or top 10 preset terms with the highest relevance.
[0158] Percentage method: Set a percentage based on the total number of preset terms in the document to determine the number of keywords, such as selecting the top 10% of the preset terms as keywords.
[0159] Threshold method: Set a relevance threshold, and only when the relevance of a preset term exceeds this threshold is it regarded as a keyword.
[0160] Keyword extraction: Extract keywords from the relevance ranking result of each document according to the determined number of keywords.
[0161] Result collation: Create a result set that includes the keyword list of each document and the relevance of each keyword.
[0162] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:
[0163] Improve the quality of keyword selection: By sorting the relevance of each preset term in each document, the terms most relevant to the document content can be preferentially selected as keywords, thereby improving the quality of keyword selection.
[0164] Enhance the document theme recognition ability: Selecting preset terms with higher relevance as keywords helps to more accurately identify and express the theme of the document, enhancing the document's theme recognition ability.
[0165] Simplify the understanding of document content: Using the keywords determined by relevance sorting can simplify the understanding of document content, enabling both readers and systems to quickly grasp the core information of the document.
[0166] Furthermore, when using the preset terms in the domain vocabulary as search terms to search in each of the documents respectively to determine the relevance of the preset terms in each of the documents, a pre-set enterprise-level search engine can be used. The preset terms in the domain vocabulary are used as search terms to search in each of the documents respectively to determine the relevance of the preset terms in each of the documents, where the enterprise-level search engine uses the ElasticSearch full-text retrieval tool.
[0167] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:
[0168] Improve search efficiency and accuracy: Using a full-text retrieval tool like ElasticSearch can quickly locate relevant content in the document, thereby improving search efficiency.
[0169] Optimize the performance of the enterprise-level search engine: ElasticSearch is a high-performance search library that can handle large amounts of data and provide fast search response times, which is very valuable for enterprise-level applications.
[0170] Enhance the precision of search: The preset terms in the domain vocabulary are constructed based on the knowledge of the enterprise's specific domain, so using them as search terms can more precisely match the relevant content in the document.
[0171] S104, convert the training set documents and the test set documents in the multiple documents into the first feature vectors of each training set document composed of the relevance of the keywords of the training set documents, and the second feature vectors of each test set document composed of the relevance of the keywords of the test set documents.
[0172] S105, train a document classifier based on the first feature vectors of each training set document, and the document classifier is constructed through a machine learning model.
[0173] In the embodiments of this specification, the machine learning model includes one or more of the Naive Bayes algorithm, decision tree algorithm, KNN nearest neighbor algorithm, center vector algorithm, and support vector machine algorithm.
[0174] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects:
[0175] Diversity: The model combination provides multiple algorithm options, and the most suitable algorithm can be selected according to different data sets and problems, thereby improving the generalization ability and adaptability of the model.
[0176] Robustness: Different algorithms have different sensitivities to data noise and outliers. Combining multiple algorithms can enhance the robustness of the model to data noise and outliers.
[0177] Accuracy: By combining the advantages of different algorithms, the prediction accuracy of the model can be improved. Each algorithm has its specific advantages, and combining them can complement each other's deficiencies.
[0178] S106. Test and evaluate the document classifier based on the second feature vector.
[0179] S107. If the document classifier passes the test and evaluation, deploy the document classifier so as to classify new documents in the classification corpus through the document classifier.
[0180] It should be noted that for the above S105 - S107, the following specific implementation solutions can be adopted:
[0181] Construction of feature vectors: For each document, use the preset vocabulary with the highest relevance as features to construct a feature vector. If there are a large number of preset vocabulary in the document, one of the following strategies can be adopted:
[0182] Fixed - length vector: Select a number of preset vocabulary as the length of the feature vector, such as the first N vocabulary with the highest relevance.
[0183] Sparse vector: Create a sparse matrix that contains the indices of all preset vocabulary, and the relevance is the corresponding value.
[0184] Feature vector of training - set documents: Use the relevance feature vector of training - set documents as input to construct the first feature vector.
[0185] Feature vector of test - set documents: Use the same method to construct the second feature vector using the relevance feature vector of test - set documents.
[0186] The specific process of training the document classifier is as follows:
[0187] Select a machine learning model: Select a suitable machine learning model according to the requirements of document classification, such as Naive Bayes, Support Vector Machine (SVM), Random Forest, deep learning models, etc.
[0188] Model training: Use the feature vectors (first feature vectors) of the training set documents to train the selected machine learning model. During the training process, parameter tuning may be required, such as cross-validation to find the optimal model parameters.
[0189] The test evaluation is as follows:
[0190] Model testing: Use the feature vectors (second feature vectors) of the test set documents to test the trained document classifier.
[0191] Evaluate the model performance: Calculate the performance metrics of the model on the test set, such as accuracy, recall rate, F1 score, etc. Evaluate the classification effect of the model according to the performance metrics.
[0192] It should be noted that the embodiments of this specification have the following beneficial effects through the above content:
[0193] Improve classification efficiency: By automatically processing documents, this method significantly improves the classification efficiency, reduces the manual operation time, and is especially suitable for processing a large number of document materials.
[0194] Reduce subjective errors: Since this method classifies based on the relevance of a preset domain vocabulary and keywords, it reduces the influence of subjective factors in the manual classification process and improves the objectivity and consistency of the classification results.
[0195] Enhance classification accuracy: The machine learning model trained by the feature vectors of the training set documents can capture the features of the documents more precisely, thereby improving the classification accuracy.
[0196] Adapt to large-scale document management: This method can process large-scale document data and meets the needs of large-scale document management through automated and model-driven classification.
[0197] It should be noted that traditional document classification mainly relies on manual operations and preset rules. This method is inefficient and error-prone when dealing with a large number of documents. With the rapid development of information technology, especially the continuous progress of natural language processing (NLP) and machine learning technologies, automated and intelligent document classification technologies have become possible.
[0198] The technical field to which the present invention belongs is information processing and data mining, especially for the automated classification processing of large-scale documents (such as documents). This field covers technologies and methods in multiple sub-fields such as natural language processing (NLP), machine learning, text mining, and knowledge management.
[0199] In the field of natural language processing, the present invention relies on advanced methods such as word segmentation technology, part-of-speech tagging, keyword extraction, and topic modeling to understand and analyze the content of document texts.
[0200] In the field of machine learning, the present invention uses classification algorithms (such as support vector machines, decision trees, neural networks, etc.) to train models, enabling the models to learn and identify the associations between keywords, topics, and classification labels in document texts.
[0201] Text mining techniques are used to extract useful information and patterns from a large number of unstructured document texts, providing data support for classification tasks.
[0202] In addition, the present invention also relates to technologies in the field of knowledge management. Through an automated classification process, it helps organizations organize, retrieve, and utilize document information more effectively.
[0203] The core objective of the present invention is to provide a document automatic classification method and system based on keyword and topic extraction, aiming to solve various problems existing in the traditional document classification process, specifically including:
[0204] (1) Improve classification efficiency: Traditional document classification mainly relies on manual operations, which are not only time-consuming and laborious but also inefficient in processing large-scale documents. The present invention can significantly shorten the time for document classification and improve work efficiency through an automated processing flow.
[0205] (2) Enhance classification accuracy: Manual classification is often affected by factors such as personal experience and knowledge level, resulting in possible biases in classification results. The present invention uses advanced natural language processing technologies and machine learning algorithms to more accurately identify document content, extract keywords and topics, thereby improving classification accuracy.
[0206] (3) Increase adaptability: Traditional classification methods are often designed for specific fields or formats and are difficult to adapt to diverse document types. The present invention can process documents in different fields and formats, with strong adaptability and flexibility.
[0207] (4) Promote intelligent management: With the rapid development of information technology, document management is moving towards intelligence and automation. The proposed present invention not only conforms to this development trend but also provides an intelligent solution for document management, promoting the improvement of document management levels.
[0208] (5) Reduce labor costs: Through automated classification, the dependence on manual classifiers can be reduced, thereby reducing the labor costs of document management.
[0209] (6) Facilitate information retrieval and utilization: Accurate classification helps to quickly locate the required documents, improve the efficiency of information retrieval, and also provides convenience for the subsequent utilization of the documents.
[0210] The steps of the embodiments of this specification are as follows:
[0211] (1) Document preprocessing
[0212] Provide standardized data input for subsequent keyword extraction and topic extraction.
[0213] Implementation method:
[0214] Remove redundant information: such as advertisements, watermarks, irrelevant text, etc.
[0215] Word segmentation: Split the document text into individual words or phrases.
[0216] Part-of-speech tagging: Tag the part of speech of each word, such as noun, verb, adjective, etc.
[0217] (2) Keyword extraction
[0218] Extract keywords from the document text that can reflect its core content and theme.
[0219] Implementation method:
[0220] Word frequency statistics: Count the occurrence frequency of each word in the document text.
[0221] TF-IDF algorithm: Calculate the TF-IDF value of each word to evaluate its importance and distinctiveness in the document.
[0222] Combine domain-specific dictionaries and rules: Use domain dictionaries and rules to filter and optimize the extracted keywords to ensure the accuracy and relevance of the keywords.
[0223] (3) Topic extraction
[0224] Identify the potential topics in the document and obtain the topic distribution of the document.
[0225] Implementation method:
[0226] Topic model selection: Select a suitable topic model according to the characteristics and requirements of the document text, such as LDA (Latent Dirichlet Allocation) or PLSA (Probabilistic Latent Semantic Analysis).
[0227] Topic modeling: Use the selected topic model to model the document text and identify the potential topics.
[0228] Topic optimization: Filter and optimize the extracted topics to ensure the accuracy and relevance of the topics.
[0229] (4) Classifier Training
[0230] Train a classifier so that it can learn the mapping relationship between keywords, topics, and document classification labels.
[0231] Implementation method:
[0232] Data preparation: Collect classified documents as training data to ensure the diversity and representativeness of the data.
[0233] Feature extraction: Extract keyword and topic information from the training data as features.
[0234] Classifier selection: Select a suitable classifier according to the complexity of the problem and the scale of the data, such as Support Vector Machine (SVM), decision tree, neural network, etc.
[0235] Training process: Use the extracted features and known document classification labels to train the classifier.
[0236] (5) Document Classification
[0237] Automatically classify the documents to be classified.
[0238] Implementation method:
[0239] Preprocessing: Perform the same preprocessing steps on the documents to be classified as on the training data.
[0240] Feature extraction: Extract keyword and topic information from the preprocessed documents as features.
[0241] Classification: Input the extracted features into the trained classifier to obtain the classification labels of the documents.
[0242] Improve classification efficiency: Significantly shorten the time for document classification through an automated processing flow.
[0243] Enhance classification accuracy: Combine keyword and topic information to more accurately identify the content of the documents, thereby improving the classification accuracy.
[0244] Enhance adaptability: This method can handle documents in different fields and different formats, with strong adaptability.
[0245] Easy to expand and optimize: As new data is added and the algorithm is continuously improved, the classifier can continuously learn and optimize to improve the classification performance.
[0246] A keyword-based document classification technology, including the following steps:
[0247] Step S1: Prepare a classification corpus, that is, the target classification system of enterprise electronic documents and the document sets corresponding to different categories, and divide the corpus into a training set and a test set;
[0248] Step S2: Construct a domain vocabulary for the enterprise;
[0249] Step S3: Use an enterprise-level search engine, take the vocabulary in the domain vocabulary as search terms, and conduct a search for each search term in the entire corpus;
[0250] Step S4: Take the top N words with the highest relevance to the document as the keywords of the document, where N is a natural number, such as 100;
[0251] Step S5: Characterize all documents as feature vectors composed of the relevance of N keywords;
[0252] Step S6: Build a classifier based on the feature vectors of the training set documents using different machine learning algorithms;
[0253] Step S7: Evaluate the constructed classifier using the test set documents, and select the optimal classifier according to the accuracy rate and recall rate of the classifier;
[0254] Step S8: Deploy the optimal classifier in the production system, and call the interface of the optimal classifier to automatically classify the newly added documents.
[0255] For the automatic classification method of electronic documents based on keyword features, 80% of the documents in the corpus are randomly selected as the training set, and 20% of the documents are used as the test set.
[0256] For the automatic classification method of electronic documents based on keyword features, the specific content of step S2 is as follows: From the formal document materials of the enterprise, including the existing knowledge base of the enterprise, relevant standard documents, business technical term documents, and business specification documents, a large number of vocabulary are discovered through word segmentation technology. Then, the vocabulary with less obvious business characteristics are preferentially deleted from the discovered large number of vocabulary, and finally a domain vocabulary is formed.
[0257] For the automatic classification method of electronic documents based on keyword features, the enterprise-level search engine uses an open-source ElasticSearch full-text retrieval tool.
[0258] For the automatic classification method of electronic documents based on keyword features, the specific content of step S4 includes:
[0259] Step S41: Perform a search for each vocabulary in the domain vocabulary, and obtain the ranking of the document in the search results;
[0260] Step S42: Calculate the relevance R between the vocabulary and the document, where R = 1 - n / m, where n is the ranking of the document in the search results, and m is the total number of documents;
[0261] Step S43: Rank according to the relevance from high to low, and obtain the top N most relevant words of the document as the keywords of the document.
[0262] An automatic classification method for electronic documents based on keyword features, wherein the machine learning algorithms include the Naive Bayes algorithm, the decision tree algorithm, the KNN nearest neighbor algorithm, the central vector algorithm, and the support vector machine algorithm.
[0263] An automatic classification method for electronic documents based on keyword features, wherein the accuracy rate and the recall rate are calculated by the following formulas: accuracy rate p = a / (a + b)×100%, recall rate r = a / (a + c)×100%, where a represents the number of test set documents input that are correctly classified into a certain category, b represents the number of test set documents input that are misclassified into a certain category by the classifier, and c represents the number of test set documents input that are wrongly excluded from a certain category by the classifier.
[0264] Corresponding to the above embodiments, Figure 2 The structure diagram of a document classification device provided by one or more embodiments of this specification includes: a corpus acquisition unit 201, a vocabulary acquisition unit 202, a keyword determination unit 203, a vector conversion unit 204, a classifier training unit 205, a classifier testing unit 206, and a classifier application unit 207.
[0265] The corpus acquisition unit 201 acquires a pre-generated classified corpus, and the classified corpus includes multiple documents of a specified enterprise, and the multiple documents include a preset training set document and a test set document;
[0266] The vocabulary acquisition unit 202 acquires the domain vocabulary of the specified enterprise, and the domain vocabulary includes preset words;
[0267] The keyword determination unit 203 obtains the keywords of each document based on the relevance between the preset words in the domain vocabulary and each document;
[0268] The vector conversion unit 204 respectively converts the training set documents and the test set documents in the multiple documents into a first feature vector of each training set document composed of the relevance of the keywords of the training set document, and a second feature vector of each test set document composed of the relevance of the keywords of the test set document;
[0269] The classifier training unit 205 trains a document classifier based on the first feature vector of each training set document, and the document classifier is constructed by a machine learning model;
[0270] A classifier test unit 206 tests and evaluates the document classifier based on the second feature vector;
[0271] A classifier application unit 207 deploys the document classifier if the document classifier passes the test and evaluation, so as to classify new documents in the classification corpus through the document classifier.
[0272] Figure 3 The following is a schematic structural diagram of a document classification device provided by one or more embodiments of this specification, including:
[0273] At least one processor; and,
[0274] A memory communicatively connected to the at least one processor; wherein,
[0275] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:
[0276] Obtain a pre-generated classification corpus, the classification corpus includes multiple documents of a specified enterprise, and the multiple documents include pre-set training set documents and test set documents;
[0277] Obtain the domain vocabulary of the specified enterprise, the domain vocabulary includes preset vocabulary;
[0278] Based on the relevance between the preset vocabulary in the domain vocabulary and each document, obtain the keywords of each document;
[0279] Convert the training set documents and the test set documents in the multiple documents into the first feature vectors of each training set document composed of the relevance of the keywords of the training set documents and the second feature vectors of each test set document composed of the relevance of the keywords of the test set documents respectively;
[0280] Based on the first feature vectors of the respective training set documents, train a document classifier, and the document classifier is constructed through a machine learning model;
[0281] Test and evaluate the document classifier based on the second feature vector;
[0282] If the document classifier passes the test and evaluation, deploy the document classifier so as to classify new documents in the classification corpus through the document classifier.
[0283] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and when the computer-executable instructions are executed by a computer, the following can be achieved:
[0284] Obtain a pre-generated classified corpus, where the classified corpus includes multiple documents of a specified enterprise, and the multiple documents include a pre-set training set of documents and a test set of documents;
[0285] Obtain the domain vocabulary of the specified enterprise, where the domain vocabulary includes pre-set vocabulary;
[0286] Based on the relevance of the pre-set vocabulary in the domain vocabulary to each document, obtain the keywords of each document;
[0287] Convert the training set of documents and the test set of documents in the multiple documents into a first feature vector of each training set of documents composed of the relevance of the keywords of the training set of documents, and a second feature vector of each test set of documents composed of the relevance of the keywords of the test set of documents respectively;
[0288] Based on the first feature vector of each training set of documents, train a document classifier, where the document classifier is constructed through a machine learning model;
[0289] Test and evaluate the document classifier based on the second feature vector;
[0290] If the document classifier passes the test and evaluation, deploy the document classifier so as to classify new documents in the classified corpus through the document classifier.
[0291] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0292] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0293] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0294] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0295] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0296] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above units can be implemented in the form of hardware or in the form of software.
[0297] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0298] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A document classification method, characterized in that: The method comprises: Acquire a pre-generated classified corpus, wherein the classified corpus includes a plurality of documents of a specified enterprise, wherein the plurality of documents include a preset training set document and a test set document; Acquire a domain vocabulary of the designated enterprise, wherein the domain vocabulary includes preset vocabulary; Based on the relevance between the preset words in the domain vocabulary and each document, obtaining the keywords of each document; Convert the training set document and the test set document in the plurality of documents into a first feature vector of each training set document composed of the relevance of the keywords of the training set document, and a second feature vector of each test set document composed of the relevance of the keywords of the test set document; Based on the first feature vector of each training set document, a document classifier is trained, where the document classifier is constructed through a machine learning model; Performing a test and evaluation on the document classifier based on the second feature vector; If the document classifier passes the test evaluation, the document classifier is deployed so as to classify the newly added documents of the classification corpus through the document classifier.
2. The method according to claim 1, characterized in that The obtaining of keywords of each document based on the relevance between the preset words in the domain vocabulary and each document includes: Using the preset words in the domain vocabulary as search words, searching each of the documents respectively, and determining the relevance of the preset words in each of the documents; Based on the relevance of the preset words to the documents, keywords of the documents are obtained.
3. The method according to claim 2, characterized in that The using the preset words in the field vocabulary as search words, searching each document respectively, and determining the relevance of the preset words in each document includes: Using each preset word in the domain vocabulary as a search word, searching each document respectively, and obtaining the ranking of each preset word in each document; Based on the ranking of each preset word in each document, the relevance of each preset word in each document is determined, and the ranking of each preset word in each document is positively correlated with the relevance of each preset word in each document.
4. The method according to claim 3, characterized in that The obtaining of keywords of each document based on the relevance of the preset vocabulary to each document includes: Sorting the relevance of each preset word in each document to obtain a relevance ranking result of each preset word in each document; Based on the relevance ranking results, a preset number of keywords for each document is determined.
5. The method according to claim 1, characterized in that The obtaining of the domain vocabulary of the specified enterprise includes: Acquire preset document materials of the designated enterprise, wherein the preset document materials include one or more of an existing knowledge base, relevant standard documents, business terminology documents, and business specification documents of the designated enterprise; The preset document material is segmented based on the word segmentation technology to obtain the domain vocabulary of the designated enterprise.
6. The method according to claim 1, characterized in that The machine learning model includes one or more of a naive Bayes algorithm, a decision tree algorithm, a KNN nearest neighbor algorithm, a center vector algorithm, and a support vector machine algorithm.
7. The method according to claim 2, characterized in that The using the preset words in the field vocabulary as search words, searching each document respectively, and determining the relevance of the preset words in each document includes: Based on a preset enterprise-level search engine, the preset words in the domain vocabulary are used as search words to search each of the documents respectively to determine the relevance of the preset words in each of the documents, wherein the enterprise-level search engine adopts the ElasticSearch full-text search tool.
8. A document classification device, characterized in that: include: A corpus acquisition unit, which acquires a pre-generated classified corpus, wherein the classified corpus includes a plurality of documents of a specified enterprise, wherein the plurality of documents include a preset training set document and a test set document; A vocabulary acquisition unit, which acquires a domain vocabulary of the designated enterprise, wherein the domain vocabulary includes preset vocabulary; A keyword determination unit, which obtains keywords of each document based on the relevance between the preset words in the field vocabulary and each document; A vector conversion unit, which converts the training set document and the test set document in the plurality of document features into a first feature vector of each training set document composed of the relevance of keywords of the training set document, and a second feature vector of each test set document composed of the relevance of keywords of the test set document; A classifier training unit, which trains a document classifier based on the first feature vector of each training set document, wherein the document classifier is constructed by a machine learning model; A classifier testing unit, which performs testing and evaluation on the document classifier based on the second feature vector; The classifier application unit deploys the document classifier if the document classifier passes the test evaluation, so as to classify the newly added documents of the classification corpus through the document classifier.
9. A document classification device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire a pre-generated classified corpus, wherein the classified corpus includes a plurality of documents of a specified enterprise, wherein the plurality of documents include a preset training set document and a test set document; Acquire a domain vocabulary of the designated enterprise, wherein the domain vocabulary includes preset vocabulary; Based on the relevance between the preset words in the domain vocabulary and each document, obtaining the keywords of each document; Convert the training set document and the test set document in the plurality of documents into a first feature vector of each training set document composed of the relevance of the keywords of the training set document, and a second feature vector of each test set document composed of the relevance of the keywords of the test set document; Based on the first feature vector of each training set document, a document classifier is trained, where the document classifier is constructed through a machine learning model; Performing a test evaluation on the document classifier based on the second feature vector; If the document classifier passes the test evaluation, the document classifier is deployed so as to classify the newly added documents of the classification corpus through the document classifier.
10. A non-volatile computer storage medium, characterized in that: Computer executable instructions are stored, and when the computer executable instructions are executed by a computer, the following can be achieved: Acquire a pre-generated classified corpus, wherein the classified corpus includes a plurality of documents of a specified enterprise, wherein the plurality of documents include a preset training set document and a test set document; Acquire a domain vocabulary of the designated enterprise, wherein the domain vocabulary includes preset vocabulary; Based on the relevance between the preset words in the domain vocabulary and each document, obtaining the keywords of each document; Convert the training set document and the test set document in the plurality of documents into a first feature vector of each training set document composed of the relevance of the keywords of the training set document, and a second feature vector of each test set document composed of the relevance of the keywords of the test set document; Based on the first feature vector of each training set document, a document classifier is trained, where the document classifier is constructed through a machine learning model; Performing a test and evaluation on the document classifier based on the second feature vector; If the document classifier passes the test evaluation, the document classifier is deployed so as to classify the newly added documents of the classification corpus through the document classifier.
Citation Information
Cited By
Open domain long text classification method and device based on topic analysis
CN121256035A
Open domain long text classification method and apparatus based on topic analysis
CN121256035B