Unstructured file classification and grading method based on rule and semantic dual mechanism

By combining rule matching and deep semantic analysis of large language models, the problem of failure to fully consider text semantics and context in the prior art is solved, and high accuracy classification rating of unstructured files is achieved.

CN120407810APending Publication Date: 2025-08-01FUJIAN FUJITSU COMM SOFTWARE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510624159.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing unstructured data classification grading methods fail to fully consider the overall semantics and context information of the text, resulting in inaccurate classification results.

Method used

Using a dual mechanism of rules and semantics, a large language model with pre-trained language model is used to conduct in-depth semantic analysis through word segmentation and rule matching, comprehensively evaluate the semantic and contextual relationships of the file to generate a detailed grading report.

Benefits of technology

It significantly improves the accuracy and adaptability of file classification and grading, can more accurately identify the actual content and type of files, and generate detailed grading reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407810A_ABST
    Figure CN120407810A_ABST
Patent Text Reader

Abstract

The invention discloses an unstructured file classification and grading method based on a rule and semantic dual mechanism. The method comprises the following steps: reading unstructured files and extracting text contents; slicing the extracted text content to obtain slice content of the independent semantic unit; performing word segmentation processing on each slice content, splitting the long text into independent vocabulary units, and performing rule matching classification to generate a word segmentation classification set; meanwhile, vectorizing the slice content, and performing deep semantic analysis on the context relationship and semantic features of the text through a pre-trained large model to generate a file content type set; evaluating the actual semantic and contextual relationship of each slice content so as to carry out sensitive grade grading to obtain the grading result of each slice; and finally grading the whole file by integrating grading results of all the slices so as to generate a detailed grading report. According to the method, the file grading accuracy can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method for classifying and grading unstructured files based on a dual mechanism of rules and semantics. Background Art

[0002] Unstructured data is known for its high diversity and complexity. This type of data does not adhere to predefined data models or structures and cannot be simply organized into rows and columns. A notable characteristic of unstructured data is that it can store a variety of different types of content, such as emails, notifications, news reports, planning plans, and even network topologies, within the same medium. This means that when processing unstructured data, it is often necessary to determine its specific type and purpose based on the actual content. Due to its lack of a fixed format, unstructured data contains rich and complex semantic information, which poses challenges for data analysis but also contains enormous potential value.

[0003] The current unstructured data classification and grading process is mainly divided into the following steps: 1. Content Reading: First, you need to obtain the text content of the file. For text files, you can directly read their content; for image files, you need to use optical character recognition (OCR) technology to convert them into text. OCR technology can recognize text information in images, providing basic text data for subsequent processing.

[0004] 2. Text preprocessing: Perform preliminary processing on the acquired text content. Word segmentation tools (such as Jieba, HanLP, and THULAC) can be used to segment the text, breaking long text into individual lexical units. Alternatively, after slicing the text, natural language processing (NLP) tools (such as Word2Vec, FastText, and TextCNN) can be used to vectorize the text slices, converting the text into numerical vectors for subsequent analysis and processing.

[0005] 3. Classification and Grading: Segmentation results are classified using rule-matching mechanisms such as keywords, dictionaries, and regular expressions, or vectorized text slices are output as corresponding classification results using activation functions. These methods can categorize text content into different categories and grade them according to pre-set rules.

[0006] However, this processing flow has certain limitations. It is prone to focusing on lower-level dimensions in the processing of file content, and different types of files are reduced to a set of classification results after processing. This processing logic may ignore the actual content and semantic information of the files, resulting in inaccurate classification and grading results. For example, low-level files may be mislabeled as high-level, or high-level files may be mislabeled as low-level. This phenomenon mainly occurs because the current processing methods do not fully consider the overall semantics and context information of the text, but only rely on local lexical or vector features for classification and grading. Summary of the Invention

[0007] The object of the present invention is to provide a method for classifying and grading unstructured files based on a dual mechanism of rules and semantics to address the limitations existing in the current unstructured data classification and grading technologies, and to improve the accuracy and adaptability of classification and grading.

[0008] The technical solution adopted by the present invention is as follows: A method for classifying and grading unstructured files based on a dual mechanism of rules and semantics, which includes the following: Step 1, file content reading and text extraction: Read the unstructured file and extract the text content; Specifically, the system reads the text content of the file. For pure text files (such as TXT, PDF, Word documents, etc.), the text information is directly extracted. If the file contains picture information (such as scanned documents, PPTs, PDF in picture format, etc.), the system will extract the picture content. Subsequently, the OCR technology is used to recognize the text in the picture and convert it into an editable text format.

[0009] Step 2, text content slicing: Slice the extracted text content, and dynamically adjust the length and boundary of the slice according to the structure and content characteristics of the text to obtain the slice content of independent semantic units; Specifically, the dimension of the slice is usually a natural paragraph or a text segment of a specific length. Slicing by natural paragraphs can preserve the semantic integrity of the text, while text segments of a specific length are suitable for long texts or scenarios that require more fine-grained analysis. During the slicing process, the system will dynamically adjust the length and boundary of the slice according to the structure and content characteristics of the text to ensure that each slice has an independent semantic unit.

[0010] Step 3, word segmentation, rule matching and semantic analysis: Perform word segmentation on the content of each slice to split the long text into individual lexical units; generate a word segmentation classification set for the word segmentation results; at the same time, vectorize the slice content and perform in-depth semantic analysis of the context relationship and semantic features of the text through a pre-trained large model to generate a file content type set; Specifically, perform word segmentation on each slice of content, splitting the long text into individual lexical units. Commonly used word segmentation tools include Jieba, HanLP, THULAC, etc. These tools can accurately identify the lexical boundaries in Chinese text.

[0011] Use rule matching mechanisms such as keywords, dictionaries, regular expressions, etc. to classify the word segmentation results. Rule matching can quickly generate a set of word segmentation classifications, providing a basis for subsequent processing.

[0012] At the same time, vectorize the slice content and perform in-depth semantic analysis on the text content through pre-trained large models. These models can understand the context relationship and semantic features of the text and generate a set of file content types. The results of semantic analysis can complement the deficiencies of rule matching, especially when dealing with complex semantics and context information.

[0013] Step 4, Comprehensive Analysis and Classification Strategy: Integrate the set of file content types and the set of word segmentation classifications, evaluate the actual semantics and context relationship of each slice of content, so as to conduct sensitive level classification to obtain the classification results of each slice; then integrate the classification results of all slices to finally classify the entire file, in order to generate a detailed classification report.

[0014] Specifically, the system will evaluate the actual semantics and context relationship of the slice content according to the results of rule matching and semantic analysis. According to the comprehensive analysis results, different classification strategies are adopted. For slices where the results of rule matching and semantic analysis are consistent, directly use their classification results for classification; for slices where the results are inconsistent, further analyze the context information, combine domain knowledge and business rules, and dynamically adjust the classification strategy. For example, some slices may contain multiple keywords, but semantic analysis shows that their actual content is not sensitive, and the system will lower their classification according to the results of semantic analysis. Integrate the classification results of all slices to finally classify the entire file. The system will generate a detailed classification report, recording the classification basis and classification results of each slice, for subsequent auditing and verification.

[0015] Furthermore, the classification report records the classification basis and classification results of each slice.

[0016] With the above technical solutions, the present invention uses a rule matching mechanism to perform word segmentation on the text and generate a word segmentation classification set. This process can quickly identify key information in the text. At the same time, with the help of an advanced large language model, in-depth semantic analysis is performed on the file content to generate an accurate set of file content types. By combining these two technologies, the system can comprehensively judge the category and level of the file from both low-dimensional and high-dimensional perspectives, thereby formulating the most suitable grading strategy. This method not only introduces high-dimensional semantic concepts but also can effectively guide low-dimensional word segmentation classification analysis and processing, significantly improving the accuracy and reliability of file grading and providing more accurate technical support for file management and information security. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments; Figure 1 It is a flowchart of the method for classifying and grading unstructured files based on the dual mechanisms of rules and semantics of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0019] As Figure 1 shown, the present invention discloses a method for classifying and grading unstructured files based on the dual mechanisms of rules and semantics, which includes the following: Step 1, file content reading and text extraction: Read the unstructured file and extract the text content; Specifically, the system reads the text content of the file. For pure text files (such as TXT, PDF, Word documents, etc.), the text information is directly extracted. If the file contains picture information (such as scanned documents, PPTs, PDF in picture format, etc.), the system will extract the picture content. Subsequently, OCR technology is used to recognize the text in the picture and convert it into an editable text format.

[0020] Step 2, text content slicing: Slice the extracted text content, and dynamically adjust the length and boundary of the slice according to the structure and content characteristics of the text to obtain the sliced content of independent semantic units; Specifically, the dimension of the slice is usually a natural paragraph or a text segment of a specific length. Slicing by natural paragraphs can preserve the semantic integrity of the text, while text segments of a specific length are suitable for long texts or scenarios that require more fine-grained analysis. During the slicing process, the system will dynamically adjust the length and boundary of the slice according to the structure and content characteristics of the text to ensure that each slice has an independent semantic unit.

[0021] Step 3, Word Segmentation, Rule Matching, and Semantic Analysis: Perform word segmentation on each slice of content, splitting the long text into individual lexical units; classify the word segmentation results to generate a word segmentation classification set; at the same time, vectorize the slice content and perform in-depth semantic analysis of the context relationship and semantic features of the text through a pre-trained large model to generate a file content type set. Specifically, perform word segmentation on each slice of content, splitting the long text into individual lexical units. Commonly used word segmentation tools include Jieba, HanLP, THULAC, etc., which can accurately identify the lexical boundaries in Chinese text.

[0022] Use rule matching mechanisms such as keywords, dictionaries, and regular expressions to classify the word segmentation results. Rule matching can quickly generate a word segmentation classification set, providing a basis for subsequent processing.

[0023] At the same time, vectorize the slice content and perform in-depth semantic analysis of the text content through a pre-trained large model. These models can understand the context relationship and semantic features of the text and generate a file content type set. The results of semantic analysis can complement the deficiencies of rule matching, especially when dealing with complex semantics and context information.

[0024] Step 4, Comprehensive Analysis and Grading Strategy: Integrate the file content type set and the word segmentation classification set, evaluate the actual semantics and context relationship of each slice of content, so as to perform sensitive level grading to obtain the grading results of each slice; then integrate the grading results of all slices to perform a final grading of the entire file to generate a detailed grading report.

[0025] Specifically, the system will evaluate the actual semantics and context relationship of the slice content according to the results of rule matching and semantic analysis. According to the comprehensive analysis results, different grading strategies will be adopted. For slices with consistent results in rule matching and semantic analysis, directly adopt their classification results for grading; for slices with inconsistent results, further analyze the context information, combine domain knowledge and business rules, and dynamically adjust the grading strategy. For example, some slices may contain multiple keywords, but semantic analysis shows that their actual content is not sensitive, and the system will lower their grading according to the results of semantic analysis. Integrate the grading results of all slices to perform a final grading of the entire file. The system will generate a detailed grading report, recording the classification basis and grading results of each slice for subsequent auditing and verification.

[0026] Furthermore, the grading report records the classification basis and grading results of each slice.

[0027] Specifically, in reality, the semantics of many words may vary or even be completely different in different contexts. For example, "She raised a little white rabbit" and "I'm a computer novice and don't understand anything". Different semantics will affect the classification results. In the existing technologies, whether using rule matching or NLP methods, most means still focus on short word segmentation without considering the context, making it difficult to accurately define and classify polysemous words. The main point of the present invention is to combine long and short texts, starting from both semantics and rules simultaneously, to achieve more accurate definition and classification of text content.

[0028] The present invention adopts the above technical solution. After tokenizing the text from a low-dimensional perspective, it classifies the tokens using rule matching. After slicing the text from a high-dimensional perspective, it uses a pre-trained large model to define the content of the text. According to the content definition of the document, different strategies are adopted for the classification set of tokens. For example, when the document type is a news report, the tolerance for personal information-related fields such as mobile phone numbers, addresses, and ID card numbers in the token classification set is increased; when the document type is a technical document, even if no obvious high-level data appears in the token classification set, the document is still marked as high-level.

[0029] The present invention introduces the semantic recognition ability of the large model, determines the actual type of the document content based on the document context, and adopts corresponding grading strategies for the classification results of the tokens according to the corresponding type. The present invention can significantly improve the accuracy of document grading.

[0030] Obviously, the described embodiments are part of the embodiments of the present application, rather than all embodiments. Without conflict, the embodiments and features in the present application can be combined with each other. Usually, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

Claims

1. A method for classifying and grading unstructured documents based on a dual mechanism of rules and semantics, characterized in that: It includes the following: Step 1, File Content Reading and Text Extraction: Read unstructured files and extract text content; Step 2, Text Content Slicing: Slice the extracted text content, and dynamically adjust the length and boundary of the slice according to the structure and content characteristics of the text to obtain the sliced content of independent semantic units; Step 3, Word Segmentation, Rule Matching, and Semantic Analysis: Perform word segmentation on each sliced content, split the long text into individual lexical units, classify the word segmentation results through rule matching to generate a word segmentation classification set; at the same time, vectorize the sliced content and perform in-depth semantic analysis of the context relationship and semantic features of the text through a pre-trained large model to generate a file content type set; Step 4, Comprehensive Analysis and Grading Strategy: Integrate the file content type set and the word segmentation classification set, evaluate the actual semantics and context relationship of each sliced content for sensitive level grading to obtain the grading results of each slice; then integrate the grading results of all slices to finally grade the entire file to generate a detailed grading report.

2. The method for classifying and grading unstructured documents based on the dual mechanisms of rules and semantics according to claim 1, wherein: In Step 1, directly extract text information for unstructured files of pure text; when the unstructured file contains picture information, after extracting the picture content, use OCR technology to recognize the text in the picture and convert it into an editable text format.

3. The method for classifying and grading unstructured documents based on the dual mechanisms of rules and semantics according to claim 1, characterized in that: In Step 2, the dimensions of slicing processing include natural paragraphs or text fragments of a specific length; slicing processing based on natural paragraphs preserves the semantic integrity of the text; text fragments of a specific length are applicable to long texts or scenarios that require fine-grained analysis.

4. The method for classifying and grading unstructured documents based on a dual mechanism of rules and semantics according to claim 1, wherein: In Step 3, word segmentation tools include Jieba, HanLP, and THULAC.

5. The method for classifying and grading unstructured documents based on the dual mechanisms of rules and semantics according to claim 1, wherein: In Step 3, use a rule matching mechanism of keywords, dictionaries, and regular expressions to classify the word segmentation results to generate a word segmentation classification set.

6. The method for classifying and grading unstructured documents based on a dual mechanism of rules and semantics according to claim 1, wherein: In Step 4, evaluate the actual semantics and context relationship of the sliced content according to the results of rule matching and semantic analysis.

7. The method for classifying and grading unstructured documents based on a dual mechanism of rules and semantics according to claim 1, characterized in that: In Step 4, for slices where the results of rule matching and semantic analysis are consistent, directly adopt the classification results for grading.

8. The method for classifying and grading unstructured documents based on a dual mechanism of rules and semantics according to claim 1, characterized in that: In Step 4, for slices where the results of rule matching and semantic analysis are inconsistent, analyze the context information and dynamically adjust the grading strategy in combination with domain knowledge and business rules to generate corresponding grading results.

9. The method for classifying and grading unstructured documents based on a dual mechanism of rules and semantics according to claim 1, characterized in that: In Step 4, the classification basis and grading results of each slice are recorded in the grading report.

Citation Information

Cited By

  • Unstructured data feature multimode extraction and identification method

    CN121705979A