Bulletin multi-label classification method and system based on rule model

Through the multi-label classification method of announcements based on the rules model, the multi-label classification of announcements of listed companies is solved, and the problems of insufficient single classification and strong subjectivity are achieved, and more accurate and objective announcement classification is achieved.

CN120179820APending Publication Date: 2025-06-20SSE INFORMATION NETWORK LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510413291.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing technology has problems of insufficient single classification and strong subjectivity when announcing a listed company, which leads to classification errors that may affect the normal progress of the information disclosure process.

Method used

The multi-label classification method of announcements based on the rules model is adopted, and the announcement files are converted, paragraph segmentation, word segmentation and annotated through the rules engine. The word2vec model and Jieba dictionary are used to combine the rule database formed by expert experience to classify the announcement contents multi-label classification.

Benefits of technology

The objective analysis and multi-label classification of the announcement content are realized, subjectivity is reduced, the accuracy and consistency of classification is improved, and the reliability of the information disclosure process is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179820A_ABST
    Figure CN120179820A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text label classification, and provides an announcement multi-label classification method and system based on a rule model, and the announcement multi-label classification method based on the rule model comprises the steps: S1, converting the format of an announcement file into a txt text format; s2, matching recognition is conducted on the primary title and the secondary title through a rule engine, a text of the announcement file is segmented into a plurality of paragraphs through paragraph rules, the paragraph rules comprise a numbering rule based on expert experience and a paragraph indentation rule, and the paragraphs comprise paragraph titles and paragraph texts; s3, performing word segmentation processing on the paragraph titles and the paragraph texts through a word segmentation library, and labeling the paragraph titles and the paragraph texts after word segmentation processing according to a rule fitting degree through a rule model; and S4, processing the paragraph title, the paragraph text and the correspondingly labeled label into a preset format and storing. After the method is introduced, the content of the whole announcement can be objectively analyzed by utilizing an automatic processing technology and combining a rule engine library in an announcement publishing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text label classification, and in particular to a method and system for multi-label classification of announcements based on a rule model. Background Art

[0002] Currently, after a listed company's announcement is written, it needs to be submitted on a professional system. When submitting, the relevant classification is selected as a unique label for the announcement, which is passed on during the subsequent announcement release process. The basis for the classification label is generally determined by the business that the announcement mainly involves.

[0003] However, with the continuous changes in the business of listed companies, many single classifications can no longer describe the characteristics of announcements well, and this manual selection method is more subjective, and it is easy to add subjective intentions when classifying, which may affect objective facts. Incorrect classification of listed company announcements may lead to abnormalities in subsequent information disclosure. Summary of the invention

[0004] In order to help solve the above technical problems, the present application provides a method and system for multi-label classification of announcements based on a rule model.

[0005] In the first aspect, the present application provides a method for multi-label classification of announcements based on a rule model, which adopts the following technical solution: A method for classifying announcements using multiple labels based on a rule model, wherein the method comprises: Step S1: Convert the format of the announcement file into txt text format; Step S2: matching and identifying the first-level title and the sub-title through the rule engine, and dividing the text of the announcement document into multiple paragraphs, wherein the paragraphs include paragraph titles and paragraph bodies. Matching and identifying the first-level title and the sub-title through the rule engine includes matching the first-level title and the sub-title content in sequence according to the priority of the regular expression by combining the Python development language and the Java development language through regular expressions, and the first matched content is used as the result of the rule engine matching and identification; Step S3: segmenting the paragraph title and paragraph text through the word segmentation library to obtain multiple phrases, and annotating the paragraph title and paragraph text after the word segmentation according to the rule fit through the rule model, including: the rule model is the word2vec open source model, each of the phrases is converted into a word segmentation vector, traversing the feature word vectors in the rule library, calculating the cosine similarity between the word segmentation vector and the feature word vector, and if the similarity exceeds a preset threshold, it is determined that the match is successful, and then annotating; Step S4: Process the paragraph title, paragraph text and corresponding annotated tags into a preset format and store them.

[0006] Preferably, step S1 includes: Converting the announcement file in doc format into an announcement file in txt text format through the poi open-source library, Converting the announcement file in docx format into an announcement file in txt text format through the dom4j open-source library, Converting the announcement file in pdf format into an announcement file in txt text format through the pdfplumber open-source library.

[0007] Preferably, step S2 includes: the first-level headings are represented by Chinese numerals, the secondary headings are represented by Arabic numerals, and the rule engine matches and identifies the first-level headings and secondary headings through Chinese numerals and Arabic numerals.

[0008] Preferably, step S3 includes: the tags include at least first-level tags and second-level tags, and the word segmentation library is the Chinese word segmentation library jieba.

[0009] Preferably, step S4 includes: the preset format is in json string format, and the paragraph headings, paragraph texts, and corresponding labeled tags are processed into json string format and stored in the database.

[0010] In a second aspect, the present application provides a multi-label classification system for announcements based on a rule model, adopting the following technical solution: A multi-label classification system for announcements based on a rule model that adopts the multi-label classification method for announcements based on a rule model as described in any one of the foregoing first aspects, wherein the multi-label classification system for announcements based on a rule model includes: An announcement submission module for receiving the announcements written by users; A format conversion module for executing step S1; A paragraph segmentation module for executing step S2; A tagging module for executing step S3; A database for storing the paragraph headings, paragraph texts, and corresponding labeled tags in the preset format of step S4.

[0011] In summary, after introducing the present application, in the announcement release process, the automated processing technology can be used to objectively analyze the content of the entire announcement in combination with the rule engine library formed by expert experience, and the multi-classification label system can be flexibly used to treat the whole and parts of the announcement differently. For announcements marked with multi-label classification, it is more convenient for market participants to interpret, and valuable content can also be refined and interpreted according to the paragraphs where the tags are located. At the same time, it also provides more effective corpus data for intelligent semantic analysis and parsing of announcements. Description of the Drawings

[0012] Figure 1 It is a flow schematic diagram of a multi-label classification method for announcements based on a rule model of this application. Specific implementation manners

[0013] The following further explains this application with reference to the accompanying drawings. The structure and principle of this application are very clear to those skilled in the art. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0014] Figure 1 It is a flow schematic diagram of a multi-label classification method for announcements based on a rule model of this application. The multi-label classification method for announcements based on a rule model of this application includes: Step S1: Convert the format of the announcement file into a txt text format; in step S1, the announcement file in doc format is converted into an announcement file in txt text format through the poi open-source library, the announcement file in docx format is converted into an announcement file in txt text format through the dom4j open-source library, and the announcement file in pdf format is converted into an announcement file in txt text format through the pdfplumber open-source library.

[0015] dom4j is an open-source Java XML processing library, widely used for XML parsing, generation, modification, and serialization. The poi (apache poi) open-source library is an open-source Java library of the apache software foundation, used for reading and writing Microsoft Office format files (such as Excel, Word, PPT, etc.). pdfplumber is an open-source Python library, used for extracting text and table data from PDF documents.

[0016] Step S2: Match and identify the first-level headings and secondary headings through a rule engine, and segment the text of the announcement file into multiple paragraphs through paragraph rules. The paragraph rules include a numbering rule based on expert experience and a paragraph indentation rule. A paragraph includes a paragraph heading and a paragraph body. In step S2, the first-level headings are represented by Chinese numerals, and the secondary headings are represented by Arabic numerals. The rule engine matches and identifies the first-level headings and secondary headings through Chinese numerals and Arabic numerals.

[0017] It should be noted here that matching and identifying the first-level headings and secondary headings through a rule engine includes using regular expressions in combination with Python and Java development languages, and sequentially matching the content of the first- and second-level headings according to the priority of regular expressions. The content that is matched first is used as the result of the rule engine's matching and identification.

[0018] Step S3: Segment the paragraph title and paragraph text through a segmentation library, and label the segmented paragraph title and paragraph text according to the rule fitness through a rule model; in Step S3, the labels at least include first-level labels and second-level labels, and the segmentation library is the Chinese segmentation library jieba.

[0019] Jieba is a Chinese word segmentation tool based on Python, which is efficient and flexible. It supports multiple word segmentation modes and can meet the word segmentation needs in different scenarios. The core function of the jieba library is to split continuous Chinese text into lexical units with semantic or grammatical meanings.

[0020] It should be noted here that labeling the segmented paragraph title and paragraph text according to the rule fitness through a rule model includes: after word segmentation, using the word2vec open-source model to convert each phrase into a vector, and at the same time traversing the feature word vector (word2vec conversion) library calculated by the rule model in turn, calculating the cosine similarity of the two word vectors in turn. If the similarity exceeds the preset threshold of more than 95%, it is considered a successful match. Word2Vec is an open-source word vector generation model widely used in the field of natural language processing.

[0021] Step S4: Process the paragraph title, paragraph text, and corresponding labeled tags into a preset format and store them. In Step S4, the preset format is the json string format. The paragraph title, paragraph text, and corresponding labeled tags are processed into the json string format and stored in the database. Json (JavaScript Object Notation) is a lightweight data exchange format, which is easy for humans to read and write, and is also easy for machines to parse and generate.

[0022] This application also provides a rule model-based announcement multi-label classification system that adopts the above-mentioned rule model-based announcement multi-label classification method. The system includes an announcement submission module, a format conversion module, a paragraph segmentation module, a tagging module, and a database. The announcement submission module is used to receive the announcements written by users. The format conversion module is used to execute Step S1. The paragraph segmentation module is used to execute Step S2. The tagging module is used to execute Step S3. The database is used to store the paragraph titles, paragraph texts, and corresponding labeled tags in the preset format of Step S4.

[0023] Specifically, after the user completes the announcement writing and submits it to the announcement submission system, a classification will be automatically given to the full text of the announcement using a rule model based on the text of the announcement's main title. The announcement file may be one of doc, docx, or pdf. Therefore, open-source libraries such as poi, dom4j, and pdfplumber will be used to convert the announcement file into a txt text format. Furthermore, a rule engine will be used to match and identify the first-level headings in Chinese numerals and the second-level headings in Arabic numerals, and the txt text will be segmented into the form of paragraph headings + paragraph text, where the deepest paragraph level does not exceed three levels.

[0024] Then, for each paragraph heading and paragraph text, segmentation is completed using the Chinese word segmentation library jieba, and the rule models in the tag library are sequentially matched. According to the rule fitting degree, one or more tags are assigned to the paragraph text related to each paragraph heading. The tag system of this embodiment can be hierarchical. For example, the first-level tag is major asset restructuring, and the second-level tags can be progress announcements, result announcements, approval announcements, etc. Finally, the paragraph headings, paragraph text, and multiple tags are packaged into the general json key:value format. The key:value format is the core component of json and is used to represent key-value pair data.

[0025] It should be noted here that in this application, for the segmentation of paragraphs, it is mainly based on paragraph rule features, such as numbering rules and paragraph indentation rules based on expert experience. The method is more general and has stronger adaptability. Moreover, the method of our segmentation is mainly used for a preliminary processing of the text structure and does not require the content segmented from the text itself to necessarily conform to the writing norms. Because the main idea in our invention is to tag the paragraph content, as long as relatively accurate tags can be assigned, some problems in the announcement release process can be solved.

[0026] In this application, whether a paragraph can completely split the text content mainly depends on expert experience, including but not limited to paragraph coding, paragraph indentation format, etc. To put it another way, even if the paragraph splitting fails due to a lack of expert experience, it will not affect the subsequent tagging process because this application will process the text again after word segmentation as long as there is no omission.

[0027] This application allows paragraph splitting to fail because the main objective of this application is to endow the text content with the attributes of one or more tags. The technical problem to be solved by this application is to objectively analyze the content of the entire announcement and assign tags. Although the paragraph splitting fails, the text content itself remains unchanged. This application can still complete the relevant tag recognition work by performing word segmentation and calculating word vector similarity. The numbering rule based on expert experience has the disadvantage of large splitting errors. However, since this application pays more attention to objectively analyzing the content of the entire announcement and the tagging is more in line with the accuracy requirements in the field of listed company announcements, this application can tolerate the disadvantage of large splitting errors in the numbering rule based on expert experience. This application does not require a complex splitting method to solve the technical problem of this application, that is, this application overcomes the technical prejudice.

[0028] However, in the prior art, first of all, there are relatively strict restrictions on the document content structure. The failure of paragraph splitting will result in the inability to find words with specific features in the corresponding paragraph, or cause the word frequency of relevant words to be abnormally low and less than the preset TF-IDF value, thus missing some keywords. These missing keywords will cause the failure of all subsequent steps. Therefore, the prior art cannot use the numbering rule based on expert experience of this application, and at the same time, it must adopt a method of accurately splitting paragraphs to process the document, increasing the complexity of the system.

[0029] This application calculates each word segmentation vector because the technical problem to be solved by this application is to perform tagging processing on the content of the full text, and coexist the tag as an attribute with the text for subsequent objective analysis.

Claims

1. A multi-label classification method for announcements based on a rule model, characterized in that: The announcement multi-label classification method based on the rule model includes: Step S1: Convert the format of the announcement file into txt text format; Step S2: matching and identifying the first-level title and the sub-title by a rule engine, and dividing the text of the announcement document into multiple paragraphs by paragraph rules, wherein the paragraph rules include numbering rules and paragraph indentation rules based on expert experience, and the paragraphs include paragraph titles and paragraph bodies. Matching and identifying the first-level title and the sub-title by the rule engine includes matching the first-level title and the sub-title content in sequence according to the priority of the regular expression by combining regular expressions with Python development language and Java development language, and the first matched content is used as the result of matching and identification by the rule engine; Step S3: segmenting the paragraph title and paragraph text through the word segmentation library to obtain multiple phrases, and annotating the paragraph title and paragraph text after the word segmentation according to the rule fit through the rule model, including: the rule model is the word2vec open source model, each of the phrases is converted into a word segmentation vector, traversing the feature word vectors in the rule library, calculating the cosine similarity between the word segmentation vector and the feature word vector, and if the similarity exceeds a preset threshold, it is determined that the match is successful, and then annotating; Step S4: Process the paragraph title, paragraph text and corresponding annotated tags into a preset format and store them.

2. The method for multi-label classification of announcements based on rule model according to claim 1 is characterized in that: The step S1 comprises: Use the poi open source library to convert the announcement file in doc format into an announcement file in txt text format. Use the dom4j open source library to convert the announcement file in docx format into an announcement file in txt text format. Convert announcement files in PDF format to TXT text format using the pdfplumber open source library.

3. The method for multi-label classification of announcements based on rule models according to claim 1 is characterized in that: The step S2 includes: the first-level title is represented by Chinese numerals, the sub-title is represented by Arabic numerals, and the rule engine matches and identifies the first-level title and the sub-title through the Chinese numerals and Arabic numerals.

4. The method for multi-label classification of announcements based on rule models according to claim 1 is characterized in that: The step S3 includes: the label at least includes a primary label and a secondary label, and the word segmentation library is a Chinese word segmentation library jieba.

5. The method for multi-label classification of announcements based on rule models according to claim 1 is characterized in that: The step S4 includes: the preset format is a json string format, and the paragraph title, paragraph text and corresponding annotated tags are processed into a json string format and stored in a database.

6. A rule-based announcement multi-label classification system using the rule-based announcement multi-label classification method according to any one of claims 1 to 5, characterized in that: The announcement multi-label classification system based on the rule model includes: Announcement submission module, used to receive announcements written by users; A format conversion module, used to execute step S1; A paragraph segmentation module, used to execute step S2; A marking module, used to execute step S3; The database is used to store the paragraph titles, paragraph texts and corresponding annotated tags in the preset format of step S4.