Multi-level label and map construction method and system for multi-text data
Through the AI large language model, the automatic classification of multi-level tags and the construction of knowledge graphs has been solved, and the problem of inefficiency of traditional manual tagging methods has been achieved, the rapid, accurate processing and in-depth mining of massive public opinion data has been achieved, and technological innovation and application have been promoted.
Patent Information
- Application Number
- CN202510428420.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional manual labeling method is inefficient and expensive, making it difficult to cope with the rapid growth and rapid change of massive public opinion data.
The AI large language model is used to automatically classify multi-level tags. Through the methods of data collection and preprocessing, intelligent tag classification, knowledge graph construction and visual display, the rapid and accurate classification and labeling of public opinion data are achieved, and a clear structure of knowledge graph is constructed.
It significantly improves data processing efficiency, builds multi-level classification standards, realizes automated construction of knowledge graphs, promotes in-depth mining of data value, and promotes technological innovation and application.
Smart Images

Figure CN120011560A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent processing of multi-text data, and in particular relates to a method and system for constructing multi-level labels and graphs for multi-text data. Background Art
[0002] With the popularization of the Internet and the development of social media, the channels for public opinion expression have become increasingly diverse, generating massive and complex public opinion data presented in text form. How to effectively process and extract valuable information from these multi-text data has become the focus of attention from all walks of life.
[0003] The traditional method is to classify public opinion data through operations such as manual labeling, and then manually analyze the classified data. However, the above manual analysis method is inefficient and costly, and it is difficult to cope with the challenge of the rapid growth of public opinion data, let alone meet the rapidly changing decision-making needs. Summary of the invention
[0004] The present invention proposes a multi-level labeling and graph construction method and system for multi-text data, which can realize the rapid and accurate classification and labeling of public opinion data, and then construct a knowledge graph with clear structure and distinct hierarchy. Through intelligent label classification and knowledge graph extraction technology, the rapid analysis and visual display of public opinion data can be realized.
[0005] To achieve the above object, the technical solution of the present invention is achieved as follows: A multi-level label and graph construction method for multi-text data, comprising: S1. Data collection and preprocessing: Collect public opinion data from multiple sources and organize and build a database; S2. Intelligent label classification: Automatically calibrate multi-level labels for public opinion data based on AI large language model; S3, knowledge graph construction: intelligently analyze public opinion data with multi-level labels, construct a knowledge graph and store it in the database; S4. Visual display: Provide multiple visualization forms of knowledge graphs for knowledge graph display, and provide individual or combined display of multi-level labels.
[0006] Furthermore, in step S1, public opinion data is collected according to special projects, and the special projects include the separate use or combined use of special project conditions.
[0007] Furthermore, in step S2, the AI big model API is called, and the public opinion data is input into the model according to the set rule text to obtain the interpreted text of the model feedback; and labels are extracted from the interpreted text.
[0008] Furthermore, in step S3, the method of intelligent analysis and construction of the knowledge graph includes: taking the special general term of the public opinion data as the root node; classifying and counting the word frequency of multi-level labels respectively, and using them as nodes connected to the root node hierarchically in order of the label levels; taking the amount of data in each node as their respective weights; and constructing the knowledge graph based on the root node, nodes, connection relationships and weights.
[0009] Furthermore, in step S4, when displaying the knowledge graph, the drawing weight of the node is determined based on the weight of the node, and the knowledge graph is drawn according to the drawing weight of the node.
[0010] On the other hand, the present invention also proposes a multi-level labeling and graph construction system for multi-text data, comprising: Collection subsystem: used for data collection and preprocessing, collecting public opinion data from multiple sources and organizing and building a database; Label subsystem: used for intelligent label classification, automatically calibrating multi-level labels for public opinion data based on the AI large language model; Construction subsystem: used for knowledge graph construction: intelligent analysis of public opinion data with multi-level labels, construction of knowledge graph and storage in the database; Display subsystem: used for visual display, providing multiple visualization forms of knowledge graphs for knowledge graph display, and providing individual or combined display of multi-level tags.
[0011] Furthermore, in the collection subsystem, public opinion data is collected according to special projects, and the special projects include the separate use or combined use of special conditions.
[0012] Furthermore, the labeling subsystem includes an AI question-answering module and an interpretation module; AI question-answering module: Calls the AI big model API, inputs public opinion data into the model according to the set rules, and obtains the interpreted text of the model feedback; Interpretation module: extract labels from the interpreted text.
[0013] Furthermore, the build subsystem includes: Root node module: The special name of public opinion data is used as the root node; Statistics module: count the frequency of words for each level of tags, and use them as nodes connected to the root node in order of the tag level; Weight module: The amount of data in each node is used as the weight of each node; Building module: Build a knowledge graph based on root nodes, nodes, connection relationships and weights.
[0014] Furthermore, in the display subsystem, when displaying the knowledge graph, the drawing weight of the node is determined based on the weight of the node, and the knowledge graph is drawn according to the drawing weight of the node.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. Significantly improve data processing efficiency: The present invention uses an AI large language model to automatically label and classify public opinion data (multi-text data). Compared with the traditional manual labeling method, the processing speed of the present invention has achieved a qualitative leap, which can greatly improve the processing efficiency of massive public opinion data. This not only shortens the data processing cycle, but also makes real-time analysis possible, providing strong support for rapid response to public opinion needs.
[0016] 2. Construct multi-level classification standards: The present invention realizes the classification perception and rapid category mining required for public opinion data. At the same time, through bottom-up autonomous classification, it can further construct multi-level classification indicator standards to form a standard indicator system that can guide classification, providing support for later accurate classification.
[0017] 3. Automatically build knowledge graphs: The present invention can automatically extract nodes, weights and connection relationships through intelligent analysis of labeled public opinion data to form a knowledge graph with a clear structure and distinct levels. This automated process not only reduces the manual burden, but also ensures the integrity and accuracy of the knowledge graph, providing a solid foundation for subsequent data analysis and application.
[0018] 4. Promote in-depth mining of data value: The implementation of this invention is not limited to data processing and knowledge graph construction. More importantly, through these technical means, the deep mining of the value of massive public opinion data can be achieved. Through in-depth analysis and mining of data, potential social problems and public needs can be discovered, providing more accurate services and solutions for the government and all sectors of society.
[0019] 5. Promote technological innovation and application: The successful implementation of this invention will promote the innovation and development of data processing and analysis technology, and provide new ideas and methods for research and application in related fields. At the same time, the application of this invention will also promote the promotion and application of advanced technologies such as AI large language models in a wider range of fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic diagram of a method according to an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of the knowledge graph of an embodiment of the present invention. Figure 1 .
[0022] Figure 3 This is a schematic diagram of the knowledge graph of an embodiment of the present invention. Figure 2 . DETAILED DESCRIPTION
[0023] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0024] The design concept of this invention is to address the current technical bottlenecks in processing massive public opinion data, especially the traditional manual labeling method, which is inefficient, costly, and difficult to cope with the challenge of rapid growth in data size. By introducing artificial intelligence technology, especially the AI large language model, the public opinion data can be quickly and accurately classified and labeled, and then a knowledge graph with a clear structure and distinct hierarchy can be constructed. Through intelligent label classification and knowledge graph extraction technology, the public opinion data can be quickly analyzed and visualized.
[0025] Based on the above design concept, the present invention proposes an efficient and automated multi-level label and graph construction method and system for multi-text data. The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.
[0026] Embodiment 1: The method of this embodiment is as follows Figure 1 As shown, including: Step 1: Data collection and preprocessing: (1) In this embodiment, public opinion data is collected from multiple sources, such as Jinyun, Weibo, Xiaohongshu, etc. The collected text data can be organized and built into a database to form attribute information including title, question, time, address, etc., and stored in the database. Public opinion data can be collected without restrictions or according to special items. The special items include setting special conditions. The special conditions can be used alone or in combination of multiple special conditions.
[0027] (2) Constructing a public opinion data entry module in the system to enter the collected public opinion data into the database. The entry module can support single entry and batch entry.
[0028] Single entry means that when each piece of public opinion data is entered into the database, the required data is filled in according to the designed form. After filling in, the entry of the public opinion data is completed.
[0029] Batch entry requires first downloading the template, organizing batches of public opinion data according to the template, and entering the data. This method is mainly for accessing other source data, making it easier to enter other source data into the database in batches.
[0030] (3) An update module is constructed in the system to keep the public opinion data updated continuously. In this embodiment, the public opinion data is updated on a weekly basis, that is, the newly acquired public opinion data is entered into the database every week.
[0031] (4) Constructing a download module in the system. The massive amount of public opinion data involves public opinions in multiple aspects, multiple regions, and at different times. It is very necessary to conduct targeted special research, such as special research or regional research. Therefore, in this embodiment, the download module sets a variety of selectable filtering conditions, including filtering by region, filtering by time, and filtering by keywords. The above conditions can be used alone or in combination to download the required data from the database.
[0032] Step 2: Smart label classification.
[0033] (1) Although the currently constructed public opinion data database can download data according to various screening conditions through the download module, the downloaded data still lacks more specific classification and is very mixed. It does not utilize the rapid perception and use of public opinion data, so labeling work is needed.
[0034] For public opinion data, one label is not easy to achieve the purpose of further understanding the details of public opinion, so it is necessary to use multi-level labels for further classification during design. In this embodiment, three-level labels are used; the three-level labels have the characteristics of clear hierarchy, classification and refinement at each level.
[0035] (2) Manual labeling: The three-level labeling can be performed by relevant researchers on the public opinion data of interest. After completion, the original public opinion data will have three more label attributes, namely Label 1, Label 2, and Label 3. The characteristic of manual labeling is that the classification is more accurate, but for the massive text data of public opinion, there are problems of large workload and low efficiency.
[0036] (3) Smart labeling: Manual labeling has the problems of low efficiency and high cost. Smart labeling is adopted in this embodiment. The smart labeling is to label the three-level labels through the AI large model. The system contains two core sub-modules, the AI question and answer module and the interpretation module. Through the AI question and answer module and the interpretation module, the first-level labels are first extracted, and then the second-level labels are extracted through the AI question and answer module and the interpretation module according to the first-level labels; and then the third-level labels are extracted through the AI question and answer module and the interpretation module according to the first-level labels and the second-level labels.
[0037] (3.1) First-level label extraction process: AI question-and-answer module: The AI question-and-answer module calls the AI big model API; inputs the public opinion data into the AI big model according to the set rules, and the AI big model will feedback an interpreted text. In this embodiment, the big model called is the Wenxinyiyan big model; its interface is called through the API, and the text data of the public opinion data is used as input; at the same time, a sentence such as "Please classify the above public opinion data" is added; after calling the API, a piece of interpreted text will be returned, and the content of the interpreted text includes the respective classification terms of each public opinion data given by the big model. For example, different public opinion data may respectively obtain classification terms such as "housing and real estate, education, urban management, consumption and commerce, government services and management, commerce and consumption, public services and management, commercial consumption" and so on.
[0038] Interpretation module: extract labels from the obtained interpreted text. The extraction process includes semantic analysis and similarity comparison of the classification terms of each public opinion data, clustering the classification terms with higher similarity, and then electing a term for each cluster as the first-level label of the public opinion data corresponding to each term in the cluster, namely Label1. The election process can refer to the official classification names of public opinion. For example, the above-mentioned "housing and real estate, education, urban management, consumption and commerce, government services and management, commerce and consumption, public services and management, commercial consumption" obtained through the AI question-answering module, through semantic analysis and similarity comparison, "consumption and commerce, commerce and consumption, commercial consumption" have a high similarity. As a cluster, the official corresponding name is "consumption and commerce", so "consumption and commerce" is used as the first-level label of the public opinion data corresponding to the three classification terms "consumption and commerce, commerce and consumption, commercial consumption".
[0039] (3.2) Secondary label extraction process: AI Q&A module: The AI Q&A module calls the AI big model API; inputs the public opinion data and the first-level label into the AI big model according to the set rule text, and the AI big model will feedback the interpreted text. When inputting, add a sentence such as "Please classify the public opinion data with the first-level label above"; after calling the API, a paragraph of interpreted text will be returned, and the content of the interpreted text includes the second-level classification terms given by the big model that are further refined than the first-level label. For example, if the first-level label of the public opinion data is "education", then the second-level classification term given by the AI Q&A module this time is "educational resources".
[0040] Interpretation module: extract labels from the interpreted text obtained this time. The extraction process is the same as that in step (3.1), including semantic analysis and similarity comparison of the secondary classification words of each public opinion data, clustering the classification words with higher similarity, and then selecting a word for each cluster as the secondary label of the public opinion data corresponding to each word in the cluster, namely Label2.
[0041] (3.3) Three-level label extraction process: AI Q&A module: The AI Q&A module calls the AI big model API; inputs the public opinion data and the first-level and second-level labels into the AI big model according to the set rule text, and the AI big model will feedback the interpreted text. When inputting, add a sentence such as "Please classify the public opinion data with the first-level and second-level labels above"; after calling the API, a paragraph of interpreted text will be returned, and the content of the interpreted text includes the third-level classification terms given by the big model that are further refined than the first-level and second-level labels. For example, the first-level label of the public opinion data is "education" and the second-level label is "educational resources", then the third-level classification term given by the AI Q&A module this time is "school facilities".
[0042] Interpretation module: extract labels from the interpreted text obtained this time. The extraction process is the same as that in step (3.1), including semantic analysis and similarity comparison of the third-level classification words of each public opinion data, clustering the classification words with higher similarity, and then selecting a word for each cluster as the second-level label of the public opinion data corresponding to each word in the cluster, namely Label3.
[0043] The above-mentioned three-level label extraction process not only realizes the classification perception and rapid category mining required for public opinion data, but also through bottom-up autonomous classification, it can further construct a three-level classification indicator standard to form a standard indicator system that can guide classification, and provide support for subsequent precise classification.
[0044] Step 3: Build the knowledge graph.
[0045] (1) Construct node (Node) and edge relationship (Edge) data tables; nodes include name (name), index (id) and weight (weight); edge relationships include starting point index (start) and ending point index (end).
[0046] (2) Root node determination: Assume that there is a public opinion topic data with a total of N data items, and the steps of intelligent label classification have been completed. The root node name is the general name of the topic of public opinion data; the index is 1, and the weight is the number of items of public opinion data.
[0047] Node.name(1)= general name of public opinion data; Node.id(1) = 1; Node.weight(1) = N; (3) Label1 knowledge graph; classify and count the labels in Label1 to obtain M1 categories, with the name (Label1name) and quantity (Label1name_sum) of each category. Assign values to the node data and connection relationship data. The connection relationship is the connection relationship between the root node and Label1.
[0048] ; (4) Label2 knowledge graph: After completing the construction of Label1 knowledge graph, the construction of Label2 knowledge graph only requires the connection between Label1 and Label2. First, the labels in Label2 are classified and counted to obtain M2 categories, with the name (Label2name) and quantity (Label2name_sum) of each category.
[0049] ; The starting index of the connection relationship requires looping for each piece of public opinion data, where the starting point is the node index of Label1 and the ending point is the node index of Label2.
[0050] (5) Label3 knowledge graph: Similar to the Label2 knowledge graph, the connection relationship between Label2 and Label3 is established on the basis of completing the construction of the label2 knowledge graph. This will not be repeated here.
[0051] Step 4: Knowledge graph display.
[0052] (1) Construct a knowledge graph object based on the node (Node) and connection relationship (Edge) data tables constructed by the knowledge graph to create the knowledge graph object.
[0053] (2) Determination of drawing weight. Drawing is based on node weight. Usually, the values of the root node and the end node are very different. In order to ensure the coordination of visualization, the formula of the drawing weight (figweight) needs to be changed, that is: Node.figweight=(log(Node.weight)+1.1)*2.
[0054] (3) Supports multiple visualization forms, such as circular layout, force-directed graph layout, and hierarchical layout; and supports the combined display of different label levels. Figure 2 , Figure 3 The knowledge graph shown.
[0055] (4) In addition, a word cloud diagram can be drawn based on the node name and weight of the node data.
[0056] The beneficial effects of this embodiment include: 1. Significantly improve data processing efficiency; 2. Construct a multi-level classification standard; 3. Automatically build knowledge graphs; 4. Promote in-depth mining of data value; 5. Promote technological innovation and application.
[0057] Embodiment 2: This embodiment proposes a multi-level labeling and graph construction system for multi-text data, including: Collection subsystem: used for data collection and preprocessing, collecting public opinion data from multiple sources and organizing and building a database; Label subsystem: used for intelligent label classification, automatically calibrating multi-level labels for public opinion data based on the AI large language model; Construction subsystem: used for knowledge graph construction: intelligent analysis of public opinion data with multi-level labels, construction of knowledge graph and storage in the database; Display subsystem: used for visual display, providing multiple visualization forms of knowledge graphs for knowledge graph display, and providing individual or combined display of multi-level tags.
[0058] In the collection subsystem, public opinion data is collected according to special projects, which include the separate use or combined use of special conditions.
[0059] The labeling subsystem includes an AI question-answering module and an interpretation module; AI question-answering module: Calls the AI big model API, inputs public opinion data into the model according to the set rules, and obtains the interpreted text of the model feedback; Interpretation module: extract labels from the interpreted text.
[0060] The build subsystem includes: Root node module: The special name of public opinion data is used as the root node; Statistics module: count the frequency of each tag and use them as nodes connected to the root node in order of the tag level; Weight module: The amount of data in each node is used as the weight of each node; Building module: Build a knowledge graph based on root nodes, nodes, connection relationships and weights.
[0061] In the display subsystem, when displaying the knowledge graph, the drawing weight of the node is determined based on the weight of the node, and the knowledge graph is drawn according to the drawing weight of the node.
[0062] The multi-level labeling and graph construction system for multi-text data proposed in this embodiment can implement the multi-level labeling and graph construction method for multi-text data proposed in Example 1, and has the same technical effect as the multi-level labeling and graph construction method for multi-text data.
[0063] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-level label and graph construction method for multi-text data, characterized in that: include: S1. Data collection and preprocessing: Collect public opinion data from multiple sources and organize and build a database; S2. Intelligent label classification: Automatically calibrate multi-level labels for public opinion data based on AI large language model; S3, knowledge graph construction: intelligently analyze public opinion data with multi-level labels, construct a knowledge graph and store it in the database; S4. Visual display: Provide multiple visualization forms of knowledge graphs for knowledge graph display, and provide individual or combined display of multi-level labels.
2. The multi-level label and graph construction method for multi-text data according to claim 1, characterized in that: In step S1, public opinion data is collected according to special projects, and the special projects include the separate use or combined use of special project conditions.
3. The multi-level label and graph construction method for multi-text data according to claim 1, characterized in that: In step S2, the AI big model API is called, and the public opinion data is input into the model according to the set rule text to obtain the interpreted text of the model feedback; and labels are extracted from the interpreted text.
4. The multi-level label and graph construction method for multi-text data according to claim 2, characterized in that: In step S3, the method of intelligent analysis and construction of knowledge graph includes: taking the special general term of public opinion data as the root node; classifying and counting the word frequency of multi-level labels respectively, and using them as nodes connected to the root node hierarchically in order of label levels; taking the amount of data in each node as their respective weights; and constructing a knowledge graph based on the root node, nodes, connection relationships and weights.
5. The method for constructing multi-level labels and graphs for multi-text data according to claim 4, characterized in that: In step S4, when displaying the knowledge graph, the drawing weight of the node is determined based on the weight of the node, and the knowledge graph is drawn according to the drawing weight of the node.
6. A multi-level labeling and graph construction system for multi-text data, characterized in that: include: Collection subsystem: used for data collection and preprocessing, collecting public opinion data from multiple sources and organizing and building a database; Label subsystem: used for intelligent label classification, automatically calibrating multi-level labels for public opinion data based on the AI large language model; Construction subsystem: used for knowledge graph construction: intelligent analysis of public opinion data with multi-level labels, construction of knowledge graph and storage in the database; Display subsystem: used for visual display, providing multiple visualization forms of knowledge graphs for knowledge graph display, and providing individual or combined display of multi-level tags.
7. The multi-level labeling and graph construction system for multi-text data according to claim 6, characterized in that: In the collection subsystem, public opinion data is collected according to special projects, which include the separate use or combined use of special conditions.
8. The multi-level labeling and graph construction system for multi-text data according to claim 6, characterized in that: The labeling subsystem includes an AI question-answering module and an interpretation module; AI question-answering module: Calls the AI big model API, inputs public opinion data into the model according to the set rules, and obtains the interpreted text of the model feedback; Interpretation module: extract labels from the interpreted text.
9. The multi-level labeling and graph construction system for multi-text data according to claim 7, characterized in that: The build subsystem includes: Root node module: The special name of public opinion data is used as the root node; Statistics module: count the frequency of words for each level of tags, and use them as nodes connected to the root node in order of the tag level; Weight module: The amount of data in each node is used as the weight of each node; Building module: Build a knowledge graph based on root nodes, nodes, connection relationships and weights.
10. The multi-level labeling and graph construction system for multi-text data according to claim 9, characterized in that: In the display subsystem, when displaying the knowledge graph, the drawing weight of the node is determined based on the weight of the node, and the knowledge graph is drawn according to the drawing weight of the node.
Citation Information
Patent Citations
Commodity futures news public opinion analysis method and system
CN110377696A
Classification processing method, model training method and related devices
CN117494051A
Diabetic nephropathy knowledge graph construction method and system based on DKD clinical data
CN118070895A
Intelligent simulation model framework construction method
CN119442365A
Industrial classification label construction method and device, medium and electronic equipment
CN119475044A