Text viewpoint automatic labeling method and device based on large model and knowledge base

By introducing automatic labeling methods of large models and knowledge bases in perspective tendency analysis, the poor quality of the data set and cross-language analysis problems are solved, and more accurate text understanding and efficient labeling process are achieved.

CN119988633AActive Publication Date: 2025-05-13NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510151779.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-19
Filing Date
2025-02-12
Publication Date
2025-05-13
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing perspective-oriented analysis technology faces problems of uneven data set quality and difficult to determine the social and cultural environment in cross-language analysis, resulting in poor analysis results.

Method used

The automatic annotation method of text perspectives based on large models and knowledge bases is adopted. By identifying entities and extended entities in the knowledge base, knowledge completion is carried out, knowledge embedded text is generated, and a large language model is used for generative annotation, attitude label is obtained, and the annotation data set is finally constructed.

Benefits of technology

By introducing knowledge related to the knowledge base, the accuracy of text understanding and analysis is improved, the workload of manual labeling is reduced, and the quality of the labeling data set and the training effect of the perspective-oriented analysis model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988633A_ABST
    Figure CN119988633A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a text viewpoint automatic labeling method and device based on a large model and a knowledge base. The text viewpoint automatic labeling method based on the large model and the knowledge base comprises the steps that knowledge base entities corresponding to entities of an original text in the knowledge base and extended entities of the knowledge base entities are recognized, the extended entities are inserted into the original text based on a predefined entity relationship for knowledge completion, and a knowledge embedded text is generated; providing a given topic entity set; creating a prompt template, filling the knowledge embedded text and a given topic entity set into the prompt template, and then performing generative labeling by using a large language model to obtain attitude tags; obtaining a tendency label based on the given topic entity and the attitude label; and constructing an annotation data set based on the original text and the tendency tag. According to the technical scheme, powerful support is provided for labeling work of the high-quality labeling data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of natural language processing, and in particular to a method and device for automatically annotating text opinions based on a large model and a knowledge base. Background Art

[0002] Opinion tendency analysis originates from artificial intelligence natural language processing technology and is a cutting-edge branch of sentiment classification technology. It can be seen as a sentiment analysis task for a specific target topic, which aims to identify and analyze the tendency towards specific objects, opinions, etc. in text data.

[0003] At present, opinion tendency analysis technology mainly uses machine learning and deep learning algorithms, which learn a large amount of labeled data through training models, so as to realize the tendency judgment of unknown data. When conducting research on opinion tendency analysis, the inventor found that the quality of existing data sets was uneven, and sometimes in the process of data annotation, a deep understanding of the text was required before annotation. At the same time, in cross-language opinion tendency analysis, it is often difficult to determine the social and cultural environment to which a specific person belongs, which brings additional challenges to opinion tendency analysis. Summary of the invention

[0004] In order to solve the problems in the related art, the embodiments of the present disclosure provide a method and device for automatically annotating text opinions based on a large model and a knowledge base.

[0005] In a first aspect, the present disclosure provides a method for automatically annotating text opinions based on a large model and a knowledge base, including:

[0006] Identify knowledge base entities and extended entities of the knowledge base entities that exist in the knowledge base and correspond to entities in the original text, insert the extended entities into the original text based on predefined entity relationships to complete the knowledge, and generate knowledge embedded text;

[0007] Provides entity set for a given topic;

[0008] Create a prompt template, embed the knowledge into the text and a given topic entity set to fill the prompt template, and then use a large language model to perform generative annotation to obtain an attitude label corresponding to each given topic entity in the given topic entity set; the prompt template requires an attitude label of support, opposition or neutrality;

[0009] Based on the given topic entity, the attitude label corresponding to the given topic entity obtains a tendency label;

[0010] A labeled data set is constructed based on the original text and the tendency labels.

[0011] In an implementation of the present disclosure, the identifying knowledge base entity corresponding to the entity of the original text and the extended entity of the knowledge base entity existing in the knowledge base includes:

[0012] Construct finite state automata;

[0013] The finite state automaton is used to identify knowledge base entities corresponding to entities in the original text and extended entities of the knowledge base entities in the knowledge base.

[0014] In one implementation of the present disclosure, constructing a finite state automaton includes:

[0015] The Aho-Corasick automaton algorithm is used to preprocess all entities in the knowledge base to construct a finite state automaton.

[0016] In an implementation of the present disclosure, inserting the extended entity into the original text based on the predefined entity relationship to complete the knowledge and generate the knowledge embedded text includes:

[0017] Constructing an insertion template according to the predefined entity relationship; the predefined entity relationship includes one or more of a character nickname, a nickname and a character name, a character name and an occupation, an event name and an honor;

[0018] The extended entity is inserted into the original text based on the insertion template to complete the knowledge and generate a knowledge embedded text.

[0019] In one implementation of the present disclosure, it further includes:

[0020] Using the labeled data set to train a viewpoint tendency analysis model;

[0021] Based on the trained opinion tendency analysis model, the opinion tendency of the new input text is analyzed to obtain the tendency label of the new input text.

[0022] In one implementation of the present disclosure, performing opinion tendency analysis on a new input text based on a trained opinion tendency analysis model to obtain an opinion tendency label of the new input text includes:

[0023] Segmenting the new input text into sentences to obtain a text sentence set;

[0024] Input the given topic entity set and text sentence set into the Bert layer to obtain entity sentence vectors and text sentence vectors;

[0025] Calculate the relationship matrix between the entity sentence vector and the text sentence vector;

[0026] The relationship matrix is ​​input into the CNN feature extraction layer for feature fusion and classification, and then the tendency label of the new input text is output through the Softmax layer.

[0027] In a second aspect, an adaptive power allocation device based on deep reinforcement learning is provided in an embodiment of the present disclosure, including:

[0028] A knowledge completion module is configured to identify knowledge base entities and extended entities of the knowledge base entities existing in the knowledge base and corresponding to entities in the original text, and insert the extended entities into the original text based on predefined entity relationships to complete the knowledge and generate knowledge embedded text;

[0029] A first acquisition module is configured to provide a given topic entity set;

[0030] A labeling module is configured to create a prompt template, embed the knowledge into the text and a given topic entity set to fill the prompt template, and then use a large language model to perform generative annotation to obtain an attitude label corresponding to each given topic entity in the given topic entity set; the prompt template requires an attitude label of support, opposition or neutrality;

[0031] A second acquisition module is configured to obtain a tendency label based on the given topic entity and the attitude label corresponding to the given topic entity;

[0032] The construction module is configured to construct a labeled data set based on the original text and the tendency label.

[0033] In one implementation of the present disclosure, the part of the knowledge completion module that identifies the knowledge base entity corresponding to the entity of the original text and the extended entity of the knowledge base entity in the knowledge base is configured as follows:

[0034] Construct finite state automata;

[0035] The finite state automaton is used to identify knowledge base entities corresponding to entities in the original text and extended entities of the knowledge base entities in the knowledge base.

[0036] In one implementation of the present disclosure, the constructing of a finite state automaton is configured as follows:

[0037] The Aho-Corasick automaton algorithm is used to preprocess all entities in the knowledge base to construct a finite state automaton.

[0038] In an implementation of the present disclosure, the part of the knowledge completion module that inserts the extended entity into the original text based on the predefined entity relationship to complete the knowledge and generate the knowledge embedded text is configured as follows:

[0039] Constructing an insertion template according to the predefined entity relationship; the predefined entity relationship includes one or more of a character nickname, a nickname and a character name, a character name and an occupation, an event name and an honor;

[0040] The extended entity is inserted into the original text based on the insertion template to complete the knowledge and generate a knowledge embedded text.

[0041] In one implementation of the present disclosure, it further includes:

[0042] A training module, configured to train a viewpoint tendency analysis model using the annotated data set;

[0043] The detection module is configured to perform opinion tendency analysis on the new input text based on the trained opinion tendency analysis model to obtain the tendency label of the new input text.

[0044] In one implementation of the present disclosure, the detection module includes:

[0045] A segmentation unit, configured to perform sentence segmentation on the new input text to obtain a text sentence set;

[0046] An acquisition unit is configured to input the given topic entity set and text sentence set into the Bert layer to obtain an entity sentence vector and a text sentence vector;

[0047] A calculation unit, configured to calculate a relationship matrix between the entity sentence vector and the text sentence vector;

[0048] The output unit is configured to input the relationship matrix into the CNN feature extraction layer for feature fusion and classification, and then output the tendency label of the new input text through the Softmax layer.

[0049] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising a memory and a processor, wherein the memory is used to store one or more computer instructions, and wherein the one or more computer instructions are executed by the processor to implement a method as described in any one of the first aspects.

[0050] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the method described in any one of the first aspects is implemented.

[0051] The technical effects provided by the embodiments of the present disclosure may include the following beneficial effects:

[0052] According to the technical solution provided by the embodiment of the present disclosure, the automatic annotation method of text opinions based on a large model and a knowledge base includes: identifying the knowledge base entities corresponding to the entities of the original text and the extended entities of the knowledge base entities in the knowledge base, inserting the extended entities into the original text based on the predefined entity relationship to complete the knowledge and generate knowledge embedded text; providing a given topic entity set; creating a prompt template, filling the prompt template with the knowledge embedded text and the given topic entity set, and then using a large language model for generative annotation to obtain the attitude label corresponding to each given topic entity in the given topic entity set; the prompt template requires the support, opposition or neutral attitude label; based on the given topic entity, the attitude label corresponding to the given topic entity obtains the tendency label; based on the original text and the tendency label, the annotation data set is constructed. The above technical solution introduces the knowledge base related knowledge embedded in the original text, and the corpus obtains richer and more accurate object and topic background information, so that the model can more accurately understand and analyze the tendency of the corpus. The application of a large language model makes the annotation process more efficient and greatly reduces the workload of manual annotation. In addition, the powerful language understanding ability of the large language model can significantly improve the quality and accuracy of the annotated data set compared to other automatic annotation models, thereby further improving the training effect of the opinion tendency analysis model. In general, the solution disclosed in this disclosure effectively combines the powerful language understanding ability of the large language model with the richness and accuracy of the knowledge base, providing strong support for the annotation work of high-quality annotated data sets.

[0053] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Other features, objectives and advantages of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments in conjunction with the accompanying drawings.

[0055] Figure 1 A flowchart of a method for automatically annotating text opinions based on a large model and a knowledge base according to an embodiment of the present disclosure is shown.

[0056] Figure 2 A schematic diagram showing a training opinion tendency analysis model according to an embodiment of the present disclosure.

[0057] Figure 3 A schematic diagram showing a calculation relationship matrix according to an embodiment of the present disclosure.

[0058] Figure 4 A structural block diagram of a device for automatically annotating text opinions based on a large model and a knowledge base according to an embodiment of the present disclosure is shown.

[0059] Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0060] Figure 6 A schematic diagram showing the structure of a computer system suitable for implementing the method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0061] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts not related to the description of the exemplary embodiments are omitted in the accompanying drawings.

[0062] In the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate the presence of features, numbers, steps, behaviors, components, parts, or a combination thereof disclosed in the present specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, behaviors, components, parts, or a combination thereof exist or are added.

[0063] It should also be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0064] At present, opinion tendency analysis technology mainly uses machine learning and deep learning algorithms, which learn a large amount of labeled data through training models, so as to realize the tendency judgment of unknown data. When conducting research on opinion tendency analysis, the inventor found that the quality of existing data sets was uneven, and sometimes in the process of data annotation, a deep understanding of the text was required before annotation. At the same time, in cross-language opinion tendency analysis, it is often difficult to determine the social and cultural environment to which a specific person belongs, which brings additional challenges to opinion tendency analysis.

[0065] Taking the above-mentioned defects into consideration, the method for automatic annotation of text opinions based on a large model and a knowledge base provided by the present disclosure, by introducing knowledge base-related knowledge and embedding it into the original text, the corpus obtains richer and more accurate object and topic background information, so that the model can more accurately understand and analyze the tendency of the corpus. The application of a large language model makes the annotation process more efficient and greatly reduces the workload of manual annotation. In addition, compared with other automatic annotation models, the powerful language comprehension ability of the large language model can significantly improve the quality and accuracy of the annotated data set, thereby further improving the training effect of the opinion tendency analysis model. In general, the scheme of the present disclosure effectively combines the powerful language comprehension ability of a large language model with the richness and accuracy of knowledge in the knowledge base, and provides strong support for the annotation work of high-quality annotated data sets.

[0066] Figure 1 A flowchart of a method for automatically annotating text opinions based on a large model and a knowledge base according to an embodiment of the present disclosure is shown.

[0067] like Figure 1 As shown, the method for automatically annotating text opinions based on a large model and a knowledge base includes the following steps S110-S150:

[0068] In step S110, a knowledge base entity corresponding to an entity in the original text and an extended entity of the knowledge base entity existing in the knowledge base are identified, and the extended entity is inserted into the original text based on a predefined entity relationship to complete the knowledge and generate a knowledge embedded text;

[0069] In step S120, a given topic entity set is provided;

[0070] In step S130, a prompt template is created, the knowledge is embedded in the text and the given topic entity set is filled into the prompt template, and then a large language model is used for generative annotation to obtain an attitude label corresponding to each given topic entity in the given topic entity set; the prompt template requires an attitude label of support, opposition or neutrality;

[0071] In step S140, based on the given topic entity, the attitude label corresponding to the given topic entity obtains a tendency label;

[0072] In step S150, a labeled data set is constructed based on the original text and the tendency labels.

[0073] In the disclosed method, the knowledge base is used to complete the entities of the original text based on the predefined entity relationships. The knowledge base includes information on people, events, etc. It can be a self-built knowledge base, or it can obtain publicly available materials and generate a knowledge base based on the materials. The disclosure does not limit this. The data stored in the knowledge base is structured data, recorded as a triple set (h, r, t); where h and t are knowledge base entities, respectively, and r is an entity relationship. The predefined entity relationships include one or more of the nicknames of characters, nicknames and character names, character names and occupations, and event names and honors received.

[0074] In the disclosed method, the given topic entity may be a custom topic entity such as celebrity gossip, concerts, basketball games, athlete transfers, etc.

[0075] The attitude labels in the present disclosure may be supportive, opposing or neutral attitude labels.

[0076] In the disclosed method, a large language model is used to generatively annotate the corpus to be annotated to form a high-quality annotated data set. The large language model can be a publicly available model, such as Qwen1.5-32B-Chat-AWQ. In the disclosed method, on the basis of selecting the above model, the inference tool vllm is also used to accelerate the inference to improve the performance and efficiency of the model.

[0077] In one implementation of the present disclosure, the step S110 of identifying the knowledge base entity corresponding to the entity of the original text and the extended entity of the knowledge base entity in the knowledge base includes:

[0078] Construct finite state automata;

[0079] The finite state automaton is used to identify knowledge base entities corresponding to entities in the original text and extended entities of the knowledge base entities in the knowledge base.

[0080] In the disclosed method, the Aho-Corasick automaton algorithm is used to preprocess all entities in the knowledge base to construct a finite state automaton.

[0081] The Aho-Corasick automaton algorithm (i.e., AC automaton) is an efficient multi-pattern string matching algorithm used to simultaneously find the occurrence positions of multiple pattern strings (substrings) in a text string. In the present disclosure, the AC automaton is used to identify the knowledge base entities corresponding to the entities of the original text and the extended entities of the knowledge base entities in the knowledge base, and then the extended entities are inserted into the original text based on the predefined entity relationship to complete the knowledge and generate knowledge embedded text.

[0082] In an implementation of the present disclosure, the step S110 of inserting the extended entity into the original text based on the predefined entity relationship to complete the knowledge and generate the knowledge embedded text includes:

[0083] Constructing an insertion template according to the predefined entity relationship; the predefined entity relationship includes one or more of a character nickname, a nickname and a character name, a character name and an occupation, an event name and an honor;

[0084] The extended entity is inserted into the original text based on the insertion template to complete the knowledge and generate a knowledge embedded text.

[0085] In the disclosed method, by creating an insertion template, the extended entity is inserted near the entity of the original text according to the predefined entity relationship, thereby obtaining the knowledge embedded text.

[0086] For example, the original sample is "A's new movie sold well at the box office on the first day of its release and became a hot topic recently";

[0087] Identify the knowledge base entity "A" and the extended entity "director" corresponding to the entity of the original sample from the knowledge base;

[0088] The insertion template constructed based on the predefined entity relationship between character name and occupation is "{character}{occupation}";

[0089] Based on the occupational information of "A" being "director", the occupational information is inserted into the front of the character, and the processed knowledge embedded text is "Director A's new movie sold well at the box office on the first day of its release and has become a hot topic recently."

[0090] In one implementation of the present disclosure, a specific prompt template in step S130 is, for example, "Based on the following content T, analyze the tendency of their opinions on the elements in P as support, opposition or neutrality, without giving the analysis process."

[0091] For the above prompt template, the knowledge is embedded in the text as content T, and filled into the prompt template with the given topic entity set P, and then the attitude label corresponding to each given topic entity is generated using the large language model. Then in step S140, based on the given topic entity, the attitude label corresponding to the given topic entity obtains the tendency label.

[0092] Taking the given topic entities of basketball games, celebrity gossip, and athlete transfers as an example, the obtained tendency labels are, for example, "basketball games, support", "celebrity gossip, neutral", and "athlete transfers, oppose".

[0093] In one implementation of the present disclosure, the method for automatically annotating text opinions based on a large model and a knowledge base further includes:

[0094] Using the labeled data set to train a viewpoint tendency analysis model;

[0095] Based on the trained opinion tendency analysis model, the opinion tendency of the new input text is analyzed to obtain the tendency label of the new input text.

[0096] In the disclosed method, the opinion tendency analysis model can be a Bert-Condition-CNN opinion tendency analysis model. In this model, the Bert layer is used to process a given topic entity set and a text sentence set (the original text is obtained by sentence segmentation) to obtain two sentence vectors, and then the Condition layer of the relationship matrix of the two sentence vectors is constructed to calculate the relationship features of each given topic entity and each text sentence, and finally the CNN feature extraction layer is used to learn the final tendency information and input it into the Softmax layer to obtain the attitude label of the text on the given topic entity, whether it is support, opposition or neutral.

[0097] Figure 2 A schematic diagram showing a training opinion tendency analysis model according to an embodiment of the present disclosure.

[0098] like Figure 2 As shown, first, the unlabeled original text is supplemented with knowledge through the knowledge base, and the knowledge information (extended entity) in the knowledge base is inserted near the entity of the original text according to the predefined entity relationship, so as to obtain the knowledge embedded text. Then, the knowledge embedded text is annotated with the large language model to obtain the tendency label of the original text, and the tendency label and the original text form a labeled data set. When the labeled data set is used to train the opinion tendency analysis model, the loss function is constructed according to the tendency label output by the model and the tendency label in the labeled data set, and the weight parameters of the Bert layer, Condition layer, and CNN feature extraction layer are adjusted. After the model converges, the trained opinion tendency analysis model is obtained.

[0099] The specific training process of the above-mentioned opinion tendency analysis model can refer to the existing technology and is not the focus of this disclosure.

[0100] In one implementation of the present disclosure, performing opinion tendency analysis on a new input text based on a trained opinion tendency analysis model to obtain an opinion tendency label of the new input text includes:

[0101] Segmenting the new input text into sentences to obtain a text sentence set;

[0102] Input the given topic entity set and text sentence set into the Bert layer to obtain entity sentence vectors and text sentence vectors;

[0103] Calculate the relationship matrix between the entity sentence vector and the text sentence vector;

[0104] The relationship matrix is ​​input into the CNN feature extraction layer for feature fusion and classification, and then the tendency label of the new input text is output through the Softmax layer.

[0105] In the disclosed method, a given topic entity set is recorded as Entities = {e1, e2, ..., e m},e i (1≤i≤m) is a given topic entity, and the text sentence set is counted as Sents={s1,s2,...,s n},s j (1≤j≤n) is a text sentence, where m and n represent the number of given topic entities contained in the given topic entity set and the number of sentences contained in the text, respectively.

[0106] Figure 3 A schematic diagram showing a calculation relationship matrix according to an embodiment of the present disclosure.

[0107] Combination Figure 3 As shown, the relationship matrix between the entity sentence vector and the text sentence vector is calculated, and the relationship matrix covers the given topic entity e i and text sentences j The implicit relationship and tendency information in this pair of text sequences can be divided into the following two steps: First, determine s j Is it around e i The first is to conduct a review, that is, whether there is an implicit relationship between the two text sequences; the second is s j For e i The tendency is to support, neutral or oppose. j With e i When there is no implication relationship, its tendency is neutral. Before calculating the relationship matrix between Entities and Sents, first i and j The output of the Bert layer is a sentence vector, denoted as u i = Bert(e i ), v j =Bert(s j ), the element x in the obtained relationship matrix ij =Score(u i ,v j ). Among them, in the calculation element c ij When the distance between the reference vectors is calculated, this disclosure will not elaborate on this.

[0108] In the disclosed method, the CNN feature extraction layer is used for the calculated c ij Perform feature fusion and classification, and calculate adjacent sequence pairs through two-dimensional convolution i ,v j The weight of the relationship feature of > on the final tendency label is calculated as follows:

[0109]

[0110] In the formula, C(m,n) is the e in the above text. i and j The relationship matrix of K(im,jn), K(im,jn) is the convolution kernel weight parameter, b is the bias term, and f is the nonlinear activation function, usually Relu, Sigmoid or Tanh. The feature matrix after convolution (whose matrix elements are M i,j ​) must be processed by the maximum pooling layer, and the pooled feature vector is fully connected for feature fusion, and then the Softmax algorithm is used to classify it. The main tasks of the fully connected layer and the Softmax layer are to fuse the feature information finally obtained, obtain the score of the feature vector for each tendency label, and output the final tendency label of <entities, sents>. This paper uses the Softmax layer as a probability conversion layer, which represents the input vector in the form of probability to complete the prediction of the tendency label.

[0111] The method for automatic text opinion annotation based on a large model and knowledge base provided by the present disclosure combines the powerful language understanding ability of a large language model with the richness and accuracy of knowledge base knowledge. Its application can not only improve the accuracy and robustness of opinion tendency analysis, but also improve annotation efficiency and reduce labor costs, providing strong support for research and application in the field of text analysis.

[0112] Figure 4 The structural block diagram of the apparatus for automatically annotating text opinions based on a large model and a knowledge base according to an embodiment of the present disclosure is shown. The apparatus can be implemented as part or all of an electronic device through software, hardware, or a combination of both.

[0113] like Figure 4 As shown, the text viewpoint automatic annotation device 400 based on a large model and a knowledge base includes:

[0114] The knowledge completion module 410 is configured to identify knowledge base entities and extended entities of the knowledge base entities that exist in the knowledge base and correspond to the entities of the original text, and insert the extended entities into the original text based on predefined entity relationships to complete the knowledge and generate knowledge embedded text;

[0115] A first acquisition module 420 is configured to provide a given topic entity set;

[0116] The labeling module 430 is configured to create a prompt template, embed the knowledge into the text and the given topic entity set to fill the prompt template, and then use the large language model to perform generative annotation to obtain an attitude label corresponding to each given topic entity in the given topic entity set; the prompt template requires an attitude label of support, opposition or neutrality;

[0117] A second acquisition module 440 is configured to obtain a tendency label based on the given topic entity and the attitude label corresponding to the given topic entity;

[0118] The construction module 450 is configured to construct a labeled data set based on the original text and the tendency label.

[0119] The text opinion automatic annotation device based on a large model and a knowledge base provided by the present disclosure, by introducing knowledge base related knowledge and embedding it into the original text, the corpus obtains richer and more accurate object and topic background information, so that the model can more accurately understand and analyze the tendency of the corpus. The application of a large language model makes the annotation process more efficient and greatly reduces the workload of manual annotation. In addition, compared with other automatic annotation models, the powerful language comprehension ability of the large language model can significantly improve the quality and accuracy of the annotated data set, thereby further improving the training effect of the opinion tendency analysis model. In general, the scheme of the present disclosure effectively combines the powerful language comprehension ability of the large language model with the richness and accuracy of the knowledge base, and provides strong support for the annotation work of high-quality annotated data sets.

[0120] In one implementation of the present disclosure, the part of the knowledge completion module that identifies the knowledge base entity corresponding to the entity of the original text and the extended entity of the knowledge base entity in the knowledge base is configured as follows:

[0121] Construct finite state automata;

[0122] The finite state automaton is used to identify knowledge base entities corresponding to entities in the original text and extended entities of the knowledge base entities in the knowledge base.

[0123] In one implementation of the present disclosure, the constructing of a finite state automaton is configured as follows:

[0124] The Aho-Corasick automaton algorithm is used to preprocess all entities in the knowledge base to construct a finite state automaton.

[0125] In an implementation of the present disclosure, the part of the knowledge completion module that inserts the extended entity into the original text based on the predefined entity relationship to complete the knowledge and generate the knowledge embedded text is configured as follows:

[0126] Constructing an insertion template according to the predefined entity relationship; the predefined entity relationship includes one or more of a character nickname, a nickname and a character name, a character name and an occupation, an event name and an honor;

[0127] The extended entity is inserted into the original text based on the insertion template to complete the knowledge and generate a knowledge embedded text.

[0128] In one implementation of the present disclosure, it further includes:

[0129] A training module, configured to train a viewpoint tendency analysis model using the annotated data set;

[0130] The detection module is configured to perform opinion tendency analysis on the new input text based on the trained opinion tendency analysis model to obtain the tendency label of the new input text.

[0131] In one implementation of the present disclosure, the detection module includes:

[0132] A segmentation unit, configured to perform sentence segmentation on the new input text to obtain a text sentence set;

[0133] An acquisition unit is configured to input the given topic entity set and text sentence set into the Bert layer to obtain an entity sentence vector and a text sentence vector;

[0134] A calculation unit, configured to calculate a relationship matrix between the entity sentence vector and the text sentence vector;

[0135] The output unit is configured to input the relationship matrix into the CNN feature extraction layer for feature fusion and classification, and then output the tendency label of the new input text through the Softmax layer.

[0136] The present disclosure also discloses an electronic device, Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0137] like Figure 5 As shown, the electronic device includes a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to an embodiment of the present disclosure.

[0138] The method for automatically annotating text opinions based on a large model and a knowledge base includes:

[0139] Identify knowledge base entities and extended entities of the knowledge base entities that exist in the knowledge base and correspond to entities in the original text, insert the extended entities into the original text based on predefined entity relationships to complete the knowledge, and generate knowledge embedded text;

[0140] Provides entity set for a given topic;

[0141] Create a prompt template, embed the knowledge into the text and a given topic entity set to fill the prompt template, and then use a large language model to perform generative annotation to obtain an attitude label corresponding to each given topic entity in the given topic entity set; the prompt template requires an attitude label of support, opposition or neutrality;

[0142] Based on the given topic entity, the attitude label corresponding to the given topic entity obtains a tendency label;

[0143] A labeled data set is constructed based on the original text and the tendency labels.

[0144] In an implementation of the present disclosure, the identifying knowledge base entity corresponding to the entity of the original text and the extended entity of the knowledge base entity existing in the knowledge base includes:

[0145] Construct finite state automata;

[0146] The finite state automaton is used to identify knowledge base entities corresponding to entities in the original text and extended entities of the knowledge base entities in the knowledge base.

[0147] In one implementation of the present disclosure, constructing a finite state automaton includes:

[0148] The Aho-Corasick automaton algorithm is used to preprocess all entities in the knowledge base to construct a finite state automaton.

[0149] In an implementation of the present disclosure, inserting the extended entity into the original text based on the predefined entity relationship to complete the knowledge and generate the knowledge embedded text includes:

[0150] Constructing an insertion template according to the predefined entity relationship; the predefined entity relationship includes one or more of a character nickname, a nickname and a character name, a character name and an occupation, an event name and an honor;

[0151] The extended entity is inserted into the original text based on the insertion template to complete the knowledge and generate a knowledge embedded text.

[0152] In one implementation of the present disclosure, it further includes:

[0153] Using the labeled data set to train a viewpoint tendency analysis model;

[0154] Based on the trained opinion tendency analysis model, the opinion tendency of the new input text is analyzed to obtain the tendency label of the new input text.

[0155] In one implementation of the present disclosure, performing opinion tendency analysis on a new input text based on a trained opinion tendency analysis model to obtain an opinion tendency label of the new input text includes:

[0156] Segmenting the new input text into sentences to obtain a text sentence set;

[0157] Input the given topic entity set and text sentence set into the Bert layer to obtain entity sentence vectors and text sentence vectors;

[0158] Calculate the relationship matrix between the entity sentence vector and the text sentence vector;

[0159] The relationship matrix is ​​input into the CNN feature extraction layer for feature fusion and classification, and then the tendency label of the new input text is output through the Softmax layer.

[0160] Figure 6 A schematic diagram showing the structure of a computer system suitable for implementing the method according to an embodiment of the present disclosure is shown.

[0161] like Figure 6 As shown, the computer system includes a processing unit, which can perform the various methods in the above-mentioned embodiments according to the program stored in the read-only memory (ROM) or the program loaded from the storage part into the random access memory (RAM). In the RAM, various programs and data required for the operation of the computer system are also stored. The processing unit, ROM and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.

[0162] The following components are connected to the I / O interface: an input part including a keyboard, a mouse, etc.; an output part including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage part including a hard disk, etc.; and a communication part including a network interface card such as a LAN card, a modem, etc. The communication part performs a communication process via a network such as the Internet. The drive is also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that the computer program read therefrom is installed into the storage part as needed. Among them, the processing unit can be implemented as a processing unit such as a CPU, a GPU, a TPU, an FPGA, an NPU, etc.

[0163] In particular, according to an embodiment of the present disclosure, the method described above can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly contained on a machine-readable medium, and the computer program includes a program code for executing the above method. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium.

[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0165] The units or modules involved in the embodiments described in the present disclosure may be implemented by software or by programmable hardware. The units or modules described may also be set in a processor, and the names of these units or modules do not constitute limitations on the units or modules themselves in some cases.

[0166] As another aspect, the present disclosure further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the electronic device or computer system in the above embodiment; or a computer-readable storage medium that exists independently and is not assembled into a device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the method described in the present disclosure.

[0167] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other.

Claims

1. A method for automatic annotation of text opinions based on a large model and a knowledge base, characterized in that: include: Identify knowledge base entities and extended entities of the knowledge base entities that exist in the knowledge base and correspond to entities in the original text, insert the extended entities into the original text based on predefined entity relationships to complete the knowledge, and generate knowledge embedded text; Provides entity set for a given topic; Creating a prompt template, embedding the knowledge into text and a given topic entity set to fill the prompt template, and then using a large language model to perform generative annotation to obtain an attitude label corresponding to each given topic entity in the given topic entity set; The prompt template requires a support, opposition or neutral attitude label; Based on the given topic entity, the attitude label corresponding to the given topic entity obtains a tendency label; A labeled data set is constructed based on the original text and the tendency labels.

2. The method for automatically annotating text opinions based on a large model and a knowledge base according to claim 1 is characterized in that: The identification of knowledge base entities corresponding to entities in the original text and the extended entities of the knowledge base entities in the knowledge base includes: Construct finite state automata; The finite state automaton is used to identify knowledge base entities corresponding to entities in the original text and extended entities of the knowledge base entities in the knowledge base.

3. The method for automatic annotation of text opinions based on a large model and a knowledge base according to claim 2 is characterized in that: The construction of the finite state automaton includes: The Aho-Corasick automaton algorithm is used to preprocess all entities in the knowledge base to construct a finite state automaton.

4. The method for automatic annotation of text opinions based on a large model and a knowledge base according to claim 1 is characterized in that: The step of inserting the extended entity into the original text based on the predefined entity relationship to complete the knowledge and generate the knowledge embedded text comprises: Constructing an insertion template according to the predefined entity relationship; the predefined entity relationship includes one or more of a character nickname, a nickname and a character name, a character name and an occupation, an event name and an honor; The extended entity is inserted into the original text based on the insertion template to complete the knowledge and generate a knowledge embedded text.

5. The method for automatic annotation of text opinions based on a large model and a knowledge base according to claim 1 is characterized in that: Also includes: Using the labeled data set to train a viewpoint tendency analysis model; Based on the trained opinion tendency analysis model, the opinion tendency of the new input text is analyzed to obtain the tendency label of the new input text.

6. The method for automatic annotation of text opinions based on a large model and a knowledge base according to claim 5 is characterized in that: The opinion tendency analysis of the new input text is performed based on the trained opinion tendency analysis model to obtain the tendency label of the new input text, including: Segmenting the new input text into sentences to obtain a text sentence set; Input the given topic entity set and text sentence set into the Bert layer to obtain entity sentence vectors and text sentence vectors; Calculate the relationship matrix between the entity sentence vector and the text sentence vector; The relationship matrix is ​​input into the CNN feature extraction layer for feature fusion and classification, and then the tendency label of the new input text is output through the Softmax layer.

7. A text opinion automatic annotation device based on a large model and a knowledge base, characterized in that: include: A knowledge completion module is configured to identify knowledge base entities and extended entities of the knowledge base entities existing in the knowledge base and corresponding to entities in the original text, and insert the extended entities into the original text based on predefined entity relationships to complete the knowledge and generate knowledge embedded text; A first acquisition module is configured to provide a given topic entity set; A labeling module is configured to create a prompt template, embed the knowledge into text and a given topic entity set to fill the prompt template, and then use a large language model to perform generative annotation to obtain an attitude label corresponding to each given topic entity in the given topic entity set; The prompt template requires a support, opposition or neutral attitude label; A second acquisition module is configured to obtain a tendency label based on the given topic entity and the attitude label corresponding to the given topic entity; The construction module is configured to construct a labeled data set based on the original text and the tendency label.

8. The text viewpoint automatic annotation device based on a large model and a knowledge base according to claim 7 is characterized in that: The part of the knowledge completion module that identifies the knowledge base entity corresponding to the entity of the original text and the extended entity of the knowledge base entity in the knowledge base is configured as follows: Construct finite state automata; The finite state automaton is used to identify knowledge base entities corresponding to entities in the original text and extended entities of the knowledge base entities in the knowledge base.

9. An electronic device, characterized in that: The method comprises a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Microblog sentiment analysis, evaluation and pushing method

    CN112612971A

  • Chinese dialogue semantic role labeling method and system

    CN114625830A

  • Dialogue pre-labeling method and system, computer equipment and storage medium

    CN116860921A

  • Controllable text generation method and device based on large language model

    CN117216193A

  • Medical named entity identification method and device based on large language model

    CN118114675A