A method, device and computer-readable medium for sentiment analysis
By introducing knowledge graphs into sentiment analysis in specific fields and assigning nodes and edges emotion factor vectors, the semantic analysis error problem caused by general data sets is solved, and more accurate sentiment analysis results are achieved.
Patent Information
- Application Number
- CN202110172800.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-08
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-02-08
AI Technical Summary
Existing sentiment analysis algorithms use common data sets in specific domains, which may lead to incorrect semantic analysis results, especially in the field of cybersecurity, where general dictionaries are difficult to accurately identify domain-specific emotional polarity.
Introduce knowledge graphs in specific fields and give emotional factor vectors to each node and edge in the knowledge graph to represent the emotional polarity of words, thereby constraining the emotional analysis process and improving the accuracy of the analysis results.
By using knowledge graphs in sentiment analysis in specific fields, texts in the field can be more accurately analyzed, the probability of error detection is reduced, and the accuracy and reliability of sentiment analysis can be improved.
Smart Images

Figure CN114912458B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of semantic analysis, and in particular, to a sentiment analysis method, device, and computer-readable medium. Background Art
[0002] Semantic analysis of text usually includes two categories: semantic analysis based on dictionaries and rules and semantic analysis based on word vectors.
[0003] Dictionary and rule-based semantic analysis relies on dictionaries and set rules. Generally, there are three types of dictionaries used for semantic analysis: positive, neutral, and negative. First, the text to be analyzed is segmented and cleaned; then, the multiple words obtained are matched with the words in the dictionary, and the sentiment polarity of the text is determined by counting the number of matches. During the whole process, rules can be set to improve accuracy.
[0004] The sentiment analysis based on word vectors converts text into vector matrices, so that various machine learning and deep learning methods can be applied during the analysis. The analysis process can also be combined with dictionaries. Similarly, first, the text to be analyzed will be segmented and cleaned; then, the segmented and cleaned text will be converted into a matrix according to the word vector conversion method. The quality of word vector conversion directly affects the accuracy of the subsequent machine learning classifier, so the choice of word vector conversion method is crucial. Commonly used word vector conversion methods include: term frequency-inverse document frequency (TF-IDF), bag-of-word, word to vector (word2vec), etc. Among them, TF-IDF is relatively easy to implement and the most widely used; the word2vec algorithm is relatively complex but the result is the best. No matter which of the above methods, it depends greatly on the text database. The better the quality of the text database, the better the result of word vector conversion. As for the selection of classifiers, machine learning classification algorithms are usually used, which have the advantages of short training time and high accuracy. In contrast, the deep learning algorithm has a long training time and high algorithm complexity, and is not as widely used as machine learning.
[0005] Whether it is a sentiment analysis based on a dictionary or a sentiment analysis based on a word vector, the selection of a text database is a crucial step. However, most of the text databases used by existing sentiment analysis algorithms are general-purpose datasets. For example, words such as "loophole", "attacks", and "severe" that often appear in some articles related to vulnerabilities are easily judged as negative. Therefore, using a general dataset in a specific scenario may lead to incorrect semantic analysis results. Summary of the invention
[0006] The embodiment of the present invention provides a sentiment analysis method, device and computer-readable medium for sentiment analysis of text in a specific field (e.g., network security). The knowledge graph used to represent the knowledge in the specific field is introduced into the sentiment analysis process, and a sentiment factor vector is assigned to each node and edge in the knowledge graph to represent the sentiment polarity of the word corresponding to the node or edge. This can effectively constrain the sentiment analysis process and avoid errors in the analysis results.
[0007] On the first aspect, a model generation method is provided, which can generate a model for sentiment analysis in a specific field. Among them, a plurality of first texts in a specific field are obtained; a group of first word vectors are generated for each first text obtained; a knowledge graph of the specific field is obtained, wherein each node and each edge in the knowledge graph has a sentiment factor vector, which is used to represent the sentiment polarity of the word represented by the node or edge in the specific field; a group of second word vectors are generated based on the knowledge graph, wherein each node and each edge of the knowledge graph corresponds to a second word vector; for each first text obtained, a group of third word vectors are generated based on a part of the knowledge graph included in the text in the knowledge graph, wherein each node and each edge in the part of the knowledge graph corresponds to a third word vector; a model is trained with each group of first training data, wherein a group of first training data takes a group of first word vectors and a group of third word vectors of a first text as input, and takes a space formed by each vector including the group of second word vectors and each sentiment factor vector corresponding to the group of second word vectors as output, so that the model is used to represent the mapping relationship between each word vector included in the text of the specific field and the group of second word vectors.
[0008] In a second aspect, a model generation device is provided, comprising a module for executing each step in the method provided in the first aspect.
[0009] In a third aspect, a sentiment analysis method is provided for sentiment analysis in a specific field. In which, a third text is obtained; a set of fifth word vectors of the third text is generated; the set of fifth word vectors is input into a model, wherein the model is used to represent the mapping relationship between each word vector included in the text of the specific field and a set of second word vectors, the set of second word vectors is generated based on the knowledge graph of the specific field, each node and each edge of the knowledge graph corresponds to a second word vector, and each has a sentiment factor vector, which is used to represent the sentiment polarity of the word represented by the node or edge in the specific field; the output of the set of fifth word vectors and the model is used as the input of a classifier to obtain the sentiment polarity of the third text, wherein the classifier is used to perform sentiment analysis on a text to obtain the sentiment polarity of the text.
[0010] In a fourth aspect, a sentiment analysis device is provided, comprising a module for executing each step in the method provided in the third aspect.
[0011] In a fifth aspect, a model generation device is provided, comprising: at least one memory configured to store computer-readable code; and at least one processor configured to call the computer-readable code to execute the steps provided in the first aspect.
[0012] In a sixth aspect, a sentiment analysis device is provided, comprising: at least one memory configured to store computer-readable code; and at least one processor configured to call the computer-readable code to execute the steps provided in the third aspect.
[0013] In a seventh aspect, a computer-readable medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the processor is caused to execute the steps provided in the first aspect or the third aspect.
[0014] By adopting the embodiment of the present invention, the knowledge graph of a specific field (such as network security) is introduced into the sentiment analysis process, and the sentiment polarity of the words corresponding to the nodes and edges is represented by setting sentiment factor vectors for the nodes and edges in the knowledge graph. In this way, in the articles and texts in these specific fields, those words that are usually considered to have negative sentiment polarity but are actually neutral in this specific field can obtain relatively accurate sentiment analysis results. For example, words such as vulnerability, attack, and serious should be considered to have neutral sentiment polarity in network security. By adding sentiment factor vectors to the edges and nodes of the knowledge graph, the sentiment characteristics of the words represented by these nodes or edges can be better refined, and the analyzed text can also be mapped more accurately, creating convenient conditions for the classification process in the subsequent sentiment analysis. By converting the text to be analyzed into a word vector and inputting it into the above-mentioned model, the knowledge of the specific field can be added to the word vector output by the model, and then further input into the classifier to obtain a more accurate classification result.
[0015] For any of the above aspects, optionally, each group of second training data can be used to train a classifier, wherein each group of second training data is used to train a classifier, wherein a group of second training data takes the group of first word vectors of a first text as input, takes the space formed by the vectors including the group of second word vectors and each of the sentiment factor vectors corresponding to the group of second word vectors as input, and takes the sentiment polarity of the first text as output, wherein the classifier is used to perform sentiment analysis on a text to obtain the sentiment polarity of the text. In this way, the classifier can learn the information of the specific field contained in the knowledge graph and determine the sentiment polarity according to the sentiment factor vector, so that the result of sentiment analysis is more accurate.
[0016] For any of the above aspects, optionally, the knowledge graph can be updated in the following manner: obtaining a second text; obtaining triple information of the second text; generating a set of fourth word vectors of the second text; inputting the set of fourth word vectors into the model; using the set of fourth word vectors and the output of the model as the input of the classifier to obtain the sentiment polarity of the second text; adding the triple information of the second text to the knowledge graph, wherein the pre-configured coefficient corresponding to the sentiment polarity of the second text is used as the sentiment factor vector of the node or edge in the knowledge graph corresponding to the triple information, wherein each sentiment polarity of the classifier is pre-configured with a coefficient, wherein the larger the coefficient, the more positive the sentiment polarity, and the smaller the coefficient, the more negative the sentiment polarity. Wherein, the triple information is identified from the text and added to the knowledge graph, and the sentiment polarity of the text is identified, and the sentiment factor vectors of the nodes and edges in the knowledge graph are updated accordingly. This multi-dimensional update can help the knowledge graph to be expanded and supplemented in many aspects, such as the coverage of the knowledge graph, the relationship between elements, the characteristics of the sentiment factor vector, and so on. On the other hand, new content can be added to the knowledge graph. The sentiment factor vectors of existing elements in the knowledge graph can also be updated, making the knowledge graph better used for subsequent sentiment analysis and the analysis results more accurate.
[0017] For any of the above aspects, optionally, the sentiment polarity of a piece of text (third text) can be determined in the following manner, wherein, a piece of third text is obtained; a set of fifth word vectors of the third text is generated; the set of fifth word vectors is input into a model, wherein the model is used to represent the mapping relationship between each word vector included in the text of the specific field and a set of second word vectors, the set of second word vectors is generated based on the knowledge graph of the specific field, each node and each edge of the knowledge graph corresponds to a second word vector, and each has a sentiment factor vector, which is used to represent the sentiment polarity of the word represented by the node or edge in the specific field; the output of the set of fifth word vectors and the model is used as the input of a classifier to obtain the sentiment polarity of the third text, wherein the classifier is used to perform sentiment analysis on a piece of text to obtain the sentiment polarity of the piece of text. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of the structure of a model generation device provided in an embodiment of the present invention.
[0019] Figure 2 It is a schematic diagram of a knowledge graph in a specific field in an embodiment of the present invention.
[0020] Figure 3 A schematic diagram of a knowledge graph with sentiment factor vectors added in an embodiment of the present invention.
[0021] Figure 4 The training process of the model in the embodiment of the present invention is shown.
[0022] Figure 5 The training process of the classifier in the embodiment of the present invention is shown.
[0023] Figure 6 The process of updating the knowledge graph in an embodiment of the present invention is shown.
[0024] Figure 7 A flow chart of a model generation method provided by an embodiment of the present invention is shown.
[0025] Figure 8 A schematic diagram of the structure of a sentiment analysis device provided in an embodiment of the present invention is shown.
[0026] Fig. 9 The process of performing sentiment analysis on text in an embodiment of the present invention is shown.
[0027] Fig.10 A flowchart of a sentiment analysis method provided by an embodiment of the present invention.
[0028] List of reference numerals:
[0029] 10: Model generation device 1001: Memory
[0030] 1002: processor 1003: communication module
[0031] 101: Model generation program 1011-1018: Program modules in the model generation program
[0032] 1011: Text acquisition module 1012: Word vector generation module
[0033] 1013: Knowledge graph acquisition module 1014: Model training module
[0034] 1015: Classifier training module 1016: Recognition module
[0035] 1017: Execution module 1018: Knowledge graph update module 20: Domain-specific knowledge graph 20': Domain-specific knowledge graph with added sentiment factor vector E1~E11: Elements in the knowledge graph (including edges and nodes)
[0036] is the sentiment factor vector of each element in the knowledge graph
[0037] 21: Part of the knowledge graph 31: First text 32: Second text 33: Third text
[0038] 41: First word vector 42: Second word vector 43: Third word vector 44: Fourth word vector
[0039] 45: Fifth word vector 51: Model 52: Classifier 60: Sentiment polarity 70: Triple information
[0040] 23: Updated Knowledge Graph
[0041] 700: Model generation method S701-S713: Method steps 80: Sentiment analysis device 8001: Memory
[0042] 8002: Processor 8003: Communication module
[0043] 801: Sentiment analysis program 8011~8013: Program modules in the sentiment analysis program
[0044] 8011: Text acquisition module 8012: Word vector generation module 8013: Execution module 1000: Sentiment analysis method S1001~S1004:Method steps DETAILED DESCRIPTION
[0045] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is only to enable those skilled in the art to better understand and implement the subject matter described herein, and is not a limitation of the scope of protection, applicability or examples set forth in the claims. The function and arrangement of the elements discussed can be changed without departing from the scope of protection of the embodiments of the present invention. Each example can omit, replace or add various processes or components as needed. For example, the described method can be performed in an order different from the described order, and each step can be added, omitted or combined. In addition, the features described relative to some examples can also be combined in other examples.
[0046] As used herein, the term "including" and its variations represent open terms, meaning "including but not limited to". The term "based on" means "based at least in part on". The terms "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other definitions may be included below, whether explicit or implicit. Unless the context clearly indicates otherwise, the definition of a term is consistent throughout the specification.
[0047] The following is a detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. It should be noted that these embodiments are only examples and should not be regarded as limiting the scope of protection of the present invention.
[0048] As mentioned above, most sentiment analysis tools and dictionaries are general, and may cause errors in the analysis results when applied to sentiment analysis in specific fields. For example, in the field of network security, texts describing network vulnerabilities usually include many negative words in the conventional sense. If they are not specially processed, it is easy to cause errors in the analysis results. In the embodiments of the present invention, by using knowledge graphs in specific fields, the texts in these specific fields can be better analyzed and the results are more accurate.
[0049] The model generation device 10 provided in the embodiment of the present invention can be implemented as a network of computer processors to execute the processing of the model generation method 700 in the embodiment of the present invention, which can realize sentiment analysis for a specific field (such as network security). Figure 1 The single computer, single-chip computer or processing chip shown includes at least one memory 1001, which includes a computer-readable medium, such as a random access memory (RAM). The model generation device 10 also includes at least one processor 1002 coupled to the at least one memory 1001. Computer executable instructions are stored in the at least one memory 1001, and when executed by the at least one processor 1002, the at least one processor 1002 can be caused to perform the steps described herein.
[0050] Figure 1 The at least one memory 1001 shown in FIG. 1 may include a model generation program 101, so that at least one processor 1002 executes the model generation method 700 described in the embodiment of the present invention. Figure 1 As shown, the model generation program 101 may include a text acquisition module 1011, a word vector generation module 1012, a knowledge graph acquisition module 1013, a model training module 1014, a classifier training module 1015, a recognition module 1016, an execution module 1017 and a knowledge graph update module 1018.
[0051] Below, refer to Figure 4 Describe the model training scheme of each module.
[0052] The text acquisition module 1011 is configured to acquire a plurality of first texts 31 in a specific field.
[0053] By giving keywords in a specific field, relevant web pages can be searched on the Internet according to the keywords. Further, tools bs4 or lxml can be used to analyze the obtained web pages and filter out irrelevant information (such as advertisements, etc.) to obtain cleaned titles and text bodies. Further, the cleaned text is segmented to obtain individual words. Among them, the tool jieba can be used to segment the text and filter out the spaced words. In this way, the first text 31 is obtained. The subsequent second text 32 and third text 33 can also use the above method when cleaning and segmenting.
[0054] The word vector generation module 1012 is configured to generate a set of first word vectors for each acquired first text 31.
[0055] The tool word2vector can be used to convert the first text 31 into a set of first word vectors 41. Each word in the text is converted into a word vector. The second text 31 and the third text 32 can also be converted from text to word vectors using the tool.
[0056] The knowledge graph acquisition module 1013 is configured to acquire a knowledge graph 20' of a specific field, wherein each node and each edge in the knowledge graph 20' has a sentiment factor vector for representing the sentiment polarity of the word represented by the node or edge in the specific field.
[0057] like Figure 2 As shown, the knowledge graph 20 is a knowledge graph of a specific field, in which each element (including nodes and edges) corresponds to a word in a text of a specific field. In the embodiment of the present invention, a sentiment factor vector is assigned to each element in the knowledge graph 20. Figure 3 As shown in the knowledge graph 20', each element E1 to E11 is assigned a sentiment factor vector
[0058] Among them, in the sentiment factor vector, the elements corresponding to a node and the node itself and other nodes can be set to non-zero, and the elements corresponding to the node and each edge can be set to zero; the elements corresponding to an edge and the edge itself are set to non-zero, and the elements corresponding to the edge and the node and the edge and other edges are set to zero. The value range of non-zero elements is (0,1), where the larger the value, the more positive the sentiment polarity; the smaller the value, the more negative the sentiment polarity, and 0.5 means that the sentiment polarity is neutral.
[0059] Specifically, for a node, the element value corresponding to the node and the node itself can represent the sentiment polarity of the word corresponding to the node in a general (non-domain-specific) text. For example, for the word "vulnerability", it can be set to 0.1, indicating that the sentiment polarity is "negative"; corresponding to the element between the node and other nodes, if there is no edge between the nodes, it can be set to 0. If there is an edge between the nodes, it can be set to the sentiment polarity value that should be assigned when the words corresponding to the two nodes and the edge appear in the text at the same time according to the knowledge of the specific domain.
[0060] Specifically, for an edge, the element value corresponding to the edge and the edge itself can be set to the value of the sentiment polarity that should be assigned when the edge and the words corresponding to the two nodes connected by the edge appear simultaneously in the text according to the knowledge of the specific field; the element corresponding to the edge and other edges can be set to 0, and the element corresponding to the edge and the node can be set to 0.
[0061] The word vector generation module 1012 is further configured to generate a set of second word vectors based on the knowledge graph 20'. Each node and each edge of the knowledge graph 20' corresponds to a second word vector 42. Here, the word2vector tool can also be used to convert the words represented by each element in the knowledge graph 20' into word vectors.
[0062] The word vector generation module 1012 is further configured to generate a set of third word vectors for each acquired first text 31 based on the part of the knowledge graph 21 included in the first text 31 in the knowledge graph 20'. Each node and each edge in the partial knowledge graph 21 corresponds to a third word vector 43.
[0063] The model training module 1014 is configured to train a model 51 with each set of first training data. The model 51 may be a Text-CNN (text-convolutional neural network) model. Figure 4 As shown, a set of first training data can be a set of first word vectors of a first text 31 and a set of third word vectors As input, to include the set of second word vectors Each sentiment factor vector corresponding to the group of second word vectors 42 The space formed by the vectors As an output, the model 51 is used to represent each word vector included in the text of the specific field and the set of second word vectors The mapping relationship between them.
[0064] The modules mentioned above cooperate with each other to train the model 51. The information of the knowledge graph 21' in the specific field is integrated into the training process (wherein, the words in the specific field are introduced through nodes and edges, and the sentiment polarity represented by some words in the specific field and multiple words when they appear at the same time is introduced through sentiment factor vectors), so that the trained model can reflect the understanding of the sentiment polarity of the text in the specific field. The mapping relationship between the text in the specific field and the knowledge and sentiment polarity in the specific field is obtained. Based on this mapping relationship, when a new text is to be subjected to sentiment analysis, the text is first input into the model to obtain the knowledge of the sentiment polarity in the specific field, and then input into the classifier to obtain accurate sentiment analysis results.
[0065] Furthermore, if Figure 5 As shown, after obtaining the mapping relationship between the text in a specific field and the knowledge and sentiment polarity in a specific field, the classifier 52 can be further trained. The classifier 52 is used to perform sentiment analysis on a section of text to obtain the sentiment polarity of the section of text. Specifically, the model generation device 10 may also include a classifier training module 1015, which is configured to: train the classifier 52 with each group of second training data, wherein a group of second training data uses a group of first word vectors 41 of a section of first text 31 as input, uses a space formed by vectors including the aforementioned group of second word vectors 42 and each sentiment factor vector corresponding to the group of second word vectors 42 as input, and uses the sentiment polarity of the section of first text 31 as output. Among them, Can Concatenate them together to form the input of classifier 52 The output of the classifier 52 is the sentiment polarity of the first text 31 .
[0066] Furthermore, if Figure 6 As shown, the knowledge graph can also be updated to expand the information in the knowledge graph about the sentiment polarity of text in a specific field.
[0067] A coefficient a, a∈(0,1) may be preconfigured for each sentiment polarity of the classifier 52. The larger the coefficient, the more positive the sentiment polarity, and the smaller the coefficient, the more negative the sentiment polarity.
[0068] During the updating process, the text acquisition module 1011 first acquires a second text 32; the model generation device 10 may also include an identification module 1016, which is configured to acquire triple information 70 of the second text 32. The triple information may represent the semantic relationship between things in a binary relationship model, that is, using a set of triple information to describe things and relationships, representing the relationship between entities or the attribute value of a certain attribute of an entity. For example, in the text "the vulnerability rating of vulnerability A is a high-risk type", the triple information may include: vulnerability A, vulnerability rating, high-risk.
[0069] Furthermore, the word vector generation module 1012 generates a set of fourth word vectors 44 for the second text 32; the execution module 1017 inputs the set of fourth word vectors into the model 51, and uses the set of fourth word vectors and the output of the model 51 as input to the classifier 52 to obtain the sentiment polarity 60 of the second text 32.
[0070] The model generation device 10 may also include a knowledge graph updating module 1018, which is configured to add the triple information 70 of the second text 32 to the knowledge graph 20' to generate a new knowledge graph 23, wherein the sentiment factor vector of the node or edge in the knowledge graph corresponding to the triple information 70 is generated based on the pre-configured coefficient corresponding to the sentiment polarity 60 of the second text 32.
[0071] Taking "The vulnerability rating of vulnerability A is a high-risk type" as an example of the second text 32, the triple information identified by the identification module 1016 includes: A vulnerability, vulnerability rating and high risk. Then the knowledge graph update module 1018 can add a head node: vulnerability, a tail node: high risk and a connection between the two nodes: vulnerability rating in the knowledge graph 20'. An optional solution is to assign the same sentiment polarity output by the classifier 52 to the head node, the tail node and the connection, and generate sentiment factor vectors for the three respectively (the values of each element in the sentiment factor vector can refer to the previous description of the sentiment factor vector). Another optional solution is that experts in a specific field assign values to the head node, the tail node and the connection respectively, and then generate sentiment factor vectors accordingly. In the case where the head node and the tail node already exist in the knowledge graph 20', the sentiment factor vectors of the head node and the tail node can use the existing ones, and the new connection between the two nodes is assigned the sentiment polarity output by the classifier 52, and the sentiment factor vector is generated accordingly.
[0072] In addition, the above modules can also be regarded as various functional modules implemented by hardware, which are used to implement various functions involved in the model generation device 10 when executing the model generation method 700, such as pre-burning the control logic of each process involved in the method into a field programmable gate array (Field-Programmable Gate Array, FPGA) chip or a complex programmable logic device (Complex Programmable Logic Device, CPLD), and these chips or devices execute the functions of the above modules. The specific implementation method can be determined according to engineering practice.
[0073] In addition, the model generation device 10 may also include a communication module 1003, which is used for communication between the model generation device 10 and other devices, such as for obtaining text, knowledge graphs, etc.
[0074] It should be mentioned that embodiments of the present invention may include devices having different Figure 1 The above architecture is merely exemplary and is used to explain the model generation method 700 provided in an embodiment of the present invention.
[0075] Combine the following Figure 7 The model generation method 700 provided by the embodiment of the present invention is described. Figure 7 As shown, method 700 may include the following steps:
[0076] - S701: Get the first text 31 of a plurality of segments of a specific field;
[0077] - S702: Generate a set of first word vectors 41 for each acquired first text 31;
[0078] - S703: Obtain a knowledge graph 20' in a specific field, wherein each node and each edge in the knowledge graph 20' has a sentiment factor vector, which is used to represent the sentiment polarity of the word represented by the node or edge in the specific field;
[0079] - S704: Generate a set of second word vectors 42 based on the knowledge graph 20', wherein each node and each edge of the knowledge graph 20' corresponds to a second word vector 42;
[0080] - S705: for each acquired first text 31, a set of third word vectors 43 is generated based on the partial knowledge graph 21 in the knowledge graph 20' included in the first text 31, wherein each node and each edge in the partial knowledge graph 21 corresponds to a third word vector 43;
[0081] -S706: Train a model 51 with each group of first training data, wherein a group of first training data takes a group of first word vectors 41 and a group of third word vectors 43 of a first text 31 as input, and takes a space formed by each vector including a group of second word vectors 42 and each sentiment factor vector corresponding to the group of second word vectors 42 as output, so that the model 51 is used to represent the mapping relationship between each word vector included in the text of a specific field and the group of second word vectors 42.
[0082] Among them, through steps S701 to S702, the model 51 is trained to obtain the mapping relationship between the text in a specific field and the knowledge and sentiment polarity in the specific field.
[0083] Furthermore, the method 700 may further include:
[0084] -S707: Train a classifier 52 with each group of second training data, wherein a group of second training data takes a group of first word vectors 41 of a section of first text 31 as input, takes a space formed by vectors including a group of second word vectors 42 and each sentiment factor vector corresponding to the group of second word vectors 42 as input, and takes the sentiment polarity of the section of first text 31 as output, wherein the classifier 52 is used to perform sentiment analysis on a section of text to obtain the sentiment polarity of the section of text.
[0085] The training of the classifier 52 is implemented through step S707.
[0086] Furthermore, the method 700 may further include:
[0087] - S708: Get a second text 32;
[0088] - S709: Obtain triple information 70 of the second text 32;
[0089] - S710: Generate a set of fourth word vectors 44 for the second text 32;
[0090] - S711: input the fourth word vector 44 of the group into the model 51;
[0091] - S712: Use the fourth word vector 44 and the output of the model as the input of the classifier 52 to obtain the sentiment polarity 60 of the second text 32;
[0092] -S713: Add the triple information 70 of the second text 32 to the knowledge graph 20', wherein the pre-configured coefficient corresponding to the sentiment polarity 60 of the second text 32 is used as the sentiment factor vector of the node or edge in the knowledge graph corresponding to the triple information 70, wherein each sentiment polarity of the classifier 52 is pre-configured with a coefficient, wherein the larger the coefficient, the more positive the sentiment polarity, and the smaller the coefficient, the more negative the sentiment polarity.
[0093] Among them, the knowledge graph is updated through steps S708 to S713.
[0094] The sentiment analysis device 80 provided in the embodiment of the present invention may be implemented as a network of computer processors to execute the sentiment analysis method 1000 in the embodiment of the present invention. Figure 8The single computer, single-chip computer or processing chip shown includes at least one memory 8001, which includes a computer-readable medium, such as a random access memory (RAM). The sentiment analysis device 80 also includes at least one processor 8002 coupled to the at least one memory 8001. Computer executable instructions are stored in the at least one memory 8001, and when executed by the at least one processor 8002, the at least one processor 8002 can perform the steps described herein.
[0095] Figure 8 The at least one memory 8001 shown in the figure may include a sentiment analysis program 801, so that at least one processor 8002 executes the sentiment analysis method 1000 described in the embodiment of the present invention. The sentiment analysis program 801 may include:
[0096] -Text acquisition module 8011, such as Fig. 9 As shown, the text acquisition module 8011 is configured to acquire a third text 33;
[0097] - a word vector generation module 8012, configured to generate a set of fifth word vectors 45 for the third text 33;
[0098] -Execution module 8013 is configured to input the group of fifth word vectors 45 into a model 51, wherein the model 51 is used to represent the mapping relationship between each word vector included in the text of the specific field and a group of second word vectors 42, the group of second word vectors 42 is generated based on the knowledge graph 20' of the specific field, each node and each edge of the knowledge graph 20' corresponds to a second word vector 42, and each has a sentiment factor vector for representing the sentiment polarity of the word represented by the node or edge in the specific field; and the output of the group of fifth word vectors 45 and the model 51 is used as the input of a classifier 52 to obtain the sentiment polarity 60 of the third text 33, wherein the classifier 52 is used to perform sentiment analysis on a piece of text to obtain the sentiment polarity of the piece of text.
[0099] In addition, the above modules can also be regarded as various functional modules implemented by hardware, which are used to realize the various functions involved in the sentiment analysis device 80 when executing the sentiment analysis method 1000, such as pre-burning the control logic of each process involved in the method into an FPGA chip or CPLD, and these chips or devices execute the functions of the above modules. The specific implementation method can be determined according to engineering practice.
[0100] In addition, the sentiment analysis device 80 may also include a communication module 8003, which is used for communication between the sentiment analysis device 80 and other devices, such as obtaining the third text 33, etc.
[0101] It should be mentioned that embodiments of the present invention may include devices having different Figure 8 The above architecture is merely exemplary and is used to explain the sentiment analysis method 1000 provided in an embodiment of the present invention.
[0102] Combine the following Fig.10 The sentiment analysis method 1000 provided by the embodiment of the present invention is described. Fig.10 As shown, method 1000 may include the following steps:
[0103] - S1001: Get a third text 33;
[0104] - S1002: Generate a set of fifth word vectors 45 for the third text 33;
[0105] - S1003: input the group of fifth word vectors 45 into a model 51, wherein the model 51 is used to represent the mapping relationship between each word vector included in the text of the specific field and a group of second word vectors 42, the group of second word vectors 42 is generated based on the knowledge graph 20' of the specific field, each node and each edge of the knowledge graph 20' corresponds to a second word vector 42, and each has a sentiment factor vector, which is used to represent the sentiment polarity of the word represented by the node or edge in the specific field;
[0106] - S1004: The output of the fifth word vector 45 and the model 51 is used as the input of a classifier 52 to obtain the sentiment polarity 60 of the third text 33, wherein the classifier 52 is used to perform sentiment analysis on a piece of text to obtain the sentiment polarity of the piece of text.
[0107] In addition, an embodiment of the present invention also provides a computer-readable medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the processor executes the aforementioned sentiment analysis method or model generation method. Embodiments of computer-readable media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the computer-readable instructions can be downloaded from a server computer or a cloud via a communication network.
[0108] In summary, the embodiments of the present invention provide a model generation method and device, a sentiment analysis method and device, and a computer-readable medium. In the sentiment analysis of texts in a specific field, a knowledge graph in a specific field is introduced, and sentiment polarity is assigned to each element by adding a sentiment factor vector in the knowledge graph, so that the sentiment analysis result of the text will be more accurate. Taking the field of network security as an example, some general concepts in articles about security vulnerabilities, such as "leak", "attack", "serious" and other words will be analyzed more accurately to reduce the probability of false detection. By assigning sentiment factor vectors to the elements in the knowledge graph, the sentiment features in these elements can be extracted, and then the sentiment tendency of the detected text can be mapped, which is convenient for the subsequent classification of sentiment polarity.
[0109] In addition, in the process of updating the knowledge graph, triple information in the new text is identified and added to the knowledge graph; and the sentiment polarity of the new text is identified, and the elements and sentiment factor vectors in the knowledge graph are updated based on the triple information and sentiment polarity. Multi-dimensional updates can help expand the knowledge graph and supplement it in many aspects. On the one hand, new content can be added to the knowledge graph; on the other hand, the sentiment factor vectors of existing elements in the knowledge graph can also be updated. This enables the knowledge graph to be better used for subsequent sentiment analysis. After such an iterative cycle, a larger and more accurate knowledge graph can be obtained, making sentiment analysis more accurate.
[0110] It should be noted that not all steps and modules in the above-mentioned processes and system structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be implemented together.
Claims
1. A model generation method (700), include: - obtaining (S701) a plurality of first texts (31) of a specific domain; - generating (S702) a set of first word vectors (41) for each acquired first text segment (31); - obtaining (S703) the knowledge graph (20') of the specific field, wherein each node and each edge in the knowledge graph (20') has a sentiment factor vector, which is used to represent the sentiment polarity of the word represented by the node or edge in the specific field; - generating (S704) a set of second word vectors (42) based on the knowledge graph (20'), wherein each node and each edge of the knowledge graph (20') corresponds to a second word vector (42); - for each acquired first text (31), generating (S705) a set of third word vectors (43) based on the partial knowledge graph (21) included in the first text (31) in the knowledge graph (20'), wherein each node and each edge in the partial knowledge graph (21) corresponds to a third word vector (43); - A model (51) is trained (S706) with each group of first training data, wherein a group of first training data has the group of first word vectors (41) and the group of third word vectors (43) of a first text (31) as input, and has a space formed by each vector including the group of second word vectors (42) and each sentiment factor vector corresponding to the group of second word vectors (42) as output, so that the model (51) is used to represent the mapping relationship between each word vector included in the text of the specific domain and the group of second word vectors (42).
2. The method according to claim 1, further comprising: include: - Training (S707) a classifier (52) with each group of second training data, wherein a group of second training data takes the group of first word vectors (41) of a first text (31) as input, takes a space formed by vectors including the group of second word vectors (42) and each of the sentiment factor vectors corresponding to the group of second word vectors (42) as input, and takes the sentiment polarity of the first text (31) as output, wherein the classifier (52) is used to perform sentiment analysis on a text to obtain the sentiment polarity of the text.
3. The method according to claim 2, further comprising: include: - obtaining (S708) a second text (32); - obtaining (S709) triple information (70) of the second text (32); - generating (S710) a set of fourth word vectors (44) for the second text (32); - inputting (S711) the set of fourth word vectors (44) into the model (51); - using (S712) the set of fourth word vectors (44) and the output of the model as inputs to the classifier (52) to obtain the sentiment polarity (60) of the second text (32); - Add (S713) the triple information (70) of the second text (32) to the knowledge graph (20'), wherein a pre-configured coefficient corresponding to the sentiment polarity (60) of the second text (32) is used as the sentiment factor vector of the node or edge in the knowledge graph corresponding to the triple information (70), wherein each sentiment polarity of the classifier (52) is pre-configured with a coefficient, wherein the larger the coefficient, the more positive the sentiment polarity, and the smaller the coefficient, the more negative the sentiment polarity.
4. A model generation program (101), It is characterized in that include: - a text acquisition module (1011), configured to acquire a plurality of first texts (31) of a specific domain; - a word vector generation module (1012), configured to generate a set of first word vectors (41) for each acquired first text segment (31); - a knowledge graph acquisition module (1013), configured to acquire the knowledge graph (20') of the specific field, wherein each node and each edge in the knowledge graph (20') has a sentiment factor vector, which is used to represent the sentiment polarity of the word represented by the node or edge in the specific field; - the word vector generation module (1012) is further configured to generate a set of second word vectors (42) based on the knowledge graph (20'), wherein each node and each edge of the knowledge graph corresponds to a second word vector (42); and for each acquired first text (31), generate a set of third word vectors (43) based on a portion of the knowledge graph (21) included in the first text (31) in the knowledge graph (20'), wherein each node and each edge in the portion of the knowledge graph (21) corresponds to a third word vector (43); - A model training module (1014), configured to train a model (51) with each group of first training data, wherein a group of first training data has the group of first word vectors (41) and the group of third word vectors (43) of a first text (31) as input, and has a space formed by each vector including the group of second word vectors (42) and each sentiment factor vector corresponding to the group of second word vectors (42) as output, so that the model (51) is used to represent the mapping relationship between each word vector included in the text of the specific domain and the group of second word vectors (42).
5. The program according to claim 4, further comprising: include: The classifier training module (1015) is configured to: - A classifier (52) is trained with each group of second training data, wherein a group of second training data takes the group of first word vectors (41) of a first text (31) as input, takes a space formed by vectors including the group of second word vectors (42) and each of the sentiment factor vectors corresponding to the group of second word vectors (42) as input, and takes the sentiment polarity of the first text (31) as output, wherein the classifier (52) is used to perform sentiment analysis on a text to obtain the sentiment polarity of the text.
6. The program according to claim 5, It is characterized in that - the text acquisition module (1011) is further configured to acquire a second text (32); - the program further comprises a recognition module (1016) configured to obtain triple information (70) of the second text (32); - the word vector generation module (1012) is further configured to generate a set of fourth word vectors (44) for the second text (32); - an execution module (1017), configured to input the set of fourth word vectors into the model (51); and use the set of fourth word vectors and the output of the model (51) as inputs of the classifier (52) to obtain the sentiment polarity of the second text (32); -The program also includes a knowledge graph updating module (1018), which is configured to add the triple information (70) of the second text (32) to the knowledge graph (20'), wherein a pre-configured coefficient corresponding to the sentiment polarity (60) of the second text (32) is used as the sentiment factor vector of the node or edge in the knowledge graph corresponding to the triple information (70), wherein each sentiment polarity of the classifier (52) is pre-configured with a coefficient, wherein the larger the coefficient, the more positive the sentiment polarity, and the smaller the coefficient, the more negative the sentiment polarity.
7. A model generation device (10), It is characterized in that include: At least one memory (1001) configured to store computer readable code; At least one processor (1002) is configured to call the computer-readable code to execute the method according to any one of claims 1 to 3.
8. A computer readable medium, It is characterized in that The computer-readable medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the processor executes the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Sentiment classification method for Chinese news title in financial field
CN110297870A