Text mining method, text mining program, and text mining device

The method enhances text mining by extracting and indexing emotional words in documents, allowing for accurate and efficient comparison of emotional trends by specifying emotional index ranges, addressing the limitations of conventional methods.

JP7818413B2Active Publication Date: 2026-02-20SCREEN HOLDINGS CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022015493
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-03
Publication Date
2026-02-20
Estimated Expiration
2042-02-03

AI Technical Summary

Technical Problem

Conventional text mining methods struggle with accurately handling words with low emotional intensity and require excessive calculations when comparing emotional trends across multiple documents, often underestimating frequently occurring emotional words.

Method used

A text mining method that extracts feature words from documents, assigns emotional indices based on a predefined dictionary, and allows users to specify the range of emotional indices for display, enabling accurate comparison of emotional trends with reduced computational effort.

Benefits of technology

Enables precise analysis of emotional trends across documents by focusing on characteristic words and adjusting emotional index ranges, avoiding underestimation and reducing computational load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007818413000001
    Figure 0007818413000001
  • Figure 0007818413000002
    Figure 0007818413000002
  • Figure 0007818413000003
    Figure 0007818413000003
Patent Text Reader

Abstract

To make it possible to compare, with a less calculation amount, emotion tendency among a plurality of documents based on proper evaluation of emotional words in the documents.SOLUTION: A text mining method according to an embodiment of the present invention comprises the steps of: receiving instructions designating, as a target document, a plurality of documents whose tendency of emotional polarity to be compared with one another and instructions designating a scope of an emotion index indicating intensity of the emotional polarity; extracting a feature word from the plurality of documents within a designated range based on the instructions; giving the emotion index to the feature word registered in a predetermined emotional word dictionary as an emotional word given with the emotion index within the predetermined range; and then displaying the extracted feature word and the given emotion index so that they can be compared with each other. In this display step, for example, the feature word given with the emotion index is allocated with a background color corresponding to the emotion index.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to text mining, and more particularly to a text mining method, a text mining program, and a text mining device for comparing trends in sentiment polarity of multiple documents. [Background technology]

[0002] In recent years, text mining, which analyzes free-form text data and extracts useful information from the analysis results, has been attracting attention. In the field of text mining, a technique is known that determines the sentiment polarity (hereinafter referred to as "sentimental tendency") of a document's text data, i.e., whether the sentiment is positive or negative toward an object, person, or content related to the document.

[0003] For example, a method is known in which an emotional word dictionary in which correspondences between words and the emotions they express (such as emotional polarity indicating whether they are positive or negative) are registered in advance is used to compare the number of words with positive emotional polarity and the number of words with negative emotional polarity contained in a document, and the emotional polarity of the document (whether the document is positive, negative, or neutral) is determined based on the comparison result (see paragraph

[0009] of Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-204226 [Patent Document 2] Japanese Patent Application Laid-Open No. 2013-246636 [Patent Document 3] Special Publication No. 2016-530651 Summary of the Invention [Problem to be solved by the invention]

[0005] The conventional method checks whether a word contained in a document is registered in an emotional dictionary, and if so, classifies the word into two categories: positive or negative, according to the dictionary. For this reason, there was no established method for appropriately handling words with low emotional intensity (emotional polarity), i.e., words that are close to neutral. Furthermore, the emotional polarity of such words needed to be adjusted depending on the content of the target document, but a simple method for doing so was not known.

[0006] Furthermore, in the above-mentioned conventional method, when comparing the emotional tendencies of multiple documents, all of the emotional words that are registered in the emotional word dictionary among the words contained in each document are tallied and the tallied results are compared. This increases the amount of calculation required for the comparison, and may result in the underestimation of emotional words that appear more frequently in one document than in other documents.

[0007] Therefore, there is a need to provide a data mining method, text mining device, etc. that can compare emotional trends between multiple documents based on appropriate evaluation of emotional words in documents with a small amount of calculation. [Means for solving the problem]

[0008] A first aspect of the present invention is a method for comparing trends in sentiment polarity among multiple documents. executed by a computer 1. A text mining method, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the characteristic words extracted by the characteristic word extraction step together with the emotion index assigned by the emotion index acquisition step for the plurality of documents designated as the target documents. 、 The instruction input step further includes a step of receiving an instruction specifying a range of feature words to be extracted from the target document; In the feature word extraction step, feature words within the range specified in the instruction input step are extracted. .

[0010] The present invention No. 2 The situation is 1. A computer-implemented text mining method for comparing sentiment polarity trends among a plurality of documents, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the feature words extracted by the feature word extraction step together with the emotion index assigned by the emotion index acquisition step for the plurality of documents designated as the target documents, The instruction input step further includes a step of receiving an instruction to specify a range of an emotion index that is an index indicating the strength of an emotion polarity, In the emotion quotient acquisition step, an emotion quotient is assigned to a feature word that is registered in the emotion word dictionary as a word that has been assigned an emotion quotient within the range specified in the instruction input step, from among the feature words extracted in the feature word extraction step.

[0011] The present invention Third This aspect of the present invention No. 2 In this situation, The instruction input step further includes a step of receiving an instruction to specify a change in the range of the emotional quotient when the extracted feature word is displayed together with the assigned emotional quotient by the display step.

[0012] The present invention Fourth The situation is 1. A computer-implemented text mining method for comparing sentiment polarity trends among a plurality of documents, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the feature words extracted by the feature word extraction step together with the emotion quotient assigned by the emotion quotient acquisition step for the plurality of documents designated as the target documents; a document sentiment index calculation step for calculating, for each of the plurality of documents designated as the target documents, a sentiment index of the document as a document sentiment index based on the feature words assigned sentiment indexes in the sentiment index acquisition step among the feature words extracted from the document in the feature word extraction step; and In the display step, a display is performed showing the document sentiment index calculated in the document sentiment index calculation step.

[0013] The present invention No. 5 The present invention provides a text mining program for comparing trends in sentiment polarity among a plurality of documents, an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the characteristic words extracted by the characteristic word extraction step together with the emotion index assigned by the emotion index acquisition step for the plurality of documents designated as the target documents by a CPU using a memory in the computer. 、 The instruction input step further includes a step of receiving an instruction specifying a range of feature words to be extracted from the target document; In the feature word extraction step, feature words within the range specified in the instruction input step are extracted. .

[0014] The present invention No. 6 The present invention provides a text mining device for comparing trends in sentiment polarity among a plurality of documents, the device comprising: an instruction input unit that receives an instruction to specify, as target documents, a plurality of documents whose semantic polarity trends are to be compared; a feature word extraction unit that extracts feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional quotient acquisition unit that assigns an emotional quotient to a feature word that is registered in a predetermined emotional word dictionary among the feature words extracted by the feature word extraction unit, the emotional quotient being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display unit that displays the characteristic words extracted by the characteristic word extraction unit together with the emotion quotient assigned by the emotion quotient acquisition unit for the plurality of documents designated as the target documents; Equipped with 、 the instruction input unit further receives an instruction specifying a range of feature words to be extracted from the target document; The characteristic word extraction unit extracts characteristic words within a range designated by the instruction input unit. .

[0015] Other aspects of the present invention will be apparent from the above aspects of the present invention and the following description of the embodiments and their modifications, and therefore will not be described here. [Effects of the Invention]

[0016] Item 1 above, No. 5 or No. 6 According to this aspect, feature words are extracted from each of a plurality of documents designated as target documents, and the extracted feature words, which are target feature words and are registered as emotive words in an emotive word dictionary, are assigned the emotional index assigned to that feature word in the emotive word dictionary. In this way, for the plurality of documents, the target feature words and the emotional indices assigned to the emotive words contained therein are displayed as the results of the emotional trend analysis for the plurality of documents. With this display, even if the plurality of documents whose emotional trends are to be compared contain feature words with weak emotive polarities, by looking at the extracted feature words and the emotional indices assigned to them, it is possible to accurately grasp the emotional trends between the plurality of documents. Furthermore, according to the first, fifth, or sixth aspect, it is possible to specify the range of characteristic words to be extracted from each of the multiple target documents, and by extracting only the more characteristic words as target characteristic words, it is possible to compare the emotional tendencies that reflect the characteristics of each of the multiple documents between the multiple documents with a smaller amount of calculation than in the past.It is also possible to avoid the problem of underestimating characteristic emotional words that appear more frequently in one of the multiple documents than in other documents.

[0018] the above No. 2According to this aspect, by specifying the range of emotional indices to be assigned to target feature words, which are feature words extracted from each of multiple target documents, it is possible to accurately compare the emotional tendencies of multiple documents that contain feature words with weak emotional polarity.

[0019] the above Third According to this aspect, when the feature words extracted as described above are displayed together with the emotional indices assigned as described above as the results of the emotional trend analysis of a plurality of documents as target documents, if an instruction to change the range of the emotional indices is received, feature words are extracted as target feature words for each of the plurality of documents based on the changed range of the emotional indices, and emotional indices are assigned to those of the target feature words that are registered as emotional words in the emotional word dictionary, and then the target feature words and the emotional indices assigned to the emotional words contained therein are displayed as the results of the emotional trend analysis of the plurality of documents.As a result, after the results of the emotional trend analysis of the plurality of documents are displayed, the user can more accurately compare the emotional trends of the plurality of documents by adjusting the specified range of the emotional indices while viewing the display.

[0020] the above Fourth According to this aspect, for each of a plurality of target documents, a document sentiment index is calculated based on the target feature words to which sentiment indexes have been assigned. In addition to comparing the sentiment indexes assigned to the feature words between the plurality of documents, it is also possible to compare the document sentiment indexes between the plurality of documents. This makes it possible to more accurately and easily compare the emotional tendencies of the plurality of documents.

[0021] The effects of other aspects of the present invention are clear from the explanation of the effects of the above aspects of the present invention and the effects of the following embodiment and its modifications, and therefore will not be explained here. [Brief explanation of the drawings]

[0022] [Figure 1]1 is a block diagram showing a configuration of a text mining device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the configuration of a computer that operates as the text mining device according to the embodiment. [Figure 3] FIG. 10 is a diagram for explaining an emotional word dictionary with an emotional index. [Figure 4] 10 is a flowchart showing the procedure of an emotional tendency analysis process executed in order for a computer to operate as the text mining device according to the embodiment. [Figure 5] FIG. 2 is a diagram showing an operation screen of the text mining device according to the embodiment. [Figure 6] FIG. 10 is a diagram for explaining extraction of feature words in the embodiment. [Figure 7] FIG. 10 is a diagram showing an example of a display showing the emotional tendencies of characteristic words extracted for each of the target documents in the embodiment. [Figure 8] FIG. 10 is a diagram showing an example of a display showing the results of emotional tendency analysis in the text mining device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0023] A text mining device according to one embodiment of the present invention will be described below with reference to the drawings. This text mining device is a device for implementing a text mining method for comparing emotional trends (emotional polarity trends) between multiple documents, and is realized by a computer executing a text mining program described below. In the following, "emotional polarity" refers to information indicating whether a document contains positive or negative opinions.

[0024] <1. Functional configuration of the text mining device> 1 is a block diagram showing the functional configuration of a text mining device 10 according to this embodiment. This text mining device 10 includes a GUI unit 11 that functions as an instruction input unit and a display unit, a text data storage unit 12, a feature word extraction unit 13, a dictionary storage unit 14 that stores an emotional word dictionary with an emotional index, a feature word emotional quotient acquisition unit 15, a document emotional quotient calculation unit 16, and a display data processing unit 17. Note that this text mining device 10 may not include one or both of the text data storage unit 12 and the dictionary storage unit 14, and may instead be configured to use one or both of the text data and the emotional word dictionary with an emotional index stored in an external storage unit via a network.

[0025] In this embodiment, the text data of a large number of documents, including target documents consisting of multiple documents whose emotional tendencies are to be compared, is stored in advance in the text data storage unit 12. When performing emotional tendency analysis processing, such as specifying target documents, the GUI unit 11 accepts user instructions (such as specifying target documents). Based on these instructions, the feature word extraction unit 13 first reads the text data of the multiple documents specified as target documents from the text data storage unit 12 and extracts feature words contained in each of the multiple documents. The feature word sentiment index acquisition unit 15 assigns the sentiment index assigned to the feature word in the sentiment word dictionary to the extracted target feature words, which are feature words registered as emotional words in the sentiment word dictionary with sentiment index in the dictionary storage unit 14. Note that this sentiment index is a numerical value indicating the strength of the sentiment polarity and will be referred to as a "word sentiment index" when distinguished from the document sentiment index described below. The document sentiment index calculation unit 16 uses the feature words to which sentiment indexes have been assigned in this way to calculate a sentiment index (document sentiment index) for each of the multiple documents using the formula described below. In this way, for each of the multiple documents designated as target documents, the target feature words and the sentiment indices (word sentiment indices) and document sentiment indices assigned to the sentiment words contained therein are obtained. The display data processing unit 17 generates display data for displaying these target feature words, word sentiment indices, and document sentiment indices so that they can be compared across the multiple documents. The GUI unit 11, as a display unit, displays a display for comparing the sentiment trends of the multiple documents based on this display data. This shows the results of the sentiment trend analysis of the target documents. By viewing this display, the user can grasp the differences in the sentiment trends among the multiple target documents, and, if necessary, can narrow the range of sentiment indices to be assigned to the feature words using the GUI unit 11, as an instruction input unit, and perform the above-mentioned sentiment trend analysis again.

[0026] <2. Hardware configuration of text mining device> FIG. 2 is a block diagram showing the configuration of a computer 20 that operates as the text mining device 10 in this embodiment using a text mining program (described later), i.e., the hardware configuration of the text mining device 10 according to this embodiment. The computer 20 shown in FIG. 2 includes a CPU 21, a main memory 22, an auxiliary storage device 23, an input operation unit 24, a display device 25, a communication interface device 26, and a recording medium reader 27. The main memory 22 may be, for example, a dynamic random access memory (DRAM). The auxiliary storage device 23 may be, for example, a hard disk or a solid-state drive. The input operation unit 24 includes, for example, a keyboard 28 and a mouse 29. The display device 25 may be, for example, a liquid crystal display (LCD). The communication interface device 26 is an interface circuit for wired or wireless communication. The recording medium reader 27 is an interface circuit for a recording medium 30 that stores a program or the like. The recording medium 30 may be, for example, a non-transitory recording medium such as a CD-ROM, a DVD-ROM, or a USB memory.

[0027] In the computer 20 configured as described above, the auxiliary storage device 23 stores, in addition to the text mining program 31 according to this embodiment, text data 32 of target documents and an emotional word dictionary 34, which is an emotional word dictionary with an emotional index, thereby realizing the text data storage unit 12 and the dictionary storage unit 14. The text mining program 31 and text data 32 may be received from a server or another computer using the communication interface device 26, for example, or may be read from the recording medium 30 using the recording medium reader 27. The emotional word dictionary 34 may also be stored in the server or another computer; in this case, the computer 20 operating as the text mining device 10 will use the emotional word dictionary 34 via the network and the communication interface device 26.

[0028] When the text mining program 31 is executed on the computer 20, the text mining program 31 and text data 32 are loaded into the main memory 22. The CPU 21 uses the main memory 22 as a working memory and executes the text mining program 31 stored in the main memory 22 to perform emotional trend analysis processing on target documents. In this emotional trend analysis processing, for each of multiple documents specified as target documents, extraction of characteristic words, acquisition of emotional indices of the characteristic words, calculation of a document emotional indices, etc. are performed (details will be described later). When the CPU 21 performs emotional trend analysis processing, the computer 20 functions as the text mining device 10. Note that the configuration of the computer 20 described above is merely an example, and the text mining device 10 can be realized using various computers.

[0029] <3. Emotional Word Dictionary with Emotion Index> The emotional tendency analysis process uses an emotional word dictionary 34, which is an emotional word dictionary with an emotional index stored in the auxiliary storage device 23. FIG. 3 is a diagram illustrating the emotional word dictionary with an emotional index used in this embodiment. In this emotional word dictionary, words indicating an emotional polarity, such as positive or negative, are collected and registered as emotional words. Furthermore, for each registered emotional word, a numerical value indicating the strength of the emotional polarity is indicated as an emotional index. This emotional index is a numerical value ranging from −1.00 to +1.00, with positive emotional words being assigned positive numerical values ​​and negative emotional words being assigned negative numerical values. For example, as shown in FIG. 3, the word “excellent” (emotional word) with a strong positive meaning is assigned an emotional index of +1.00, and the word “horrible” (emotional word) with a strong negative meaning is assigned an emotional index of −1.00. Several methods are known for creating an emotional word dictionary with an emotional index, such as vectorizing (quantifying) words and then calculating the similarity with known emotional words. In this embodiment, data of an emotional word dictionary with emotional indexes created by any known method is stored in advance in the auxiliary storage device 23 as the emotional word dictionary 34.

[0030] <4. Emotional tendency analysis processing> As described above, emotional tendency analysis processing is performed on target documents by the CPU 21 of the computer 20 executing the text mining program 31. Figure 4 is a flowchart showing the procedure for this emotional tendency analysis processing. In this embodiment, the CPU 21 executes the text mining program 31, causing the computer 20 to operate as shown in Figure 4.

[0031] First, instructions for specifying the target document, the range of characteristic words, and the range of the emotional quotient (word emotional quotient) are accepted (step S10). Specifically, an operation screen such as that shown in FIG. 5 is displayed on the display device 25, and the user operates the operation screen to specify the target document, the range of characteristic words, and the range of the emotional quotient using the keyboard 28 or mouse 29 of the input operation unit 24, and clicks the "OK" button 260 on the operation screen. This causes the computer 20, acting as the text mining device 10, to receive input information indicating the specified target document, the range of characteristic words, and the range of the emotional quotient. In the example shown in FIG. 5, the range of the emotional quotient can be specified by operating a slider 250 having a first knob 251 and a second knob 252. That is, by setting the positions of the first knob 251 and the second knob 252 on the slider 250, two ranges can be specified for the emotional index: a negative emotional index range (negative emotional index range) ranging from "-1.00" to the negative value indicated by the position of the first knob 251, and a positive emotional index range (positive emotional index range) ranging from the positive value indicated by the position of the second knob 252 to "+1.00." Note that multiple documents whose emotional tendencies are to be compared are specified as the target documents. In the following explanation, it is assumed that review documents (documents containing user impressions, critiques, opinions, etc. for each model of the product) for models A, B, and C of a certain product are specified as the target documents.

[0032] In step S10 above, the specification of the range of characteristic words is premised on the assumption that a numerical value indicating the characteristic degree of each word is used when extracting characteristic words from each document designated as a target document (details will be described later). The specification of the range of characteristic words is performed by specifying how many words are to be extracted as characteristic words in each document designated as a target document, in descending order of characteristic degree.

[0033] After receiving instructions to specify the target documents, the range of feature words, and the range of emotion quotients in this way, the text data 32 of the multiple documents specified as the target documents is first read from the auxiliary storage device 23 into the main memory 22 (step S12). Next, using this text data 32, feature words within the specified range are extracted as target feature words from each of the multiple target documents (step S14).

[0034] Figure 6 shows an example of feature word extraction when review documents for models A, B, and C of a certain product are specified as target documents and the top 10 feature words in descending order of characteristic degree are specified as the range of feature words. In Figure 6, the 10 feature words in descending order of characteristic degree are shown for each of the review documents for models A, B, and C, along with numerical values ​​indicating their characteristic degrees.

[0035] In the example shown in Figure 6, the Jaccard coefficient is used as a numerical value indicating the distinctiveness of a word. If the target documents are review documents for model A, model B, and model C, respectively, and are denoted by Da, Db, and Dc, the Jaccard coefficient Jxw of word w in document Dx is calculated by the following steps (p1) to (p4) (x=a, b, c). (p1) The number Nw of sentences containing word w among all sentences contained in documents Da, Db, and Dc is calculated. (p2) The number of sentences Nx contained in document Dx is calculated. (p3) The number Nxw of sentences that contain word w among the sentences contained in document Dx is calculated. (p4) The Jaccard coefficient Jxw of the word w in the document Dx is calculated using the following formula. Jxw=Nxw / (Nw+Nx-Nxw) …(1)

[0036] Generally, when a plurality of documents D1, D2, . . . , Dn are specified as target documents, the Jaccard coefficient Jkw of a word w in a document Dk (1≦k≦n) among them is expressed by the following formula. Jkw=|Sw∩Sk| / |Sw∪Sk| …(2) Here, Sw denotes a set whose elements are all sentences that contain word w among all sentences contained in documents D1, D2, . . . , Dn, and Sk denotes a set whose elements are all sentences contained in document Dk.

[0037] Instead of using the Jaccard coefficient, it is possible to use the number of sentences containing a word in document Dk (hereinafter referred to as "document occurrence count") as a numerical value indicating the distinctiveness of a word in document Dk (1 ≦ k ≦ n) among multiple documents D1, D2, ..., Dn serving as target documents. Using this document occurrence count can lead to the following problems. A word wp that appears frequently in all documents D1, D2, ..., Dn cannot be said to have a high distinctiveness in any of documents Dk (1 ≦ k ≦ n), but has a large number of document occurrences. Furthermore, a word wq that appears in a certain document Dk (1 ≦ k ≦ n) but is almost completely absent in other documents Dj (j ≠ k and 1 ≦ j ≦ n) may be considered to have a high distinctiveness even if the number of sentences containing the word wq in document Dk (document occurrence count) is not large. However, if the document occurrence count is small beyond a certain level, the word cannot be considered a distinctive word in document Dk. In contrast, when the Jaccard coefficient is used, the Jaccard coefficient calculated for these two words wp and wq in document Dk using the above formula (2) becomes sufficiently small, and neither of these two words wp and wq is extracted as a feature word.

[0038] Once target feature words have been extracted from each of the multiple target documents in the above manner, attention is then focused on one of the target documents that has not yet been focused on (step S15). Note that when step S15 is executed for the first time after the start of the emotional tendency analysis process, all of the multiple documents designated as target documents are in a state of not yet being focused on. As described above, if review documents for model A, model B, and model C of a certain product are designated as target documents, one of the review documents for model A, model B, and model C will be the document of interest.

[0039] Next, each of the target feature words extracted from the target document is searched in the emotional word dictionary 34 with emotional indices, and those of the target feature words that are registered in the emotional word dictionary 34 as words (emotional words) with emotional indices within a specified range are assigned an emotional index (step S16). FIG. 7 shows an example display of the emotional tendencies of feature words extracted from each of the review documents for model A, model B, and model C designated as target documents. This display example is provided for ease of explanation and constitutes the main part of the actual display example shown in FIG. 8, which will be described later. In this display example, feature words assigned with emotional indices are assigned a background color whose color varies depending on whether the feature word is positive or negative (whether the assigned emotional indices are positive or negative) and whose intensity corresponds to the emotional indices assigned to the feature word. For example, positive feature words are assigned a blue background color whose intensity corresponds to their emotional indices, and negative feature words are assigned a red background color whose intensity corresponds to their emotional indices. As described above, when a review document of model A is the target document, the background color of the feature words shown for "model A" in Fig. 7 indicates the emotional index assigned based on the emotional word dictionary 34. No background color is assigned to feature words that are not registered in the emotional word dictionary 34 among the target feature words.

[0040] Next, the sentiment index of the target document is calculated based on the feature words assigned sentiment indices among the target feature words extracted from the target document (step S18). That is, the document sentiment index Ctx, which is the sentiment index of the target document, is calculated using the following formula. Ctx=(Naf-Nng) / (Naf+Nng) …(3) Here, Naf is the number of feature words extracted from the target document that have been assigned a positive sentiment index, i.e., the number of positive feature words that appear in the target document. Nng is the number of feature words extracted from the target document that have been assigned a negative sentiment index, i.e., the number of negative feature words that appear in the target document. As can be seen from the above formula (3), the document sentiment index Ctx ranges from -1 to +1.

[0041] It is then determined whether the target documents include any uninterested documents (step S20), and if any uninterested documents are included, the process returns to step S15. Steps S15 to S20 are then repeated until the target documents no longer include any uninterested documents, and if it is determined in step S20 that the target documents do not include any uninterested documents, the process proceeds to step S22. As described above, if review documents for model A, model B, and model C of a certain product are specified as target documents, steps S15 to S20 are executed for each of the review documents for model A, model B, and model C, and then the process proceeds to step S22.

[0042] At the time of proceeding to step S22, feature words within the specified range have been extracted as target feature words for each of the review documents for model A, model B, and model C designated as target documents (see FIG. 6), and those feature words among the target feature words that are registered in the emotional word dictionary 34 as emotional words with emotional indices within the specified range have been assigned emotional indices (see FIG. 7), and the document emotional quotient Ctx has been calculated based on the target feature words to which emotional indices have been assigned. In step S22, display data is generated to display the target feature words thus obtained for the multiple target documents, the emotional indices assigned to the emotional words contained therein, and the document emotional quotient Ctx so that the multiple documents can be compared (step S22). In other words, display data is generated to compare the emotional tendencies of the review documents for model A, model B, and model C based on the target feature words obtained for each of the target documents, the emotional indices assigned to the emotional words contained therein, and the document emotional quotient Ctx.

[0043] Next, using the display data generated in this way, the emotional tendencies of multiple documents designated as target documents are displayed on the display device 25 in a comparative manner (step S24). This shows the results of the emotional tendency analysis of the target documents. For example, the results of the emotional tendency analysis of target documents consisting of review documents for models A, B, and C are displayed as shown in FIG. 8. In the display example of FIG. 8, the target feature words assigned with emotional indices (word emotional indices) are given background colors with a color and intensity corresponding to the emotional indices. In addition, the names of the documents designated as target documents, "Model A," "Model B," and "Model C," are each given a background color with a color and intensity corresponding to the document emotional indices Ctx. Furthermore, in this display example, a slider 250, the same as the slider 250 (see FIG. 5) displayed in step S10 for specifying the emotional index range, is displayed together with the results of the emotional tendency analysis of the target documents.

[0044] The slider 250 shown in Fig. 8 is also configured to be operable by the user using the mouse 29. The user can change the range of emotional quotients of the feature words to be extracted from the target document by operating the slider 250 while viewing the display of the results of the emotional trend analysis of the target document as shown in Fig. 8. That is, while the results of the emotional trend analysis of the target document are being displayed, the computer 20 waits until the slider 250 is operated (step S26), and when the slider 250 is operated by the user, all documents specified as target documents (all review documents for model A, model B, and model C) are returned to an unfocused state (step S28), and the process returns to step S12.

[0045] Thereafter, with the specified range of the emotion quotient changed, steps S15 to S20 are executed for each of the multiple documents specified as target documents (here, review documents for models A, B, and C) in the same manner as above, and then the process proceeds to step S22. After this, when steps S22 and S24 are executed, the results of the emotional tendency analysis of the target documents for the specified range after the emotion quotient has been changed are displayed in a format similar to that shown in FIG. 8. As described above, the computer 20 waits in this display state until the slider 250 is operated. During this wait, if the range of the emotion quotient is further changed by operating the slider 250, the process returns to step S12 and the same process as above is performed. Note that during this wait, if an end instruction is received by interrupt processing, the emotional tendency analysis process shown in FIG. 4 is terminated.

[0046] As can be seen from the above explanation, in this embodiment, steps S10, S24, and S26, which perform processing related to the input operation unit 24 and the display device 25, realize the GUI unit 11 as an instruction input unit and display unit, step S14 realizes the feature word extraction unit 13, step S16 realizes the feature word sentiment index acquisition unit 15, and step S18 realizes the document sentiment index calculation unit 16.

[0047] <5. Effects> According to the present embodiment described above, feature words are extracted from each of multiple documents (target documents) for which emotional tendencies are to be compared, and those extracted as target feature words that are registered as emotional words in the emotional word dictionary 34 are assigned the emotional index assigned to that feature word in the emotional word dictionary 34. In this way, for each of the multiple documents, the target feature words and the emotional indices assigned to the included emotional words are displayed as the results of the emotional trend analysis for the multiple documents designated as target documents. With this display (see steps S22 and S24 in FIG. 4, and FIGS. 7 and 8), even if the multiple documents for which emotional tendencies are to be compared contain feature words with weak emotional polarity, the emotional tendencies of the multiple documents can be accurately grasped by viewing the extracted feature words and the emotional indices assigned to them, which range from -1 to +1.

[0048] Furthermore, according to this embodiment, it is possible to specify the range of emotional indices assigned to target feature words, which are feature words extracted from each of the multiple documents (see step S10 in FIG. 4 and FIG. 5). That is, by operating the slider 250 shown in FIG. 5, two ranges consisting of the negative emotional indices and the positive emotional indices can be specified as the emotional indices, as described above, and no emotional indices can be assigned to neutral feature words with weak emotional polarity among the target feature words. This allows for accurate comparison of the emotional tendencies of multiple documents containing feature words with weak emotional polarity. In this embodiment, the document emotional indices Ctx are calculated for each of the multiple documents using the above-described formula (3) based on the target feature words assigned emotional indices within the specified range (see step S16 in FIG. 4). Therefore, in addition to comparing the emotional indices assigned to the feature words among the multiple documents, it is also possible to compare the document emotional indices Ctx among the multiple documents (see FIG. 8). This allows for more accurate and easier comparison of the emotional tendencies of multiple documents.

[0049] Furthermore, according to this embodiment, since it is possible to specify the range of characteristic words to be extracted from each of the multiple target documents (see steps S10 and S14 in FIG. 4 and FIG. 5), by extracting only the more characteristic words (for example, the top 10 words in descending order of characteristic degree) as target characteristic words, it is possible to compare the emotional trends that reflect the characteristics of each of the multiple documents between the multiple documents with a smaller amount of calculation than in the past.In addition, it is possible to avoid the problem of underestimating characteristic emotional words that appear more frequently in one of the multiple documents than in other documents.

[0050] Furthermore, according to this embodiment, the display device 25 showing the results of the emotional tendency analysis of the plurality of documents also displays a slider 250 for specifying the range of the emotional quotient (FIG. 8), so that the user can change the range of the emotional quotient while looking at the results of the emotional tendency analysis, and display the results of the emotional tendency analysis of the plurality of documents based on the changed range of the emotional quotient (steps S26 → S28 → S15 → ... → S24 in FIG. 4). As a result, the user can once display the results of the emotional tendency analysis of the plurality of documents, and then adjust the specified range of the emotional quotient while looking at the display, thereby more accurately comparing the emotional tendencies of the plurality of documents.

[0051] <6. Variations> The present invention is not limited to the above-described embodiment, and various modifications can be made without departing from the scope of the present invention.

[0052] For example, in the above embodiment, when extracting feature words from each of multiple documents designated as target documents for comparing emotional tendencies, the range of feature words to be extracted from the target documents is specified by the range (minimum and maximum values) of the Jaccard coefficient, which is a numerical value indicating the distinctiveness of the feature words. However, the range of feature words may be specified using other numerical values ​​indicating the distinctiveness of the feature words, without being limited to the Jaccard coefficient. For example, the Dice coefficient or Simpson coefficient may be used instead of the Jaccard coefficient, and the range of feature words to be extracted may be specified by a numerical value indicating the distinctiveness based on the TF-IDF (Term Frequency - Inverse Document Frequency) method.

[0053] In the above embodiment, the display of the results of the emotional trend analysis for multiple documents specified as target documents shows the feature words extracted for each of the multiple documents and numerical values ​​indicating their distinctiveness, as well as the feature word emotional quotient (word emotional quotient) and document emotional quotient (document emotional quotient) Ctx, which are displayed using background colors of colors and densities corresponding to the emotional quotients, as shown in Fig. 8. However, the display format of the results of the emotional trend analysis is not limited to this, and the feature word emotional quotient and document emotional quotient may be displayed in other formats, such as numerical values ​​or bar graphs. [Explanation of symbols]

[0054] 10...Text mining device 11...GUI section 12...Text data storage unit 13...Feature word extraction section 14...Dictionary memory section 15...Feature word emotion index acquisition section 16...Document sentiment index calculation section 17...Display data processing unit 20...Computer 21...CPU 22...Main memory 23…Auxiliary storage device 24...Input operation section 25…Display device 30...Recording media 31...Text mining program 32...Text data 34...Emotional Word Dictionary 250...Slider

Claims

1. 1. A computer-implemented text mining method for comparing sentiment polarity trends among a plurality of documents, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the characteristic words extracted by the characteristic word extraction step together with the emotion index assigned by the emotion index acquisition step for the plurality of documents designated as the target documents; Equipped with The instruction input step further includes a step of receiving an instruction specifying a range of feature words to be extracted from the target document; A text mining method in which, in the feature word extraction step, feature words within the range specified in the instruction input step are extracted.

2. A computer-implemented text mining method for comparing trends in sentiment polarity among a plurality of documents, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the characteristic words extracted by the characteristic word extraction step together with the emotion index assigned by the emotion index acquisition step for the plurality of documents designated as the target documents; Equipped with The instruction input step further includes a step of receiving an instruction to specify a range of an emotion index that is an index indicating the strength of an emotion polarity, In the emotion index acquisition step, an emotion index is assigned to a feature word that is registered in the emotion word dictionary as a word that has been assigned an emotion index within the range specified in the instruction input step, from among the feature words extracted in the feature word extraction step.

3. 3. The text mining method according to claim 2, wherein the instruction input step further includes a step of receiving an instruction to specify a change in the range of the emotional quotient when the extracted feature words are displayed together with the assigned emotional quotient by the display step.

4. A computer-implemented text mining method for comparing trends in sentiment polarity among a plurality of documents, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the feature words extracted by the feature word extraction step together with the emotion quotient assigned by the emotion quotient acquisition step for the plurality of documents designated as the target documents; a document sentiment index calculation step of calculating, for each of the plurality of documents designated as the target documents, a sentiment index of the document as a document sentiment index based on the feature words to which sentiment indexes have been assigned in the sentiment index acquisition step among the feature words extracted from the document in the feature word extraction step; Equipped with A text mining method, wherein the display step displays the document sentiment index calculated in the document sentiment index calculation step.

5. 5. The text mining method according to claim 4, wherein in the document sentiment index calculation step, the document sentiment index Ctx is calculated for each of the plurality of documents designated as the target documents by the following formula: Ctx=(Naf-Nng) / (Naf+Nng) Here, Naf is the number of occurrences of positive feature words in the document, and Nng is the number of occurrences of negative feature words in the document.

6. 6. The text mining method according to claim 4, wherein in the display step, the names of the plurality of documents designated as the target documents are displayed against a background color that varies depending on whether the emotional quotient of the document is positive or negative, and that has a density that corresponds to the emotional quotient of the document.

7. A computer-implemented text mining method for comparing trends in sentiment polarity among a plurality of documents, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the characteristic words extracted by the characteristic word extraction step together with the emotion index assigned by the emotion index acquisition step for the plurality of documents designated as the target documents; Equipped with A text mining method in which, in the display step, characteristic words assigned with emotional indices in the emotional index acquisition step are displayed against a background color that differs depending on whether the emotional index of the characteristic word is positive or negative, and has a density that corresponds to the emotional index of the characteristic word.

8. A text mining program for comparing trends in sentiment polarity among a plurality of documents, an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the characteristic words extracted by the characteristic word extraction step together with the emotion index assigned by the emotion index acquisition step for the plurality of documents designated as the target documents; The CPU executes the program using the memory. The instruction input step further includes a step of receiving an instruction specifying a range of feature words to be extracted from the target document; A text mining program in which, in the feature word extraction step, feature words within the range specified in the instruction input step are extracted.

9. A text mining program for comparing trends in sentiment polarity among multiple documents, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the characteristic words extracted by the characteristic word extraction step together with the emotion index assigned by the emotion index acquisition step for the plurality of documents designated as the target documents; The CPU executes the program using the memory. The instruction input step further includes a step of receiving an instruction to specify a range of an emotion index that is an index indicating the strength of an emotion polarity, A text mining program in which, in the emotion quotient acquisition step, an emotion quotient is assigned to a feature word that is registered in the emotion word dictionary as a word that has been assigned an emotion quotient within the range specified in the instruction input step, from among the feature words extracted in the feature word extraction step.

10. 10. The text mining program according to claim 9, wherein the instruction input step further includes a step of receiving an instruction to specify a change in the range of the emotion quotient when the extracted feature words are displayed together with the assigned emotion quotient by the display step.

11. A text mining program for comparing trends in sentiment polarity among multiple documents, comprising: an instruction input step of receiving an instruction to designate a plurality of documents to be compared in terms of semantic polarity trends as target documents; a feature word extraction step of extracting feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional index acquisition step of assigning an emotional index to a feature word registered in a predetermined emotional word dictionary from among the feature words extracted by the feature word extraction step, the emotional index being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display step of displaying the feature words extracted by the feature word extraction step together with the emotion quotient assigned by the emotion quotient acquisition step for the plurality of documents designated as the target documents; a document sentiment index calculation step of calculating, for each of the plurality of documents designated as the target documents, a sentiment index of the document as a document sentiment index based on the feature words to which sentiment indexes have been assigned in the sentiment index acquisition step among the feature words extracted from the document in the feature word extraction step; The CPU executes the program using the memory. The display step displays the document sentiment index calculated in the document sentiment index calculation step.

12. A text mining device for comparing trends in sentiment polarity among a plurality of documents, comprising: an instruction input unit that receives an instruction to specify, as target documents, a plurality of documents whose semantic polarity trends are to be compared; a feature word extraction unit that extracts feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional quotient acquisition unit that assigns an emotional quotient to a feature word that is registered in a predetermined emotional word dictionary among the feature words extracted by the feature word extraction unit, the emotional quotient being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display unit that displays the characteristic words extracted by the characteristic word extraction unit together with the emotion quotient assigned by the emotion quotient acquisition unit for the plurality of documents designated as the target documents; Equipped with the instruction input unit further receives an instruction specifying a range of feature words to be extracted from the target document; The feature word extraction unit extracts feature words within a range specified by the instruction input unit.

13. A text mining device for comparing trends in sentiment polarity among a plurality of documents, comprising: an instruction input unit that receives an instruction to specify, as target documents, a plurality of documents whose semantic polarity trends are to be compared; a feature word extraction unit that extracts feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional quotient acquisition unit that assigns an emotional quotient to a feature word that is registered in a predetermined emotional word dictionary among the feature words extracted by the feature word extraction unit, the emotional quotient being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display unit that displays the characteristic words extracted by the characteristic word extraction unit together with the emotion quotient assigned by the emotion quotient acquisition unit for the plurality of documents designated as the target documents; Equipped with the instruction input unit further receives an instruction to specify a range of an emotion index which is an index indicating the strength of an emotional polarity; The emotional quotient acquisition unit assigns an emotional quotient to a characteristic word that is registered in the emotional word dictionary as a word that has been assigned an emotional quotient within the range specified by the instruction input unit, from among the characteristic words extracted by the characteristic word extraction unit.

14. 14. The text mining device according to claim 13, wherein the instruction input unit further receives an instruction to specify a change in a range of the emotion quotient when the extracted feature word is displayed by the display unit together with the assigned emotion quotient.

15. A text mining device for comparing trends in sentiment polarity among a plurality of documents, comprising: an instruction input unit that receives an instruction to specify, as target documents, a plurality of documents whose semantic polarity trends are to be compared; a feature word extraction unit that extracts feature words from each of the plurality of documents based on text data of the plurality of documents designated as the target documents; an emotional quotient acquisition unit that assigns an emotional quotient to a feature word that is registered in a predetermined emotional word dictionary among the feature words extracted by the feature word extraction unit, the emotional quotient being a numerical value indicating the strength of emotional polarity of the feature word in the emotional word dictionary; a display unit that displays, for the plurality of documents designated as the target documents, the feature words extracted by the feature word extraction unit together with the emotion quotient assigned by the emotion quotient acquisition unit; a document sentiment index calculation unit that calculates, for each of the plurality of documents designated as the target documents, a sentiment index of the document as a document sentiment index based on the feature words that have been assigned sentiment indices by the sentiment index acquisition unit among the feature words extracted from the document by the feature word extraction unit; Equipped with The display unit displays the document sentiment index calculated by the document sentiment index calculation unit.

Citation Information

Patent Citations

  • Device, method and program for processing information

    JP2003157255A

  • Device, and method for supporting input of evaluation information and program for executing the same method

    JP2010271757A

  • System and method for classifying text feeling polarities based on sentence sequence

    JP2011204226A

  • Document analysis apparatus and program

    JP2012093966A

  • Evaluation expression polarity determination device, method and program

    JP2013246636A