Apparatus, method and program for supporting text analysis and recording medium

The text analysis support device processes text data by classifying and labeling partial texts, addressing the limitation of existing systems to analyze text data, thereby enhancing student ability analysis with text-based insights.

JP2025110632APending Publication Date: 2025-07-29NEC SOLUTION INNOVATORS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024004576
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing student ability analysis systems cannot effectively analyze text data such as questionnaire contents, limiting their ability to process and understand student abilities accurately.

Method used

A text analysis support device that includes a text data acquisition unit, classification unit, group word extraction unit, and label setting unit to process text data into analyzable data by classifying partial texts into groups, assigning group identification information, extracting group characteristic words, and setting labels for each partial text.

Benefits of technology

Enables the conversion of text data into analyzable data, allowing for more comprehensive student ability analysis by incorporating text-based insights into the evaluation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025110632000001_ABST
    Figure 2025110632000001_ABST
Patent Text Reader

Abstract

To provide an apparatus, a method and a program for supporting text analysis that allow for processing text data into analyzable data, and a recording medium.SOLUTION: An apparatus for supporting text analysis 10 comprises a text-data acquiring section that acquires text data including at least one partial text, a classifying section that classifies, for each partial text, the partial text into at least one group based on a word included in the partial text and gives group identifying information to the partial text, a group-word extracting section that extracts at least one group feature word based on an in-group word included in the group and a word included in the text data, a label setting section that sets, for each partial text, a set of the group identifying information and the group feature word as a label, and an output section that outputs for-analysis text data obtained by tying the label to the partial text.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a text analysis support device, a text analysis support method, a text analysis support program, and a recording medium. [Background technology]

[0002] Universities and other organizations are increasingly using IT (information technology) to manage students, staff, and other resources. A student ability analysis system is known that allows students to accurately understand their own abilities and promotes optimal instruction by faculty and staff (see Patent Document 1). The student ability analysis system in Patent Document 1 stores ability category data containing multiple ability category names indicating each student's ability category; subject data containing subject names indicating each subject and multiple ability category names related to each subject; grade data containing the student's name, subject names indicating the subject, and grades indicating the evaluation scores for the subject; and extracurricular activity data containing the student's name, ability category names for the extracurricular activities, and extracurricular activity evaluation points indicating the evaluation scores for the extracurricular activities. An objective evaluation data calculation processor calculates objective evaluation points that objectively indicate the student's evaluation based on the subject data, grade data, and extracurricular activity data, and creates and outputs objective evaluation data containing the student's name, ability category names for the objective evaluation, and the objective evaluation points. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2012-194507 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the system described in Patent Document 1 has a problem in that while it is possible to analyze data that has been converted into numerical values in advance when analyzing student abilities, it is unable to analyze text data such as questionnaire contents.

[0005] Therefore, the present disclosure aims to provide a text analysis support device, a text analysis support method, a text analysis support program, and a recording medium that can process text data into analyzable data.

Means for Solving the Problems

[0006] To achieve the above object, the text analysis support device of the present disclosure includes a text data acquisition unit, a classification unit, a group word extraction unit, a label setting unit, and an output unit, wherein the text data acquisition unit acquires text data including at least one partial text, the classification unit classifies each of the partial texts into at least one group based on the words included in the partial text, and assigns group identification information to the partial text, the group word extraction unit extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data, the label setting unit sets, for each of the partial texts, a pair of the group identification information and the group characteristic word as a label, and the output unit outputs analysis text data in which the label is associated with the partial text.

[0007] The text analysis support method of the present disclosure includes a text data acquisition step, a classification step, a group word extraction step, a label setting step, and an output step, wherein the text data acquisition step acquires text data including at least one partial text, the classification step classifies each of the partial texts into at least one group based on the words included in the partial text, and assigns group identification information to the partial text, the group word extraction step extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data, In the label setting step, for each of the partial texts, a pair of the group identification information and the group characteristic word is set as a label. In the output step, analysis text data in which the label is associated with the partial text is output, and each step is executed by a computer, which is a text analysis support method.

[0008] The text analysis support program of the present disclosure includes a text data acquisition procedure, a classification procedure, a group word extraction procedure, a label setting procedure, and an output procedure. The text data acquisition procedure acquires text data including at least one partial text. The classification procedure classifies each of the partial texts into at least one group based on the words included in the partial text, and assigns group identification information to the partial text. The group word extraction procedure extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data. The label setting procedure sets, for each of the partial texts, a pair of the group identification information and the group characteristic word as a label. The output procedure outputs analysis text data in which the label is associated with the partial text. It is a text analysis support program for causing a computer to execute each of the procedures.

[0009] The recording medium of the present disclosure includes a text data acquisition procedure, a classification procedure, a group word extraction procedure, a label setting procedure, and an output procedure. The text data acquisition procedure acquires text data including at least one partial text. The classification procedure classifies each of the partial texts into at least one group based on the words included in the partial text, and assigns group identification information to the partial text. The group word extraction procedure extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data. The label setting procedure sets, for each partial text, a pair of the group identification information and the group characteristic word as a label. The output procedure outputs analysis text data in which the label is associated with the partial text. It is a computer-readable recording medium that records a text analysis support program for causing a computer to execute each of the above procedures.

Advantages of the Invention

[0010] According to the present disclosure, text data can be processed into analyzable data.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0012] Next, embodiments of the present disclosure will be described with reference to the drawings. The present disclosure is not limited to the following embodiments. In the following figures, the same parts are denoted by the same reference numerals. Also, the descriptions of the respective embodiments can be mutually referred to unless otherwise specified, and the configurations of the respective embodiments can be combined unless otherwise specified.

[0013] [Embodiment 1] The text analysis support device of this embodiment will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of an example of the text analysis support device 10 of this embodiment. As shown in FIG. 1, the text analysis support device 10 (hereinafter also referred to as "this device 10") includes a text data acquisition unit 11, a classification unit 12, a group word extraction unit 13, a label setting unit 14, and an output unit 15. Although not shown, this device 10 may include, for example, a storage unit.

[0014] This device 10 may be, for example, one device including the above-mentioned respective units, or the above-mentioned respective units may be devices connectable via a communication line network. Further, this device 10 can be connected to an external device described later via a communication line network. The communication line network is not particularly limited, and a known network can be used. For example, it may be wired or wireless. Examples of the communication line network include an Internet line, WWW (World Wide Web), a telephone line, a LAN (Local Area Network), a SAN (Storage Area Network), a DTN (Delay Tolerant Networking), an LPWA (Low Power Wide Area), an L5G (local 5G), and the like. Examples of the wireless communication include Wi-Fi (registered trademark), Bluetooth (registered trademark), local 5G, LPWA, and the like. The wireless communication may be in a form in which each device directly communicates (Ad Hoc communication), infrastructure communication, indirect communication via an access point, or the like. This device 10 may be incorporated into a server as a system, for example. Further, this device 10 may be, for example, a personal computer (PC, for example, a desktop type or a notebook type) installed with the program of the present disclosure, a smartphone, a tablet terminal, or the like. Furthermore, this device 10 may be in a form such as cloud computing or edge computing, for example, in which at least one of the above-mentioned respective units is on a server and the other above-mentioned respective units are on a terminal.

[0015] FIG. 2 illustrates a block diagram of the hardware configuration of the present apparatus 10. The present apparatus 10 includes, for example, a central processing unit (CPU, GPU, etc.) 101, a memory 102, a bus 103, a storage device 104, an input device 105, an output device 106, a communication device 107, and the like. Each part of the present apparatus 10 is interconnected via the bus 103 by respective interfaces (I / F).

[0016] The central processing unit 101 operates in cooperation with other components by a controller (system controller, I / O controller, etc.) and is responsible for overall control of the present apparatus 10. In the present apparatus 10, for example, the program of the present disclosure and other programs are executed by the central processing unit 101, and various information is read and written. Specifically, for example, the central processing unit 101 causes the text analysis support apparatus 10 (hereinafter also referred to as "the present apparatus 10") to function as a text data acquisition unit 11, a classification unit 12, a group word extraction unit 13, a label setting unit 14, and an output unit 15. The present apparatus 10 may include other arithmetic units such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an APU (Accelerated Processing Unit) as an arithmetic unit, or a combination thereof.

[0017] The bus 103 can be connected to an external device, for example. Examples of the external device include an external storage device (external database, etc.), a printer, an external input device, an external display device, an audio output device such as a speaker, an external imaging device such as a camera, and various sensors such as an acceleration sensor, a geomagnetic sensor, and a direction sensor. The present apparatus 10 can be connected to an external network (the communication line network) by, for example, a communication device 107 connected to the bus 103, and can also be connected to other devices such as a user's terminal via the external network.

[0018] The memory 102 includes, for example, a main memory (primary storage device). When the central processing unit 101 performs processing, for example, the memory 102 reads various operation programs such as the program of the present disclosure stored in the storage device 104 described later, and the central processing unit 101 receives data from the memory 102 and executes the program. The main memory is, for example, a RAM (Random Access Memory). Also, the memory 102 may be, for example, a ROM (Read Only Memory).

[0019] The storage device 104 is also referred to as, for example, a so-called auxiliary storage device with respect to the main memory (primary storage device). As described above, the storage device 104 stores an operation program including the program of the present disclosure. The storage device 104 may be, for example, a combination of a recording medium and a drive for reading and writing to the recording medium. The recording medium is not particularly limited, and may be, for example, an internal type or an external type, and examples include an HD (Hard Disk), CD-ROM, CD-R, CD-RW, MO, DVD, flash memory, memory card, etc. The storage device 104 may be, for example, a hard disk drive (HDD) in which a recording medium and a drive are integrated, and a solid state drive (SSD). When the device 10 includes the storage unit, for example, the storage device 104 functions as the storage unit.

[0020] In the device 10, the memory 102 and the storage device 104 can also store various information such as log information, information acquired from an external database (not shown) or an external device, information generated by the device 10, and information used when the device 10 executes processing. In this case, the memory 102 and the storage device 104 may store, for example, information of the user of the device described above. Note that at least some of the information may be stored in an external server other than the memory 102 and the storage device 104, or may be distributed and stored in a plurality of terminals using blockchain technology or the like.

[0021] The device 10 further includes, for example, an input device 105 and an output device 106. The input device 105 includes, for example, pointing devices such as touch panels, track pads, and mice; keyboards; imaging means such as cameras and scanners; card readers such as IC card readers and magnetic card readers; voice input means such as microphones; and the like. The output device 106 includes, for example, display devices such as LED displays and liquid crystal displays; voice output devices such as speakers; printers; and the like. In the present embodiment, the input device 105 and the output device 106 are separately configured, but the input device 105 and the output device 106 may be integrally configured like a touch panel display.

[0022] Next, an example of the text analysis support method of the present embodiment will be described based on the flowchart of FIG. 3. The text analysis support method of the present embodiment can be implemented as follows, for example, using the text analysis support device 10 shown in FIG. 1 or FIG. 2. Note that the text analysis support method of the present embodiment is not limited to the use of the text analysis support device 10 of FIG. 1 or FIG. 2. FIG. 3 is a flowchart showing an example of the processing by the text analysis support device 10.

[0023] The text data acquisition unit 11 acquires text data including at least one partial text (S1, text data acquisition step). The text data is not particularly limited and may be, for example, sentence data in a natural language or a set of a plurality of sentences. The partial text is, for example, sentence data included in the text data. The sentence included in the partial text may be, for example, a grammatically correct sentence or a non-sentence. The non-sentence is not particularly limited and includes, for example, a part of a sentence (such as only a noun phrase), a list, and the like. The text data may be, for example, data in a format including rows and columns. In this case, in the text data, for example, the partial text may be stored row by row. The format of the text data is not particularly limited, and examples include, but are not limited to, CSV (Comma Separated Value), TSV (Tab Separated Values), Fixed Width Format, Markdown Table: Markdown format, etc. The text data may include, for example, text data in the free description field of a questionnaire. Further, the text data may be, for example, transcribed data of voice data or the like. The text data acquisition unit 11 may acquire, for example, the text data stored in the storage unit of the present device 10, or may acquire the text data from outside the present device 10. The text data acquisition unit 11 may store the text data in, for example, the memory 102 or the storage device 104.

[0024] A specific example of the text data acquired by the text data acquisition unit 11 will be described with reference to Table 1. The following Table 1 is an example of text data in which questionnaire responses regarding school classes by students are described, but the present disclosure is not limited or restricted in any way by the following examples. As shown in the following Table 1, the text data includes partial texts (questionnaire response contents) numbered from 1 to 10.

Table 1

[0025] The classification unit 12 classifies the partial text into at least one group based on the words included in the partial text for each partial text, and assigns group identification information to the partial text (S2, classification step). The classification unit 12 first divides, for example, the partial text included in the text data into words included in the partial text (word segmentation process). Next, the classification unit 12 converts the divided words into numerical values (numerical conversion process). Then, the classification unit 12 classifies the words into at least one group based on the numerical values (grouping process). The classification unit 12 can perform the classification process using, for example, known natural language processing techniques. The group identification information is, for example, information that can identify the classified group and is preferably numerical data. When the group identification information is numerical data, the group identification information is also referred to as a group identification number.

[0026] The word segmentation process can be performed, for example, by processing the partial text with known natural language processing software and performing morphological analysis. The natural language processing software is not particularly limited, and examples include morphological analysis engines such as MeCab, Chasen, JUMAN, and Kytea. At this time, the classification unit 12 may perform, for example, the removal of so-called stop words in addition to morphological analysis. The classification unit 12 can obtain a group of divided words as an output result by inputting the partial text into these open source softwares, for example.

[0027] The numerical conversion process can be performed by, for example, a method for evaluating the importance of words included in a document, such as TF-IDF (Term Frequency-Inverse Document Frequency). TF-IDF can be performed by using open source softwares such as scikit-learn, NLTK (Natural Language Toolkit), Gensim, and SparkMLib, for example. The classification unit 12 can obtain the evaluation value of the word group as a numerical value as an output result by inputting the words included in the text data and the words included in the partial text into these open source softwares, for example.

[0028] The grouping process can be implemented, for example, using a known algorithm. The known algorithm is not particularly limited, and examples include algorithms for cluster analysis such as K-means included in scikit-learn, algorithms for topic modeling such as LDA (Latent Dirichlet Allocation) included in Gensim; etc., but are not limited thereto. The classification unit 12 can classify the partial text included in the text data into at least one group by inputting the evaluation values of the word groups into these open-source softwares. The number of groups in the grouping process is not particularly limited, and for example, any number can be specified according to the purpose of analysis, the type of input text data, etc. The number of groups can be arbitrarily set to a value that is, for example, 1 or more and not more than the number of the partial texts. As a specific example, when the text data is the text in the free description column of a questionnaire, the number of groups can be set to, for example, 5, 4, or 3, etc.

[0029] The group word extraction unit 13 extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data (S3, group word extraction step). For example, the group word extraction unit 13 can compare the words within the group with the words included in the text data, estimate the feature amount of the words within the group, and extract the group characteristic word based on the feature amount. The feature amount can be implemented by, for example, an importance evaluation method of words included in a document such as TF-IDF (Term Frequency-Inverse Document Frequency). TF-IDF can be implemented by using open source software such as scikit-learn, NLTK (Natural Language Toolkit), Gensim, SparkMLib, etc. The group word extraction unit 13 can, for example, input the words within the group included in the group and the words included in the text data into these open source software, and obtain the evaluation value of the group of words within the group as a numerical value (feature amount) as the output result. Then, the group word extraction unit 13 extracts at least one group characteristic word based on the feature amount for each group, for example. The group word extraction unit 13 may, for example, extract the top n words (n is an integer of 1 or more) in descending order of the feature amount for each group, or extract the words whose feature amount exceeds a threshold value, or a combination of these (for example, extract the top n words whose feature amount exceeds the threshold value). The number of group characteristic words extracted by the group word extraction unit 13 is not particularly limited and can be any number. As an example, there are the top 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 in descending order of the feature amount.

[0030] The group word extraction unit 13 may, for example, estimate a corrected feature amount obtained by correcting the feature amount with reference to the word weight data, and extract the group feature words based on the corrected feature amount. The word weight data is, for example, information on coefficients based on an index of the influence degree of words included in the sentence, and arbitrary values can be set. The word weight data may be stored, for example, in the storage unit of the present apparatus 10. In this case, the group word extraction unit 13 can estimate the corrected feature amount, for example, by referring to the word weight data and multiplying the coefficient by the feature amount of the words within the group.

[0031] The label setting unit 14 sets, for each of the partial texts, a pair of the group identification information and the group feature words as a label (S4, label setting step). The label setting unit 14 may, for example, record, for each group, the pair of the extracted group feature words and the group identification information in association with each other in the storage unit. Thereby, for the partial texts included in the text data, grouping reflecting the content of each partial text and quantification with numerical values (group identification numbers) associated with the groups become possible.

[0032] Then, the output unit 15 outputs analysis text data in which the label is associated with the partial text (S5, output step). For example, when the text data is data in a format including rows and columns like a CSV file and the partial texts are stored row by row, the output unit 15 may, for example, add a label column to a column adjacent to the row in which the partial text is stored and store the label in the added label column. The output unit 15 may output, for example, to the output device (for example, a display) 106 of the present apparatus 10, or may output to the outside of the present apparatus 10 via the communication device 107 connected to the bus 103. In the latter case, the output unit 15 can output the analysis text data to, for example, an external text analysis device.

[0033] Using Table 2 below, a specific example of the processing by the group word extraction unit 13, the label setting unit 14, and the output unit 15 is shown. Table 2 below is data when cluster analysis is adopted in the grouping process of the group word extraction unit 13. As shown in Table 2 below, the analysis text data is the text data shown in Table 1 with labels that are the analysis results attached thereto. [Table 2]

[0034] Also, using Table 3 below, another specific example of the processing by the group word extraction unit 13, the label setting unit 14, and the output unit 15 is shown. Table 3 below is data when topic modeling is adopted in the grouping process of the group word extraction unit 13. As shown in Table 3 below, the analysis text data is the text data shown in Table 1 with labels that are the analysis results attached thereto. [Table 3]

[0035] The text analysis support device of the present disclosure includes the following: a text data acquisition unit acquires text data including at least one partial text; a classification unit classifies each partial text into at least one group based on the words contained in the partial text; a group identification unit assigns group identification information to the partial text; a group word extraction unit extracts at least one group characteristic word based on the group-specific words contained in the group and the words contained in the text data; a label setting unit sets a pair of the group identification information and the group characteristic word as a label for each partial text; and an output unit outputs analytical text data in which the label is associated with the partial text. Therefore, the text analysis support device of the present disclosure can, for example, classify partial texts contained in text data into predetermined patterns. Furthermore, the text analysis support device of the present disclosure can, for example, group texts containing multiple texts (partial texts) with different content into any number of groups. Therefore, for example, by referencing group identification numbers corresponding to groups of partial texts contained in the text data, the text data can be used for analysis. Furthermore, without being limited to this, for example, by outputting the analytical text data to an analytical tool capable of text-based analysis, efficient analysis becomes possible using text data labeled with group characteristic words.

[0036] [Embodiment 2] The second embodiment is another example of a text analysis support device of the present disclosure.

[0037] FIG. 4 is a block diagram showing a configuration example of the text analysis support device 10A. As shown in FIG. 4, in addition to the configuration of the text analysis support device 10 of the first embodiment, the text analysis support device 10A includes a sentiment analysis unit 16. The hardware configuration of the text analysis support device 10A is the same as that of the text analysis support device 10 in FIG. 2, except that the central processing unit 101 has the configuration of the text analysis support device 10A in FIG. 4 instead of the configuration of the text analysis support device 10 in FIG. 1. Hereinafter, the processing of the sentiment analysis unit 16 will be described. The processing of the sentiment analysis unit 16 may be appropriately inserted at an arbitrary position in the flowchart of FIG. 3 described in the first embodiment, or may be executed independently of each step of the flowchart of FIG. 3, as shown in FIG. 5.

[0038] The sentiment analysis unit 16 analyzes, for example, the sentiment of the partial text included in the text data (S11, sentiment analysis step). The sentiment analysis unit 16 is a process of analyzing the sentiment (also referred to as polarity) included in the partial text. The sentiment may be, for example, a qualitative evaluation (e.g., positive, neutral, negative) or a quantitative evaluation. The quantitative evaluation may be, for example, a simple numerical scale or a combination with the qualitative evaluation. The combination of the quantitative evaluation and the qualitative evaluation is, for example, an evaluation such as "positive x points, neutral y points, negative z points (x, y, and z are integers of 1 or more)". The processing by the sentiment analysis unit 16 can be implemented, for example, using a known sentiment analysis algorithm. Examples of known sentiment analysis algorithms include Transformers and NaVie Bayes.

[0039] The label setting unit 14 may, for example, set the emotion as an emotion label in the partial text (S12), and the output unit 15 may, for example, output analysis text data in which the emotion label is associated with the partial text (S13). At this time, for example, if the text data is data in a format including rows and columns like a CSV file and the partial text is stored row by row, the output unit 15 may, for example, add an emotion column adjacent to the row in which the partial text is stored and store the emotion label in the added emotion column.

[0040] A specific example of the processing by the emotion analysis unit 16, the label setting unit 14, and the output unit 15 is shown using Table 4 below. As shown in Table 4 below, for the text data shown in Table 1, an emotion label which is an emotion analysis result is attached to the analysis text data.

Table 4

[0041] The text analysis support device of the present embodiment can analyze the emotion of the partial text included in the text data by the emotion analysis unit, set the emotion as an emotion label in the partial text by the label setting unit, and output analysis text data in which the emotion label is associated with the partial text by the output unit. Therefore, according to the text analysis support device of the present embodiment, for example, not only classification based on characteristic words included in text data but also the emotion of the text inputter can be added to the analysis data. Therefore, according to the text analysis support device of the present embodiment, for example, more accurate text analysis becomes possible.

[0042] [Embodiment 3] The text analysis support program of the present embodiment is a program for causing a computer to execute each step of the text analysis support method described above. Specifically, the text analysis support program of the present embodiment is a program for causing a computer to execute a text data acquisition procedure, a classification procedure, a group word extraction procedure, a label setting procedure, and an output procedure.

[0043] The text data acquisition procedure acquires text data including at least one partial text, The classification procedure classifies the partial text into at least one group based on the words included in the partial text for each partial text, and assigns group identification information to the partial text, The group word extraction procedure extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data, The label setting procedure sets, for each partial text, a combination of the group identification information and the group characteristic word as a label, The output procedure outputs analysis text data in which the label is associated with the partial text.

[0044] Also, the text analysis support program of the present embodiment can also be said to be a program that causes a computer to function as a text data acquisition procedure, a classification procedure, a group word extraction procedure, a label setting procedure, and an output procedure.

[0045] The text analysis support program of this embodiment can incorporate the descriptions in the text analysis support device and the text analysis support method of the present disclosure. For example, each of the above procedures can be read as "processing" instead of "procedure". Also, the text analysis support program of this embodiment may be recorded on a computer-readable recording medium, for example. The recording medium is, for example, a non-transitory computer-readable storage medium. The recording medium is not particularly limited, and examples include random access memory (RAM), read-only memory (ROM), hard disk (HD), flash memory (e.g., solid state drive (SSD), USB flash memory, SD / SDHC card, etc.), optical disk (e.g., CD-R / CD-RW, DVD-R / DVD-RW, BD-R / BD-RE, etc.), magneto-optical disk (MO), floppy (registered trademark) disk (FD), and the like. Further, the text analysis support program of this embodiment (also referred to as a programming product or a text analysis support program product, for example) may be in a form distributed from an external computer, for example. The "distribution" may be, for example, distribution via a communication network or distribution via a device connected by wire. The text analysis support program of this embodiment may be installed and executed on the distributed device, or may be executed without installation.

[0046] As described above, the present disclosure has been described with reference to the embodiments. However, the present disclosure is not limited to the above embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. And each embodiment can be combined with other embodiments as appropriate.

[0047] <Supplementary Note> Some or all of the above embodiments can be described as follows in the supplementary note, but are not limited thereto. (Supplementary Note 1) It includes a text data acquisition unit, a classification unit, a group word extraction unit, a label setting unit, and an output unit. The text data acquisition unit acquires text data including at least one partial text, The classification unit classifies each of the partial texts into at least one group based on words included in the partial text, and assigns group identification information to the partial text, The group word extraction unit extracts at least one group characteristic word based on words within the group included in the group and words included in the text data, The label setting unit sets, for each of the partial texts, a pair of the group identification information and the group characteristic word as a label, The output unit outputs analysis text data in which the label is associated with the partial text, a text analysis support device. (Appendix 2) The group word extraction unit compares words within the group with words included in the text data, estimates a feature amount of the words within the group, and extracts the group characteristic word based on the feature amount. The text analysis support device according to claim 1. (Appendix 3) The group word extraction unit estimates a corrected feature amount obtained by correcting the feature amount with reference to word weight data, and extracts the group characteristic word based on the corrected feature amount. The text analysis support device according to claim 2. (Appendix 4) including an emotion analysis unit, The emotion analysis unit analyzes the emotion of the partial text included in the text data, The label setting unit sets the emotion as an emotion label for the partial text, The output unit outputs analysis text data in which the emotion label is associated with the partial text. The text analysis support device according to any one of claims 1 to 3. (Appendix 5) The text data is data in a format including rows and columns, and the partial texts are stored row by row, 5. The text analysis support device according to claim 1, wherein the output unit adds a column adjacent to the row in which the partial text is stored, and stores the label in the added column. (Appendix 6) 6. The text analysis support device according to claim 1, wherein the text data includes text data in a free-form comment field of a questionnaire. (Appendix 7) The method includes a text data acquisition step, a classification step, a group word extraction step, a label setting step, and an output step, the text data acquiring step acquires text data including at least one partial text; the classifying step classifies each of the partial texts into at least one group based on words contained in the partial text, and assigns group identification information to the partial text; The group word extraction step extracts at least one group characteristic word based on group words included in the group and words included in the text data, the label setting step sets a pair of the group identification information and the group characteristic word as a label for each of the partial texts; The output step outputs analytical text data in which the label is associated with the partial text, and each of the steps is executed by a computer. (Appendix 8) 8. The text analysis support method according to claim 7, wherein the group word extraction step compares the words in the group with words contained in the text data to estimate features of the words in the group, and extracts the group characteristic words based on the features. (Appendix 9) 9. The text analysis support method according to claim 8, wherein the group word extraction step estimates corrected feature quantities obtained by correcting the feature quantities with reference to word weight data, and extracts the group characteristic words based on the corrected feature quantities. (Appendix 10) A sentiment analysis step is included. The sentiment analysis step analyzes the sentiment of the partial text included in the text data, The label setting step sets the sentiment as a sentiment label for the partial text, The output step outputs analysis text data in which the sentiment label is associated with the partial text. The text analysis support method according to any one of claims 7 to 9. (Appendix 11) The text data is data in a format including rows and columns, and the partial text is stored row by row. The output step adds a column adjacent to the row in which the partial text is stored, and stores the label in the added column. The text analysis support method according to any one of claims 7 to 10. (Appendix 12) The text data includes text data in the free description column of a questionnaire. The text analysis support method according to any one of claims 7 to 11. (Appendix 13) Including a text data acquisition procedure, a classification procedure, a group word extraction procedure, a label setting procedure, and an output procedure. The text data acquisition procedure acquires text data including at least one partial text. The classification procedure classifies the partial text into at least one group based on the words included in the partial text for each partial text, and assigns group identification information to the partial text. The group word extraction procedure extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data. The label setting procedure sets a combination of the group identification information and the group characteristic word as a label for each partial text. The output procedure outputs analysis text data in which the label is associated with the partial text. A text analysis support program for causing a computer to execute each of the above procedures. (Appendix 14) The group word extraction procedure according to claim 13 of the text analysis support program compares the in-group words with the words included in the text data, estimates the feature amount of the in-group words, and extracts the group feature words based on the feature amount. (Appendix 15) The group word extraction procedure according to claim 14 of the text analysis support program estimates a corrected feature amount obtained by correcting the feature amount with reference to word weight data, and extracts the group feature words based on the corrected feature amount. (Appendix 16) Including a sentiment analysis procedure, The sentiment analysis procedure analyzes the sentiment of the partial text included in the text data, The label setting procedure sets the sentiment as a sentiment label for the partial text, The output procedure outputs analysis text data in which the sentiment label is associated with the partial text according to any one of claims 13 to 15 of the text analysis support program. (Appendix 17) The text data is data in a format including rows and columns, and the partial text is stored row by row. The output procedure adds a column adjacent to the row in which the partial text is stored, and stores the label in the added column according to any one of claims 13 to 16 of the text analysis support program. (Appendix 18) The text data includes text data in the free description column of the questionnaire according to any one of claims 13 to 17 of the text analysis support program. (Appendix 19) Including a text data acquisition procedure, a classification procedure, a group word extraction procedure, a label setting procedure, and an output procedure. The text data acquisition procedure acquires text data including at least one partial text. The classification procedure classifies the partial text into at least one group based on the words included in the partial text for each partial text, and assigns group identification information to the partial text. The group word extraction procedure extracts at least one group feature word based on the words within the group included in the group and the words included in the text data. The label setting procedure sets, for each partial text, a pair of the group identification information and the group feature word as a label. The output procedure outputs analysis text data in which the label is associated with the partial text. A computer-readable recording medium recording a text analysis support program for causing a computer to execute each of the above procedures. (Appendix 20) The group word extraction procedure according to claim 19, wherein the group word extraction procedure compares the words within the group with the words included in the text data, estimates a feature amount of the words within the group, and extracts the group feature word based on the feature amount. (Appendix 21) The recording medium according to claim 20, wherein the group word extraction procedure estimates a corrected feature amount obtained by correcting the feature amount with reference to word weight data, and extracts the group feature word based on the corrected feature amount. (Appendix 22) Including a sentiment analysis procedure. The sentiment analysis procedure analyzes the sentiment of a partial text included in the text data. The label setting procedure sets the sentiment as a sentiment label for the partial text. The recording medium according to any one of claims 19 to 21, wherein the output procedure outputs analysis text data in which the sentiment label is associated with the partial text. (Appendix 23) The text data is data in a format including rows and columns, and the partial text is stored row by row. The recording medium according to any one of claims 19 to 22, wherein the output procedure adds a column adjacent to the row in which the partial text is stored, and stores the label in the added column. (Appendix 24) The recording medium according to any one of claims 19 to 23, wherein the text data includes text data in a free description field of a questionnaire.

Industrial Applicability

[0048] The text analysis support device of the present disclosure can classify, for example, partial texts included in text data into a predetermined pattern. Therefore, for example, by referring to a group identification number corresponding to a group of partial texts included in the text data, the text data can be used for analysis. Accordingly, the present disclosure is widely useful in various fields that utilize data analysis.

Explanation of Signs

[0049] 10 Text analysis support device 11 Text data acquisition unit 12 Classification unit 13 Group word extraction unit 14 Label setting unit 15 Output unit 101 CPU 102 Memory 103 Bus 104 Storage device 105 Input device 106 Output device 107 Communication device

Claims

1. A text analysis support device including a text data acquisition unit, a classification unit, a group word extraction unit, a label setting unit, and an output unit, wherein the text data acquisition unit acquires text data including at least one partial text, the classification unit classifies each of the partial texts into at least one group based on words included in the partial text, and assigns group identification information to the partial text, the group word extraction unit extracts at least one group feature word based on words included in the group and words included in the text data, the label setting unit sets, for each of the partial texts, a combination of the group identification information and the group feature word as a label, and the output unit outputs analysis text data in which the label is associated with the partial text.

2. The text analysis support device according to claim 1, wherein the group word extraction unit compares words included in the group with words included in the text data, estimates a feature amount of the words included in the group, and extracts the group feature word based on the feature amount.

3. The text analysis support device according to claim 2, wherein the group word extraction unit estimates a corrected feature amount obtained by correcting the feature amount with reference to word weight data, and extracts the group feature word based on the corrected feature amount.

4. including a sentiment analysis unit, wherein the sentiment analysis unit analyzes the sentiment of a partial text included in the text data, the label setting unit sets the sentiment as a sentiment label for the partial text, and the output unit outputs analysis text data in which the sentiment label is associated with the partial text.

5. wherein the text data is data in a format including rows and columns, the partial texts are stored row by row, and the output unit adds a column adjacent to the row in which the partial text is stored, and stores the label in the added column.

6. The text analysis support device according to any one of claims 1 to 3, wherein the text data includes text data in a free description field of a questionnaire.

7. including a text data acquisition step, a classification step, a group word extraction step, a label setting step, and an output step, The text data acquisition step acquires text data including at least one partial text, The classification step classifies each of the partial texts into at least one group based on the words included in the partial text, and assigns group identification information to the partial text, The group word extraction step extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data, The label setting step sets, for each of the partial texts, a pair of the group identification information and the group characteristic word as a label, The output step outputs analysis text data in which the label is associated with the partial text, and each of the steps is executed by a computer, a text analysis support method.

8. The group word extraction step compares the words within the group with the words included in the text data, estimates a feature amount of the words within the group, and extracts the group characteristic word based on the feature amount, the text analysis support method according to Claim 7.

9. Including a text data acquisition procedure, a classification procedure, a group word extraction procedure, a label setting procedure, and an output procedure, The text data acquisition procedure acquires text data including at least one partial text, The classification procedure classifies each of the partial texts into at least one group based on the words included in the partial text, and assigns group identification information to the partial text, The group word extraction procedure extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data, The label setting procedure sets, for each of the partial texts, a pair of the group identification information and the group characteristic word as a label, The output procedure outputs analysis text data in which the label is associated with the partial text, A text analysis support program for causing a computer to execute each of the procedures.

10. Including a text data acquisition procedure, a classification procedure, a group word extraction procedure, a label setting procedure, and an output procedure, The text data acquisition procedure acquires text data including at least one partial text, The classification procedure classifies the partial text into at least one group based on the words included in the partial text for each of the partial texts, and assigns group identification information to the partial text. The group word extraction procedure extracts at least one group characteristic word based on the words within the group included in the group and the words included in the text data. The label setting procedure sets, for each of the partial texts, a pair of the group identification information and the group characteristic word as a label. The output procedure outputs analysis text data in which the label is associated with the partial text. A computer-readable recording medium recording a text analysis support program for causing a computer to execute each of the above procedures.

Citation Information

Patent Citations

  • Students' ability analysis system

    JP2012194507A