Information Processing Apparatus, Information Processing Method, and Information Processing Program
The information processing apparatus addresses the challenge of estimating text readability by calculating the hiragana ratio in text and adjusting difficulty levels, resulting in improved learning experiences by matching text difficulty with user comprehension.
Patent Information
- Application Number
- JP2021092761
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-02
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2041-06-02
AI Technical Summary
Existing difficulty level estimation devices for text, particularly those analyzing picture books, struggle with accurate analysis when the text contains a high proportion of hiragana, affecting the estimation of readability.
An information processing apparatus that acquires text information, calculates the ratio of hiragana used in the text, and estimates the readability of the text based on this ratio, allowing for the adjustment of text difficulty levels according to user learning situations.
Enables accurate estimation and adjustment of text readability, improving the learning experience by ensuring that text difficulty aligns with the user's comprehension level.
Smart Images

Figure 0007698866000001 
Figure 0007698866000002 
Figure 0007698866000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Conventionally, there has been a technique for obtaining the difficulty level of text. The difficulty level estimation device described in Patent Document 1 performs morphological analysis on the text of a picture book that is the target for obtaining the difficulty level, and calculates the appearance frequency of each word constituting the text. Further, the difficulty level estimation device stores the occurrence probability of each word included in the text to which the difficulty level class is assigned for each difficulty level class of the text. The difficulty level estimation device estimates the likelihood that the text of the picture book belongs to the difficulty level class based on the calculated appearance frequency of each word and the stored occurrence probability. Then, the difficulty level estimation device estimates the difficulty level class with the highest likelihood among the estimated likelihoods as the text of the picture book.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The above-described difficulty level estimation device estimates the difficulty level class of the text of a picture book based on performing morphological analysis on the text of the picture book. However, morphological analysis may not be able to perform accurate analysis when there are relatively many hiragana in the text. That is, when the text of the picture book that is the target for estimating the difficulty level contains relatively many hiragana, the above-described difficulty level estimation device may not be able to accurately estimate the difficulty level class of the text.
[0005] Also, the difficulty level of text is related to the readability of that text. As an example, even when the number of Chinese characters contained in the text relatively increases as a user (e.g., an elementary school student) learns Chinese characters in an elementary school etc. (according to the grade), the user can read relatively difficult text. For this reason, it is required to estimate the difficulty level of the text, that is, the readability of the text, so that the user can read text according to the learning situation (comprehension level).
[0006] An object of the present invention is to provide an information processing apparatus, an information processing method, and an information processing program capable of obtaining the readability of text.
Means for Solving the Problems
[0007] An information processing apparatus according to one aspect includes a first acquisition unit that acquires text information in which text is recorded, a second acquisition unit that acquires the ratio of hiragana used in the text based on the text information acquired by the first acquisition unit, a calculation unit that calculates the readability of the text recorded in the text information based on the ratio of hiragana acquired by the second acquisition unit, and an output control unit that controls to output the readability of the text calculated by the calculation unit.
[0008] In an information processing method according to one aspect, a computer executes a first acquisition step of acquiring text information in which text is recorded, a second acquisition step of acquiring the ratio of hiragana used in the text based on the text information acquired by the first acquisition step, a calculation step of calculating the readability of the text recorded in the text information based on the ratio of hiragana acquired by the second acquisition step, and an output control step of controlling to output the readability of the text calculated by the calculation step.
[0009] An information processing program according to one aspect causes a computer to implement a first acquisition function for acquiring text information in which text is recorded, a second acquisition function for acquiring the ratio of hiragana used in the text based on the text information acquired by the first acquisition function, a calculation function for calculating the readability of the text recorded in the text information based on the ratio of hiragana acquired by the second acquisition function, and an output control function for controlling to output the readability of the text calculated by the calculation function.
Effect of the Invention
[0010] An information processing apparatus according to one aspect acquires text information in which text is recorded, acquires the ratio of hiragana used in the text based on the text information, calculates the readability of the text recorded in the text information based on the ratio of hiragana, and controls to output the calculated readability of the text, so that the readability of the text can be obtained. An information processing method and an information processing program according to one aspect can achieve the same effects as the information processing apparatus according to the above-described one aspect.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Mode for Carrying Out the Invention
[0012] Hereinafter, an embodiment of the present invention will be described.
[0013] [Outline of Information Processing Apparatus 1] First, the outline of the information processing apparatus 1 according to an embodiment will be described. FIG. 1 is a diagram for explaining the information processing apparatus 1 according to an embodiment.
[0014] The information processing apparatus 1 may be a computer such as a server, a desktop, a laptop, and a tablet. The information processing apparatus 1 acquires the readability of the text recorded in the text information 100 based on the text information 100 generated from various contents. The text information 100 may be information in which characters such as Chinese characters, hiragana, and katakana are recorded, for example. The readability of the text can also be considered as the difficulty level of the text, for example. In this case, the information processing apparatus 1 may acquire the readability of the text based on the ratio of the hiragana ratio described in the text. The hiragana ratio may be, for example, the ratio of hiragana to the characters (Chinese characters and hiragana, etc.) in the entire text or a part (predetermined range) of the text. Or, the hiragana ratio may be, for example, the ratio of Chinese characters and hiragana to the characters in the entire text or a part (predetermined range) of the text.
[0015] [Details of Information Processing Apparatus 1] Next, the details of the information processing apparatus 1 according to an embodiment will be described. FIG. 2 is a block diagram for explaining the information processing apparatus 1 according to an embodiment.
[0016] The information processing apparatus 1 includes a communication unit 21, a storage unit 22, a display unit 23, and a control unit 10. The communication unit 21, the storage unit 22, and the display unit 23 may be an embodiment of an "output unit". The control unit 10 may be configured by, for example, an arithmetic processing unit of the information processing apparatus 1. The control unit 10 may function as, for example, a first acquisition unit 11, a second acquisition unit 12, a third acquisition unit 13, a fourth acquisition unit 14, a calculation unit 15, a rewriting unit 16, a first specifying unit 17, a second specifying unit 18, and an output control unit 19.
[0017] The communication unit 21 can transmit and receive information to and from, for example, a device (external device) (not shown) outside the information processing apparatus 1. The communication unit 21 can communicate, for example, with external devices such as a server, a terminal (e.g., desktop, laptop, tablet, and smartphone, etc.), an OCR (Optical Character Recognition) device, and a voice acquisition device. The server and the terminal generate, for example, a file (text information) in which a sentence generated by the own device is described, and store a file (text information) in which sentences generated by the own device and other devices are described. When an image is generated based on a medium on which sentences are described, such as a book, a picture book, or a magazine, the OCR device recognizes the characters recorded in the image and generates text information. The voice acquisition device acquires voice with a microphone or the like, performs voice recognition based on the acquired voice, and generates text information.
[0018] The storage unit 22 stores, for example, various information, programs, and the like. The storage unit 22 may store text information, for example. In this case, the storage unit 22 may store, for example, text information generated by the information processing apparatus 1, or may store text information acquired from outside the information processing apparatus 1. The storage unit 22 may store various information generated by the information processing apparatus 1, such as information regarding the readability of the text calculated by the calculation unit 15, based on the control of the output control unit 19 described later. The storage unit 22 may store information acquired from outside the information processing apparatus 1, for example.
[0019] The display unit 23 displays, for example, various characters, symbols, images, and the like. The display unit 23 may display characters, symbols, images, and the like based on various information generated by the information processing apparatus 1, such as the readability of the text calculated by the calculation unit 15, based on the control of the output control unit 19 described later. The display unit 23 may display characters, symbols, images, and the like based on information acquired from outside the information processing apparatus 1, for example.
[0020] The first acquisition unit 11 acquires text information in which text is recorded. The first acquisition unit 11 may acquire, for example, the text information stored in the storage unit 22. Alternatively, the first acquisition unit 11 may acquire text information via the communication unit 21, for example. The text information may be generated based on a medium on which sentences such as books, picture books, magazines, etc. are described, or may be generated based on sentences such as daily newspapers, monthly reports, reports, and reports. Further, the text information may be generated based on the voice or the like uttered by the user. That is, the text information may be information in which various sentences are recorded, for example.
[0021] The second acquisition unit 12 acquires the ratio of hiragana used in the text based on the text information acquired by the first acquisition unit 11. The second acquisition unit 12 acquires, for example, the ratio of hiragana in the entire text or a part (predetermined range) of the text recorded in the text information. The predetermined range of the text may be, for example, an arbitrary range in the entire text, that is, text of a predetermined number of characters at an arbitrary position in the entire text. That is, the ratio of hiragana in the predetermined range of the text may be, for example, the ratio of hiragana to the text (characters) of a predetermined number of characters at an arbitrary position in the entire text. As an example, the predetermined number of characters may be about 500 to 3000 characters, and as the number of characters increases, it becomes possible to more accurately calculate the readability of the text by the calculation unit 15 described later. The ratio of hiragana may be, for example, the number of hiragana characters with respect to the characters (total number of characters or number of characters in the predetermined range) of the entire text or a part (predetermined range) of the text.
[0022] Further, the second acquisition unit 12 may acquire, for example, the ratio of Chinese characters to hiragana from among some of the texts recorded in the text information. That is, the ratio of hiragana used in the above-described text may be the ratio of Chinese characters to hiragana used in the text. A part of the text may be, for example, the text within the above-described predetermined range. The ratio of Chinese characters to hiragana may be, for example, the ratio of the number of Chinese characters to the number of hiragana with respect to the total number of characters (the total number of all characters or the number of characters within a predetermined range) of the entire text or a part (predetermined range) of the text.
[0023] The third acquisition unit 13 may acquire the average length of one sentence of the text based on the text information acquired by the first acquisition unit 11. In this case, when the second acquisition unit 12 acquires the ratio of hiragana (the ratio of Chinese characters and hiragana) from a part of the text (text within a predetermined range) as described above, the third acquisition unit 13 may acquire the average length of one sentence of that part of the text (text within a predetermined range). The average length of the text may be, for example, the arithmetic mean length of the sentences constituting the entire text or a part (predetermined range). Also, the average length of the text may be, for example, the average number of characters per sentence. Further, when acquiring the length of one sentence, the third acquisition unit 13 may, as an example, acquire the length (the number of characters in one sentence) without including symbols such as parentheses ("(", ")", "[", and "]"), question marks and exclamation marks ("?", "!", and "!?") and full stops (".") (delimiter characters). Therefore, the third acquisition unit 13 may acquire the average length (number of characters) of one sentence from a plurality of sentences based on the length (number of characters) of one sentence that does not include symbols such as delimiter characters.
[0024] The fourth acquisition unit 14 may perform morphological analysis based on the text information acquired by the first acquisition unit 11, and may acquire at least one of the Chinese word rate, Japanese native word rate, verb rate, and particle rate in the text based on the result of the analysis. That is, the fourth acquisition unit 14 performs morphological analysis on the text based on the text information. In this case, for example, when the second acquisition unit 12 described above acquires the ratio of hiragana (the ratio of kanji and hiragana) from a part of the text (text within a predetermined range), the fourth acquisition unit 14 may perform morphological analysis on that part of the text (text within a predetermined range). The fourth acquisition unit 14 may acquire at least one of the ratio of Chinese words in the text, the ratio of Japanese words in the text, the ratio of verbs in the text, and the ratio of auxiliary words in the text, for example, based on the result of the morphological analysis. In this case, the fourth acquisition unit 14 may acquire, for example, the number of Chinese words in the text as the ratio of Chinese words, the number of Japanese words in the text as the ratio of Japanese words, the number of verbs in the text as the ratio of verbs, and the number of auxiliary words in the text as the ratio of auxiliary words. As a specific example, the fourth specifying unit uses various known morphological analysis engines (for example, MeCab, etc.) to perform word segmentation on the text based on the text information, and acquires at least one of the Chinese word ratio, Japanese word ratio, verb ratio, and auxiliary word ratio.
[0025] The calculation unit 15 acquires the readability (text difficulty) of the text recorded in the text information based on the ratio of hiragana acquired by the second acquisition unit 12. The readability of the text is, for example, information that quantitatively indicates the degree of readability of the text. The calculation unit 15 may, for example, presume that the text is more readable when the ratio of hiragana acquired by the second acquisition unit 12 is relatively large (there are more hiragana). In other words, the calculation unit 15 may, for example, presume that the text is more difficult to read when the ratio of hiragana acquired by the second acquisition unit 12 is relatively small (there are fewer hiragana).
[0026] In this case, the calculation unit 15 may obtain the readability of the text according to the age (for example, the school year of schools such as elementary school, junior high school, and high school). When a threshold value of the ratio of hiragana for estimating the readability of the text according to the age is set, if the ratio of hiragana obtained by the second acquisition unit 12 is equal to or higher than the threshold value (if the number of hiragana is larger than the threshold value), the calculation unit 15 may estimate that the text is easy to read. In other words, if the ratio of hiragana obtained by the second acquisition unit 12 is less than the threshold value (if the number of hiragana is smaller than the threshold value), the calculation unit 15 may estimate that the text is difficult to read.
[0027] Note that a plurality of the above-described threshold values may be set. As an example, when two threshold values (the first threshold value > the second threshold value) are set, if the ratio of hiragana obtained by the second acquisition unit 12 is equal to or higher than the first threshold value, the calculation unit 15 may estimate that the text is easier to read. As another example, if the ratio of hiragana obtained by the second acquisition unit 12 is less than the first threshold value and equal to or higher than the second threshold value, the calculation unit 15 may estimate that the text is easy to read. As another example, if the ratio of hiragana obtained by the second acquisition unit 12 is less than the second threshold value, the calculation unit 15 may estimate that the text is difficult to read.
[0028] The calculation unit 15 calculates the readability of the text recorded in the text information based on the ratio of hiragana obtained by the second acquisition unit 12. Further, the calculation unit 15 may calculate the readability of the text based on, for example, the ratio of kanji and hiragana obtained by the second acquisition unit 12. The calculation unit 15 may substitute the ratio of hiragana (the ratio of kanji and hiragana) obtained by the second acquisition unit 12 into a preset arithmetic expression to calculate the readability of the text. Although the calculation unit 15 can use various arithmetic expressions as the arithmetic expression, for example, the readability of the text may be calculated using Expression (1) described later. Here, as for calculating the readability of the text, for example, the calculation unit 15 may include obtaining the readability of the text as described above, that is, estimating the readability using a threshold value or the like.
[0029] The calculation unit 15 may calculate the readability of the text according to the age based on the ratio of hiragana (the ratio of kanji and hiragana) acquired by the second acquisition unit 12. For example, the calculation unit 15 may estimate the readability of the text according to the age by classifying the value (the value indicating the readability of the text) obtained by substituting the ratio of hiragana (the ratio of kanji and hiragana) into an arithmetic expression (for example, Expression (1) described later). The calculation unit 15 may perform classification using, for example, a preset threshold value. As an example, when the age is 8 years old (the school year is the third grade of elementary school), it is set that the text is easy to read if it is equal to or greater than the third threshold value and less than the fourth threshold value. When the age is 9 years old (the school year is the fourth grade of elementary school), it is set that the text is easy to read if it is equal to or greater than the fourth threshold value and less than the fifth threshold value. When the age is 10 years old (the school year is the fifth grade of elementary school), it is set that the text is easy to read if it is equal to or greater than the fifth threshold value and less than the sixth threshold value (the sixth threshold value > the fifth threshold value > the fourth threshold value > the third threshold value). In such an example, for example, when the value obtained by substituting into an arithmetic expression (for example, Expression (1)) is equal to or greater than the fourth threshold value and less than the fifth threshold value, the calculation unit 15 may estimate that the text is easy to read (appropriate) for a child of 9 years old (the school year is the fourth grade of elementary school), estimate that the text is difficult to read (difficult) for a child of 8 years old (the school year is the third grade of elementary school), and estimate that the text is easy to read (easy) for a child of 10 years old (the school year is the fifth grade of elementary school).
[0030] In addition to or in addition to calculating the readability of the text according to the age as described above, the calculation unit 15 may calculate the readability of the text according to the age by using at least one of the difficulty information regarding the difficulty of the vocabulary used in the text and the content of the Chinese characters to be learned according to the school grade (learning Chinese character information). The difficulty information may be, for example, information associating the vocabulary with the difficulty of the vocabulary. The learning Chinese character information may be, for example, information associating the school grades of primary schools, middle schools, high schools, etc. with the Chinese characters learned in each grade. In this case, the calculation unit 15 may calculate the readability of the text according to the age (grade) based on, for example, the ratio of hiragana (the ratio of Chinese characters and hiragana) acquired by the second acquisition unit 12 and at least one of the difficulty information and the learning Chinese character information.
[0031] Further, the calculation unit 15 may calculate the readability of the text based on the ratio acquired by the second acquisition unit 12 and the average length of one sentence acquired by the third acquisition unit 13. That is, the calculation unit 15 may calculate the readability of the text by using, for example, the average length of one sentence acquired by the third acquisition unit 13 in addition to the ratio of hiragana (the ratio of Chinese characters and hiragana) acquired by the second acquisition unit 12. The calculation unit 15 may, for example, substitute the ratio of hiragana (the ratio of Chinese characters and hiragana) acquired by the second acquisition unit 12 and the average length of one sentence acquired by the third acquisition unit 13 into a preset arithmetic expression to calculate the readability of the text. The calculation unit 15 may use various arithmetic expressions as the arithmetic expression, for example, and may calculate the readability of the text by using, for example, Expression (1) described later.
[0032] The calculation unit 15 may calculate the readability of the text based on the ratio acquired by the second acquisition unit 12, the average length of one sentence acquired by the third acquisition unit 13, and at least one of the Chinese word rate, Japanese native word rate, verb rate, and particle rate acquired by the fourth acquisition unit 14. That is, for example, in addition to the ratio of hiragana (the ratio of Chinese characters and hiragana) acquired by the second acquisition unit 12 and the average length of one sentence acquired by the third acquisition unit 13, the calculation unit 15 may calculate the readability of the text using at least one of the Chinese word rate, Japanese native word rate, verb rate, and particle rate acquired by the fourth acquisition unit 14. The calculation unit 15 may, for example, substitute the ratio of hiragana (the ratio of Chinese characters and hiragana) acquired by the second acquisition unit 12, the average length of one sentence acquired by the third acquisition unit 13, and at least one of the Chinese word rate, Japanese native word rate, verb rate, and particle rate acquired by the fourth acquisition unit 14 into a preset arithmetic expression to calculate the readability of the text. The calculation unit 15 may use various arithmetic expressions as the arithmetic expression. For example, the readability of the text may be calculated using the formula (1) described later.
[0033] If the ratio (the ratio of hiragana (the ratio of Chinese characters and hiragana)) acquired by the second acquisition unit 12 is defined as a, the average length of one sentence acquired by the third acquisition unit 13 is defined as b, the Chinese word rate acquired by the fourth acquisition unit 14 is defined as c, the Japanese native word rate is defined as d, the verb rate is defined as e, and the particle rate is defined as f, the readability YL of the text may be calculated by the following formula (1). YL = 60 * a + 0.24 * b + 126 * c + 42 * d + 145 * e + 44 * f …(1)
[0034] Each term of the above formula (1) may increase or decrease according to the values obtained by the above-described second to fourth acquisition units 12, 13, and 14. That is, as an example, when the calculation unit 15 calculates the readability of the text using only the ratio a acquired by the second acquisition unit 12, the readability YL of the text may be calculated by the following formula (2) based on formula (1). YL = 10 * a …(2)
[0035] Similarly, as an example, when the calculation unit 15 calculates the readability of the text using the ratio a acquired by the second acquisition unit 12 and the average length b of one sentence acquired by the third acquisition unit 13, the readability YL of the text may be calculated by the following formula (3) based on formula (1). YL = 60 * a + 0.24 * b …(3)
[0036] Similarly, as an example, when the calculation unit 15 calculates the readability of the text using the ratio a acquired by the second acquisition unit 12, the average length b of one sentence acquired by the third acquisition unit 13, and the Chinese character rate c acquired by the fourth acquisition unit 14, the readability YL of the text may be calculated by the following formula (4) based on formula (1). YL = 60 * a + 0.24 * b + 126 * c …(4)
[0037] The coefficients of the above-mentioned formulas (1) to (4) are not limited to the above-mentioned example and may be appropriately set values.
[0038] Note that the calculation unit 15 may, for example, use machine learning, deep learning, etc. to set the coefficients of the above-mentioned formula (1) etc. and introduce new terms to improve the accuracy of the readability YL. In this case, the calculation unit 15 may, for example, construct a learning model adapted to the education of the user (as an example, children's reading education, etc.) using information obtained from various services. Also, the calculation unit 15 is not limited to the example of calculating the readability YL of the text using the arithmetic formula as described above. For example, the readability YL may be calculated using machine learning, deep learning, etc. As an example, the calculation unit 15 may calculate the readability YL based on a learning model generated by learning by associating the text with the readability of the text and the information acquired by the second to fourth acquisition units 12, 13, 14. Also, the calculation unit 15 may, for example, use natural language processing technology, etc. to perform a difficulty determination incorporating the complexity of the flow of the entire text. Further, the calculation unit 15 may evaluate an index such as psychological readability by reflecting the undulations of the entire text. In this case, the calculation unit 15 may perform positive / negative determination in units such as words and read the undulations of the story in time series.
[0039] The information processing apparatus 1 can use the readability YL of the text acquired as described above for various purposes. The control unit 10 may rewrite the text according to the reading ability of the user (for example, a child) (into a plurality of patterns for each readability YL). As an example, the control unit 10 may be used for reference books, textbooks, articles, and the like. The control unit 10 may gradually increase the reading comprehension ability of the user, for example, by publishing media such as books for each readability YL. The control unit 10 may support the writing activity of the user (for example, a writer or a writer) by acquiring in real time the readability YL (difficulty level) of the text (text) created by the user. In this case, the control unit 10 can be applied to various uses of various texts in general, not limited to writing, and can be expanded to various services for children and various services for the elderly. Further, the control unit 10 may, for example, convert the user's voice into text and analyze it to acquire the habits of the user's speaking style and the vocabulary used, and give advice (for brushing up the presentation) so that a presentation that is easy for other users to understand can be made. The control unit 10 can also improve the writing ability of the user by using the readability YL. In this case, the control unit 10 can determine, for example, the complex logical structure and the level of the vocabulary used in the text, and can encourage further improvement. Further, the control unit 10 can also determine elements such as the sentence structure in the text and can encourage improvement. The control unit 10 may, for example, provide reading guidance to the user by using the readability level YL. In this case, the control unit 10 may create a problem (test) that can measure by separating a complex group of elements that constitute the user's reading ability. Also, the control unit 10 can lead to an approach that can eliminate the content that is difficult for the user in the shortest time by measuring the strength not only by the score of the problem (test) but also by the user's ability. Hereinafter, a specific example of the control unit 10 when using the readability YL of the above-described text for various applications will be described.
[0040] The rewriting unit 16 may rewrite the text recorded in the text information acquired by the first acquisition unit 11 so as to have a level different from the readability level calculated by the calculation unit 15 based on the first correspondence information in which the text and the readability level corresponding to the text are associated. That is, the rewriting unit 16 may rewrite the text based on the first correspondence information so as to increase (or decrease) the readability level calculated by the calculation unit 15. As a specific example, there is a textbook used by fourth-grade elementary school children, and when the children feel the textbook is difficult, the rewriting unit 16 may rewrite the text based on the textbook to a level that can be read by third-grade elementary school children based on the first correspondence information. Similarly, as a specific example, there is a textbook used by fourth-grade elementary school children, and when the children feel the textbook is easy, the rewriting unit 16 may rewrite the text based on the textbook to a level that can be read by fifth-grade elementary school children based on the first correspondence information.
[0041] The first correspondence information may record, for example, Chinese characters, idioms, expressions, and learning contents (for example, words, etc.) learned according to the age (school grade). When the rewriting unit 16 rewrites a textbook used by a fourth - grade elementary school student to a level that can be read by a third - grade elementary school student, for example, as in the above - mentioned example, based on the first correspondence information, for the Chinese characters learned in the fourth grade of elementary school, they are rewritten into hiragana, and for the Chinese characters learned in the first to third grades of elementary school, they are maintained. Thus, the textbook (text) used by the fourth - grade elementary school student may be rewritten to the third - grade elementary school level. Similarly, when the rewriting unit 16 rewrites a textbook used by a fourth - grade elementary school student to the fifth - grade elementary school level, for example, as in the above - mentioned example, based on the first correspondence information, for the Chinese characters learned in the fifth grade of elementary school, they are rewritten from hiragana to Chinese characters, and for the Chinese characters learned in the first to fourth grades of elementary school, they are maintained. Thus, the textbook (text) used by the fourth - grade elementary school student may be rewritten to the fifth - grade elementary school level.
[0042] The first specifying unit 17 may specify advice for a user to raise the readability level calculated by the calculation unit 15 based on the second correspondence information that associates a text, the readability level corresponding to the text, and advice for the user corresponding to the readability level. As a specific example, when a fourth - grade elementary school student writes a composition and the readability of the written text is calculated by the calculation unit 15 to be at the third - grade elementary school level, the first specifying unit 17 may specify advice based on the second correspondence information so as to raise the level of the composition of the child to the fourth - grade elementary school level.
[0043] The second correspondence information may record, for example, Chinese characters, idioms, expressions, and learning contents (such as words, etc.) learned according to age (school grade), and advice for raising the level according to age (school grade). In this case, the advice may be, for example, when raising the readability level corresponding to the third - grade elementary school student to the readability level corresponding to the fourth - grade elementary school student, if the Chinese characters learned in the fourth grade of elementary school in the text are written in hiragana, the content may be to teach the Chinese characters of the hiragana. The advice may be, for example, in a form of displaying an article on the display unit 23 of the terminal used by the user, and a form of outputting voice from the speaker of the terminal used by the user, etc.
[0044] That is, for example, when the readability level of an article written by a fourth - grade elementary school student is at the third - grade elementary school level and kanji at the fourth - grade elementary school level are written in hiragana in the article, the first specifying unit 17 may specify advice to prompt rewriting the hiragana into kanji based on the second corresponding information.
[0045] Note that the first specifying unit 17 is not limited to the example of giving advice regarding composition (creation of an article) as described above, and may give various advice based on the result (text) of recognizing the voice spoken by the user. In this case, the first specifying unit 17 may, for example, give advice on the user's presentation.
[0046] The second specifying unit 18 may specify a problem corresponding to the readability level calculated by the calculation unit 15 based on third corresponding information associating a readability level with a problem corresponding to the readability level. As a specific example, when giving a problem (test) to a fourth - grade elementary school student, the second specifying unit 18 may specify a problem based on the third corresponding information so as to be able to give a problem (test) suitable for a fourth - grade elementary school student.
[0047] The third corresponding information may record, for example, kanji, idiomatic expressions, and turns of phrase to be learned according to age (school grade), and problems described in kanji, idiomatic expressions, and turns of phrase corresponding to age (school grade). That is, for example, when the readability level of the text of a fourth - grade elementary school student is at the third - grade elementary school level, the second specifying unit 18 specifies a problem that is a problem of the content learned by fourth - grade elementary school students but is described in kanji and the like learned by third - grade elementary school students based on the third corresponding information.
[0048] The output control unit 19 controls the output unit to output the readability of the text calculated by the calculation unit 15. The output unit may be, for example, the communication unit 21, the storage unit 22, the display unit 23, etc. That is, the output control unit 19 may control the communication unit 21 to transmit, for example, information regarding the readability of the text calculated by the calculation unit 15 to an external device (such as a server and a terminal, etc.). The output control unit 19 may control the storage unit 22 to store, for example, information regarding the readability of the text calculated by the calculation unit 15. The output control unit 19 may control the display unit 23 to display, for example, the readability of the text calculated by the calculation unit 15.
[0049] The output control unit 19 may control the output unit to output the text rewritten by the rewriting unit 16. That is, the output control unit 19 may control the communication unit 21, the storage unit 22, and the display unit 23 (output unit) to output the text rewritten by the rewriting unit 16 in the same manner as the output regarding the "readability of the text" described above.
[0050] The output control unit 19 may control the output unit to output the advice specified by the first specifying unit 17. That is, the output control unit 19 may control the communication unit 21, the storage unit 22, and the display unit 23 (output unit) to output the advice specified by the first specifying unit 17 in the same manner as in the above-described case.
[0051] The output control unit 19 may control the output unit to output the problem specified by the second specifying unit 18. That is, the output control unit 19 may control the communication unit 21, the storage unit 22, and the display unit 23 (output unit) to output the problem specified by the second specifying unit 18 in the same manner as in the above-described case.
[0052] Here, as a specific example, the control unit 10 (for example, the first acquisition unit 11, the second acquisition unit 12, the third acquisition unit 13, the fourth acquisition unit 14, the calculation unit 15, the rewriting unit 16, the first specifying unit 17, the second specifying unit 18, and the output control unit 19) may be configured to output the readability of the text and the like based on the text information by using an API (Application Programing Interface). That is, the control unit 10 performs morphological analysis and analysis of the text based on text information (information of text described in Japanese) with a predetermined number of characters or more, or voice (voice information) or image (image information) that can be converted into text information, and outputs various data regarding the readability YL of the text and the difficulty level of the sentence including it. That is, as an example, the readability YL, the details of various parameters used for calculating the readability YL, and the vocabulary difficulty level, etc. are output. In addition to this, the control unit 10 may also output data such as the composition and characteristics of the sentence, which do not fall within the category of the difficulty level of simple text. For example, the control unit 10 may output the flow of the sentence expressed as time-series data of numerical values, various tendencies such as the "solidity" of the sentence, and the genre of the sentence (for example, broad classification such as story and non-fiction, and detailed genres such as "love", "adventure", and "school stories", etc.).
[0053] [Information Processing Method] Next, an information processing method according to an embodiment will be described.
[0054] First, an information processing method according to an embodiment, which is a process of obtaining the readability of text, will be described. FIG. 3 is a first flowchart for explaining an information processing method according to an embodiment.
[0055] In step ST101, the first acquisition unit 11 acquires text information in which the text is recorded.
[0056] In step ST102, the second acquisition unit 12 acquires the ratio of hiragana used in the text based on the text information acquired in step ST101. In this case, the second acquisition unit 12 may acquire the ratio of kanji and hiragana used in the text. Here, the second acquisition unit 12 may acquire the ratio of hiragana (the ratio of kanji and hiragana) based on all or part of the text based on the text information.
[0057] In step ST103, the third acquisition unit 13 acquires the average length of one sentence of the text based on the text information acquired in step ST101. In this case, the third acquisition unit 13 may acquire the average length of one sentence, for example, from within the range of the text for which the hiragana ratio (the ratio of kanji and hiragana) is acquired in step ST102.
[0058] In step ST104, the fourth acquisition unit 14 performs morphological analysis based on the text information acquired in step ST101, and acquires at least one of the Chinese word ratio, Japanese native word ratio, verb ratio, and particle ratio in the text based on the result of the analysis. In this case, the fourth acquisition unit 14 may acquire at least one of the Chinese word ratio, Japanese native word ratio, verb ratio, and particle ratio, for example, from within the range of the text for which the hiragana ratio (the ratio of kanji and hiragana) is acquired in step ST102, or from within the range of the text for which the average length of one sentence is acquired in step ST103.
[0059] In step ST105, the calculation unit 15 calculates the readability of the text recorded in the text information acquired in step ST101 based on the ratio (hiragana ratio or the ratio of kanji and hiragana) acquired in step ST102. In this case, the calculation unit 15 may calculate the readability of the text according to the age based on the ratio acquired in step ST102. Further, the calculation unit 15 may calculate the readability of the text based on the ratio acquired in step ST102 and the average length of one sentence acquired in step ST103. Further, the calculation unit 15 may calculate the readability of the text based on the ratio acquired in step ST102, the average length of one sentence acquired in step ST103, and at least one of the Chinese word ratio, Japanese native word ratio, verb ratio, and particle ratio acquired in step ST104. If the ratio acquired in step ST102 is a, the average length of one sentence acquired in step ST103 is b, the Chinese word ratio acquired in step ST104 is c, the Japanese native word ratio is d, the verb ratio is e, and the particle ratio is f, the calculation unit 15 may calculate the readability YL of the text by the following formula (5). YL = 60*a + 0.24*b + 126*c + 42*d + 145*e + 44*f …(5)
[0060] In step ST106, the output control unit 19 controls the output unit to output the readability of the text calculated in step ST105. The output unit may be, for example, the communication unit 21, the storage unit 22, the display unit 23, etc.
[0061] Next, an information processing method according to an embodiment will be described, which is a process of outputting various information using the readability of text. FIG. 4 is a second flowchart for explaining an information processing method according to an embodiment.
[0062] The information processing apparatus 1 can generate various information by performing the following processes of steps ST201 to ST203 using the readability YL of the text calculated in step ST104 of FIG. 3, and output the information in step ST204. Hereinafter, a specific description will be given.
[0063] In step ST201, the rewriting unit 16 rewrites the text recorded in the text information acquired in step ST101 of FIG. 3 based on the first correspondence information so that the level is different from the level of the readability YL calculated in step ST104. Here, the first correspondence information is, for example, information in which a text is associated with a readability level corresponding to the text.
[0064] In step ST202, the first specifying unit 17 specifies advice for the user to increase the level of the readability YL calculated in step ST104 based on the second correspondence information. Here, the second correspondence information is, for example, information in which a text is associated with a readability level corresponding to the text and advice for the user corresponding to the readability level.
[0065] In step ST203, the second specifying unit 18 specifies a problem (test) corresponding to the level of the readability YL calculated in step ST104 based on the third correspondence information. Here, the third correspondence information is information associating, for example, a readability level and a problem corresponding to the readability level.
[0066] Note that the processes of steps ST201 to ST203 described above are not limited to the example where all of them are performed in this order. For example, all or part of the processes of steps ST201 to ST203 may be performed in an arbitrary order.
[0067] In step ST204, the output control unit 19 controls the output unit to output the results of the processes of step ST201, step ST202, and step ST203, respectively. That is, the output control unit 19 may control the output unit to output, for example, the text rewritten in step ST201. The output control unit 19 may control the output unit to output, for example, the advice specified in step ST202. The output control unit 19 may control the output unit to output, for example, the problem specified in step ST203. The output unit may be, for example, the communication unit 21, the storage unit 22, the display unit 23, or the like.
[0068] In this specification, the term "information" is used, but the term "information" can be replaced with "data", and the term "data" can be replaced with "information".
[0069] Each part of the information processing apparatus described above may be realized as a function of a computer's arithmetic processing unit or the like. That is, the first acquisition unit, the second acquisition unit, the third acquisition unit, the fourth acquisition unit, the calculation unit, the rewriting unit, the first specifying unit, the second specifying unit, and the output control unit (control unit) of the information processing apparatus may be realized as a first acquisition function, a second acquisition function, a third acquisition function, a fourth acquisition function, a calculation function, a rewriting function, a first specifying function, a second specifying function, and an output control function (control function) by a computer's arithmetic processing unit or the like, respectively. The information processing program can cause a computer to implement each of the functions described above. The information processing program may be recorded, for example, on a non-transitory computer-readable recording medium such as an external memory or an optical disk. Also, as described above, each part of the information processing apparatus may be implemented by an arithmetic processing unit or the like of a computer. The arithmetic processing unit or the like is configured by, for example, an integrated circuit or the like. For this reason, each part of the information processing apparatus may be implemented as a circuit constituting the arithmetic processing unit or the like. That is, the first acquisition unit, the second acquisition unit, the third acquisition unit, the fourth acquisition unit, the calculation unit, the rewriting unit, the first specifying unit, the second specifying unit, and the output control unit (control unit) of the information processing apparatus may be implemented as a first acquisition circuit, a second acquisition circuit, a third acquisition circuit, a fourth acquisition circuit, a calculation circuit, a rewriting circuit, a first specifying circuit, a second specifying circuit, and an output control circuit (control circuit) constituting an arithmetic processing unit or the like of a computer. Also, the communication unit, the storage unit, and the display unit (output unit) of the information processing apparatus may be implemented, for example, as a communication function, a storage function, and a display function (output function) including functions such as an arithmetic processing unit. Also, the communication unit, the storage unit, and the display unit (output unit) of the information processing apparatus may be implemented as a communication circuit, a storage circuit, and a display circuit (output circuit) by being configured by, for example, an integrated circuit or the like. Also, the communication unit, the storage unit, and the display unit (output unit) of the information processing apparatus may be configured as a communication device, a storage device, and a display device (output device) by being configured by, for example, a plurality of devices.
[0070] [Aspects and Effects of the Present Embodiment] Next, an aspect of the present embodiment and the effects exhibited by each aspect will be described. Note that the effects described below are examples, and the effects exhibited by each aspect are not limited to those described below.
[0071] (Aspect 1) An information processing apparatus according to one aspect includes a first acquisition unit that acquires text information in which text is recorded, a second acquisition unit that acquires the ratio of hiragana used in the text based on the text information acquired by the first acquisition unit, a calculation unit that calculates the readability of the text recorded in the text information based on the ratio of hiragana acquired by the second acquisition unit, and an output control unit that controls to output the readability of the text calculated by the calculation unit. Accordingly, the information processing apparatus can acquire the readability of the text.
[0072] (Aspect 2) In an information processing apparatus according to one aspect, the second acquisition unit may acquire the ratio of kanji and hiragana from among a part of the text recorded in the text information, and the calculation unit may calculate the readability of the text according to the age based on the ratio of kanji and hiragana acquired by the second acquisition unit. It is considered that the kanji learned according to the age are different. That is, it is considered that the number of kanji learned increases as the age increases. Therefore, the information processing apparatus can acquire the readability of the text according to the age by considering the ratio of kanji and hiragana in the text.
[0073] (Aspect 3) An information processing apparatus according to one aspect includes a third acquisition unit that acquires the average length of one sentence of the text based on the text information acquired by the first acquisition unit, and the calculation unit may calculate the readability of the text based on the ratio acquired by the second acquisition unit and the average length of one sentence acquired by the third acquisition unit. Generally, a user may feel that it is difficult to read text as the length of one sentence becomes longer. In particular, for children such as elementary school students, the possibility of feeling that it is difficult to read text increases as the length of one sentence becomes longer. The information processing apparatus can acquire the readability of the text by considering the average length of one sentence in addition to the ratio of hiragana (the ratio of kanji and hiragana) in the text.
[0074] (Aspect 4) An information processing apparatus according to one aspect performs morphological analysis based on text information acquired by a first acquisition unit, and includes a fourth acquisition unit that acquires at least one of a Chinese character ratio, a Japanese word ratio, a verb ratio, and a particle ratio in the text based on the result of the analysis. The calculation unit may calculate the readability of the text based on the ratio acquired by the second acquisition unit, the average length of one sentence acquired by the third acquisition unit, and at least one of the Chinese character ratio, Japanese word ratio, verb ratio, and particle ratio acquired by the fourth acquisition unit. Generally, it is considered that the readability of a text changes according to the Chinese character ratio, Japanese word ratio, verb ratio, and particle ratio of the text. The information processing apparatus can acquire the readability of the text by considering at least one of the Chinese character ratio, Japanese word ratio, verb ratio, and particle ratio in addition to the ratio of hiragana (the ratio of Chinese characters and hiragana) in the text and the average length of one sentence.
[0075] (Aspect 5) In an information processing apparatus according to one aspect, when the ratio acquired by the second acquisition unit is a, the average length of one sentence acquired by the third acquisition unit is b, the Chinese character ratio acquired by the fourth acquisition unit is c, the Japanese word ratio is d, the verb ratio is e, and the particle ratio is f, the readability YL of the text may be calculated by the following formula (6). YL = 60 * a + 0.24 * b + 126 * c + 42 * d + 145 * e + 44 * f …(6) The information processing apparatus can calculate the readability of the text by various calculation formulas. For example, the readability of the text can be acquired based on the above-described formula (6).
[0076] (Aspect 6) An information processing apparatus according to one aspect includes a rewriting unit that rewrites the text recorded in the text information acquired by the first acquisition unit so that the level of readability calculated by the calculation unit is different from the level of readability corresponding to the text based on the first correspondence information associated with the text and the readability level corresponding thereto. The output control unit may control to output the text rewritten by the rewriting unit. Generally, when calculating the readability of text as described above, there may be a desire to change the readability level of the text according to the user. For example, in a school or the like, even within the same grade, there may be cases where the learning levels of children are different. That is, for example, there may be a case where child A in the same class can read and understand a textbook, but for another child B, the textbook is difficult to read and difficult to understand. By including a rewriting unit, the information processing apparatus can change the level of readability of the text to another level of readability (increase or decrease the level of readability). That is, the information processing apparatus can output one or texts with different readability levels from one text. Therefore, as an example, the information processing apparatus can output text according to the level of a child from one textbook, and can assist the learning of the child.
[0077] (Aspect 7) An information processing apparatus according to one aspect includes a first specifying unit that specifies advice for a user who raises the level of readability calculated by the calculation unit based on the second correspondence information associating the text, the readability level corresponding to the text, and the advice for the user corresponding to the readability level. The output control unit may control to output the advice specified by the first specifying unit. As an example, when a child in a school or the like writes a composition, they may not know how to write the composition because they are not good at it, and the Chinese characters and expressions used in the text may be at a level lower than the grade to which the child belongs. The information processing device can, for example, obtain a readability level based on a text created by a user and output advice for improving the level of the text according to that level. Therefore, as an example, the information processing device can output advice according to the level of a child and assist the child's learning.
[0078] (Aspect 8) An information processing device of one aspect includes a second specifying unit that specifies a problem corresponding to the readability level calculated by a calculation unit based on third correspondence information associating a readability level with a problem corresponding to the readability level, and an output control unit may control to output the problem specified by the second specifying unit. Generally, in a school or the like, there are cases where the learning levels of children are different even in the same grade. In such a case, when a problem (test) C is given to children A and B in the same class, child A can understand the problem C, but for child B, the problem C may be difficult to read and they may not be able to answer it. The information processing device can, for example, prepare third correspondence information regarding problems corresponding to a plurality of readability levels for one task content, and when obtaining the readability level of the text created by the user as described above, can output a problem corresponding to that user based on the obtained readability level and the third correspondence information. Therefore, as an example, the information processing device can output problems according to the level of a child from one textbook and assist the child's learning.
[0079] (Aspect 9) In an information processing method according to one aspect, a computer executes a first acquisition step of acquiring text information in which text is recorded, a second acquisition step of acquiring the ratio of hiragana used in the text based on the text information acquired in the first acquisition step, a calculation step of calculating the readability of the text recorded in the text information based on the ratio of hiragana acquired in the second acquisition step, and an output control step of controlling to output the readability of the text calculated in the calculation step. The information processing method can achieve the same effects as the information processing apparatus according to the above-described one aspect.
[0080] (Aspect 10) An information processing program according to one aspect causes a computer to realize a first acquisition function of acquiring text information in which text is recorded, a second acquisition function of acquiring the ratio of hiragana used in the text based on the text information acquired by the first acquisition function, a calculation function of calculating the readability of the text recorded in the text information based on the ratio of hiragana acquired by the second acquisition function, and an output control function of controlling to output the readability of the text calculated by the calculation function. The information processing program can achieve the same effects as the information processing apparatus according to the above-described one aspect.
Explanation of Signs
[0081] 1 Information processing apparatus 10 Control unit 11 First acquisition unit 12 Second acquisition unit 13 Third acquisition unit 14 Fourth acquisition unit 15 Calculation unit 16 Rewriting unit 17 First specifying unit 18 Second specifying unit 19 Output control unit 21 Communication unit 22 Storage unit 23 Display unit
Claims
1. A first acquisition unit that acquires text information in which text is recorded; A second acquisition unit that acquires the ratio of hiragana to kanji used in the text based on the text information acquired by the first acquisition unit; Based on at least one of the ratio of hiragana to kanji acquired by the second acquisition unit, the difficulty information associating the vocabulary used in the text with the difficulty of the vocabulary, and the learning kanji information associating the school years of elementary school, junior high school, and high school with the kanji learned in each school year, a calculation unit that calculates the readability of the text recorded in the text information; An output control unit that controls to output the readability of the text calculated by the calculation unit; An information processing apparatus comprising:
2. The calculation unit calculates the readability of the text according to the age based on the ratio of kanji to hiragana acquired by the second acquisition unit The information processing apparatus according to claim 1.
3. Comprising a third acquisition unit that acquires the average length of one sentence of the text based on the text information acquired by the first acquisition unit, The calculation unit calculates the readability of the text based on the ratio acquired by the second acquisition unit and the average length of one sentence acquired by the third acquisition unit. The information processing apparatus according to claim 1 or 2.
4. Comprising a fourth acquisition unit that performs morphological analysis based on the text information acquired by the first acquisition unit and acquires at least one of the Chinese word rate, Japanese native word rate, verb rate, and particle rate in the text based on the result of the analysis, The calculation unit calculates the readability of the text based on the ratio acquired by the second acquisition unit, the average length of one sentence acquired by the third acquisition unit, and at least one of the Chinese word rate, Japanese native word rate, verb rate, and particle rate acquired by the fourth acquisition unit. The information processing apparatus according to claim 3.
5. When the ratio acquired by the second acquisition unit is a, the average length of one sentence acquired by the third acquisition unit is b, the Chinese word rate acquired by the fourth acquisition unit is c, the Japanese native word rate is d, the verb rate is e, and the particle rate is f, the calculation unit calculates the readability YL of the text by the following formula YL = 60 * a + 0.24 * b + 126 * c + 42 * d + 145 * e + 44 * f The information processing apparatus according to claim 4.
6. Based on the first correspondence information associating text with a readability level corresponding to the text, a rewriting unit is provided to rewrite the text recorded in the text information acquired by the first acquisition unit so as to have a level different from the readability level calculated by the calculation unit. The output control unit controls to output the text rewritten by the rewriting unit. The information processing apparatus according to any one of claims 1 to 5.
7. Based on the second correspondence information associating text with a readability level corresponding to the text and advice for the user corresponding to the readability level, a first specifying unit is provided to specify advice for the user to raise the readability level calculated by the calculation unit. The output control unit controls to output the advice specified by the first specifying unit. The information processing apparatus according to any one of claims 1 to 6.
8. Based on the third correspondence information associating a readability level with a problem corresponding to the readability level, a second specifying unit is provided to specify a problem corresponding to the readability level calculated by the calculation unit. The output control unit controls to output the problem specified by the second specifying unit. The information processing apparatus according to any one of claims 1 to 7.
9. A computer A first acquisition step of acquiring text information in which text is recorded; A second acquisition step of acquiring the ratio of hiragana to kanji used in the text based on the text information acquired in the first acquisition step; Based on at least one of the ratio of hiragana to kanji acquired in the second acquisition step, the difficulty information associating the vocabulary used in the text with the difficulty of the vocabulary, and the learning kanji information associating the school years of elementary school, junior high school, and high school with the kanji learned in each school year, a calculation step of calculating the readability of the text recorded in the text information; An output control step of controlling to output the readability of the text calculated in the calculation step; An information processing method for executing.
10. In a computer, A first acquisition function of acquiring text information in which text is recorded; A second acquisition function of acquiring the ratio of hiragana to kanji used in the text based on the text information acquired by the first acquisition function; Based on at least one of the ratio of hiragana to kanji obtained by the second acquisition function, the difficulty information associating the vocabulary used in the text with the difficulty of the vocabulary, and the learning kanji information associating the school years of elementary school, junior high school, and high school with the kanji learned in each school year, a calculation function for calculating the readability of the text recorded in the text information; An output control function for controlling to output the readability of the text calculated by the calculation function; An information processing program for realizing the above.
Citation Information
Patent Citations
Document editing device
JP1993274306A
Composition problem preparing method and device therefor
JP1996152843A
System and method for education by correspondence
JP2001331089A
Text legibility evaluation system and text legibility evaluation method
JP2009032240A
Information processing apparatus, information processing method and program
JP2014194637A