Content processing device and program

The content processing device automates the extraction of word relationships in educational content, allowing for efficient and relevant content association and delivery in learning services.

JP7851151B2Active Publication Date: 2026-04-24NIPPON HOSO KYOKAI
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NIPPON HOSO KYOKAI
Filing Date
2022-02-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Conventional methods for assigning content relationships in learning services require manual intervention, lacking an automated approach to determine correspondences between educational content items.

Method used

A content processing device and program that automatically extracts important words from document titles and bodies, identifies basic and development words, and establishes relationships between them to facilitate automatic content association.

Benefits of technology

Enables the automatic extraction and presentation of related content based on pre-determined relationships, enhancing the efficiency and relevance of content delivery in learning services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007851151000001
    Figure 0007851151000001
  • Figure 0007851151000002
    Figure 0007851151000002
  • Figure 0007851151000003
    Figure 0007851151000003
Patent Text Reader

Abstract

To provide a content processing apparatus and program for automatically extracting information representing a relationship between multiple pieces of content, on the basis of content data.SOLUTION: A content processing apparatus includes an importance word extraction unit and a basic and advanced relationship extraction unit. The important word extraction unit extracts, based on document data including a title and a body, important words out of words included in at least one of the title and the body. The basic and advanced relationship extraction unit decides, out of the important words extracted by the important word extraction unit, as an advanced word, an important word selected based on a requirement that the word is included in the title, decides, as a basic word, an important word not selected as the advanced word, and extracts a pair of the decided basic word and the decided advanced word, as a basic and advanced relationship.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a content processing apparatus and a program.

Background Art

[0002] In the field of education, learning services utilizing Internet technology have attracted attention. In this learning service, users can view educational content by accessing a server device or the like. Also, in the learning service, it has been pointed out that it is necessary to provide content according to the understanding and progress of users (learners), rather than simply following a uniform learning order.

[0003] As a conventional technique, many methods for automatically presenting learning content to users (learners and instructors) have been studied.

[0004] For example, Patent Document 1 discloses a method of determining a learning area based on the curriculum guidelines, and presenting teaching materials to be recommended to a user based on the grade, subject, and textbook unit within that area.

[0005] Also, Non-Patent Document 1 shows a method of describing the relationship between learning content and other video content in the data format of RDF (Resource Description Framework) that supports the Semantic Web, and presenting the relationship explicitly.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Non-Patent Documents

[0007]

Non-Patent Document 1

[0008] However, using the conventional technologies described above presented a problem: the correspondence between content had to be manually assigned. In other words, it is desirable to be able to assign correspondences automatically.

[0009] This invention was made in consideration of these problems, and aims to provide a content processing device and program that can automatically extract information representing the relationships between multiple content items based on content data. [Means for solving the problem]

[0010] [1] In order to solve the above problems, a content processing device according to one aspect of the present invention includes: an important word extraction unit that extracts important words from among words included in at least one of the title and the body of the document based on document data including a title and body of the document; and a basic development relationship extraction unit that selects important words from among the important words extracted by the important word extraction unit, on the condition that they are words included in the title, as development words, determines important words that were not selected as development words as basic words, and extracts pairs of the determined basic words and the determined development words as basic development relationships.

[0011] [2] In another aspect of the present invention, in the above-described content processing apparatus, the important word extraction unit extracts important words from each of the plurality of document data, each of which includes a title and body text, wherein the plurality of document data are ordered, and the basic development relationship extraction unit determines as development words important words selected from among the important words extracted by the important word extraction unit, on the condition that the words included in the title are newly appearing words based on the order of the document data.

[0012] [3] Another aspect of the present invention further comprises the following in the content processing apparatus described above: a basic development relationship data storage unit that stores the basic development relationships extracted by the basic development relationship extraction unit; a metadata storage unit that stores information identifying content and keywords relating to the content as metadata for the content; and a related content search unit that, for a specified main content, searches the basic development relationship data storage unit using a main content keyword which is a keyword associated with the main content, and determines the content associated with the related content keyword as related content corresponding to the main content, by determining the development word corresponding to the basic word when the main content keyword is a basic word, and the basic word corresponding to the development word when the main content keyword is a development word.

[0013] [4] Another aspect of the present invention further comprises a content display unit that displays main content on a screen, requests the related content search unit to search for related content by specifying the main content, and displays the related content determined by the related content search unit as a search result together with the main content on the screen.

[0014] [5] In addition, in one aspect of the present invention, in the above-described content processing apparatus, the related content search unit determines the related content as an advanced content for the main content if (1) the main content keyword related to the main content corresponds to the basic word and the related content keyword related to the related content corresponds to the advanced word, and (2) the related content as basic content for the main content if the main content keyword related to the main content corresponds to the advanced word and the related content keyword related to the related content corresponds to the basic word, and the content presentation unit, when presenting the related content, indicates information that distinguishes whether the related content is basic content or advanced content.

[0015] [6] In addition, one aspect of the present invention further comprises a document conversion unit that generates a composite sentence, which is data obtained by concatenating the title and the body, based on the document data including the title and the body, and the important word extraction unit extracts the important words by performing an analysis process on the composite sentence.

[0016] [7] Furthermore, a content processing device according to one aspect of the present invention includes: a basic-development relationship data storage unit that stores pairs of basic words and developed words as basic-development relationships; a metadata storage unit that stores information identifying content and keywords relating to the content as metadata for the content; and a related content search unit that, for a specified main content, searches the basic-development relationship data storage unit using a main content keyword which is a keyword associated with the main content, and determines the content associated with the related content keyword as related content corresponding to the main content, by determining the developed word corresponding to the basic word as a related content keyword when the main content keyword is a basic word, and the basic word corresponding to the developed word as a related content keyword when the main content keyword is a developed word. According to this configuration, the content processing device can search for related content related to the main content based on pre-extracted basic development relationships.

[0017] [8] Another aspect of the present invention is a content processing device described above, wherein the basic development relationship is information represented as a pair of determined basic words and determined development words, wherein the important word extraction unit extracts important words from words included in at least one of the title and the body of the document based on document data including the title and the body of the document, the important word extraction unit determines the important words selected from the extracted important words to be words included in the title as development words, and the important words not selected as development words as basic words.

[0018] [9] Another aspect of the present invention is a program for causing a computer to function as a content processing device as described in any of [1] to [8] above. [Effects of the Invention]

[0019] According to the present invention, a content processing apparatus can extract the relationship between a basis and a development between concepts based on document data including pairs of a title and a body text. This relationship can be used for content presentation.

Brief Description of the Drawings

[0020] [Figure 1] It is a block diagram showing a schematic functional configuration of a content processing system according to an embodiment of the present invention. [Figure 2] It is a block diagram showing a more detailed functional configuration of a basis-development relationship extraction unit according to the embodiment. [Figure 3] It is a schematic diagram showing an example of one piece of document data acquired by a document acquisition unit of a basis-development relationship extraction unit according to the embodiment. [Figure 4] It is a schematic diagram showing another example of one piece of document data acquired by a document acquisition unit of a basis-development relationship extraction unit according to the embodiment. [Figure 5] It is a schematic diagram showing document conversion processing by a document conversion unit according to the embodiment and the relationship between its input / output data. [Figure 6] It is a schematic diagram showing an example of document data after a document conversion unit 202 according to the embodiment performs conversion processing. [Figure 7] It is a schematic diagram showing an example of a basis-development relationship list generated and output by a basis-development relationship list output unit according to the embodiment. [Figure 8] It is a schematic diagram showing another example of a basis-development relationship list generated and output by a basis-development relationship list output unit according to the embodiment. [Figure 9] It is a flowchart showing a procedure of processing for a basis-development relationship extraction unit according to the embodiment to extract a basis-development relationship. [Figure 10] It is a flowchart showing a procedure of processing for a content processing system according to the embodiment to present content. [Figure 11] It is a schematic diagram showing an example of the configuration of a screen presented by a content presentation unit according to the embodiment. [Figure 12]This block diagram shows an example of the internal configuration of a device for configuring a content processing system according to the same embodiment. [Modes for carrying out the invention]

[0021] Next, this embodiment will be described. In this embodiment, the relationships between words and phrases appearing in document data are extracted, and these relationships are associated with keywords assigned to the content to determine the relationship between one piece of content and other pieces of content.

[0022] Specifically, the relationships between terms used in this embodiment are basic-development relationships. Basic-development relationships represent the relationship between basic concepts (terms) and advanced concepts (terms). In other words, in this embodiment, related content corresponding to a specific main content is searched based on basic-development relationships. In extracting basic-development relationships, when document data consists of title and body text pairs, the basic structure of the teaching material is utilized, where "newly learned terms exist in the title and are explained by terms that appear in the body text." That is, among the important words in the document data of the teaching material, important words included in the title are estimated to be advanced words, and important words not included in the title are estimated to be basic words, thereby extracting basic-development relationships between terms.

[0023] In the following description, "words" refers to words, idioms, and phrases (such as noun phrases). The processing related to words in the embodiments may be replaced with processing related to idioms and phrases.

[0024] In this embodiment, relationships between words corresponding to the learning order are automatically extracted from a document. These relationships are those between basic words and advanced words. In other words, in this embodiment, pairs of words in a basic-advanced relationship are automatically extracted from educational materials such as textbooks and supplementary materials for educational content. Then, the basic-advanced relationships between words are structured, and based on the relationships of keywords between content, the system can present the aforementioned "other content" while clearly indicating whether a particular piece of content (for example, the content the user is currently viewing) is basic or advanced.

[0025] Figure 1 is a block diagram illustrating the schematic functional configuration of a content processing system according to this embodiment. As shown in the figure, the content processing system 1 includes a basic development relationship extraction unit 101, a structuring unit 102, a metadata storage unit 103, a content presentation unit 104, a related content search unit 105, and a structured data storage unit 110. Each of these functional units can be implemented, for example, by a computer and a program. Each functional unit also has storage means as needed. The storage means is, for example, a variable in the program or memory allocated by the execution of the program. Alternatively, non-volatile storage means such as a magnetic hard disk drive or a solid-state drive (SSD) may be used as needed. Furthermore, at least some of the functions of each functional unit may be implemented as a dedicated electronic circuit instead of a program.

[0026] The content processing system 1 may also be called a "content processing device." The content processing system 1 can be configured with one device (computer) or with multiple devices. For example, as shown in the figure, the functions of the basic development relationship extraction unit 101, the structuring unit 102, and the structured data storage unit 110 may be configured as a relationship extraction device 11. Alternatively, the functions of the metadata storage unit 103, the content presentation unit 104, and the related content search unit 105 may be configured as a content presentation device 21.

[0027] In this case, the relationship extraction device 11 extracts relationships between concepts based on the input document data and stores the relationship information in the structured data storage unit 110. In other words, the relationships between concepts are the relationships between words. Furthermore, the relationships between concepts extracted by the relationship extraction device 11 are "basic development relationships." These "basic development relationships" will be explained further later. The content presentation device 21 then presents content based on the relationships, referring to the relationship information stored in the structured data storage unit 110. Specifically, the content presentation device 21 presents (displays, plays) a certain content while simultaneously displaying information about other content related to that content. The relationships here are, for example, the "basic development relationships" mentioned above. In other words, the content presentation device 21 presents information about content that corresponds to basic relationships or development relationships (for example, clickable icons) from the perspective of the content being presented at that time.

[0028] The functions of each component of the content processing system 1 are as follows:

[0029] The basic development relationship extraction unit 101 acquires document data and extracts relationships between concepts based on that document data. Specifically, the basic development relationship extraction unit 101 extracts basic development relationships between concepts and outputs a basic development relationship list, which is data representing those relationships. The basic development relationship extraction unit 101 passes the generated basic development relationship list to the structuring unit 102. A more detailed explanation of the basic development relationship extraction unit 101's functional configuration will be provided later with reference to Figure 2.

[0030] The structuring unit 102 generates structured data based on the list of basic development relationships output by the basic development relationship extraction unit 101. This structured data represents the basic development relationships between concepts (between words) and is expressed in RDF format, for example. RDF stands for "Resource Description Framework." Data in RDF format is also suitable for representing directed graphs, and the list of basic development relationships output by the basic development relationship extraction unit 101 is equivalent to a set of such directed graphs. The structuring unit 102 then writes the generated structured data to the structured data storage unit 110.

[0031] An example of structured data generated by the structuring unit 102 is as follows: When the word "distributive law" is used as the base word, one of its derived words is "factorization." Also, when the word "factorization" is used as the base word, its derived words are "cross multiplication" and "quadratic equation." The structured data generated by the structuring unit 102 can represent any number of correspondences between base words and derived words, not limited to the example given here. In other words, the structured data (RDF data) generated by the structuring unit 102 systematizes the base development relationships between many words.

[0032] The metadata storage unit 103 stores metadata related to content that may be presented. This metadata may also be expressed in RDF format, for example. The metadata stored by the metadata storage unit 103 is data that associates information identifying the content with keywords related to the content. Specifically, the metadata may include identification information, content name, URL (Uniform Resource Location), and one or more keywords for each piece of content. The identification information is information (ID) for uniquely identifying the content. For example, a URL may be used as the identification information. The content name is text data that represents the content. For example, the content name may be "High School Mathematics / Linear Equations". The URL is location information for accessing the content. For example, the URL is a string of characters that identifies the content server device and represents its location within that device. Keywords are words that represent the characteristics of the content. The metadata storage unit 103 may also store metadata related to the content other than those described above.

[0033] The content presentation unit 104 has the function of displaying content on the user's device. The content presentation unit 104 presents the main content, which is the content requested by the user, and related content, which is content related to the main content. Based on the URLs of the main content and related content, the content presentation unit 104 retrieves and presents the content data from the content server device.

[0034] The content display unit 104 displays the main content on the screen and requests the related content search unit 105 to search for related content based on the main content. The related content search unit 105 then displays the related content determined as a search result on the screen together with the main content. When displaying related content, the content display unit 104 may indicate information that distinguishes whether the related content is basic content or advanced content.

[0035] Specifically, the content presentation unit 104 passes information of the main content specified by the user's device to the related content search unit 105, and requests the related content search unit 105 to search for related content related to that main content. Then, the content presentation unit 104 uses the search results information returned by the related content search unit 105 to present the related content. When presenting related content, the content presentation unit 104 clearly indicates the relationship between that related content and the original main content. Specifically, the content presentation unit 104 distinguishes between related content that has foundational content for the main content and related content that has expanded content for the main content, and presents them in a way that makes their relationship clear. An example of the presentation screen will be explained later with reference to Figure 11.

[0036] The related content search unit 105 searches for related content related to a given content based on a request from the content presentation unit 104. For example, for a specified main content, the related content search unit 105 searches the structured data storage unit 110 (basic and developmental relationship data storage unit) using the main content keyword, which is a keyword associated with the main content. If the main content keyword is a basic word, the developmental word corresponding to that basic word is used as the related content keyword. If the main content keyword is a developmental word, the basic word corresponding to that developmental word is used as the related content keyword. The related content search unit 105 then determines the content associated with those related content keywords as related content corresponding to the main content.

[0037] The related content search unit 105 performs the following (1) and / or (2): (1) If the main content keyword related to the main content corresponds to the basic term and the related content keyword related to the related content corresponds to the advanced term, the related content shall be determined to be advanced content for the main content. (2) If the main content keyword related to the main content corresponds to the development term and the related content keyword related to the related content corresponds to the base term, the related content shall be determined to be the base content for the main content.

[0038] Specifically, the related content search unit 105 receives the URL of the main content from the content presentation unit 104. Based on this URL, the related content search unit 105 retrieves keywords for the main content by searching the metadata storage unit 103. Then, based on those keywords, the related content search unit 105 retrieves the base words and advanced words corresponding to those keywords by searching the structured data storage unit 110. Based on each of those base words and advanced words, the related content search unit 105 retrieves information about other content (related content of the main content) that has those base words or advanced words as keywords by searching the metadata storage unit 103. The related content search unit 105 retrieves, for example, the URL of the relevant related content. The related content search unit 105 returns the obtained related content information (for example, URL) to the requesting content presentation unit 104. The related content search unit 105 distinguishes between the information of related content corresponding to the base and the information of related content corresponding to the advanced and passes them to the content presentation unit 104.

[0039] The number of basic words and advanced words obtained by the related content search unit 105 may be multiple. In the above process, the related content search unit 105 may, after obtaining basic words and advanced words respectively, obtain information on related content based on the number of basic words or advanced words it has as keywords, in descending order. This is because related content with many basic words is considered to be more likely to be basic content, and related content with many advanced words is considered to be more likely to be advanced content.

[0040] The structured data storage unit 110 stores the structured data generated by the structuring unit 102. In other words, the structured data storage unit 110 holds information about the basic-development relationships between words. The structured data storage unit 110 may also be called the "basic-development relationship data storage unit." In other words, the structured data storage unit 110 stores the basic-development relationships extracted by the basic-development relationship extraction unit 101.

[0041] Note that the configuration in Figure 1 assumes the existence of a content server device and a user device. The content server device holds content and provides content (video data, etc.) in response to requests from the content processing system 1. The content presentation unit 104 can retrieve specific content from the content server device, for example, by specifying the URL of the content. The user device is a device operated by the user, such as a PC, tablet terminal, or smartphone. The content presentation unit 104 presents the content on the screen of the user device.

[0042] Figure 2 is a block diagram showing a more detailed functional configuration of the basic development relationship extraction unit 101. As shown in the figure, the basic development relationship extraction unit 101 is composed of a document acquisition unit 201, a document conversion unit 202, an important word extraction unit 203, a basic development relationship estimation unit 204, and a basic development relationship list output unit 205.

[0043] The document acquisition unit 201 acquires document data from an external source. The document data is educational data (teaching material data) and includes a title and body text. The teaching material data consists of multiple sections (units, etc.), and each section has a pair of titles and body texts. The document acquisition unit 201 passes the pair of titles and body texts as one document data item to the document conversion unit 202. In other words, the data acquired from an external source by the document acquisition unit 201 includes multiple document data items. These multiple document data items included in one set of document data are pre-ordered. This order may correspond to, for example, the page order in a textbook. Alternatively, this order may correspond to, for example, the broadcast order of broadcast content. Alternatively, this order may correspond to, for example, the presentation order of content presented on the web.

[0044] An example of a single document data (i.e., a pair of title and body text) will be explained later, referring to Figures 3 and 4.

[0045] The document conversion unit 202 generates a composite sentence, which is data that concatenates the title and the body, based on the document data, which includes the title and the body. Specifically, the document conversion unit 202 performs conversion processing on the document data passed from the document acquisition unit 201. The data before the conversion processing is, as described above, the title and body data passed from the document acquisition unit 201. The data after the conversion processing by the document conversion unit 202 is the title and composite sentence data. Here, the composite sentence is data that concatenates the title and the body. Note that the title, body, and composite sentence are each text (a sequence of characters) data. The document conversion unit 202 passes the converted data (i.e., the title and composite sentence pairs) to the important word extraction unit 203. The conversion processing by the document conversion unit 202 will be explained further later with reference to Figure 5.

[0046] The important word extraction unit 203 extracts important words from at least one of the words contained in the title and the body of the document, based on the document data which includes the title and the body of the document. The important word extraction unit 203 may extract the important words for each document data, based on multiple document data which each include the title and the body of the document. The important word extraction unit 203 may extract important words by performing an analysis process on the synthesized sentence.

[0047] In this embodiment, specifically, the important word extraction unit 203 receives the data after the above conversion process (pairs of title and composite sentence) from the document conversion unit 202 and extracts important words from the title and the composite sentence, respectively. Specifically, the important word extraction unit 203 extracts nouns (which may include noun phrases; the same applies hereinafter) from the title and the composite sentence by performing morphological analysis or the like. Existing technologies can be used for the process of extracting nouns. The important word extraction unit 203 also extracts important words from the nouns extracted from the composite sentence. At this time, values ​​such as TF-IDF can be used as the importance of the nouns. TF-IDF is a value calculated as the product of the TF (term frequency) value and the IDF (inverse document frequency) value. TF is a numerical value that represents how often a certain word appears in a given document. IDF is a numerical value that represents how rarely documents containing a certain word exist among all documents. Here, "all documents" includes not only the title-to-compound sentence pair in question, but also other title-to-compound sentence pairs. In other words, the important word extraction unit 203 calculates a TF-IDF value for each noun extracted from the compound sentence, and extracts nouns in a predetermined range (for example, a predetermined number) with high TF-IDF values ​​as important words. The important word extraction unit 203 passes the information of the extracted important words, along with the data of the title-to-compound sentence pair, to the basic development relationship estimation unit 204.

[0048] The basic-development relationship estimation unit 204 estimates pairs of words that have a specific relationship based on the data passed from the important word extraction unit 203. A specific relationship is, for example, a basic-development relationship. A basic-development relationship is as follows: When the concept represented by word A is the basis of the concept represented by word B, then word A is the basis of word B. Furthermore, when word A is the basis of word B, word B is always an evolution of word A.

[0049] The basic-development relationship estimation unit 204 determines, as development words, important words selected from the important words extracted by the important word extraction unit 203, based on the necessary condition that they are included in the title, and determines, as basic words, important words not selected as development words, and extracts pairs of the determined basic words and determined development words as basic-development relationships. The multiple document data to be processed by the basic-development relationship estimation unit 204 may be ordered. In this case, the basic-development relationship estimation unit 204 may determine, as development words, important words selected from the important words extracted by the important word extraction unit 203, based on the condition that the words included in the title are newly appearing words based on the order of the document data. The procedure for determining whether or not important words included in the title are newly appearing words will be explained later with reference to the flowchart in Figure 9.

[0050] The method by which the Basic-Development Relationship Estimation Unit 204 estimates the basic-development relationship is as follows: Among the important words extracted by the important word extraction unit 203, newly introduced important words included in the title are development words, and the other important words are basic words. The Basic-Development Relationship Estimation Unit 204 estimates the pairs of basic words and development words determined in this way as basic-development relationships. In other words, the pairs of basic words and development words are word pairs that represent basic-development relationships.

[0051] The Basic and Advanced Relationship Estimation Unit 204 determines whether a given important word is new or has appeared before based on its order in the set of document data (the group of teaching material documents). For example, in high school mathematics teaching materials, the item "Linear Inequalities" is followed by the item "Various Linear Inequalities." Then, in the text of the item "Various Linear Inequalities," advanced content such as "Systems of Linear Inequalities" may be explained. The order of appearance in the teaching materials matches the order of the document data passed to the Basic and Advanced Relationship Extraction Unit 101. In this way, by determining whether an important word is new or has appeared before based on a predetermined order of document data (the determination in step S007 of Figure 9, which will be described later), it becomes possible to eliminate reversed pairs of words in the basic and advanced relationship.

[0052] The basic development relationship list output unit 205 generates a basic development relationship list, which is a list of word pairs output by the basic development relationship estimation unit 204, and passes it to the structuring unit 102. In other words, the basic development relationship list output unit 205 outputs the basic development relationship list as information representing the estimation results of relationships between concepts. The data that constitutes the basic development relationship list will be explained later with reference to Figures 7 and 8.

[0053] Next, with reference to Figures 3 and 4, we will describe examples of document data that the content processing system 1 processes. In other words, Figures 3 and 4 each show examples of data used as information for extracting fundamental development relationships.

[0054] Figure 3 is a schematic diagram showing an example of a single document data item acquired by the document acquisition unit 201 of the basic development relationship extraction unit 101. This single document data item includes a pair of title data and body data. As shown in the figure, in this example, the title is "Linear Equations". The body text corresponding to this title is "Reviewing the method of solving linear equations. In order to reliably solve linear equations, ... (omitted)". In other words, this document data corresponds to a section of a mathematics textbook.

[0055] Figure 4 is a schematic diagram showing another example of a single document data item acquired by the document acquisition unit 201 of the basic development relationship extraction unit 101. This single document data item includes a pair of title data and body data. As shown in the figure, in this example, the title is "Considering the Relationship Between Force and Motion ~ The Law of Inertia". The body text corresponding to this title is "When no force acts on an object, or when forces act on it but those forces are in equilibrium, a stationary object remains stationary... (omitted)". In other words, this document data corresponds to a section of a physics textbook.

[0056] Figure 5 is a schematic diagram showing the processing by the document conversion unit 202 (document conversion processing) and the relationship between its input and output data. As shown in the figure, the data input to the document conversion unit 202 is the document data 301 before conversion. The document data 301 before conversion is a series of multiple document data 302. The document data 302 have an overall order. Each document data 302 contains a pair of title 303 and body text 304. On the other hand, the data output from the document conversion unit 202 is the document data 401 after conversion. The document data 401 after conversion is a series of multiple document data 402. The document data 402 corresponds to the input document data 302 and inherits the order (overall order) of the input document data 302. Each document data 402 contains a pair of title 403 and composite text 405. The data of the composite text 405 is data obtained by concatenating the title 303 and body text 304 from the input. Note that the data for title 403 is identical to the data for title 303 on the input side.

[0057] In other words, as shown in Figure 5, the document conversion unit 202 generates a composite text 405 by concatenating the title 303 and body text 304 of each document data 302. The document data 402 (after conversion) output by the document conversion unit 202 has pairs of titles 403 and composite texts 405.

[0058] Figure 6 is a schematic diagram showing an example of document data converted by the document conversion unit 202. The converted document data shown in Figure 6 is data obtained by converting from the original document data explained in Figure 3. As mentioned above, a composite sentence is data that concatenates the title and the body text. In the data exemplified in Figure 6, the first part of the composite sentence, "linear equation," is the title data. The part that follows in the composite sentence, "Let's review the solution method for linear equations... (omitted)," is the body text data.

[0059] Next, we will explain the data in the basic-development relationship list, referring to Figures 7 and 8. As shown in Figures 7 and 8, the basic-development relationship list is tabular data that represents the correspondence between basic words and advanced words.

[0060] Figure 7 is a schematic diagram showing an example of a basic-development relationship list generated and output by the basic-development relationship list output unit 205. The basic-development relationship list in Figure 7 was generated based on the document data exemplified in Figure 3. The basic-development relationship list in Figure 7 has five pairs of basic words and development words. In all five pairs, the development word is "linear equation". This development word was determined as a development word by the basic-development relationship estimation unit 204 from the title "linear equation" in the document data shown in Figure 3. On the other hand, in these five pairs, the basic words are "equation", "both sides", "transposition", "left side", and "right side", respectively. These basic words are important words extracted from the text of the document data shown in Figure 3 and were determined as basic words by the basic-development relationship estimation unit 204.

[0061] Figure 8 is a schematic diagram showing another example of a basic-development relationship list generated and output by the basic-development relationship list output unit 205. The basic-development relationship list in Figure 8 was generated based on the document data exemplified in Figure 4. The basic-development relationship list in Figure 8 has five pairs of basic words and developed words. In all five pairs, the developed word is "law of inertia." This developed word is an important word extracted from the title "Considering the Relationship Between Force and Motion ~ Law of Inertia" in the document data shown in Figure 4, and was determined as a developed word by the basic-development relationship estimation unit 204. On the other hand, in these five pairs, the basic words are "force," "motion," "object," "uniform linear motion," and "velocity," respectively. Of these five basic words, three—"object," "uniform linear motion," and "velocity"—are important words extracted from the main text of the document data shown in Figure 4. Also, two—"force" and "motion"—are important words extracted from the title of the document data shown in Figure 4. However, the two words "force" and "motion" are not new important words in the document data (Figure 4), but rather important words that have already appeared in other document data. Even if an important word is extracted from the title, if it is a word that has appeared before, it is determined to be a basic word rather than a development word (see also the flowchart in Figure 9, which will be explained later). These five basic words are words that were determined to be basic words by the basic-development relationship estimation unit 204.

[0062] Figure 9 is a flowchart showing the procedure for the extraction of basic and developmental relationships by the basic and developmental relationship extraction unit 101. The following describes the process of extracting basic and developmental relationships according to this flowchart. Note that this flowchart shows the process corresponding to one document data (a pair of title and body text). When multiple document data exist, they are ordered as described above, and the process in this flowchart is executed for each document data according to that order.

[0063] In step S001, the document acquisition unit 201 acquires one document data. As mentioned above, the document data acquired by the document acquisition unit 201 is educational material data and includes pairs of title and body data.

[0064] In step S002, the document conversion unit 202 performs a process to convert the document data input in step S001. As a result of this conversion process, composite sentence data is generated. As mentioned above, the composite sentence data is data that concatenates the title and the body text.

[0065] In step S003, the important word extraction unit 203 extracts important words. Important words are extracted from both the title and the synthesized sentence. Important words may be nouns or noun phrases. The importance of a word may be determined by, for example, the TF-IDF value, as described above. The important word extraction unit 203 extracts words with relatively high importance as important words.

[0066] In step S004, the basic development relationship estimation unit 204 determines whether or not important words (extracted in step S003) are present in the title. If important words are present in the title (step S004: YES), the process proceeds to the next step S005. If important words are not present in the title (step S004: NO), the processing for that pair (title and body text pair) is terminated.

[0067] Step S005 marks the start of a loop. The end of this loop is in step S010 below. The condition for executing the process in between (steps S006 to S009) is that there are undecided important words among the important words extracted in step S003. In other words, this loop is executed for each important word until there are no more undecided important words among the important words extracted from the compound sentence.

[0068] In step S006, the basic development relationship estimation unit 204 determines whether or not a particular important word is included in the title. If the important word is included in the title (step S006: YES), the unit proceeds to the next step S007. If the important word is not included in the title (step S006: NO), the unit jumps to step S009.

[0069] If the process proceeds to step S007, the basic development relationship estimation unit 204 determines whether the important word is a new important word or not. The title-to-compound sentence pair is generated based on the title-to-body text pair. As mentioned above, the title-to-body text pair is ordered. When processing the title-to-body text pair according to that order, if the important word currently being processed has not appeared in a previous title, then that important word is new. If the important word currently being processed has appeared in a previous title, then that important word is not new (it has appeared before). Note that when determining whether the important word currently being processed is new or has appeared before, instead of using whether it appeared in a previous title as the criterion, the criterion may be whether it appeared in either a previous title or the body text. If the important word currently being processed is new (step S007: YES), the process proceeds to step S008. If the important word currently being processed is not new, i.e., it has appeared before (step S007: NO), the process proceeds to step S009.

[0070] If the process proceeds to step S008 (from step S007), the basic development relationship estimation unit 204 extracts the important words currently being processed as development words. After processing in this step, the process proceeds to step S010.

[0071] If the process proceeds to step S009 (from step S006 or S007), the basic development relationship estimation unit 204 extracts the important words currently being processed as basic words. After processing in this step, the process proceeds to step S010.

[0072] As described above, in the processing from steps S006 to S009, the basic-development relationship estimation unit 204 determines whether the important word being processed is a development word or a basic word.

[0073] Step S010 is the endpoint of the loop, as mentioned above. From step S010, we return to step S005 to process the next important word (the important word extracted from the compound sentence).

[0074] Once the processing from steps S005 to S010 (repeating for each important word) is complete, the process moves on to step S011.

[0075] In step S011, the basic-development relationship list output unit 205 creates and outputs a list representing basic-development relationships. The basic-development relationship list is a list of data representing a set of pairs of basic words and developed words. The basic-development relationship list contains pairs of basic words and developed words, which are extracted as basic words (as described in step S009) and words extracted as developed words (as described in step S008) based on a given title-text pair. Based on a title-text pair, zero, one, or more pairs of basic words and developed words may be generated. When the processing in step S011 is completed, the basic-development relationship extraction unit 101 completes the entire processing for that document data.

[0076] Figure 10 is a flowchart showing the processing steps for content processing system 1 (content presentation device 21) to present content. Here, content processing system 1 presents content based on already extracted basic development relationships. The following explanation will describe the processing steps for content presentation with reference to this flowchart.

[0077] In step S101, the content presentation unit 104 requests the main content from the content server device based on a request from the user and retrieves the main content. The content presentation unit 104 makes the request to the content server device using, for example, the URL (uniform resource locator) of the content requested by the user. The content presentation unit 104 also requests information about content related to the main content (related content) from the related content search unit 105. When making this request, the content presentation unit 104 passes information identifying the main content to the related content search unit 105. The requested related content information includes the URL information of that related content.

[0078] In step S102, the related content search unit 105 first obtains keywords related to the main content specified by the content presentation unit 104 from the metadata storage unit 103. Specifically, the related content search unit 105 obtains keywords related to the main content by searching the metadata storage unit 103 using the identification information of the main content as a key. The keywords related to the main content obtained here may be singular or plural. Then, the related content search unit 105 searches for base words and advanced words corresponding to the keywords of the main content obtained in this step. That is, the related content search unit 105 obtains base words and advanced words corresponding to the keywords of the main content by searching the structured data storage unit 110. When searching using multiple keywords, base words and advanced words are obtained corresponding to each of those multiple keywords. In other words, as a result of the search process in this step, the related content search unit 105 obtains zero or more base words and zero or more advanced words.

[0079] In step S103, the related content search unit 105 determines whether one or more basic words exist in the search results obtained in step S102. If one or more basic words exist in the search results (step S103: YES), the process proceeds to the next step S104. If no basic words exist (step S103: NO), the process jumps to step S105.

[0080] If the process proceeds to step S104, the related content search unit 105 retrieves information from the metadata storage unit 103 about content (referred to as "basic content") that uses the basic word, which is the search result from step S102, as a keyword. Specifically, the related content search unit 105 retrieves information about content that uses the obtained basic word as a keyword by searching the metadata storage unit 103 using the basic word as a key. Here, the related content search unit 105 retrieves the identification information and URL for each of the zero or more basic content items. This URL may be conveniently referred to as the "basic content URL". The related content search unit 105 passes the identification information and URL of the basic content to the content presentation unit 104 as a search result.

[0081] In step S105, the related content search unit 105 determines whether there is one or more related terms in the search results obtained in step S102. If there is one or more related terms in the search results (step S105: YES), the process proceeds to the next step S106. If there are no related terms (step S105: NO), the process jumps to step S107.

[0082] If the process proceeds to step S106, the related content search unit 105 retrieves information from the metadata storage unit 103 about content (referred to as "developmental content") that uses the developmental term, which is the search result from step S102, as a keyword. Specifically, the related content search unit 105 retrieves information about content that uses the developmental term as a keyword by searching the metadata storage unit 103 using the obtained developmental term as a key. Here, the related content search unit 105 retrieves the identification information and URL for each of the zero or more developmental content items. This URL may be conveniently referred to as the "developmental content URL". The related content search unit 105 passes the identification information and URL of the developmental content to the content presentation unit 104 as a search result.

[0083] In step S107, the content presentation unit 104 retrieves the corresponding content (basic content and advanced content) as related content from the content server device based on the basic content URL and advanced content URL obtained in steps S104 and S106, respectively.

[0084] In step S108, the content presentation unit 104 presents the main content and related content (basic content and advanced content) obtained through the above process on the screen of the user's device. After that, the content processing system 1 completes the entire process shown in this flowchart.

[0085] Figure 11 is a schematic diagram showing an example of the screen configuration presented by the content presentation unit 104. The content presentation unit 104 displays the screen shown in this figure on the display of the user's device. In the figure, 501 is the display area. The display area 501 may be, for example, a part or all of the display screen of a liquid crystal display device. The display area 501 includes the main content display area 502, the basic content display area 510, and the advanced content display area 520. The main content display area 502 is the area for the content presentation unit 104 to display the main content. The basic content display area 510 is the area for the content presentation unit 104 to display the basic content. The advanced content display area 520 is the area for the content presentation unit 104 to display the advanced content. The basic content and advanced content together are called related content. That is, related content is displayed as content related to the main content.

[0086] The content display unit 104 can display thumbnail images of multiple basic content items in the basic content display area 510. All of the thumbnail images 511 in the figure are thumbnail images of basic content. The content display unit 104 can also display thumbnail images of multiple advanced content items in the advanced content display area 520. All of the thumbnail images 521 in the figure are thumbnail images of advanced content.

[0087] In the illustrated example, the title of the main content is "Function". The content display unit 104 displays the main content video in the main content display area 502. Thumbnail images 511 and 521 can be selected by user operation (click operation or touch operation, etc.). If either of the thumbnail images 511 or 521 is selected, the related content corresponding to the selected thumbnail image will then be displayed as the main content in the main content display area 502.

[0088] Thus, when the content processing system 1 of this embodiment displays the main content specified by the user, it can present information on related content (such as thumbnail images and titles) related to the main content based on pre-stored basic development relationship information.

[0089] Figure 12 is a block diagram showing an example of the internal configuration of a device for constituting content processing system 1. The device may consist of one or more devices. Each device can be implemented using a computer. As shown in the figure, the computer consists of a central processing unit 901, RAM 902, input / output ports 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be implemented using existing technology. The central processing unit 901 executes instructions contained in programs read from RAM 902, etc. The central processing unit 901 writes data to RAM 902, reads data from RAM 902, and performs arithmetic and logical operations according to each instruction. RAM 902 stores data and programs. Each element contained in RAM 902 has an address and can be accessed using that address. RAM stands for "Random Access Memory". Input / output ports 903 are ports for the central processing unit 901 to exchange data with external input / output devices, etc. Input / output devices 904 and 905 are input / output devices. Input / output devices 904 and 905 exchange data with the central processing unit 901 via input / output port 903. Bus 906 is a common communication channel used within the computer. For example, the central processing unit 901 reads and writes data to RAM 902 via bus 906. Also, for example, the central processing unit 901 accesses input / output ports via bus 906.

[0090] At least some of the functions of the content processing system 1 in the above-described embodiment can be realized by a computer and a program. In this case, the program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be loaded into a computer system and executed. Here, "computer system" includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, DVD-ROMs, USB memory, and storage devices such as hard disks built into a computer system. In other words, "computer-readable recording medium" may be a non-transitory computer-readable recording medium. Moreover, "computer-readable recording medium" may also include those that temporarily and dynamically hold programs, such as communication lines when transmitting programs via networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside a computer system that acts as a server or client in such a case. Furthermore, the above-mentioned program may be for realizing some of the functions described above, and may also be able to realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0091] As explained above, according to this embodiment, word pairs can be extracted based on document data representing the content of, for example, textbooks or supplementary materials for educational content. Word pairs represent the relationships between words (concepts). Furthermore, these words can be associated with keywords of the content. In other words, the extracted word pairs represent the relationship between one piece of content and another piece of content. That is, according to this embodiment, it is possible to determine other content (related content) that has a predetermined relationship corresponding to a certain piece of content (main content). Therefore, the content processing system of this embodiment can identify and present related content together when presenting the main content. In addition, the content processing system can present related content while clearly indicating the nature of the relationship (basic relationship or developmental relationship).

[0092] Modifications of the above embodiment may be implemented. Next, several modifications will be described. In addition, if possible, multiple modifications may be combined and implemented.

[0093] [Example 1] In the above embodiment, the basic development relationship extraction unit 101 selected important terms from among the important terms extracted by the important term extraction unit 203 as development terms, based on the condition that the terms included in the title were newly appearing terms based on the order of the document data. However, any important terms included in the title may be selected as development terms regardless of whether they are newly appearing or not.

[0094] [Differentiation 2] In the content processing system 1 shown in Figure 1, only the function of the relationship extraction device 11 may be performed. In this case, the relationship extraction device 11 extracts basic development relationships and stores the resulting data in the structured data storage unit 110. In this case, an external device can refer to and use the basic development relationship data stored in the structured data storage unit 110.

[0095] [Difference 3] Of the content processing system 1 shown in Figure 1, only the functions of the content presentation device 21 may be implemented. In this case, the content presentation device 21 can present content while referring to pre-held data on basic and developmental relationships. In other words, in this case, the basic and developmental relationships referred to by the content presentation device are information that is represented as pairs of determined basic and determined developmental words, based on document data including the title and body text. The important word extraction unit extracts important words from among the extracted important words, selecting those that are included in the title as necessary conditions, and determining the important words that were not selected as developmental words as basic words.

[0096] Although embodiments of this invention have been described in detail above with reference to the drawings, the specific configuration is not limited to these embodiments and includes designs and the like that do not depart from the spirit of this invention. [Industrial applicability]

[0097] The present invention can be used, for example, in the analysis of document data or in the presentation of content based on the analysis results. However, the scope of use of the present invention is not limited to those exemplified herein. [Explanation of Symbols]

[0098] 1. Content Processing System (Content Processing Device) 11. Relational Extraction Device 21. Content display device 101 Extraction section of basic development relationships 102 Structuring part 103 Metadata storage unit 104 Content Presentation Section 105 Related Content Search Section 110 Structured data storage unit (basic development-related data storage unit) 101 Extraction section of basic development relationships 201 Document Acquisition Department 202 Document Conversion Department 203 Important word extraction part 204 Basic Development Relationship Estimation Section 205 Basic Development Relationship List Output Section 501 Display area 502 Main content display area 510 Basic content display area 511 Thumbnail Images (Basic Content) 520 Advanced content display area 521 Thumbnail Images (Advanced Content) 901 Central Processing Unit 902 RAM 903 Input / Output Ports 904,905 Input / Output Devices 906 Bus

Claims

1. A key word extraction unit extracts important words from at least one of the words contained in the title and the body of the document, based on document data including the title and the body of the document. The basic-development relationship extraction unit selects from the important words extracted by the important word extraction unit the important words that are included in the title as development words, determines the important words that were not selected as development words as basic words, and extracts pairs of the determined basic words and the determined development words as basic-development relationships. A basic development relationship data storage unit stores the basic development relationship extracted by the basic development relationship extraction unit, A metadata storage unit that stores information identifying the content and keywords related to the content as metadata for the content, A related content search unit searches the basic development relationship data storage unit using the main content keyword, which is a keyword associated with the specified main content, and determines the content associated with the related content keyword as the related content corresponding to the main content, by determining the development word corresponding to the basic word as the related content keyword when the main content keyword is the basic word, and the basic word corresponding to the development word as the related content keyword when the main content keyword is the development word, for the specified main content, A content processing device equipped with the following features.

2. The aforementioned key word extraction unit extracts key words from each of the document data, based on a plurality of document data, each of which includes a title and body text. The basic development relationship extraction unit, for each document data, determines, as development words, important words selected from the important words extracted by the important word extraction unit, provided that they are included in the title, and determines, as base words, the important words not selected as development words, and extracts pairs of the determined base words and the determined development words as basic development relationships. The aforementioned multiple document data are ordered, The basic development relationship extraction unit determines the important terms selected from the important terms extracted by the important term extraction unit as development terms, provided that the terms included in the title are newly appearing terms based on the order of the document data. The content processing apparatus according to claim 1.

3. A content display unit that displays the main content on the screen, requests the related content search unit to search for related content by specifying the main content, and displays the related content determined by the related content search unit as a search result together with the main content on the screen. The content processing apparatus according to claim 1 or 2, further comprising:

4. The aforementioned related content search unit, (1) If the main content keyword related to the main content corresponds to the basic term and the related content keyword related to the related content corresponds to the advanced term, the related content shall be determined as advanced content for the main content, (2) If the main content keyword related to the main content corresponds to the development term and the related content keyword related to the related content corresponds to the base term, the related content shall be determined to be the base content for the main content. The content presentation unit, when presenting the related content, displays information to distinguish whether the related content is the basic content or the advanced content. The content processing apparatus according to claim 3.

5. A document conversion unit generates a composite sentence, which is data formed by concatenating the title and the body, based on the document data including the title and the body. Furthermore, The aforementioned important word extraction unit extracts the important words by performing analysis processing on the synthesized sentence. A content processing device according to any one of claims 1 to 4.

6. A basic-development relationship data storage unit that stores pairs of basic words and developed words as basic-development relationships, A metadata storage unit that stores information identifying the content and keywords related to the content as metadata for the content, A related content search unit searches the basic development relationship data storage unit using the main content keyword, which is a keyword associated with the specified main content, and determines the content associated with the related content keyword as the related content corresponding to the main content, by determining the development word corresponding to the basic word as the related content keyword when the main content keyword is the basic word, and the basic word corresponding to the development word as the related content keyword when the main content keyword is the development word, for the specified main content, Equipped with, The aforementioned basic-development relationship is information represented as pairs of determined basic words and determined developed words. This relationship is obtained when the important word extraction unit extracts important words from at least one of the words contained in the title and the text, based on document data including the title and the text, and then determines the important words selected by the important word extraction unit as developed words if they are included in the title, and determines the important words not selected as developed words as basic words. Content processing device.

7. Content processing apparatus according to any one of claims 1 to 6, A program that makes a computer function.

Citation Information

Patent Citations

  • Lexical analysis method, lexical analysis program, and lexical analysis system

    JP2005149331A

  • Educational teaching material navigation system

    JP2015018159A

  • Related word extraction device and program

    JP2015130111A

  • Related document extraction device

    JP2017211687A

  • Content presenting apparatus and program

    JP2021176081A