Generation apparatus, generation method, and generation program
The generating apparatus addresses the challenge of insufficient market trend and VoC understanding in B2B manufacturing by generating and visualizing co-occurrence networks with similarity and cross-field comparisons, enhancing product development efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HITACHI LTD
- Filing Date
- 2022-11-30
- Publication Date
- 2026-05-07
AI Technical Summary
In B2B manufacturing, insufficient understanding of market trends and corporate Voice of the Customer (VoC) hinders product development efficiency, as existing systems lack comprehensive analysis of external information and fail to provide insights beyond keyword-based search results or co-occurrence networks.
A generating apparatus that searches document data based on keywords, generates co-occurrence networks by connecting co-occurring words, modifies conditions, updates the network, calculates similarity, and outputs the results in a displayable format, enabling deeper insights through similarity visualization and cross-field comparisons.
Enhances user awareness by providing visualized co-occurrence networks that prompt insights and discrepancies, facilitating improved product development through enhanced understanding of market trends and customer needs.
Smart Images

Figure 0007854925000001 
Figure 0007854925000002 
Figure 0007854925000003
Abstract
Description
Technical Field
[0001] The present invention relates to a generation device, a generation method, and a generation program for generating information.
Background Art
[0002] The digital transformation (DX) of the engineering chain (EC, abbreviated) in the B2B (Business to Business) manufacturing industry is progressing mainly in the management of the product life cycle and product data from design to mass production preparation processes. In the upstream of EC (market research department, product planning department, research and development department, design department), external information including the issues and needs of customer companies (defined as enterprise VoC (Voice Of Customer)) is collected and reflected in the development of products and services. However, the DX in the upstream of EC is at an intermediate stage. Therefore, external information including enterprise VoC of customer and potential customer companies is manually collected and analyzed from databases of patents and papers outside the company and exhibitions.
[0003] Patent Document 1 discloses a self-producing information processing system that continuously provides new information leading to user awareness and discovery. This self-producing information processing system is an information processing system that collects and outputs information, and includes means for inputting first information, means for collecting second information related to the first information, means for selecting third information from the second information, means for outputting the second information or the third information, means for collecting second information with the third information as new first information, means for merging existing second information and new second information at a predetermined ratio, means for selecting new third information from the merged second information, and means for outputting the merged second information or the new third information, and operates recursively.
[0004] Patent Document 2 discloses an idea generation support program. This idea generation support program performs the following processes: morphological analysis of first information data consisting of multiple natural language sentences limited to a specific subject and extraction of multiple first terms; extraction of multiple first terms in the first information data according to their frequency of occurrence in each of multiple topics using a latent Dirichlet allocation method; and morphological analysis of second information data consisting of multiple natural language sentences not limited to a specific subject and extraction of second terms that co-occur with multiple first terms in each of multiple topics. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] International Publication No. 2016 / 027372 [Patent Document 2] Japanese Patent Publication No. 2022-117931 [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] In B2B manufacturing, there is a problem in that insufficient understanding of market trends and corporate Voice of the Customer (VoC) makes it difficult to improve product development efficiency. Patent Document 1, mentioned above, is limited to a process that returns search results including surrounding areas for a search keyword entered by the user. Patent Document 2, mentioned above, is limited to a process that displays a co-occurrence network created from keywords included in two information groups.
[0007] The present invention aims to provide information that prompts users to become aware of something. [Means for solving the problem]
[0008] A generating apparatus comprising one aspect of the invention disclosed herein is a generating apparatus having a processor for executing a program and a storage device for storing the program, wherein the processor includes a search process for searching document data from an information source based on search keywords, a generating process for generating a co-occurrence network by connecting co-occurring words in each sentence within the document data retrieved by the search process based on conditions relating to the increase or decrease of the number of words, a modification process for changing the conditions, and an update process for updating the co-occurrence network generated by the generating process based on the conditions modified by the modification process. A calculation process that calculates the similarity between the search keyword and the words in the co-occurrence network, and an output process that outputs the co-occurrence network in a displayable format based on the similarity of the words in the co-occurrence network calculated by the calculation process, It is characterized by performing the following. A generating apparatus representing another aspect of the invention disclosed in this application is a generating apparatus having a processor for executing a program and a storage device for storing the program, wherein the processor performs a search process for searching document data from an information source based on search keywords; a generating process for generating a co-occurrence network by connecting words that co-occur in each sentence within the document data retrieved by the search process based on conditions relating to an increase or decrease in the number of words; a modification process for changing the conditions; an update process for updating the co-occurrence network generated by the generating process based on the conditions modified by the modification process; a comparison process for comparing a first word in the co-occurrence network with a second word in another co-occurrence network; and an output process for relating the first word and the second word based on the comparison results of the comparison process and outputting the co-occurrence network and the other co-occurrence network in a displayable format. [Effects of the Invention]
[0009] According to a typical embodiment of the present invention, it is possible to provide information that prompts awareness in the user. Problems, configurations, and effects other than those mentioned above will be clarified by the following description of the embodiments. [Brief explanation of the drawing]
[0010] [Figure 1] Figure 1 is an explanatory diagram showing an example of the configuration of the generation system. [Figure 2] Figure 2 is a block diagram showing an example of the hardware configuration of the generation device. [Figure 3] Figure 3 is an explanatory diagram showing an example of a concern database. [Figure 4] Figure 4 is a block diagram showing an example of the functional configuration of the generating device. [Figure 5] Figure 5 is a flowchart showing an example of the visualization information generation process procedure by the generation device. [Figure 6] Figure 6 is a flowchart showing a detailed example of the co-occurrence network generation process (step S502). [Figure 7] Figure 7 is an explanatory diagram showing the search results of step S602. [Figure 8] Figure 8 is an explanatory diagram showing the morphological analysis results. [Figure 9] FIG. 9 is an explanatory diagram showing a first example of the display of the co-occurrence network. [Figure 10] FIG. 10 is an explanatory diagram showing a second example of the display of the co-occurrence network. [Figure 11] FIG. 11 is a flowchart showing a detailed example of the processing procedure of the similarity visualization process (step S504). [Figure 12] FIG. 12 is an explanatory diagram showing an example of the display of the similarity visualization. [Figure 13] FIG. 13 is a flowchart showing a detailed example of the processing procedure of the cross-field comparison process (step S505). [Figure 14] FIG. 14 is an explanatory diagram showing an example of the display of the cross-field comparison.
MODE FOR CARRYING OUT THE INVENTION
[0011] <Configuration Example of the Generation System> FIG. 1 is an explanatory diagram showing a configuration example of the generation system. The generation system 100 includes a generation device 101 and a terminal 102. The generation device 101 and the terminal 102 are communicably connected via a network 103 such as the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network). The generation device 101 is connected to a search result DB (Data Base) 104. The search result DB 104 is a database that stores search results of users of a plurality of terminals 102 with respect to an in-company database group 110 and an out-company database group 120. The terminal 102 has a browser function and displays the display information on a display screen when receiving visualization information from the generation device 101.
[0012] Each user of the plurality of terminals 102 is, for example, an employee within the same company.
[0013] Further, the generation device 101 and the terminal 102 are communicably connected to the in-company database group 110 and the out-company database group 120 via the network 103. The in-company database group 110 and the out-company database group 120 hold various document data.
[0014] The in-house database group 110 is a set of databases held within the company to which each user of the plurality of terminals 102 belongs. Specifically, for example, the in-house database group 110 has a research and development / prototype information DB 111, a sales and marketing information DB 112, a production information DB 113, a quality information DB 114, a sales information DB 115, and an area of interest DB 116.
[0015] The research and development / prototype information DB 111 has an in-house research and development DB and an in-house product prototype DB related to product prototyping within the company. Specifically, for example, the research and development DB stores, for example, research number, research theme name, research period, budget, name of the responsible researcher, researcher identification number, research plan (research fund name, budget, achievement goals, research and development items, milestones), current progress, weekly reports, meeting materials, materials used, equipment and laboratories, deliverables (in-house research reports, patent applications, external publications), and search keywords.
[0016] Specifically, for example, the product prototype DB stores, for example, prototype product name, prototype number, prototype details, prototype time, name of the person in charge of prototyping, prototyping person identification number, research theme name and research number from which the prototype originated, prototyping plan (prototyping items, milestones), and prototype evaluation results.
[0017] The sales and marketing information DB 112 is a database related to information obtained in sales activities, and specifically stores, for example, customer information relationships, sales strategy relationships, claim relationships, inquiry relationships from customers and potential customers, and market research relationships. The customer information relationship includes, for example, customer name, name of the person in charge on the customer side, name of the salesperson, salesperson identification number, project management information (project name, progress status, estimate, delivery date, salesperson comment), customer list, and customer questionnaire content.
[0018] Sales strategy-related information includes, for example, the annual sales plan (target sales, achievement status, milestones). Complaint-related information includes, for example, complaint information (complaint details, name of the person handling the complaint, response details). Customer and potential customer inquiry-related information includes, for example, inquiry information (inquiry details, name of the person handling the inquiry, response details, analysis results of the inquiry information). Market research-related information includes, for example, market research results (research content, name of the researcher, analysis results, marketing strategy based on the research results).
[0019] The production information DB113 is a database of information related to the production of a product. Specifically, it stores information such as the product production plan, the procurement of materials and components that make up the product, product production information, product number, product lot number, bill of materials, production process details, production start date, production end date, production location, equipment used, and the names and identification numbers of the workers in charge of each production process.
[0020] The Quality Information DB114 is a database that stores information about product quality. Specifically, it stores information such as quality defects in the manufacturing process, quality defects in pre-shipment quality inspections, quality defects, malfunctions, and accidents reported by customers after sale, analysis results of quality defects, malfunctions, and accidents and the corresponding actions taken, information about the research and development process on which the product is based, the prototype number on which the product is based, and the research number.
[0021] Sales Information DB115 is a database containing sales information for a business or product. Specifically, it stores, for example, the business or product name, sales data for each business or product (sales amount, customer name, sales period), the name of the research theme related to the business or product, the research number, and the implementation period.
[0022] The Concerns DB116 is a database that stores the user's concerns. Details about the Concerns DB116 will be described later.
[0023] The external database group 120 is a collection of databases maintained outside the company to which each user of the multiple terminals 102 belongs. Specifically, for example, the external database group 120 includes a journal database 121, a patent database 122, a newspaper database 123, a statistics database 124, an SNS (Social Networking Service) database 125, other company websites 126, and government websites 127.
[0024] The journal DB121 is a database of papers published in various academic journals. Specifically, it stores information such as the name of the journal, publication date, title, content of the article, author names, author affiliations, keywords derived from the content of the article, citations, and reference information.
[0025] Patent DB122 is a database of published domestic and international patents, specifically storing information such as patent title, application number, application date, publication number, publication date, inventor name, applicant name, examination request information, patent grant information, information on foreign applications, patent number, and registration date.
[0026] Newspaper DB123 is a database of articles published in various newspapers. Specifically, it stores information such as the name of the newspaper, the publication date, the article content, keywords derived from the article content, and the name of the article's author.
[0027] Statistical DB124 is a database of statistical information provided by governments, public institutions, and industry associations. Specifically, it stores information such as the Ministry of Economy, Trade and Industry's "Industrial Statistics," the Statistics Bureau of the Ministry of Internal Affairs and Communications' "Population Census," and the Japan Automobile Manufacturers Association's "Four-Wheeled Vehicle Production Figures."
[0028] SNSDB125 is a database of information posted on various social networks, and specifically stores information provided by social network operators, whether for a fee or free of charge, in accordance with the Personal Information Protection Act.
[0029] Third-party websites 126 are information obtained from the websites of other companies or organizations and are stored on the sites of those other companies or organizations. Specifically, third-party websites 126 store, for example, information about products produced by the company (product name, model number, specifications, price, delivery date), the company's or organization's management strategy, management plan, performance and financial information, financial results, ESG (Environment, Social, Governance) information, and technological development information of the other company or organization.
[0030] Government websites 127 are information obtained from the websites of the government or government-affiliated public entities and stored on the websites of the government or government-affiliated public entities. Government websites 127 store information on policies, budgets, policy-based research and development, laws, regulations, guidelines, and various statistical information.
[0031] <Example hardware configuration of generation device 101> Figure 2 is a block diagram showing an example of the hardware configuration of the generation device 101. The generation device 101 includes a processor 201, a storage device 202, an input device 203, an output device 204, and a communication interface (communication IF) 205. The processor 201, storage device 202, input device 203, output device 204, and communication IF 205 are connected by a bus 206. The processor 201 controls the generation device 101. The storage device 202 serves as the work area for the processor 201. The storage device 202 is a non-temporary or temporary recording medium that stores various programs and data. Examples of storage devices 202 include ROM (Read Only Memory), RAM (Random Access Memory), HDD (Hard Disk Drive), and flash memory. The input device 203 inputs data. Examples of input devices 203 include a keyboard, mouse, touch panel, numeric keypad, scanner, microphone, and sensor. The output device 204 outputs data. Output devices 204 include, for example, displays, printers, and speakers. The communication IF 205 connects to the network 103 and sends and receives data.
[0032] <Interest Database 116> Figure 3 is an explanatory diagram showing an example of the Interest DB 116. The Interest DB 116 has a user ID 301 and an interest 302. The user ID 301 is identification information that uniquely identifies each user of multiple terminals 102. The interest 302 is entered from terminal 102. The interest 302 is an arbitrary string that indicates something the user is interested in, for example, "Ukraine conflict", "ramen", "children's extracurricular activities", etc. The interest 302 becomes the title of its co-occurrence network.
[0033] <Example of functional configuration of the generation device 101> Figure 4 is a block diagram showing an example of the functional configuration of the generation device 101. The generation device 101 includes an input unit 401, an output unit 402, a control unit 403, and a storage unit 404. The input unit 401, the output unit 402, and the control unit 403 are specifically realized, for example, by causing the processor 201 to execute a program stored in the storage device 202 shown in Figure 2. The storage unit 404 is specifically realized, for example, by the storage device 202 shown in Figure 2.
[0034] The input unit 401 receives data input from the terminal 102, the internal database group 110, the external database group 120, and the search results DB 104, which are received via the communication IF 205.
[0035] The output unit 402 outputs data from the control unit 403 and the storage unit 404 to the terminal 102, the internal database group 110, the external database group 120, and the search result DB 104 using the communication IF 205.
[0036] The control unit 403 includes an information collection unit 431, an information analysis unit 432, and an information generation unit 433.
[0037] The information gathering unit 431 has a search function 431A and an external website crawling function 431B. The search function 431A is a function that searches the internal database group 110 and the external database group 120 by inputting search keywords. The external website crawling function 431B is a function that automatically collects information from specified websites among the websites of other companies 126 and government websites 127.
[0038] The information analysis unit 432 has a data analysis function 432A, a similarity calculation function 432B, and a medical interview function 432C. The data analysis function 432A is a function that analyzes the data collected by the information collection unit 431. Specifically, for example, the data analysis function 432A performs natural language analysis such as morphological analysis and dependency parsing on the input text, and calculates co-occurrence relationships.
[0039] The similarity calculation function 432B is a function that calculates the similarity between words that have undergone morphological analysis, for example. The medical interview function 432C is a function that checks for any deficiencies or excesses in the information collected by the information collection unit 431, as well as any omissions or errors in the user's information.
[0040] The generation unit 433 generates information (visualization information) to be visualized on the terminal 102 using the visualization information generation function 433A.
[0041] The memory unit 404 stores the information collection results 441, the information analysis results 442, the dictionary 443, and the user information 444. The information collection results 441 are the information collected by the information collection unit 431.
[0042] The information analysis results 442 include the natural language analysis results, co-occurrence relationships, word similarity, and interview results from the interview function 432C performed by the information analysis unit 432. The information analysis results 442 also include size information 442A.
[0043] Size information 442A is a parameter that determines the size of the visualization information generated by the generation unit 433. Details of size information 442A will be described later.
[0044] Dictionary 443 contains words and tags assigned to those words. Tags indicate the category of a word. Categories include, for example, the part of speech, synonyms, related words, and antonyms of that word. The set of tags constitutes tag information 443A.
[0045] User information 444 is information that associates user ID 301 with the user's personal information. User ID 301 in user information 444 is used to match it with user ID 301 from terminal 102.
[0046] <Visualization Information Generation Processing Procedure> Figure 5 is a flowchart showing an example of the visualization information generation process procedure by the generation device 101. The generation device 101 receives input of user ID 301 from terminal 102 and performs user authentication (step S501). Specifically, for example, the generation device 101 determines whether the user ID 301 from terminal 102 matches any of the user IDs 301 in the user information 444. If there is a matching user ID 301, the user is authenticated and the process proceeds to step S502.
[0047] Next, the generation device 101 executes the co-occurrence network generation process (step S502). The co-occurrence network generation process (step S502) is a process in which the control unit 403 generates a co-occurrence network, which is an example of visualization information. A co-occurrence network is a network composed of co-occurring word groups. The co-occurring word groups are obtained from sentences within information sources that hit the search keywords.
[0048] In the co-occurrence network generation process (step S502), the number of co-occurring words in the co-occurring word group within the co-occurrence network can be increased or decreased by changing the size information 442A. In other words, by visualizing the co-occurrence network before and after the increase or decrease in the number of words, it is possible to give the user an opportunity to gain insights (sensory discoveries, flashes of insight, changes in interpretation or understanding). Details of the co-occurrence network generation process (step S502) will be described later in Figures 6 to 10.
[0049] Next, the generation device 101 determines whether the user has selected similarity visualization or cross-field comparison (step S503). If similarity visualization is selected (step S503: similarity visualization), the generation device 101 executes the similarity visualization process (step S504). On the other hand, if cross-field comparison is selected (step S503: cross-field comparison), the generation device 101 executes the cross-field comparison process (step S505).
[0050] Similarity visualization is the process of visualizing the similarity between a search keyword and other words within a co-occurrence network. By visualizing this similarity within the co-occurrence network, it is possible to provide users with a trigger to notice any discrepancies between the search keyword and other words. Details of the similarity visualization process (step S504) will be described later in Figures 11 and 12.
[0051] Cross-field comparison involves comparing the co-occurrence network obtained from a search keyword with the co-occurrence network of a different field from the theme of the search keyword. By comparing the two co-occurrence networks, the comparison results can be provided to the user as a starting point for gaining insights. Details of the cross-field comparison process (step S505) will be described later in Figures 13 and 14.
[0052] <Co-occurrence network generation process (step S502)> Figure 6 is a flowchart showing a detailed example of the co-occurrence network generation process (step S502).
[0053] The generation device 101 receives input from terminal 102 of a string indicating a theme that will become the title of the co-occurrence network and a search keyword (step S601). The search keyword may be the same as the theme, a word included in the theme, or a word not included in the theme.
[0054] Next, the generating device 101 searches for information sources using the search keywords (step S602). The information sources are the search destinations for the search keywords, and may be, for example, the in-house database group 110 (excluding the concern DB 116) mentioned above, the external database group 120, the in-house database group 110 (excluding the concern DB 116) and the external database group 120, or one or more databases within the in-house database group 110 and the external database group 120 (excluding the concern DB 116).
[0055] Figure 7 is an explanatory diagram showing the search results of step S602. The search result 700 is displayed on terminal 102. The search result 700 includes a number 701, a source 702, a title 703, and a date 704 on which the information was generated.
[0056] Number 701 is a number assigned in ascending order. Number 701 can be selected by operating terminal 102. Information source 702 is the search destination for the search keyword mentioned above, for example, the internal database group 110 and external database group 120 mentioned above, which contain the document data group. Title 703 is the title or summary contained in the searched database. For example, it is the title set for the document data contained in the searched database, or, if the document data is a web page, for example, the string enclosed in the title tag.
[0057] The information generation date 704 refers to the date on which the search results were generated. Specifically, this could be, for example, the date the document data included in the searched database was uploaded, the creation date recorded in the document data, or, if the document data is a web page, the publication date described in the web page, URL (Uniform Resource Locator), or source code.
[0058] Returning to Figure 6, the generation device 101 accepts the selection of the searched document data and extracts the document data (step S603). Specifically, for example, when the generation device 101 accepts the selection of number 701 from terminal 102, the generation device 101 extracts the document data identified in the row of the selected number 701 from the information source 702.
[0059] The generation device 101 performs morphological analysis, a type of natural language processing, on the text within the extracted document data in step S603, counts the occurrence count of each extracted word, and extracts words that satisfy the conditions regarding word occurrence count (step S604). Note that particles are not included in the part of speech of the words to be counted. The part of speech of the words to be counted can be set in advance.
[0060] Figure 8 is an explanatory diagram showing the morphological analysis results. The morphological analysis result 800 has a number 801, an extracted word 802, a part of speech 803, and an occurrence count 804. The number 801 is a number assigned in ascending order. The extracted word 802 is a word decomposed by morphological analysis. The part of speech 803 is a group that classifies the extracted word 802 according to grammatical criteria. The occurrence count 804 is the number of times the extracted word 802 appeared in the extracted document data. Note that sa-verbs (combinations of nouns and "suru") may be classified as a single word.
[0061] Returning to Figure 6, the generation device 101 counts the number of times each extracted word 802 in step S604 co-occurs with other extracted words 802 in the same sentence within the extracted document data (the numerator of equation (1)) (step S605).
[0062] The generating device 101 uses the number of co-occurrences calculated in step S605 to calculate the co-occurrence probability for word pairs of co-occurring extracted words 802 and other extracted words 802 using the above formula (1) (step S606).
[0063] Next, the generation device 101 sets conditions regarding the increase or decrease in the number of words as size information 442A (step S607). The conditions regarding the increase or decrease in the number of words are conditions for increasing or decreasing the number of nodes that represent words in the generated co-occurrence network. Specifically, for example, the conditions regarding the increase or decrease in the number of words include at least one of the following: conditions regarding the number of word occurrences and conditions regarding the co-occurrence relationship of co-occurring word pairs.
[0064] The condition regarding word occurrences defines the range of occurrences, which is the number of times a word appears in the document data. For example, if the range is "5 or more and 10 or less," words that appear between 5 and 10 times will be extracted. Also, if "minimum occurrence = 5," words that appear 5 or more times will be extracted. And if "maximum occurrence = 10," words that appear between 10 and 0 times will be extracted.
[0065] The condition for the co-occurrence relationship of co-occurring word pairs is that the co-occurrence probability of co-occurring word pairs in each sentence within the document data falls within a predetermined probability range, thereby generating a co-occurrence network. Co-occurrence relationship is the relationship between two words appearing in the same sentence, specifically, for example, the co-occurrence probability. Co-occurrence probability is the probability that two words appear in the same sentence, and can be calculated, for example, by the following formula (1).
[0066] Co-occurrence probability = (Number of sentences in which word A and word B appear) / (Number of sentences in which word A or word B appears) ... (1)
[0067] Therefore, the predetermined probability range may be an absolutely evaluated range, such as word pairs that co-occur with a co-occurrence probability of X% or higher, or a relatively evaluated range, such as word pairs that co-occur with a co-occurrence probability from 1st to Nth. Note that X and N are predetermined.
[0068] The generation device 101 generates a co-occurrence network based on the extracted words 802 and size information 442A whose co-occurrence probabilities calculated in step S606 satisfy the conditions regarding the co-occurrence relationship of the co-occurring word pairs (step S608). The generation device 101 outputs the generated co-occurrence network to the terminal 102 (step S609). As a result, the co-occurrence network is displayed on the terminal 102.
[0069] The generator 101 determines from the terminal 102 whether there has been a change in the conditions regarding the increase or decrease in the number of words (step S610). If there has been a change in the conditions regarding the increase or decrease in the number of words from the terminal 102 (step S610: Yes), the process returns to step S607. On the other hand, if there has been no change in the conditions regarding the increase or decrease in the number of words from the terminal 102 (step S610: No), the generator 101 determines whether there has been a change in the extracted document data (step S611). If there has been a change in the extracted document data, the process returns to step S604. If there has been no change in the extracted document data, the process proceeds to step S503.
[0070] Figure 9 is an explanatory diagram showing example 1 of the co-occurrence network display. Figure 10 is an explanatory diagram showing example 2 of the co-occurrence network display. Terminal 102 displays the display screen 900. The display screen 900 has a network display area 901, a first condition setting area 902, a second condition setting area 903, a similarity visualization selection button 904, and a different field comparison selection button 905.
[0071] The network display area 901 displays the co-occurrence network. Specifically, for example, in Figure 9, the co-occurrence network 911 is displayed, and in Figure 10, the co-occurrence network 1011 is displayed. In co-occurrence networks 911 and 1011, the circular shapes are nodes representing extracted words 802, and the line segments connecting the circular shapes are links indicating the co-occurrence relationships of extracted words 802.
[0072] For example, in co-occurrence network 1011, nodes A and B have a co-occurrence relationship, nodes B and C have a co-occurrence relationship, nodes A and D have a co-occurrence relationship, nodes A and E have a co-occurrence relationship, and nodes D and E have a co-occurrence relationship. Note that the size of a node indicates its frequency of occurrence. That is, the more frequently the extracted word corresponding to a node appears, the larger the node.
[0073] Furthermore, when a node is specified by terminal 102, the generation device 101 may display the location in the document data where the extracted word corresponding to that node appears. Similarly, when two nodes connected by a link are specified by terminal 102, the generation device 101 may display the location in the document data where a sentence in which a pair of extracted words corresponding to those two nodes appear simultaneously appears.
[0074] The first condition setting area 902 is an area where the user can set conditions regarding the number of occurrences of words (for example, the minimum number of occurrences) by operating the terminal 102. The second condition setting area 903 is an area where the user can set conditions regarding the co-occurrence relationship of co-occurring word pairs (for example, up to the top N) by operating the terminal 102. The similarity visualization selection button 904 is an area where the user can select similarity visualization by operating the terminal 102.
[0075] When the similarity visualization selection button 904 is pressed, the process proceeds to the similarity visualization process (step S504). The different field comparison selection button 905 is an area where the user can select a comparison to a different field by operating the terminal 102. When the different field comparison selection button 905 is pressed, the process proceeds to the different field comparison process (step S505).
[0076] In the network display area 901, the number of extracted words 802 (nodes) decreases as the minimum occurrence count increases, and increases as the minimum occurrence count decreases. Furthermore, in the conditions concerning the co-occurrence relationship of co-occurring word pairs, the number of extracted words 802 (nodes) increases as the lower limit, i.e., N, increases, and decreases as N decreases.
[0077] For example, in the state shown in Display Example 1 in Figure 9, if in step S610 the user changes the value of the first condition setting area 902 from "10" to "5" and changes the value of the second condition setting area 903 from "up to the top 50" to "up to the top 100", the display becomes Display Example 2 in Figure 10. In other words, the co-occurrence network 911 is updated to a co-occurrence network 1011 with more extracted words 802.
[0078] In the above explanation, the value of the first condition setting area 902 was decreased, and the value of the second condition setting area 903 (lower limit N) was increased. However, even if only one of these is changed, the co-occurrence network 911 will be updated to a co-occurrence network with more extracted words 802.
[0079] Furthermore, in the above explanation, the value of the first condition setting area 902 was made smaller and the value (lower limit) of the second condition setting area 903 was made larger. However, if the value of the first condition setting area 902 is made larger and the value of the second condition setting area 903 (lower limit N) is made smaller, the co-occurrence network will be updated to have fewer extracted words 802.
[0080] For example, in the state shown in Display Example 2 of Figure 10, if in step S610 the user changes the value of the first condition setting area 902 from "5" to "10" and changes the value of the second condition setting area 903 from "up to the top 100" to "up to the top 50", the result will be Display Example 1 of Figure 9. In other words, the co-occurrence network 1011 is updated to a co-occurrence network 911 with fewer extracted words 802.
[0081] In the above explanation, the value of the first condition setting area 902 was increased and the value of the second condition setting area 903 (lower limit N) was decreased. However, even if only one of these is changed, the co-occurrence network 1011 will be updated to a co-occurrence network with fewer extracted words 802.
[0082] <Similarity Visualization Process (Step S504)> Figure 11 is a flowchart showing a detailed example of the similarity visualization process (step S504). The similarity visualization process (step S504) is executed when the similarity visualization selection button 904 is pressed (step S503: similarity visualization).
[0083] Figure 12 is an explanatory diagram showing an example of the similarity visualization display. The similarity visualization display screen 1200 is updated from the display screen 900 when the similarity visualization selection button 904 is pressed. The similarity visualization display screen 1200 displays a similarity display slider 1201 and a similarity legend 1202 in the network display area 901. The similarity display slider 1201 is a user interface that allows the user on terminal 102 to select whether to display or hide the similarity. In the example in Figure 12, the state in which the similarity display is selected is shown.
[0084] Similarity Legend 1202 is a legend that shows how the similarity of a node to a search keyword is represented by the intensity of the color within the node. The double-circle node (G in Figure 12) represents the search keyword. White within a node indicates that it is not being visualized. In the example in Figure 12, nodes A, D, F, H, and I are shown to be visualized. Note that the representation of similarity is not limited to the intensity of color, as long as it is a representation that the user can recognize. For example, text or shapes indicating the type of similarity may be added to the node.
[0085] Furthermore, co-occurrence network 1211 is a co-occurrence network to which similarity visualization has been applied to co-occurrence network 1011.
[0086] In Figure 11, the generation device 101 sets the visualization targets from the extracted word group (step S1101). The visualization targets are the extracted words whose similarity to the search keyword is to be visualized. The generation device 101 sets one or more nodes specified by the user operation of terminal 102 as visualization targets.
[0087] Furthermore, the generation device 101 may automatically set the extracted words as targets for visualization. The automatically set extracted words may be all extracted words, or extracted words narrowed down to specific categories. For example, if a specific category is the same part of speech as the search keyword, the automatically set extracted words will be words of the same part of speech as the search keyword. Also, if a specific category is synonyms, related words, or antonyms of the search keyword, the automatically set extracted words will be synonyms, related words, or antonyms of the search keyword.
[0088] The generation device 101 vectorizes the search keywords and the visualization targets (step S1102). Specifically, for example, the generation device 101 generates word vectors for each of the search keywords and visualization targets using Word2Vec, which is a word vectorization technique.
[0089] The generation device 101 calculates the similarity between the search keyword and the visualization target (step S1103). Specifically, for example, the generation device 101 calculates the inter-vector distance between the word vector of the search keyword and the word vector of the visualization target. The shorter the inter-vector distance, the more similar the search keyword and the visualization target are. Therefore, the generation device 101 calculates the similarity between the search keyword and the visualization target based on the inter-vector distance. For example, the generation device 101 may use the reciprocal of the inter-vector distance as the similarity. In this case, the higher the similarity, the more similar the search keyword and the visualization target are. If cosine similarity is used as the inter-vector distance, the closer the value is to 1, the higher the similarity, and the closer it is to 0, the lower the similarity.
[0090] The generation device 101 outputs the similarity of the objects to be visualized in a displayable format (step S1104). Specifically, for example, the generation device 101 draws the intensity of the nodes to be visualized based on the similarity according to the similarity legend 1202.
[0091] The generation device 101 determines whether or not there is an input to terminate similarity visualization (step S1105). If there is no input to terminate similarity visualization (step S1105: No), the generation device 101 determines whether or not there is an input to change the visualization target (step S1106). If there is an input to change the visualization target (step S1106: Yes), the process returns to step S1101. If there is no input to change the visualization target (step S1106: No), the process returns to step S1105. If there is an input to terminate similarity visualization (step S1105: Yes), the similarity visualization process (step S504) ends.
[0092] Thus, the similarity visualization process (step S504) visualizes the similarity between the search keyword and the extracted words. Therefore, the user can visually identify which extracted words are similar to the search keyword and which are not, and thus identify any sense of incongruity the user may have.
[0093] In Figure 12, similarity visualization was applied to co-occurrence network 1211. However, the generation device 101 may apply similarity visualization only to co-occurrence network 1211 where no search keywords exist, or it may apply similarity visualization only to co-occurrence network 1212 where search keywords exist.
[0094] Furthermore, the generator 101 may tag extracted words 802 with a similarity of 1 or higher to the search keyword in the dictionary 443 as synonyms, or tag extracted words 802 with a similarity of 2 or higher, which is greater than the 1st threshold, as synonyms.
[0095] Similarly, the generator 101 may tag extracted words 802 with a similarity of at least a first threshold in the dictionary 443 with the search keyword as a synonym, or tag extracted words 802 with a similarity of at least a second threshold greater than the first threshold with the search keyword as a synonym. In this way, the generator 101 can perform learning on the dictionary 443.
[0096] <Comparison processing across different fields (Step S505)> Figure 13 is a flowchart showing a detailed example of the inter-field comparison process (step S505). The inter-field comparison process (step S505) is executed when the inter-field comparison selection button 905 is pressed (step S503: inter-field comparison).
[0097] Figure 14 is an explanatory diagram showing an example of a display for comparing different fields. The field comparison display screen 1400 is updated from the display screen 900 when the field comparison selection button 905 is pressed. The field comparison display screen 1400 includes a network display area 901, a theme display area 1401, a field network display area 1402, a concern display area 1403, and a comparison display slider 1404.
[0098] In the example shown in Figure 14, the network display area 901 displays the co-occurrence network 1011 shown in Figure 10. The theme display area 1401 displays the string T1 (hereinafter referred to as theme T1) which represents the theme entered in step S601 in order to generate the co-occurrence network 1011.
[0099] The alternative field network display area 1402 is a display area that displays a comparison co-occurrence network for a concern 302 in a field different from theme T1. In the example in Figure 14, the comparison co-occurrence network 1411 is displayed. Note that, to distinguish it from the comparison co-occurrence network 1411, co-occurrence networks 911 and 1011 may be referred to as the source co-occurrence networks 911 and 1011. The concern 302 may belong to the same user who entered theme T1, or to a different user. Furthermore, a concern 302 in a different field is, for example, a concern 302 that is dissimilar to theme T1.
[0100] The concern display area 1403 displays a string T2 (hereinafter referred to as "concern T2") indicating a concern 302 in another field. If there are multiple concerns T2 in another field, the generator 101 may randomly select a concern T2 in another field, or it may display multiple concerns T2 in a pull-down format in the concern display area 1403 for the user to select. The generator 101 obtains a comparison co-occurrence network 1411 for the selected concern T2 in another field and displays it in the other field network display area 1402.
[0101] The comparison display slider 1404 is a user interface that allows the user on terminal 102 to select whether to display or hide comparisons of different fields. In the example in Figure 14, the state in which comparisons of different fields are selected is shown.
[0102] Returning to Figure 13, the generation device 101 obtains a comparison co-occurrence network related to the user's concerns 302 (step S1301). Specifically, for example, the generation device 101 calculates the similarity between the theme T1 entered by the user and the concerns 302 associated with the user's user ID 301 in the concerns DB 116. For example, doc2vec is used to calculate the similarity between texts. The generation device 101 then extracts concerns 302 whose text similarity is below a predetermined threshold as concerns 302 in a different field. It is also possible to use concerns 302 associated with a user ID 301 different from the user in question.
[0103] The generation device 101 searches for information sources using the extracted interests 302 from other fields instead of the search keywords in step S602, and generates the comparison co-occurrence network 1411 by performing the same processing as in steps S603 to S606. In this way, the comparison co-occurrence network 1411 is obtained.
[0104] The generation device 101 may also store the previously acquired comparison co-occurrence networks for each concern 302 in the storage unit 404, and when a concern 302 in a different field is extracted, it may read the comparison co-occurrence network associated with that concern 302 from the storage unit 404. The comparison co-occurrence network 1411 can also be acquired by this method.
[0105] The generation device 101 detects the compatibility between co-occurrence networks (step S1302). Compatibility between co-occurrence networks refers to the presence or absence of identical or similar nodes between the source co-occurrence network and the comparison target co-occurrence network. In Figure 14, for example, the generation device 101 searches the comparison target co-occurrence network 1411 for nodes that are identical or similar to the search keyword G in the source co-occurrence network 1011, that is, nodes whose similarity to the search keyword G is equal to or greater than a predetermined value.
[0106] In the example in Figure 14, node G is identified in the comparison target co-occurrence network 1411. The generation device 101 associates the search keyword G in the source co-occurrence network 1011 with node G in the comparison target co-occurrence network 1411 using compatibility-related information 1421. The compatibility-related information 1421 is, for example, a line segment connecting the search keyword G in the source co-occurrence network 1011 with node G in the comparison target co-occurrence network 1411.
[0107] The generation device 101 detects the unexpectedness between co-occurrence networks (step S1303). The unexpectedness between co-occurrence networks is the presence or absence of an unexpected extracted word 802 in the comparison source co-occurrence network and the comparison target co-occurrence network. An unexpected extracted word 802 is an extracted word 802 in the comparison source co-occurrence network that is similar to an extracted word 802 in the comparison source co-occurrence network that is dissimilar to an extracted word 802 in the comparison target co-occurrence network that is dissimilar to an interest T2 in a different field.
[0108] Specifically, for example, the generator 101 searches for extracted words 802 that are dissimilar to the subject of interest T2 in another field that defines the comparison co-occurrence network. For example, if the subject of interest T2 in another field is a single noun or a combination of nouns, word2vec is used to search for dissimilar extracted words 802; otherwise, doc2vec is used. The generator 101 identifies extracted words 802 whose similarity is below a predetermined threshold as dissimilar extracted words 802. In the example in Figure 14, node R is identified as a dissimilar extracted word 802.
[0109] The generator 101 searches the source co-occurrence network for nodes that are identical or similar to the dissimilar extracted word 802, i.e., nodes whose similarity is equal to or greater than a predetermined value. In the example in Figure 14, node H in the source co-occurrence network 1011 is identified as a word similar to node R, which is the dissimilar extracted word 802. The generator 101 associates node H in the source co-occurrence network 1011 with node R in the target co-occurrence network 1411 using surprise-related information 1422. The surprise-related information 1422 is, for example, a line segment connecting node H in the source co-occurrence network 1011 with node R in the target co-occurrence network 1411.
[0110] The generation device 101 detects the relationship of expressions between specific word pairs that co-occur between co-occurrence networks (step S1304). The relationship of expressions between specific word pairs that co-occur between co-occurrence networks is the relationship of expressions between a specific word pair that co-occurs in the source co-occurrence network and a specific word pair that co-occurs in the target co-occurrence network.
[0111] A specific pair of co-occurring words is a pair of words that have a grammatical connection and co-occur. Grammatical connections include, for example, the relationship between a modifier and its non-modifier, the relationship between a subject and a predicate, and the relationship between a predicate and an object.
[0112] For example, if the co-occurring extracted word pairs are a combination of adjective and noun in the source sentence (e.g., "red" and "automobile"), then they represent a modifier-unmodifier relationship. Note that it does not matter whether the modifier extracted word 802 (red) modifies the unmodifier extracted word 802 (automobile) in the source sentence.
[0113] However, in the co-occurrence network generation process (step S502), the generation device 101 may count the number of times a modifier extracted word 802 (red) modifies a non-modifier extracted word 802 (automobile), and if the counted number of co-occurrences is greater than or equal to a predetermined number, it may identify it as a specific word pair that co-occurs.
[0114] Furthermore, for example, if a co-occurring extracted word pair is a combination of a word indicating the subject and a word indicating the predicate in the source sentence (e.g., "ramen" and "delicious"), then it represents a subject-predicate relationship. Note that it does not matter whether the extracted word 802 (ramen) and the extracted word 802 (delicious) in the source sentence constitute the subject and predicate.
[0115] However, in the co-occurrence network generation process (step S502), the generation device 101 may count the number of times the extracted subject word 802 (ramen) and the extracted predicate word 802 (delicious) co-occur, and if the counted number of co-occurrences is greater than or equal to a predetermined number, it may identify them as a specific word pair that co-occurs.
[0116] Furthermore, for example, if a co-occurring extracted word pair is a combination of a word indicating a predicate and a word indicating an object in the source sentence (e.g., "sell" and "beverage"), then it represents a predicate-object relationship. Note that it does not matter whether the extracted word 802 (sell) and the extracted word 802 (beverage) constitute a predicate and its object in the source sentence.
[0117] However, the generating device 101 may count the number of times a predicate extracted word 802 (sell) and an object extracted word 802 (beverage) co-occur in the co-occurrence network generation process (step S502) when the predicate and its object are composed of the predicate and its object, and if the counted number of co-occurrences is greater than or equal to a predetermined number, it may identify them as a specific word pair that co-occurs.
[0118] Furthermore, the relationship of expressions refers to the matching of expressions between specific word pairs that co-occur in the source co-occurrence network and specific word pairs that co-occur in the target co-occurrence network. Examples of expressions include facilitative expressions, suppressive expressions, affirmative expressions, and negative expressions.
[0119] A facilitating expression is an expression in which one of the extracted words 802 in a specific co-occurring word pair contains a facilitating word that promotes the other extracted word 802. For example, if the specific co-occurring word pair is "demand" (subject) and "increase" (predicate), then "increase" is the facilitating word. Other examples of facilitating words include predicates such as "expand" and "rise."
[0120] An inhibitory expression is an expression in which one of the extracted words 802 in a specific co-occurring word pair contains an inhibitory word that suppresses the other extracted word 802. For example, if the specific co-occurring word pair is "demand" (subject) and "decrease" (predicate), then "decrease" is the inhibitory word. Other examples of inhibitory words include predicates such as "shrink" and "decline."
[0121] Affirmative expressions are those in which one of the extracted words 802 in a specific co-occurring word pair contains a facilitator that affirms the other extracted word 802. For example, if the specific co-occurring word pair is "accept" (predicate) and "refugee" (object), then "accept" is the affirmative word. Other examples of affirmative words include predicates such as "permit," "tolerate," "agree," "delicious," and "good."
[0122] A negative expression is an expression in which one of the extracted words 802 in a specific pair of co-occurring words contains a facilitator that negates the other extracted word 802. For example, if the specific pair of co-occurring words is "prohibit" (predicate) and "park on the street" (object), then "prohibit" is the negative word. Other examples of negative words include predicates such as "refuse," "deny," "oppose," "bad," and "wrong."
[0123] In dictionary 443, the relevant words are assumed to have been pre-assigned tags such as facilitators, inhibitors, affirmatives, and negatives. Some words may be tagged with both facilitators and affirmatives, and some may be tagged with both inhibitors and negatives.
[0124] In the example in Figure 14, a specific word pair 1431 within the source co-occurrence network 1011 is represented by the combination of nodes B and C. For example, node B is "crime rate" and node C is "rise". Similarly, a specific word pair 1432 within the target co-occurrence network 1411 is represented by the combination of nodes P and Q. For example, node P is "demand" and node Q is "increase".
[0125] In this case, since nodes C and Q are facilitated expressions, the expressions match between the specific word pairs 1431 and 1432. Therefore, the generator 101 associates the specific word pair 1431 in the source co-occurrence network 1011 with the specific word pair 1432 in the comparison target co-occurrence network 1411 using the expression relationship information 1423. The expression relationship information 1423 is, for example, a line segment connecting the specific word pairs 1431 and 1432.
[0126] The generation device 101 outputs the detection results from steps S1302 to S1304 to the terminal 102 (step S1305). As a result, the terminal 102 displays the cross-field comparison display screen 1400 shown in Figure 14. Note that it is sufficient for at least one of steps S1302 to S1304 to be executed. Which of these steps to be executed can be set in advance.
[0127] The generator 101 determines whether or not there is an input to end the inter-field comparison (step S1306). If there is no input to end the inter-field comparison (step S1306: No), the generator 101 determines whether or not there is an input to change concern 302 (step S1307). If there is an input to change concern 302 (step S1307: Yes), the process returns to step S1301. If there is no input to change concern 302 (step S1307: No), the process returns to step S1306. If there is an input to end the inter-field comparison (step S1306: Yes), the inter-field comparison process (step S505) ends.
[0128] Thus, according to the embodiment described above, by performing the co-occurrence network generation process (step S502), it is possible to give the user an opportunity to gain insights (sensory discoveries, flashes of inspiration, changes in interpretation or understanding).
[0129] Furthermore, by performing the similarity visualization process (step S504), the discrepancies between the search keyword and other words can be presented to the user as a trigger for gaining insight. Additionally, by performing the cross-field comparison process (step S505), the comparison results of the two co-occurrence networks can be presented to the user as a trigger for gaining insight.
[0130] In Figure 5, the generation device 101 performs either the similarity visualization process (step S504) or the cross-field comparison process (step S505) after step S503. However, the cross-field comparison process (step S505) may be performed after the similarity visualization process (step S504), or the similarity visualization process (step S504) may be performed after the cross-field comparison process (step S505).
[0131] Furthermore, the generation device 101 may output query information to the user of terminal 102 prompting them to select nodes to be used for similarity calculation in the similarity visualization process (step S504). For example, query information such as "Your search keyword is XX, but are there any words that seem out of place?" may be output to prompt the user to take notice.
[0132] Furthermore, if terminal 102 specifies a node within the co-occurrence network through user operation, the tag attached to the word indicated by the specified node may be displayed. This can provide the user with information that prompts awareness.
[0133] It should be noted that the present invention is not limited to the embodiments described above, but includes various modifications and equivalent configurations within the spirit of the attached claims. For example, the embodiments described above are described in detail to make the present invention easier to understand, and the present invention is not necessarily limited to having all of the described configurations. Furthermore, some of the configurations of one embodiment may be replaced with those of another embodiment. Furthermore, some of the configurations of one embodiment may be added to those of another embodiment. Furthermore, some of the configurations of each embodiment may be added, deleted, or replaced with other configurations.
[0134] Furthermore, each of the aforementioned configurations, functions, processing units, and processing means may be implemented in hardware, for example, by designing them as integrated circuits, or they may be implemented in software by having a processor interpret and execute programs that realize each function.
[0135] Information such as programs, tables, and files that implement each function can be stored in memory, hard disks, SSDs (Solid State Drives), or on recording media such as IC (Integrated Circuit) cards, SD cards, and DVDs (Digital Versatile Discs).
[0136] Furthermore, the control lines and information lines shown are those deemed necessary for explanation purposes and do not necessarily represent all control lines and information lines required for implementation. In reality, it can be assumed that almost all components are interconnected. [Explanation of Symbols]
[0137] 100 generation systems 101 Generator 102 terminals 103 Network 110 Internal database group 120 External Databases 201 Processor 202 Storage Devices 401 Input section 402 Output section 403 Control Unit 404 Storage section 431 Information Gathering Department 432 Information Analysis Department 433 Generation part 702 Source of information 911,1011,1211 co-occurrence network 1411 Comparison Co-occurrence Network
Claims
1. A generating apparatus having a processor for executing a program and a storage device for storing the program, The aforementioned processor, A search process that retrieves document data from information sources based on search keywords, A process for generating a co-occurrence network by connecting co-occurring words in each sentence within the document data retrieved by the search process, comprising a generation process for generating the co-occurrence network based on conditions for increasing or decreasing the number of words in the co-occurrence network, A modification process to change the aforementioned conditions, An update process that updates the co-occurrence network generated by the generation process based on the conditions changed by the modification process, A calculation process that calculates the similarity between the search keyword and the words in the co-occurrence network, Based on the similarity of words within the co-occurrence network calculated by the calculation process, an output process is performed to output the co-occurrence network in a displayable format. A generating device characterized by performing the following actions.
2. A generating apparatus comprising a processor for executing a program and a storage device for storing the program, The aforementioned processor, A search process that retrieves document data from information sources based on search keywords, A process for generating a co-occurrence network by connecting co-occurring words in each sentence within the document data retrieved by the search process, comprising a generation process for generating the co-occurrence network based on conditions for increasing or decreasing the number of words in the co-occurrence network, A modification process to change the aforementioned conditions, An update process that updates the co-occurrence network generated by the generation process based on the conditions changed by the modification process, A comparison process that compares the first word in the aforementioned co-occurrence network with the second word in another co-occurrence network, Output processing that associates the first word and the second word based on the comparison results of the comparison process and outputs the co-occurrence network and the other co-occurrence networks in a displayable format, A generating device characterized by performing the following actions.
3. The generating apparatus according to claim 2, In the comparison process, the processor detects a specific second word that is similar to the first word, In the output processing, the processor associates the first word and the specific second word and outputs the co-occurrence network and the other co-occurrence networks in a displayable format. A generating apparatus characterized by the following features.
4. The generating apparatus according to claim 2, In the comparison process, the processor detects a specific second word that is similar to the first word and dissimilar to the titles of the other co-occurrence networks. In the output processing, the processor associates the first word and the specific second word and outputs the co-occurrence network and the other co-occurrence networks in a displayable format. A generating apparatus characterized by the following features.
5. The generating apparatus according to claim 2, In the comparison process, the processor identifies a first word pair consisting of the first word and a third word in the co-occurrence network that co-occurs with the first word and has a grammatical connection to the first word, identifies a second word pair consisting of the second word and a fourth word in the other co-occurrence network that co-occurs with the second word and has a grammatical connection to the second word, and detects the relationship between the expressions of the first word pair and the second word pair. In the output processing, the processor associates the first word pair and the second word pair and outputs the co-occurrence network and the other co-occurrence networks in a displayable format. A generating apparatus characterized by the following features.
6. A generating apparatus according to claim 1 or 2, The aforementioned condition defines the range of word occurrences within the co-occurrence network. A generating apparatus characterized by the following features.
7. A generating apparatus according to claim 1 or 2, The aforementioned condition is that the co-occurrence network is generated using words whose co-occurrence probability is within a predetermined probability range. A generating apparatus characterized by the following features.
8. A generating apparatus according to claim 1 or 2, The aforementioned conditions include a condition that defines the range of the number of occurrences of words within the co-occurrence network, and a condition that the co-occurrence network is generated using words whose co-occurrence probability of co-occurring word pairs falls within a predetermined probability range. A generating apparatus characterized by the following features.
9. A generation method performed by a generation device having a processor for executing a program and a storage device for storing the program, The aforementioned processor, A search process that retrieves document data from information sources based on search keywords, A process for generating a co-occurrence network by connecting co-occurring words in each sentence within the document data retrieved by the search process, comprising a generation process for generating the co-occurrence network based on conditions for increasing or decreasing the number of words in the co-occurrence network, A modification process to change the aforementioned conditions, An update process that updates the co-occurrence network generated by the generation process based on the conditions changed by the modification process, A calculation process that calculates the similarity between the search keyword and the words in the co-occurrence network, Based on the similarity of words within the co-occurrence network calculated by the calculation process, an output process is performed to output the co-occurrence network in a displayable format. A generation method characterized by performing the following.
10. A generation method performed by a generation device having a processor for executing a program and a storage device for storing the program, The aforementioned processor, A search process that retrieves document data from information sources based on search keywords, A process for generating a co-occurrence network by connecting co-occurring words in each sentence within the document data retrieved by the search process, comprising a generation process for generating the co-occurrence network based on conditions for increasing or decreasing the number of words in the co-occurrence network, A modification process to change the aforementioned conditions, An update process that updates the co-occurrence network generated by the generation process based on the conditions changed by the modification process, A comparison process that compares the first word in the aforementioned co-occurrence network with the second word in another co-occurrence network, Output processing that associates the first word and the second word based on the comparison results of the comparison process and outputs the co-occurrence network and the other co-occurrence networks in a displayable format, A generation method characterized by performing the following.
11. The processor includes: A search process that retrieves document data from information sources based on search keywords, A process for generating a co-occurrence network by connecting co-occurring words in each sentence within the document data retrieved by the search process, comprising a generation process for generating the co-occurrence network based on conditions for increasing or decreasing the number of words in the co-occurrence network, A modification process to change the aforementioned conditions, An update process that updates the co-occurrence network generated by the generation process based on the conditions changed by the modification process, A calculation process that calculates the similarity between the search keyword and the words in the co-occurrence network, Based on the similarity of words within the co-occurrence network calculated by the calculation process, an output process is performed to output the co-occurrence network in a displayable format. A generation program characterized by causing the execution of a specific action.
12. The processor includes: A search process that retrieves document data from information sources based on search keywords, A process for generating a co-occurrence network by connecting co-occurring words in each sentence within the document data retrieved by the search process, comprising a generation process for generating the co-occurrence network based on conditions for increasing or decreasing the number of words in the co-occurrence network, A modification process to change the aforementioned conditions, An update process that updates the co-occurrence network generated by the generation process based on the conditions changed by the modification process, A comparison process that compares the first word in the aforementioned co-occurrence network with the second word in another co-occurrence network, Output processing that associates the first word and the second word based on the comparison results of the comparison process and outputs the co-occurrence network and the other co-occurrence networks in a displayable format, A generation program characterized by causing the execution of a specific action.
Citation Information
Patent Citations
Method and device for supporting document retrieval and document retrieving service using the method and device
JP1998074210A
Retrieval support method for document data base and storage medium where program thereof is stored
JP2000010986A
Document analyzer and document analysis method
JP2021093080A
Thinking support program and method
JP2022117931A
Autopoietic information processing system and method
WO2016027372A1