Generation system, generation method, and generation program
The generation system addresses the challenge of generating ideas that consider unfamiliar customer and social issues by combining known and unfamiliar domain words, facilitating effective idea generation for planners in manufacturing industries.
Patent Information
- Application Number
- JP2023199637
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-06
AI Technical Summary
Planners in manufacturing industries often struggle to generate ideas that consider customer issues and social issues outside their familiar technical domains, leading to biased information and limited idea generation.
A generation system that accesses both a first database containing words from a known domain and a second database containing words from an unfamiliar domain, searches for document data related to a search keyword combining words from both domains, extracts relevant words, and generates combinations of known and unfamiliar domain words to facilitate idea generation.
The system enables planners to easily generate new ideas by linking familiar technical strengths with unfamiliar customer and social issues, thereby overcoming the limitations of traditional idea generation methods.
Smart Images

Figure 2025085925000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a generation system, a generation method, and a generation program for generating data. [Background technology]
[0002] In manufacturing industries that operate business models such as B2B (Business to Business) and B2B2C (Business to Business to Customer), there is now a demand for research and development that takes into account "needs" such as customer issues and societal issues. Such manufacturing industries analyze the information they collect in their own way, gain insights from the results of that analysis, and then decide on measures to take.
[0003] Furthermore, the following Patent Document 1 discloses an idea generation support device that presents words that are related to an input keyword based on a new relationship that goes beyond the known relationship between concepts. In this idea generation support device, a first concept similarity calculation unit calculates a first concept similarity indicating the similarity between the target word and each of the words based on a partial order relationship between a formal concept including the target word as an extension and a formal concept including each of the words as an extension, which are extracted from a first DB, a second concept similarity calculation unit calculates a second concept similarity indicating the similarity between the target word and each of the words based on an inclusion relationship between the concept of the target word and the concept of each of the words, which are selected from a second DB, and a related word selection unit selects related words from each of the words based on an evaluation value calculated from the first concept similarity and the second concept similarity, and displays the related words on a screen.
[0004] Furthermore, the following Patent Document 2 discloses an information processing device capable of acquiring and presenting unexpected information that is difficult for a user to find by himself from among commonly recognized related documents, regardless of category. This information processing device acquires a phrase that is linked to an input keyword based on link structure data of a document group, and calculates the degree of unexpectedness between a document related to the input keyword and a document related to the acquired phrase. Then, based on the calculated degree of unexpectedness, the information on the unexpectedness of the input keyword is acquired and output. The degree of unexpectedness can be calculated using dissimilarity between document vectors generated using a TF-IDF value, which is a feature value based on the frequency of occurrence of each word in a document. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2013-125454 A [Patent Document 2] JP 2017-91270 A Summary of the Invention [Problem to be solved by the invention]
[0006] However, when formulating hypotheses on which decisions are based, planners often end up sticking to ideas specific to their known fields, and therefore tend not to be able to consider ideas from fields other than the known field.
[0007] For example, in areas where planners are familiar with their own core technologies, they have a wealth of knowledge and can easily come up with ideas. On the other hand, in areas where planners are unfamiliar with, such as customer issues or social issues, the information collected may be biased or they may not know how to use their own core technologies to solve the issues. For this reason, planners tend not to be able to come up with ideas from areas that are unfamiliar to them.
[0008] In addition, the technology of Patent Document 1 can only display words for which essential concepts and ordinary concepts have been calculated, and the hints for ideas are limited. Therefore, the idea domain is abstract and unfamiliar to the user, making it difficult for the user to come up with ideas. In addition, the technology of Patent Document 2 can only display keywords selected by the user that have a link relationship set in advance in the database.
[0009] An object of the present invention is to facilitate idea generation support. [Means for solving the problem]
[0010] A generation system which is one aspect of the invention disclosed in the present application is a generation system having a processor which executes a program and a storage device which stores the program, and is capable of accessing a first database having a first word belonging to a first domain and a second database having a second word belonging to a second domain different from the first domain, and is characterized in that the processor executes a search process which searches an information source for document data including a search keyword consisting of the second word and an additional keyword, an extraction process which extracts a third word from the document data searched by the search process, a generation process which generates a combination of the first word and the third word extracted by the extraction process, and a first output process which outputs the combination generated by the generation process in a displayable manner. Effect of the Invention
[0011] According to the representative embodiment of the present invention, it is possible to facilitate idea generation support. Problems, configurations and effects other than those described above will become apparent from the following description of the embodiment. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 is an explanatory diagram showing an example of hypothesis formulation support. [Diagram 2] FIG. 2 is an explanatory diagram showing an example of a network configuration. [Diagram 3]FIG. 3 is a block diagram illustrating an example of a hardware configuration of the generation system. [Figure 4] FIG. 4 is an explanatory diagram showing an example of the configuration of the first area. [Diagram 5] FIG. 5 is an explanatory diagram showing an example of the configuration of the second area. [Figure 6] FIG. 6 is a block diagram illustrating an example of a functional configuration of the generation system. [Figure 7] FIG. 7 is a flowchart illustrating an example of a hypothesis formulation support process performed by the generation system. [Figure 8] FIG. 8 is an explanatory diagram showing an example of the word selection screen. [Figure 9] FIG. 9 is a flowchart showing a detailed example of the process steps of the third region data creation process (step S704). [Figure 10] FIG. 10 is an explanatory diagram showing an example of the information source designation screen. [Figure 11] FIG. 11 is an explanatory diagram showing the relationship between the second domain words, the additional keywords, and the third domain words. [Figure 12] FIG. 12 is an explanatory diagram illustrating an example of the third region table. [Figure 13] FIG. 13 is an explanatory diagram showing an example of the third region data display screen. [Figure 14] FIG. 14 is an explanatory diagram showing another example of the third region data display screen 1300. As shown in FIG. [Figure 15] FIG. 15 is a flowchart showing a detailed example of the process steps of the third area word selection process (step S707). [Figure 16] FIG. 16 is an explanatory diagram showing a selection screen example 1 in the third area word selection process (step S707). [Figure 17] FIG. 17 is an explanatory diagram showing a selection screen example 2 in the third area word selection process (step S707). [Figure 18] FIG. 18 is an explanatory diagram showing a selection screen example 3 in the third area word selection process (step S707). [Figure 19]FIG. 19 is a flowchart showing a detailed example of the process of generating an idea area (step S708). [Figure 20] FIG. 20 is an explanatory diagram showing an example of a selection screen in the idea area generation process (step S708). [Figure 21] FIG. 21 is an explanatory diagram showing an example of an idea area. [Figure 22] FIG. 22 is an explanatory diagram showing an example of the idea area display screen. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] <Figure 1 Example of hypothesis formulation support> FIG. 1 is an explanatory diagram showing an example of hypothesis formulation support. FIG. 1 shows the processing contents executed by the generation system according to the present embodiment. First, the first area 101 and the second area 102 will be explained. Note that information on the first area 101 and the second area 102 is set in advance. Furthermore, in a manufacturing industry that develops business models such as B2B and B2B2C, a planner who formulates measures is, for example, a member (executive or employee) of the manufacturing industry, and serves as a user who operates the generation system. However, the planner is not limited to a member of the manufacturing industry. Furthermore, the user is not limited to a planner or a member of the manufacturing industry.
[0014] The first area 101 is an area (known area) related to a field known to a planner who plans measures in a manufacturing industry that deploys business models such as B2B and B2B2C. Specifically, for example, the first area 101 is a known area related to a core technology that is a strength (seed) of the company.
[0015] For example, the first area 101 includes, as seeds, a technology ST that is a core technology (for example, synthetic fiber technology) and a product SP (for example, environmentally friendly plastic: "environmentally friendly plastic").
[0016] Technology ST has, for example, functions ft1 (e.g., "prevention of deterioration" A1), ft2 ("biodegradation" A2), .... When functions ft1, ft2, ... are not distinguished, they are written as function ft. "Prevention of deterioration" A1, "biodegradation" A2, ... are words that are components of Technology ST.
[0017] A product SP has, for example, characteristics cp1 (e.g., “plant-based material” A3), cp2 (e.g., “biodegradable” A4), etc. “Plant-based material” A3, “biodegradable” A4, etc. are words that are components of the product SP. When there is no need to distinguish between characteristics cp1, cp2, etc., they are written as characteristic cp.
[0018] These words A1, A2, A3, A4, . . . are referred to as first domain words. When there is no need to distinguish between the first domain words A1, A2, A3, A4, . . . , they are written as first domain words A.
[0019] The first domain 101 has n dictionaries SD1 to SDn (n is an integer equal to or greater than 1). When the dictionaries SD1 to SDn are not distinguished from one another, they are referred to as dictionaries SD. The dictionary SD is a dictionary that collects similar first domain words A.
[0020] Next, the second area 102 will be described. The second area 102 is an area different from the first area 101. The second area 102 is, for example, an area related to a field that is unfamiliar to the planner in the first area 101, and includes, for example, a problem N1 (needs) regarding customers and society. Taking "marine pollution" as an example of problem N1, words related to "marine pollution" include keywords B1 (for example, "microplastics"), B2 (for example, "leaked chemicals"), ....
[0021] These keywords B1, B2, . . . which are components of the second domain 102 are called second domain words. When the second domain words B1, B2, . . . are not to be distinguished from one another, they are written as second domain words B.
[0022] The additional keywords C1, C2, C3, C4, ... are words used to collect information together with the second area word B. When there is no need to distinguish between the additional keywords C1, C2, C3, C4, ..., they are written as additional keywords C. The additional keywords C are also information that is set in advance, like the first area 101 and the second area 102.
[0023] The additional keywords C are, for example, words for collecting information by logical thinking about the second domain words B. In FIG. 1, examples of the additional keywords C are listed as "cause" C1, "issue" C2, "impact" C3, "measure" C4, and so on.
[0024] Specifically, this additional keyword C is used to clarify the content (xxx, yyy) of the second domain word B from the perspective of, for example, "What is the additional keyword C1 (cause) that becomes xxx due to the second domain word B1 (microplastic)?" or "Why does the second domain word B2 (leaking chemical substance) have the additional keyword C3 (impact) on yyy?"
[0025] Additional keywords C may be words that are dug deeper from a PEST (Politics, Economy, Society, Technology) perspective. In the case of politics, for example, words such as laws, regulations, standards, industry standards, export restrictions, or words related to these are selected as additional keywords C.
[0026] In the case of economy, for example, words such as cost, supply chain, value chain, market forecast, economic forecast, and market size, or words related to these, are selected as additional keywords C.
[0027] For example, in the case of society, words such as ELSI, future prediction, SDGs, GX, and environment, or words related to these, are selected as additional keywords C.
[0028] In the case of technology, for example, words such as specifications, performance, quality, raw materials, and materials, or words related to these, are selected as additional keywords C.
[0029] The generation system searches the information source using an AND condition that combines the second domain word B and the additional keyword C as a search condition. If there are multiple second domain words B in the search condition, the multiple second domain words B may be combined with an AND condition or with an OR condition. Similarly, if there are multiple additional keywords C, the multiple additional keywords C may be combined with an AND condition or with an OR condition.
[0030] The information source is document data stored in the generation system or document data stored in another device communicatively connected to the generation system via a network. Document data is data that includes text.
[0031] The generation system collects sentences retrieved from the information source according to the above search conditions and stores them as the third domain 103. The third domain 103 constitutes elements that serve as hints for coming up with new research themes, for example. The generation system breaks down the collected sentences into words using morphological analysis. The generation system extracts, for each sentence, the following words D1 (synthetic fiber), D2 (material), D3 (deterioration), D4 (biodegradability), D5 (improvement), D6 (human body), ... that are nouns and are not additional keywords C from among the broken down word group. The words D1 to D6, ... are referred to as third domain words. When the third domain words D1 to D6 are not distinguished, they are written as third domain words D. The third domain words D are words that connect the first domain words A and the second domain words B.
[0032] The generation system generates an idea area 104 and outputs it visibly. The idea area 104 is an area generated by combining a first domain word AX selected from the first domain words A in the first area 101 and a third domain word DX selected from the third domain words D in the third area 103. Specifically, the idea area is an area that supports the idea of solving the problem in the second area 102 associated with the third domain word DX based on information on the technology ST and product SP of the first area 101 to which the first domain word A belongs. Since the idea area 104 is an area that associates the first area 101 familiar to the user with the second area 102 unfamiliar to the user, the user can easily get a new idea by referring to the first domain words AX and the third domain words DX that constitute the idea area 104.
[0033] <Figure 2 Network configuration example> 2 is an explanatory diagram showing an example of a network configuration. A network system 200 includes a generation system 201 and a terminal 202. The generation system 201 and the terminal 202 are communicatively connected via a network 203 such as the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network).
[0034] The generation system 201 is configured with one or more computers. The generation system 201 is connected to a search result DB (Data Base) 204. The search result DB 204 is a database that stores search results of the in-house database group 210 and the external database group 220 by users of multiple terminals 202.
[0035] The terminal 202 has a browser function, and displays the display information on a display screen when visualization information is received from the generation system 201. The user may operate the generation system 201 directly, or may operate the generation system 201 via the terminal 202.
[0036] Furthermore, the generation system 201 and the terminal 202 are communicatively connected to an in-house database group 210 and an external database group 220 via a network 203. The in-house database group 210 and the external database group 220 hold various document data.
[0037] The in-house database group 210 is a collection of databases held within a company to which each user of the multiple terminals 202 belongs. Specifically, the in-house database group 210 includes, for example, a dictionary DB 211, a technical information DB 212, and a product information DB 213.
[0038] The dictionary DB 211 is a database that stores the dictionary SD shown in FIG.
[0039] The technical information DB 212 is a database that stores information (technical information) related to the seeds of technology ST. Specifically, for example, the technical information DB 212 stores, as technical information, a research number, a research theme name, a research period, a budget, a name of a researcher in charge, a researcher identification number, a research plan (a research fund name, a budget, an achievement goal, a research and development item, a milestone), a current progress, a weekly report, meeting materials, materials used, an apparatus and a laboratory, an outcome (an internal research report, a patent application, an external presentation), and a search keyword.
[0040] The product information DB 213 stores information (product information) related to the product SP that is the seed. Specifically, for example, the product information stored in the product information DB 213 includes product prototype information, product production information, product quality information, and product sales information.
[0041] Specifically, the product prototype information includes, for example, the prototype name, prototype number, prototype content, prototype time, the name of the prototype person in charge, the prototype person's identification number, the name of the research topic and research number on which the prototype was based, the prototype plan (prototype items, milestones), and the prototype evaluation results.
[0042] Specifically, product production information includes, for example, information regarding the production plan for the product, information regarding the procurement of materials and parts that make up the product, information regarding the production of the product, product number, product lot number, bill of materials, production process details, production start date, production end date, production location, equipment used, and the names and worker identification numbers of workers responsible for each production process.
[0043] Specifically, product quality information includes, for example, information regarding quality defects in the manufacturing process, information regarding quality defects in quality inspections before shipment, information regarding quality defects, failures, and accidents reported by customers after the sale, analysis results of the quality defects, failures, and accidents and responses thereto, information regarding the research and development process on which the product is based, the prototype number on which the product is based, and the research number.
[0044] Specifically, the product sales information includes, for example, the product name, sales data for each product (sales amount, customer name, sales period), the name of a research topic related to the product, the research number, and the implementation period.
[0045] The external database group 220 is a collection of databases held outside the company to which each user of the multiple terminals 202 belongs. Specifically, the external database group 220 includes, for example, a journal DB 221, a patent DB 222, a newspaper DB 223, a statistics DB 224, a social networking service (SNS) DB 225, and a website 226.
[0046] The journal DB221 is a database of papers published in various academic journals, and specifically stores, for example, the name of the journal in which the paper was published, the publication date, the title, the content of the text, the author's name, the author's affiliation, keywords created from the content of the paper, cited literature, and reference information.
[0047] The patent DB222 is a database of published domestic and foreign patents, and specifically stores, for example, the patent name, application number, application date, publication number, publication date, inventor name, applicant name, examination request information, patent decision information, information regarding foreign applications, patent number, and registration date.
[0048] The newspaper DB 223 is a database relating to articles published in various newspapers, and specifically stores, for example, the name of the newspaper in which the article was published, the publication date, the article content, keywords created from the article content, and the name of the article creator.
[0049] The statistical DB224 is a database of statistical information provided by the government, public institutions, and industry associations, and specifically stores information regarding, for example, the Ministry of Economy, Trade and Industry's "Industrial Statistics," the Ministry of Internal Affairs and Communications' Statistics Bureau's "Popular Census," and the Japan Automobile Manufacturers Association's "Four-wheel Vehicle Production Volume."
[0050] The SNSDB 225 is a database of information posted on various social networks, and specifically stores information provided by social network operating companies, for a fee or free of charge, based on the Personal Information Protection Act, for example.
[0051] The website 226 includes other company homepages and government homepages. The other company homepages are information obtained from the homepages of other companies or various organizations, and are stored on the sites of the other companies or various organizations. Specifically, the other company homepages store, for example, information about the products produced by the company (product name, model number, specifications, price, delivery date), the business strategy, business plan, performance and financial information, settlement information, information about ESG (Environment, Social, Governance) of the company or organization, and technological development information of the other company or organization.
[0052] Government websites are information obtained from websites of governments or government-affiliated public organizations, and are stored on the websites of governments or government-affiliated public organizations. Government websites store information on policies, budgets, research and development based on policies, information on laws, regulations, guidelines, and various statistical information.
[0053] <Figure 3 Example of hardware configuration of generation system 201> FIG. 3 is a block diagram showing an example of a hardware configuration of the generating system 201. The generating system 201 includes a processor 301, a storage device 302, an input device 303, an output device 304, and a communication interface (communication IF) 305. The processor 301, the storage device 302, the input device 303, the output device 304, and the communication IF 305 are connected by a bus 306. The processor 301 controls the generating system 201. The storage device 302 is a working area for the processor 301. The storage device 302 is a non-transient or temporary recording medium that stores various programs and data. Examples of the storage device 302 include a ROM (Read Only Memory), a RAM (Random Access Memory), a HDD (Hard Disk Drive), and a flash memory. The input device 303 inputs data. Examples of the input device 303 include a keyboard, a mouse, a touch panel, a numeric keypad, a scanner, a microphone, and a sensor. The output device 304 outputs data. The output device 304 may be, for example, a display, a printer, or a speaker. The communication IF 305 is connected to the network 203 and transmits and receives data.
[0054] The generation system 201 is configured with one or more computers. Therefore, the generation system 201 has one or more processors 301, storage devices 302, input devices 303, output devices 304, and communication IFs 305. The generation system 201 may also include a terminal 202.
[0055] <Figure 4 Configuration example of the first area 101> 4 is an explanatory diagram showing an example of the configuration of the first area 101. As shown in FIG.
[0056] The dictionary DB211 has fields of a dictionary number 411 and a first domain word 412. The dictionary number 411 is an identification number that uniquely identifies the dictionary SD. The first domain word 412 is a first domain word A collected from the first domain words 423 of the technical information DB212 and the first domain words 433 of the product information DB213.
[0057] The technical information DB 212 has fields of a technology number 421, technology information 422, and a first domain word 423. The technology number 421 is an identification number that uniquely identifies a technology ST. The technology information 422 is information related to a technology ST that is a seed. The first domain word 423 is a first domain word A that indicates a function ft.
[0058] The product information DB 213 has fields of a product number 431, product information 432, and a first domain word 433. The product number 431 is an identification number that uniquely identifies the product SP. The product information 432 is information about the product SP that is a seed. The first domain word 433 is a first domain word A that indicates the characteristic cp.
[0059] In FIG. 4, “biodegradation” A2, which is a first domain word 423 indicating a function ft2 of synthetic fiber technology ST1 in technical information DB212, and “biodegradability” A4, which is a first domain word 433 indicating a characteristic cp2 of environmentally friendly plastic “environmentally friendly plastic” SP1 in product information DB213, are similar words, and therefore are collected as first domain words 412 in dictionary SD1.
[0060] <Figure 5: Example of configuration of the second area 102> 5 is an explanatory diagram showing an example of the configuration of the second area 102. As shown in FIG. 2, the second area 102 includes a search result DB 204. As described above, the search result DB 204 is a database in which the users of the multiple terminals 202 store search results for the in-house database group 210 and the external database group 220. The search result DB 204 is realized by the storage device 302.
[0061] The search result DB 204 has fields including a search result number 501, a field 502, a problem 503, and a second domain word 504. The search result number 501 is an identification number that uniquely identifies a search result. When the search result numbers R1, R2, ... are not distinguished, they are written as a search result number R. The field 502 is a viewpoint for identifying the second domain 102 that is different from the first domain 101. The problem 503 is a problem that needs to be solved in the field 502. When the individual problems that need to be solved are not distinguished, they are written as a problem N. The second domain word 504 is a second domain word B that is searched from the in-house database group 210 and the external database group 220 using the field 502 and the problem 503 as search conditions.
[0062] <Figure 6 Example of functional configuration of generation system 201> 6 is a block diagram showing an example of a functional configuration of the generation system 201. The generation system 201 has an input unit 601, an output unit 602, a control unit 603, and a storage unit 604. Specifically, the input unit 601, the output unit 602, and the control unit 603 are realized by, for example, causing the processor 301 to execute a program stored in the storage device 302. Specifically, the storage unit 604 is realized by, for example, the storage device 302.
[0063] The input unit 601 accepts data input from the input device 303 or the communication IF 305 , and transfers the data to the output unit 602 , the control unit 603 , or the storage unit 604 .
[0064] The output unit 602 outputs data from the input unit 601 , the control unit 603 , or the storage unit 604 to the output device 304 or the communication IF 305 .
[0065] The control unit 603 has an information collection unit 631, an information analysis unit 632, and a generation unit 633. The information collection unit 631 has a function of collecting information by searching the in-house database group 210 and the external database group 220 by crawling or scraping. The information collection unit 631 stores the information collection result 641 in the storage unit 604.
[0066] The information analysis unit 632 is a function that analyzes the information collection result 641. The information analysis unit 632 has a morphological analysis function 632A, a co-occurrence probability calculation function 632B, and a similarity calculation function 632C. The information analysis unit 632 stores the information analysis result 642 in the storage unit 604.
[0067] The morphological analysis function 632A is a function that performs morphological analysis on the sentence obtained in the information collection result 641 and breaks it down into words. The co-occurrence probability calculation function 632B is a function that calculates the co-occurrence probability that two words appear in a sentence. Specifically, for example, the co-occurrence probability is calculated by the following formula (1).
[0068] Co-occurrence probability = (number of sentences in which two words W1 and W2 appear) / (number of sentences in which either word W1 or word W2 appears) (1)
[0069] The similarity calculation function 632C is a function that calculates the similarity between words using, for example, Word2Vec.
[0070] The generating unit 633 generates information (visualized information) by visualizing the information collection result 641 and the information analysis result 642.
[0071] The storage unit 604 stores an information collection result 641, an information analysis result 642, a dictionary DB 211, and user information 643. The information collection result 641 includes a search result DB 204, an additional keyword group 641A, and a document data group 641B. The additional keyword group 641A is a set of additional keywords C. The document data group 641B is a set of document data constituting a search result obtained by combining second domain words B and additional keywords C. The information analysis result 642 is a calculation result by a morphological analysis function 632A, a co-occurrence probability calculation function 632B, and a similarity calculation function 632C.
[0072] The user information 643 is information about a user, and includes, for example, a user name, a user ID, a password, an affiliation, and a job title. The user information 643 is used for authentication when logging in to the generation system 201.
[0073] The generation system 201 is configured with one or more computers. When the generation system 201 is configured with a plurality of computers, the input unit 601, the output unit 602, the control unit 603, and the storage unit 604 may be distributed.
[0074] Specifically, for example, the input unit 601 and the output unit 602 are implemented in respective computers, and the control unit 603 and the storage unit 604 are distributed among a plurality of computers. Any of the information collecting unit 631, the information analyzing unit 632, and the generating unit 633 may be implemented in a computer different from the other units. Similarly, any of the morphological analysis function 632A, the co-occurrence probability calculation function 632B, and the similarity calculation function 632C in the information analyzing unit 632 may be implemented in a computer different from the other functions.
[0075] Furthermore, any one of the information collection results 641, the information analysis results 642, the dictionary DB 211, and the user information 643 in the storage unit 604 may be stored in a computer different from that in which the other information is stored.
[0076] <Figure 7 Hypothesis planning support processing procedure> FIG. 7 is a flowchart showing an example of a procedure for a hypothesis formulation support process performed by the generation system 201.
[0077] (Step S701) The generation system 201 executes user authentication. Specifically, for example, the generation system 201 receives a pair of a user ID and a password from the input unit 601, and determines whether or not the pair of the user ID and the password matches a pair of the user ID and the password in the user information 643. If they match, the generation system 201 permits the user to use the generation system 201, and enables the execution of steps S702 to S709.
[0078] (Step S702) The generation system 201 selects the second domain words B only when the user is authenticated (step S701). Specifically, for example, the generation system 201 selects the second domain words 504 from the search result DB 204. The generation system 201 may automatically select the second domain words 504 based on a given selection condition, or may accept the second domain words 504 selected by a user operation via the input unit 601. Hereinafter, the selected second domain words 504 are referred to as second domain words BX.
[0079] (Step S703) The generation system 201 selects an additional keyword C. Specifically, for example, the generation system 201 selects the additional keyword C from the additional keyword group 641A. The generation system 201 may automatically select the additional keyword C based on a given selection condition, or may accept the additional keyword C selected by a user operation via the input unit 601. Hereinafter, the selected additional keyword is represented as CX. Here, a specific example of steps S702 and S703 will be described.
[0080] <Figure 8 Word selection screen> 8 is an explanatory diagram showing an example of a word selection screen. A word selection screen 800 displays a second area word list 801, an additional keyword list 802, an additional keyword input area 803, a decision button 804, and a reset button 805.
[0081] The second domain word list 801 is list information in which a selection 810 is added to the search result number 501, the field 502, the topic 503, and the second domain word 504 in the search result DB 204. The selection 810 is an interface that displays a state in which the second domain word 504 is selected via the input unit 601 by a user operation. An "o" indicates that the second domain word 504 is selected. In the example of FIG. 8, it is indicated that "microplastics" B1 is selected.
[0082] The added keyword list 802 is list information in which a selection 820 has been added to the added keyword group 641A in the information collection result 641. The selection 820 is an interface that displays a state in which the added keyword group 641A has been selected via the input unit 601 by a user operation. An "o" indicates that the added keyword C has been selected. In the example of FIG. 8, it is indicated that "Cause" C1, "Issue" C2, and "Impact" C3 have been selected.
[0083] The additional keyword input area 803 is an area where the user can input an additional keyword C that is not displayed in the additional keyword list 802. The example in Fig. 8 shows that "regulation" and "trend" have been input.
[0084] The decision button 804 is a user interface for deciding, through a user operation, the selection 810, the selection 820, and the input to the additional keyword input area 803. The reset button 805 is a user interface for resetting, through a user operation, the selection 810, the selection 820, and the input to the additional keyword input area 803. Fig. 8 shows a state in which the decision button 804 is pressed.
[0085] When the enter button 804 is pressed, the second domain word 504 in the selection 810 is accepted (step S702). In this case, "microplastics" B1 becomes the second domain word BX.
[0086] Similarly, when the decision button 804 is pressed, the additional keyword group 641A in the selection 820 is accepted, and the additional keyword C input in the additional keyword input area 803 is accepted (step S703). In this case, "cause" C1, "issue" C2, "impact" C3, "regulation" and "trend" become the additional keywords CX. Also in this case, "regulation" and "trend" are added to the additional keyword group 641A.
[0087] When the reset button 805 is pressed, the selection 810, the selection 820, and the additional keyword input area 803 are reset, and the user can make a new selection and re-input.
[0088] (Step S704) Returning to Fig. 7, the generation system 201 executes the third region data creation process. Specifically, for example, when the decision button 804 is pressed, the generation system 201 executes the third region data creation process. The third region data is the third region word D in the third region 103 and data required for its generation. For example, the third region data includes the searched document data, the words resolved by the morphological analysis function 632A, the co-occurrence probability calculated by the co-occurrence probability calculation function 632B, and the similarity calculated by the similarity calculation function 632C.
[0089] <FIG. 9 Third region data creation process (step S704)> FIG. 9 is a flowchart showing a detailed example of the process steps of the third region data creation process (step S704).
[0090] (Step S901) The generation system 201 obtains the second domain words BX and the additional keywords CX.
[0091] (Step S902) The generation system 201 specifies information sources to be searched from among the in-house database group 210 and the external database group 220. Specifically, for example, the generation system 201 may automatically specify information sources based on given selection conditions, or may accept information sources specified by a user operation via the input unit 601. Here, a specific example of step S902 will be described.
[0092] <Figure 10 Source selection screen> 10 is an explanatory diagram showing an example of an information source designation screen. The information source designation screen 1000 displays an information source designation button 1001, a non-designated information source button 1002, an information source list 1003, an information source detail display area 1004, an information source input area 1005, an execute button 1006, and a reset button 1007.
[0093] The information source designation button 1001 is a user interface that accepts the designation of an information source through the input unit 601 by a user operation. Fig. 10 shows a state in which the information source designation button 1001 has been pressed and accepted.
[0094] The no information source designation button 1002 is a user interface that does not accept the designation of an information source by a user operation via the input unit 601. When the designation of an information source is not accepted, the generation system 201 searches for accessible sites on the network 203.
[0095] The information source list 1003 is a list of information sources and has the following fields: information source name 1031, URL (Uniform Resource Locator) 1032, first history 1033, second history 1034, and specification 1035.
[0096] The information source name 1031 is the name of the information source. The URL 1032 is information specifying the information source. The first history 1033 is information indicating whether or not any department in the organization to which the user belongs has specified the information source in the past. If it is "○", it indicates a history of the organization specifying the information source in the past. The second history 1034 is information indicating whether or not a specific user has specified the information source in the past. If it is "○", it indicates a history of the specific user specifying the information source in the past.
[0097] Designation 1035 is an interface that displays the state in which URL 1032 has been designated through input unit 601 by user operation. "◯" indicates that URL 1032 has been designated. In the example of Fig. 10, it is shown that URL 1032 with information source name 1301 of "C database" has been designated.
[0098] The information source detail display area 1004 is an area for displaying details of an information source that has been previously specified in the first history 1033 and the second history 1034. In the example of FIG. 10, in the information source list 1003, the first history 1033 is "○" for the URL 1032 of the information source name 1301 "C database", so the information source name 1031 "C database", its URL 1032, and the detailed history 1040 of the first history 1033 are displayed. The detailed history 1040 of the first history 1033 includes the organization name (for example, C department C section) and the date when the information source was specified. Although not shown, the detailed history 1040 of the second history 1034 includes a specific user name and the date when the information source was specified.
[0099] The generation system 201 may automatically recommend the URL 1032. For example, the generation system 201 may automatically recommend the URL 1032 based on the total number of accesses in a predetermined period of time in the past, or may randomly select and automatically recommend the URL 1032.
[0100] The information source input area 1005 is an area that accepts input of an information source via the input unit 601 by user operation. The user can input an information source that is not displayed in the information source list 1003 in the information source input area 1005. An information source name 1051 and a URL 1052 can be input in the information source input area 1005. When the information source name 1051 is input, the generation system 201 identifies an information source on the network 203 by the information source name 1051 and searches for the website. When the URL 1052 is input, the generation system 201 searches for the website specified by the URL 1052 as an information source.
[0101] The execute button 1006 is a user interface for executing a search in an information source specified on the information source specification screen 1000 by a user operation. If the specify information source button 1001 is pressed, the generation system 201 executes a search in the information sources specified in the information source list 1003 and the information source input area 1005, and if the no information source specified button 1002 is pressed, the generation system 201 executes a search in accessible sites on the network 203. Fig. 10 shows a state in which the specify information source button 1001 and the execute button 1006 are pressed.
[0102] The reset button 1007 is a user interface for resetting the designation and input on the information source designation screen 1000 by a user operation.
[0103] (Step S903) 9, the generation system 201 searches the information source specified in step S902 using the search keyword when the execute button 1006 is pressed. The search keyword is a combination of the second domain word BX acquired in step S901 and the additional keyword CX.
[0104] Specifically, for example, the generation system 201 searches for information sources using an AND condition that combines a second domain word BX and an additional keyword CX as a search condition. In addition, when there are multiple second domain words BX in the search condition, the multiple second domain words BX may be combined with an AND condition or may be combined with an OR condition. Similarly, when there are multiple additional keywords CX, the multiple additional keywords CX may be combined with an AND condition or may be combined with an OR condition.
[0105] (Step S904) Returning to FIG. 9, the generation system 201 extracts the document data group 641B found in step S903 from the information source, and stores it in the storage unit 604 as data of the third area 103.
[0106] (Step S905) The generation system 201 performs morphological analysis on each sentence in the document data group 641B extracted in step S904, and extracts the third domain words D. Specifically, for example, the generation system 201 extracts, from each sentence, a word that is a noun and is not an additional keyword CX from the morphologically analyzed word group as the third domain word D. Here, the relationship between the second domain word BX, the additional keyword CX, and the third domain word D in steps S903 to S905 will be specifically described.
[0107] <Figure 11 Relationship between the second domain word BX, the additional keyword CX, and the third domain word D> FIG. 11 is an explanatory diagram showing the relationship between the second domain words BX, the additional keywords CX, and the third domain words D. As shown in FIG.
[0108] The second domain 102 has second domain words 504 (B1, B2, ...) of the search result DB 204. The generation system 201 generates the search keyword 1100 by combining a selected second domain word BX (for example, "microplastics" B3) selected from the second domain words 504 and an additional keyword CX (for example, "cause" C1, "issue" C2, and "impact" C3) selected from the additional keyword group 641A (step S903).
[0109] The search keyword 1100 is, for example, a logical AND of a selected second domain word BX and an additional keyword CX, and the additional keyword CX is a logical OR of "cause" C1, "issue" C2, and "effect" C3.
[0110] The information sources 1130A (A newspaper), 1130B (B industry association), 1130C (C database), . . . have document data groups 1131, 1132, 1133, .
[0111] In step S902, the generation system 201 accepts a designation of an information source. In Fig. 11, it is assumed that the generation system 201 accepts a designation of the information source 1130C (C database) from the information sources 1130A (A newspaper), 1130B (B industry association), 1130C (C database), ... as shown in Fig. 10.
[0112] The generation system 201 searches the specified information source 1130C by the search keyword 1100 (step S903), extracts the corresponding document data group from the document data group 1133, and stores it in the third area DB 1110 (step S904). The extracted document data group is called extracted document data group 1111X. The extracted document data group 1111X is stored in the storage unit 604 as document data group 641B in the information collection result 641.
[0113] The third domain DB 1110 stores a third domain word group 1111D, a third domain table 1112, a third domain word appearance count list 1113, and a third domain inter-word co-occurrence probability list 1114 in addition to an extracted document data group 1111X.
[0114] The third domain word group 1111D is a set of the third domain words D extracted in step S905. The third domain table 1112 is created based on the search keyword 1100, the extracted document data group 1111X, and the third domain word group 1111D. The third domain table 1112 will be described later with reference to FIG.
[0115] The third domain word occurrence count list 1113 is list information of the occurrence count of the third domain word D in the extracted document data group 1111X, and is created in step S906. The third domain word co-occurrence probability list 1114 is list information of the co-occurrence probability between the third domain words D in each sentence of the extracted document data group 1111X, and is created in step S907.
[0116] <Figure 12 Third area table 1112> 12 is an explanatory diagram showing an example of the third domain table 1112. The third domain table 1112 has, as fields, a file number 1200, an extracted second domain word 1201, an additional keyword 1202, a specified information source 1203, a search result 1204, and a morphological analysis result 1205.
[0117] The file number 1200 is an area for storing an identification number that uniquely identifies a file. A file is, for example, a set of entries of the common extracted second domain word BX.
[0118] The extracted second domain word 1201 is an area for storing the extracted second domain word BX. The additional keyword 1202 is an area for storing the additional keyword CX. The designated information source 1203 is an area for storing information on the information source designated in step S902 (for example, the URL 1032).
[0119] The search result 1204 is an area for storing the results of searching the specified information source with the search keyword 1100 in step S903. Specifically, for example, the search result 1204 is a sentence including the extracted second domain word BX and the additional keyword 1202. The morphological analysis result 1205 is an area for storing the third domain word D obtained as a result of morphologically analyzing the search result 1204 in step S905.
[0120] Here, the description of the third region data creation process (step S704) is completed. Next, the display of the third region data (step S705) will be described.
[0121] (Step S705) 7, the generation system 201 executes display of the third region data. The display of the third region data (step S705) will be specifically described.
[0122] <Fig. 13, Fig. 14 Third area data display screen> 13 is an explanatory diagram showing an example of a third region data display screen. The third region data display screen 1300 is displayed when the third region data creation process (step S704) is completed. The third region data display screen 1300 has a file number selection 1301, an execution button 1302, a stop button 1303, a reset button 1304, a co-occurrence network selection 1305, a list selection 1306, an analysis result display area 1307, a third region word selection button 1308, and a third region word non-selection button 1309.
[0123] The file number selection 1301 is a user interface that accepts the selection of the file number 1200 by user operation via the input unit 601. Based on the entry of the selected file number 1200, an analysis result regarding the number of occurrences or the co-occurrence probability is extracted.
[0124] The execute button 1302 is a user interface for starting extraction of an analysis result related to the number of occurrences or the co-occurrence probability through a user operation via the input unit 601. Fig. 13 shows a state in which the execute button 1302 is selected.
[0125] The stop button 1303 is a user interface for stopping ongoing extraction of the occurrence count or co-occurrence probability through the input unit 601 by a user operation.
[0126] The reset button 1304 is a user interface for resetting the file number selection 1301, the co-occurrence network selection 1305, the list selection 1306, and the analysis result display area 1307 via the input unit 601 by a user operation.
[0127] Co-occurrence network selection 1305 is a user interface that accepts the selection of displaying a co-occurrence network via the input unit 601 by user operation.
[0128] The list selection 1306 is a user interface that accepts a selection to display the third domain word appearance count list 1113 through a user operation via the input unit 601. Fig. 13 shows a state in which the list selection 1306 has been accepted.
[0129] The analysis result display area 1307 is an area that displays the analysis results regarding the occurrence count or co-occurrence probability performed in steps S906 and S907 in the detailed processing procedure of the third domain data creation process (step S704). In Fig. 13, the execute button 1302 is selected and the list selection 1306 is accepted, so the third domain word occurrence count list 1113 is displayed. The third domain word occurrence count list 1113 has third domain words 1333 and their occurrence counts 1332, but when displayed as the analysis result display area 1307, the ranking 1331 and checks 1334 are also displayed.
[0130] The rank 1331 is a number in descending order (highest to lowest) of the occurrence count 1332, starting from 1. The occurrence count 1332 is the number of third domain words 1333 that appear in the extracted document data group 1111X. The third domain words 1333 are the third domain words D that appear in the extracted document data group 1111X. The check 1334 is an item that accepts whether or not confirmation is required by the user. When selected by the user via the input unit 601, "○" is displayed. In FIG. 13, it is shown that confirmation is required for the third domain words 1333, "synthetic fiber" D1 and "material" D2.
[0131] The third area word selection button 1308 is a user interface for starting the execution of the third area word selection process (step S707). When the third area word selection button 1308 is pressed (step S706: Yes), the process proceeds to step S707.
[0132] The third area word non-selection button 1309 is a user interface for not executing the third area word selection process (step S707). When the third area word non-selection button 1309 is pressed (step S706: No), the process proceeds to step S708.
[0133] 13, the third domain word appearance count list 1113 is displayed, but the third domain word co-occurrence probability list 1114 may be displayed instead of or together with the third domain word appearance count list 1113. Although not shown, the third domain word co-occurrence probability list 1114 displays the descending order of co-occurrence probability starting from 1, the co-occurrence probability, the third domain word pairs in a co-occurrence relationship, and a check 1334.
[0134] Fig. 14 is an explanatory diagram showing another example of the third domain data display screen 1300. Unlike Fig. 13, Fig. 14 shows a state in which co-occurrence network selection 1305 has been accepted instead of list selection 1306. In this case, co-occurrence network diagram information 1401 and third domain word appearance count list 1402 are displayed in analysis result display area 1307.
[0135] The co-occurrence network diagram information 1401 includes a co-occurrence network 1410, a minimum occurrence number input 1411, and a lowest rank input 1412. The co-occurrence network 1410 is a network diagram in which the third domain words D are nodes (circles) and the co-occurrence relationships between the third domain words D are links (lines). The larger the occurrence number of the third domain words D, the larger the node is displayed. In FIG. 14, the thickness of the link is constant, but the thicker the link may be, the higher the co-occurrence probability between the third domain words D. In addition, "microplastics" B1 is a second domain word BX constituting the search keyword 1100, but is also a third domain word D. Similarly, "cause" is an additional keyword CX constituting the search keyword 1100, but is also a third domain word D.
[0136] The minimum occurrence number input 1411 is a user interface that accepts input of the minimum occurrence number through the input unit 601 by user operation. Fig. 14 shows a state in which "10" is input as the minimum occurrence number. That is, the third domain words D displayed as nodes in the co-occurrence network 1410 are third domain words D with an occurrence number of 10 or more.
[0137] The lowest rank input 1412 is a user interface that accepts input of the lowest rank through a user operation via the input unit 601. Fig. 14 shows a state in which "50" has been input as the lowest rank. That is, the co-occurrence relations displayed as nodes and links in the co-occurrence network 1410 are third domain word pairs whose co-occurrence probabilities are ranked from 1st to "50th" in descending order of co-occurrence probability.
[0138] The third domain word appearance count list 1402 is a list indicating the number of appearances of the third domain words 1333 displayed as nodes in the co-occurrence network 1410 in the third domain word appearance count list 1113 .
[0139] (Step S706) Returning to Fig. 7, the generation system 201 judges whether or not the selection instruction of the third-area word D is accepted. Specifically, for example, when the generation system 201 accepts the input of the third-area word selection button 1308 through the input unit 601 by the user operation (step S706: Yes), it proceeds to step S707. In this case, the generation system 201 executes the third-area word selection process (step S707), and executes the idea area generation process (step S708) using the third-area word DX selected in the third-area word selection process (step S707).
[0140] On the other hand, when the generation system 201 receives the input of the third area word non-selection button 1309 through the input unit 601 by the user operation (step S706: No), it proceeds to the ideation area generation process (step S708). That is, in this case, the third area word selection process (step S707) is not executed, so the generation system 201 selects all the third area words D and executes the ideation area generation process (step S708) for all the selected third area words D in a brute force manner.
[0141] (Step S707) The generation system 201 executes a third domain word selection process (step S707). The third domain word selection process (step S707) is a process of selecting a specific third domain word D (hereinafter, a third domain word DX) that is a junction between the first domain 101 and the second domain 102 from the third domain word group 1111D.
[0142] <FIG. 15 Third-area word selection process (step S707)> FIG. 15 is a flowchart showing a detailed example of the process steps of the third area word selection process (step S707).
[0143] (Step S1501) The generation system 201 determines whether any of the additional keywords CX (for example, any of "Cause" C1, "Issue" C2, and "Impact" C3) has been specified. The selected additional keyword CX is referred to as the specified additional keyword CX. If any of the additional keywords CX has not been selected (step S1501: No), the process proceeds to step S1507. If any of the additional keywords CX has been selected (step S1501: Yes), the process proceeds to step S1502.
[0144] (Step S1502) The generation system 201 selects document data including the specified additional keyword CX from the extracted document data group 1111X, and proceeds to step S1503. The selected document data is referred to as selected document data.
[0145] (Step S1503) The generation system 201 judges whether the selection of the third domain word D is performed based on the number of occurrences or the co-occurrence probability. If it is judged that the selection of the third domain word D is performed based on the number of occurrences (step S1503: number of occurrences), the process proceeds to step S1504. If it is judged that the selection of the third domain word D is performed based on the co-occurrence probability (step S1503: co-occurrence probability), the process proceeds to step S1511.
[0146] (Step S1504) The generation system 201 counts the number of occurrences of the third domain word D in the selected document data, and proceeds to step S1505.
[0147] (Step S1505) The generation system 201 sets a condition regarding the ranking of the number of occurrences of the third domain word D. The condition regarding the ranking is, for example, the minimum number of occurrences, the lowest ranking in descending order of the number of occurrences starting from 1, and can be changed by the user's operation input. After that, the process proceeds to step S1506.
[0148] (Step S1506) The generation system 201 selects, as third domain words DX, the third domain words D that appear in the selected document data from the first place to the place set in step S1505, and then proceeds to step S708.
[0149] (Step S1507) The generation system 201 judges whether the user selects the third domain word D based on the frequency of occurrence. If it is judged that the user selects the third domain word D based on the frequency of occurrence (step S1507: Yes), the process proceeds to step S1508. If it is not judged that the user selects the third domain word D based on the frequency of occurrence (step S1507: No), the process proceeds to step S1510.
[0150] (Step S1508) As in step S1505, the generation system 201 sets a condition regarding the ranking of the number of occurrences of the third domain word D. After that, the process proceeds to step S1509.
[0151] (Step S1509) The generation system 201 refers to the third domain word appearance count list 1113 and selects the third domain words D that appear in the extracted document data group 1111X from the first place to the place set in step S1508 as the third domain words DX. Then, the process proceeds to step S708.
[0152] (Step S1510) The generation system 201 selects the third domain word D designated by the user from any one of the third domain word group 1111D as the third domain word DX, and then proceeds to step S708.
[0153] (Step S1511) As in step S1505, the generation system 201 sets a condition regarding the order of the co-occurrence probability of the third domain word D. After that, the process proceeds to step S1512.
[0154] (Step S1512) The generation system 201 selects, from the selected document data, the third domain words D that have a co-occurrence probability with the designated additional keyword CX from the first place to the set place as the third domain words DX, and then proceeds to step S708.
[0155] <FIGS. 16 to 18: Selection screen in third-area word selection process (step S707)> 16 is an explanatory diagram showing a selection screen example 1 in the third area word selection process (step S707). In the selection screen 1600, a designation button 1601 is a user interface for designating any one additional keyword CX via the input unit 601 by a user operation. A no designation button 1602 is a user interface for not designating any one additional keyword CX via the input unit 601 by a user operation. When the designation button 1601 is pressed, an additional keyword designation 1603 is displayed.
[0156] The additional keyword designation 1603 is a user interface for accepting designation of an additional keyword CX via the input unit 601 by a user operation. The additional keyword designation 1603 has an additional keyword 1202 and a designation area 1630. The designation area 1630 is a user interface for accepting designation of any of the additional keywords 1202 via the input unit 601 by a user operation. In the designation area 1630, a "circle" indicates that the additional keyword 1202 has been designated. In the example of FIG. 16, "cause" C1 is designated.
[0157] When the designation button 1601 is pressed and the designation is accepted in the designation area 1630, the generation system 201 determines that one of the additional keywords CX has been designated (step S1501: Yes), and executes step S1502.
[0158] The occurrence number selection button 1604 is a user interface for selecting a third domain word D by the occurrence number through the input unit 601 by a user operation. The co-occurrence probability selection button 1605 is a user interface for selecting a third domain word D by the co-occurrence probability through the input unit 601 by a user operation.
[0159] When the occurrence number selection button 1604 is pressed, the generation system 201 determines that the selection of the third domain word D is to be performed by the occurrence number (step S1503: occurrence number), and executes step S1504. FIG. 16 shows a state in which the occurrence number selection button 1604 is pressed.
[0160] When the co-occurrence probability selection button 1605 is pressed, the generation system 201 determines that the selection of the third domain word D is to be performed based on the co-occurrence probability (step S1503: co-occurrence probability), and executes step S1511.
[0161] The ranking setting 1606 is a user interface that accepts the setting of the lower limit of the ranking of the occurrence count or the co-occurrence probability by a user operation via the input unit 601. When the setting of a value (an integer equal to or greater than 1) is accepted in the ranking setting 1606, a condition regarding the ranking of the occurrence count or the co-occurrence probability of the third domain word D is set (steps S1505, S1511).
[0162] Specifically, for example, when the occurrence count selection button 1604 is pressed, the third domain words D from 1st in the number of occurrences to the value of the rank setting 1606 become selection candidates in step S1506, and when the co-occurrence probability selection button 1605 is pressed, the third domain words D from 1st in the co-occurrence probability to the value of the rank setting 1606 become selection candidates in step S1512.
[0163] When the occurrence number selection button 1604 is pressed and a value is input in the ranking setting 1606, a third domain word list 1607 is displayed.
[0164] The third domain word list 1607 has a rank 1331, an occurrence count 1332, a third domain word 1333, and a selection state 1670. The selection state 1670 indicates whether the third domain word 1333 corresponds to the third domain word D from the first place in the occurrence count to the rank of the value of the rank setting 1606. "○" indicates that it corresponds, and "×" indicates that it does not correspond. That is, the generation system 201 selects the third domain word D whose selection state 1670 is "○" as the third domain word DX (step S1506).
[0165] 16, when the co-occurrence probability button 1605 is pressed and a value is input to the rank setting 1606, a co-occurrence probability list 1608 is displayed that has third domain words D whose co-occurrence probabilities with the specified additional keyword CX range from 1 to the value of the rank setting 1606. Although not shown, the co-occurrence probability list 1608 displays the descending order of co-occurrence probability starting from 1, the value of the co-occurrence probability with the additional keyword CX, the third domain words in a co-occurrence relationship with the additional keyword CX, and the selection state.
[0166] 17 is an explanatory diagram showing a selection screen example 2 in the third-area word selection process (step S707) Fig. 17 shows a display state of the selection screen 1600 when the reject button 1602 is pressed.
[0167] When the reject button 1602 is pressed, the generation system 201 determines that the specification of the additional keyword CX has been rejected (step S1501: No), and executes step S1507.
[0168] The select button 1701 is a user interface for selecting a third domain word D based on the frequency of occurrence through a user operation via the input unit 601. The no-select button 1702 is a user interface for not selecting a third domain word D based on the frequency of occurrence through a user operation via the input unit 601. Fig. 17 shows a state in which the select button 1701 is pressed.
[0169] The ranking setting 1703 is a user interface that accepts the setting of the lower limit of the ranking of the number of occurrences by a user operation via the input unit 601. When the setting of a value (an integer of 1 or more) is accepted in the ranking setting 1703, a condition regarding the ranking of the number of occurrences of the third domain word D is set (step S1508).
[0170] When the select button 1701 is pressed and a value is input in the rank setting 1703, the third domain word appearance count list 1113 is displayed.
[0171] The third domain word appearance count list 1113 has a rank 1331, an appearance count 1332, and a third domain word 1333, but in the selection screen 1600, in addition to the rank 1331, the appearance count 1332, and the third domain word 1333, a selection state 1740 is also displayed. The selection state 1740 indicates whether the third domain word 1333 corresponds to the third domain word D from the first place in the appearance count to the rank of the value of the rank setting 1606. "○" indicates that it corresponds, and "×" indicates that it does not correspond. That is, the generation system 201 selects the third domain word D whose selection state 1740 is "○" as the third domain word DX (step S1509).
[0172] 18 is an explanatory diagram showing a selection screen example 3 in the third-domain word selection process (step S707) Fig. 18 shows a display state of the selection screen 1600 when the reject button 1602 is pressed.
[0173] When the No designation button 1602 is pressed, the generation system 201 determines that the designation of the additional keyword CX is denied (step S1501: No), and executes step S1507. Fig. 18 shows a state in which the No selection button 1702 is pressed.
[0174] When the no selection button 1702 is pressed, the third region table 1112 is displayed.
[0175] The third domain table 1112 has the extracted second domain words 1201, the additional keywords 1202, the designated information source 1203, the search results 1204, and the morphological analysis results 1205, but in the selection screen 1600, in addition to the extracted second domain words 1201, the additional keywords 1202, the designated information source 1203, the search results 1204, and the morphological analysis results 1205, a selection 1800 is also displayed. The selection 1800 is a user interface that accepts the selection of the morphological analysis results 1205. "○" indicates that the morphological analysis results 1205 are selected as the third domain words DX. That is, the generation system 201 selects the morphological analysis results 1205 for which the selection 1800 is "○" as the third domain words DX (step S1510).
[0176] (Step S708) 7, the generation system 201 executes the idea area generation process (step S708) to generate the idea area 104.
[0177] <Figure 19: Ideation area generation process (step S708)> FIG. 19 is a flowchart showing a detailed example of the process of generating an idea area (step S708).
[0178] (Step S1901) When the third domain word DX is selected in the third domain word selection process (step S707), the generation system 201 acquires the third domain word DX and proceeds to step S1903.
[0179] (Step S1902) If the selection instruction of the third domain word DX is not accepted in step S706 (step S706: No), the generation system 201 acquires the third domain word group 1111D from the third domain DB 1110 and proceeds to step S1903. Each third domain word D of the third domain word group 1111D is also hereinafter referred to as a third domain word DX.
[0180] (Step S1903) The generation system 201 acquires the first domain word group in the column of the first domain word 412 from the dictionary DB 211, and proceeds to step S1904.
[0181] (Step S1904) The generation system 201 judges whether the third domain word DX or a word co-occurring with the third domain word DX is selected as the target for calculating the similarity with each of the first domain words AX of the first domain word group. If the third domain word DX is selected (step S1904:DX), the process proceeds to step S1905. If the word co-occurring with the third domain word DX is selected (step S1904:Co-occurrence), the process proceeds to step S1910.
[0182] (Step S1905) The generation system 201 calculates the similarity between each of the first domain words AX in the first domain word group and the third domain word DX, and then proceeds to step S1906.
[0183] (Step S1906) The generation system 201 sets a condition regarding the ranking of similarity. The condition regarding the ranking of similarity is, for example, the lowest ranking of similarity starting from 1, and can be changed by a user's operational input. After that, the process proceeds to step S1907.
[0184] (Step S1907) The generation system 201 determines whether the selection criterion for similarity is descending order or ascending order of similarity. If it is descending order (step S1907: descending order), the process proceeds to step S1908. If it is ascending order (step S1907: ascending order), the process proceeds to step S1909.
[0185] (Step S1908) The generation system 201 selects a combination of the first domain word AX and the third domain word DX in descending order of similarity from 1st place to the order set in step S1906. The higher the similarity, the more effective and immediate the idea will be. After this, the process proceeds to step S709.
[0186] (Step S1909) The generation system 201 selects combinations of the first domain word AX and the third domain word DX in ascending order of similarity from 1st to the order set in step S1906. The lower the similarity, the more unexpected the idea. Then, the process proceeds to step S709.
[0187] (Step S1910) The generation system 201 refers to the third domain word co-occurrence probability list 1114 calculated in step S907, and extracts the third domain word D (hereinafter, the third domain co-occurrence word EX) based on the high co-occurrence probability with the third domain word DX. Specifically, for example, the generation system 201 extracts the third domain word D whose co-occurrence probability with the third domain word DX is equal to or higher than a predetermined probability or is equal to or higher than a predetermined rank in descending order as the third domain co-occurrence word EX. Then, the process proceeds to step S1911.
[0188] (Step S1911) The similarity calculation function 632C of the generation system 201 calculates the similarity between each of the first domain words AX of the first domain word group and the third domain co-occurring words EX, and proceeds to step S1912. Specifically, for example, the similarity calculation function 632C calculates the inter-vector distance between the word vector of the first domain word AX and the word vector of the third domain co-occurring word EX. The similarity between the first domain word AX and the third domain co-occurring word EX is in a proportional or inversely proportional relationship with the inter-vector distance. When cosine similarity is used as the inter-vector distance, the closer the value is to 1, the higher the similarity, and the closer the value is to 0, the lower the similarity. Note that a method other than cosine similarity can also be used to calculate the inter-vector distance.
[0189] (Step S1912) The generation system 201 sets a condition regarding the ranking of similarities, as in step S1906. The condition regarding the ranking of similarities is, for example, the lowest ranking of similarities starting from 1, and can be changed by a user's operational input. Then, the process proceeds to step S1913.
[0190] (Step S1913) As in step S1907, the generation system 201 determines whether the selection criterion for the similarity is descending or ascending order of similarity. If it is descending order (step S1913: descending order), the process proceeds to step S1914. If it is ascending order (step S1913: ascending order), the process proceeds to step S1915.
[0191] (Step S1914) The generation system 201 selects combinations of the first domain word AX and the third domain co-occurring word EX from the first place to the place set in step S1912 in descending order of similarity, and proceeds to step S709.
[0192] (Step S1915) The generation system 201 selects combinations of the first domain word AX and the third domain co-occurring word EX in ascending order of similarity from 1st place to the rank set in step S1912, and proceeds to step S709.
[0193] <Figure 20 Selection screen in the idea area generation process (step S708)> 20 is an explanatory diagram showing an example of a selection screen in the idea area generation process (step S708). The selection screen 2000 is displayed when the execution of the third area word selection process (step S707) is completed. The selection screen 2000 has a third area word display 2001, a third area word selection instruction button 2002, a co-occurring word selection instruction button 2003, a ranking setting 2004, a descending order selection button 2005, an ascending order selection button 2006, a decision button 2007, and a reset button 2008.
[0194] The third domain word display 2001 displays the third domain word DX selected in the selection 1800 from the morphological analysis result 1205 selected in the third domain word selection process (step S707). In the example of Fig. 20, "synthetic fiber" D1 and "biodegradable" D4 are displayed as the third domain word DX.
[0195] The third region word selection instruction button 2002 is a user interface for receiving an instruction to select a third region word DX through the input unit 601 by a user operation. Fig. 20 shows a state in which the third region word selection instruction button 2002 is pressed. When the third region word selection instruction button 2002 is pressed (step S1904:DX), the generation system 201 executes step S1905.
[0196] The co-occurring word selection instruction button 2003 is a user interface for receiving an instruction to select a word co-occurring with the third area word DX by a user operation via the input unit 601. When the co-occurring word selection instruction button 2003 is pressed (step S1904: co-occurrence), the generation system 201 executes step S1910.
[0197] The ranking setting 2004 is a user interface that accepts the setting of the lower limit of the ranking of similarity through a user operation via the input unit 601. When the setting of a value (an integer equal to or greater than 1) is accepted in the ranking setting 2004, a condition regarding the ranking of similarity between words is set (steps S1906, S1912).
[0198] The descending order selection button 2005 is a user interface that accepts the selection of descending order by user operation via the input unit 601. Fig. 20 shows a state in which the descending order selection button 2005 is pressed. When the descending order selection button 2005 is pressed (step S1907: descending order, step S1913: descending order), the generation system 201 executes steps S1908 and S1914.
[0199] The ascending order selection button 2006 is a user interface that accepts the selection of ascending order by user operation via the input unit 601. When the ascending order selection button 2006 is pressed (step S1907: ascending order, step S1913: ascending order), the generation system 201 executes steps S1909 and S1915.
[0200] The decision button 2007 is a user interface that accepts the decision of the idea area 104 with the selection content on the selection screen 2000 through the input unit 601 by the user's operation. Fig. 20 shows a state in which the decision button 2007 is pressed. When the decision button 2007 is pressed, the generation system 201 executes step S709.
[0201] The reset button 2008 is a user interface that accepts a reset of the selection contents on the selection screen 2000 through the input unit 601 by a user operation. When the reset button 2008 is pressed, the generating system 201 resets the information selected on the selection screen 2000 (pressing a button, inputting a value).
[0202] [Figure 21 Idea area 104] 21 is an explanatory diagram showing an example of the idea area 104. When the third-area word selection instruction button 2002 is pressed, the idea area 104 is an area specified by a combination of the first-area word AX and the third-area word DX that has similarity to the first-area word AX.
[0203] In this case, the idea area 104 is an area that associates the first area 101, which is familiar to the user, with the second area 102, which is unfamiliar to the user, so that the user can easily come up with new ideas by referring to the first area words AX and the third area words DX that constitute the idea area 104A.
[0204] When the co-occurrence word selection instruction button 2003 is pressed, the idea area 104 is an area specified by a combination of a first domain word AX, a third domain word DX having similarity to the first domain word AX, and a third domain co-occurrence word EX having similarity to the first domain word AX and having a co-occurrence relationship with the third domain word DX.
[0205] In this case, the idea area 104 becomes an area that associates the first area 101 familiar to the user with the second area 102 unfamiliar to the user, so that the user can easily get new ideas by referring to the first area words AX, the third area words DX, and the third area co-occurring words EX that constitute the idea area 104B.
[0206] (Step S709) Returning to Fig. 7, the generation system 201 outputs the ideation area 104 in a displayable manner. Specifically, for example, the generation system 201 displays the ideation area 104 on a display, which is an example of the output device 304, or transmits the ideation area 104 to the terminal 202 via the communication IF 305 to display the ideation area 104 on the display of the terminal 202. This completes the series of processes.
[0207] <Figure 22 Idea area display screen> 22 is an explanatory diagram showing an example of an idea area display screen. The idea area display screen 2200 displays an idea area list 2201. The idea area list 2201 has a first area 101 (technology ST or product SP, first area word AX), a third area 103 (third area co-occurring word EX, third area word DX), additional keyword CX, and a second area 102 (second area word BX, task N) for each idea area number 2202. The idea area number 2202 is a number that uniquely identifies the idea area 104.
[0208] For example, when referring to an entry with the idea area number 2202 of "1", the idea area 104 is a combination of "biodegradation" A2, which is the first domain word AX, and "biodegradable" D4, which is the third domain word DX. "Biodegradable" D4, which is the third domain word DX, is a component that connects "biodegradation" A2, which is the first domain word AX known to the user, and "microplastic" B1, which is the second domain word BX of the second domain 102 different from the first domain 101. Therefore, by referring to "biodegradation" A2 and "biodegradable" D4 that constitute the idea area 104, the user can easily obtain a new idea such as "Attempt to solve the problem N "environmental pollution", which is the source of the second domain word "microplastic" B1, by using synthetic fiber technology, which is the source of the first domain word "biodegradation" A2".
[0209] Also, when referring to the entry with the idea area number 2202 of "3", the idea area 104 is a combination of "biodegradable" A4, which is the first domain word AX, "deterioration" D3, which is the third domain word DX, and "biodegradable" D4, which is the third domain co-occurring word EX. "Deterioration" D3, which is the third domain word DX, is a component that connects "biodegradable" A4, which is the first domain word AX known to the user, and "microplastic" B1, which is the second domain word BX of the second domain 102 different from the first domain 101.
[0210] Moreover, such a component "deterioration" D3 is in a co-occurring relationship with "biodegradable" D4, which is a third domain co-occurring word EX. Therefore, by referring to "biodegradable" A4, "biodegradable" D4 (EX), and "deterioration" D3 (DX) that constitute the idea domain 104, the user can easily obtain a new idea such as "trying to solve the problem N "environmental pollution" that is the source of the second domain word "microplastic" B1 by using environmentally friendly plastic products and related technologies that are the source of the first domain word "biodegradable" A4."
[0211] As described above, according to this embodiment, a mechanism can be provided that makes it easier for users to come up with ideas by linking the strengths (seeds) of the company with the issues (needs) of customers and society.
[0212] The present invention is not limited to the above-described embodiments, and includes various modified examples and equivalent configurations within the spirit of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the configurations described. Also, a part of the configuration of one embodiment may be replaced with a configuration of another embodiment. Also, a configuration of another embodiment may be added to a configuration of one embodiment. Also, a part of the configuration of each embodiment may be added, deleted, or replaced with another configuration.
[0213] Furthermore, each of the aforementioned configurations, functions, processing units, processing means, etc. may be realized in hardware, for example by designing some or all of them as an integrated circuit, or may be realized in software by a processor interpreting and executing a program that realizes each function.
[0214] Information such as programs, tables, files, etc. that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or in a recording medium such as an IC (Integrated Circuit) card, an SD card, or a DVD (Digital Versatile Disc).
[0215] In addition, the control lines and information lines shown are those considered necessary for the explanation, and do not necessarily show all the control lines and information lines necessary for implementation. In reality, it can be considered that almost all components are connected to each other. [Explanation of symbols]
[0216] 101 First area 102 Second area 103 Third area 104 Idea Area 200 Network Systems 201 Generator System 202 Terminal 601 Input section 602 Output section 603 Control Unit 604 Storage section 631 Information Gathering Department 632 Information Analysis Department 632A Morphological analysis function 632B Co-occurrence probability calculation function 632C Similarity calculation function 633 Generation part 641 Information Collection Results 641A Additional Keywords 641B Document Data Group 642 Information analysis results 801 Second Area Word List 802 Additional Keyword List 1100 Search Keywords 1111D Third Area Word Group 1111X Extracted document data set 1112 Third Area Table 1113 Third domain word occurrence count list 1114 Third Domain Word Co-occurrence Probability List 1130A~1130C Information source 1131,1132,1133 Document data set A, AX 1st domain word B, BX Second domain words C, CX Additional Keywords D, DX Third Area Words EX Third Area Co-occurring Words
Claims
1. A generating system having a processor for executing a program and a storage device for storing the program, a first database having first words belonging to a first domain and a second database having second words belonging to a second domain different from the first domain; The processor, a search process for searching an information source for document data including a search keyword composed of the second word and an additional keyword; an extraction process for extracting a third word from the document data searched by the search process; a generating process for generating combinations of the first words and the third words extracted by the extracting process; a first output process for outputting the combination generated by the generation process in a displayable manner; A generating system comprising:
2. 2. The production system of claim 1, The processor, execute a first designation process for accepting designation by a user of access information for accessing the information source; In the search process, the processor searches for the document data from the information source for which the access information is specified by the first specification process. A generating system comprising:
3. 3. The production system of claim 2, The processor, executing a second output process for displayably outputting information about other users or organizations that have accessed the information source; A generating system comprising:
4. 3. The production system of claim 2, The processor, executing a second output process for displayably outputting access information for accessing each of the plurality of information sources based on the number of past accesses; A generating system comprising:
5. 2. The production system of claim 1, The processor, executing a selection process for selecting a specific third word from the document data based on the third word extracted by the extraction process; In the generating process, the processor generates a combination of the first word and the specific third word selected by the selecting process. A generating system comprising:
6. 6. The production system of claim 5, The processor, execute a second designation process of accepting designation of a specific additional keyword from among the specific additional keywords; In the selection process, the processor selects the specific third word from specific document data that includes the specific additional keyword designated by the second designation process among the document data. A generating system comprising:
7. 7. A generating system according to claim 6, comprising: In the selection process, the processor selects the specific third word from the specific document data based on the number of occurrences of the third word in the specific document data. A generating system comprising:
8. 7. A generating system according to claim 6, comprising: In the selection process, the processor selects the specific third word from the specific document data based on a co-occurrence probability of the third word and the specific additional keyword co-occurring in the same sentence in the specific document data. A generating system comprising:
9. 6. The production system of claim 5, In the selection process, the processor selects the specific third word from specific document data that includes the additional keyword among the document data. A generating system comprising:
10. 10. The production system of claim 9, In the selection process, the processor selects the specific third word from the specific document data based on the number of occurrences of the third word in the specific document data. A generating system comprising:
11. 6. The production system of claim 5, In the selection process, the processor selects the specific third word from the document data based on the number of occurrences of the third word extracted in the extraction process. A generating system comprising:
12. 6. The production system of claim 5, In the generating process, the processor generates the combination based on a similarity between the first word and a specific third word. A generating system comprising:
13. 6. The production system of claim 5, In the generation process, the processor generates combinations of the first words and words that co-occur with the specific third words in the document data based on a similarity between the first words and words that co-occur with the specific third words. A generating system comprising:
14. A generation method executed by a generation system having a processor that executes a program and a storage device that stores the program, comprising: a first database having first words belonging to a first domain and a second database having second words belonging to a second domain different from the first domain; The processor, a search process for searching an information source for document data including a search keyword composed of the second word and an additional keyword; an extraction process for extracting a third word from the document data searched by the search process; a generating process for generating combinations of the first words and the third words extracted by the extracting process; a first output process for outputting the combination generated by the generation process in a displayable manner; A generating method comprising:
15. a processor having access to a first database having first words belonging to a first domain and a second database having second words belonging to a second domain different from the first domain; a search process for searching an information source for document data including a search keyword composed of the second word and an additional keyword; an extraction process for extracting a third word from the document data searched by the search process; a generating process for generating combinations of the first words and the third words extracted by the extracting process; a first output process for outputting the combination generated by the generation process in a displayable manner; A generating program for causing a user to execute the above steps.
Citation Information
Patent Citations
Idea support device, method and program
JP2013125454A
Information processing device, information processing system, and program
JP2017091270A
Cited By
Text information processing method and electronic equipment
CN121166941A