Program, information processing system, data processing system, and data processing method

The information processing system addresses the challenge of analyzing trends across multiple data sources by creating cumulative graphs and using an iMap to identify keywords nearing commercialization, enhancing the accuracy of trend detection in technological development.

JP2025145718APending Publication Date: 2025-10-03RICOH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024046043
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies lack a method for analyzing a common analysis target contained in multiple data sources, such as academic papers and patent documents, to accurately identify trends and signs of technological development, particularly distinguishing between early-stage research and pre-commercialization phases.

Method used

An information processing system that acquires data from multiple sources, creates cumulative graphs and scatter plots, and uses an iMap to analyze the transition patterns of keyword frequencies, enabling the identification of keywords nearing commercialization by comparing iMap coordinates and link lengths.

Benefits of technology

Enables accurate identification of keywords indicating technological development nearing commercialization, providing a higher accuracy in detecting trends and future hot topics by analyzing common analysis targets across multiple data sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025145718000001_ABST
    Figure 2025145718000001_ABST
Patent Text Reader

Abstract

To analyze common analytical targets contained in multiple data sources.SOLUTION: The present invention is provided with a data processing unit for extracting one or more common analytical targets from first data and second data acquired by a data acquisition unit, a graph creation unit for creating a first graph 305 and a second graph 307 representing cumulative values over a period for values related to the analytical targets in the first data and the second data, an area calculation unit for calculating a third area using a first area formed by a graph created by the graph creation unit with the x-axis and a second area formed by a quadratic function of the graph with the x-axis, and a scatter plot creation unit for arranging a first data point associated with the first area and the third area calculated based on the first graph and a second data point associated with the first area and the third area calculated based on the second graph on a scatter plot 210 with the x-axis as the first area and the y-axis as the third area.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a program, an information processing system, a data processing system, and a data processing method. [Background technology]

[0002] There are times when you want to extract various analytical targets, such as keywords contained in time-series data, and analyze recent topics. For example, analyzing the frequency of keyword appearances is a method for discovering trends, trends, and fashions of the times. Analyzing the frequency of appearances of each keyword makes it easier to discover topics, themes, popularity, and interesting events that have recently increased in various fields.

[0003] There is a known technology for extracting social trends from a collection of text (see, for example, Patent Document 1). Patent Document 1 discloses a system that determines the topicality (degree of rapid rise) of each term in a target document by comparing the frequency of appearance of the term in the distant past with the frequency of appearance in the recent past, and presents documents that are highly relevant to terms with a high degree of rapid rise to the user. Summary of the Invention [Problem to be solved by the invention]

[0004] However, the prior art does not describe a technique for analyzing a common analysis target contained in multiple data sources.

[0005] In view of the above-mentioned problems, the present invention provides a technique for analyzing a common analysis target contained in a plurality of data sources. [Means for solving the problem]

[0006] In view of the above-mentioned problems, the present invention provides an information processing system including: a data acquisition unit that acquires, from a first data source and a second data source, first data and second data, respectively, with one or more of year, month, day, hour, minute, and second associated with each other; a data processing unit that extracts one or more common analysis targets from the first data and the second data acquired by the data acquisition unit; a graph creation unit that creates a first graph and a second graph, which are cumulative values ​​over a period of values ​​related to the analysis target of the first data and values ​​related to the analysis target of the second data; an area calculation unit that calculates a third area using a first area formed by the graph created by the graph creation unit and its x-axis, and a second area formed by a square function of the graph and its x-axis; and a scatter plot creation unit that arranges, in a scatter plot with the first area as the x-axis and the third area as the third area, first data points that associate the first area calculated based on the first graph with the third area, and second data points that associate the first area calculated based on the second graph with the third area. [Effects of the Invention]

[0007] The present invention can analyze a common analysis target contained in multiple data sources. [Brief explanation of the drawings]

[0008] [Figure 1] This is a diagram explaining the process of extracting multiple keywords from papers in a specific technical field and detecting changes and signs of these changes. [Figure 2] FIG. 1 is a diagram showing a schematic diagram of a technology and a business cycle using that technology over time. [Figure 3] This is an example of a chart showing the hype cycle and the cumulative frequency of keywords that are on the rise in papers. [Figure 4] FIG. 1 illustrates the relationship between the Hype Cycle and a word frequency graph of keywords across multiple data sources. [Figure 5] FIG. 10 is a diagram illustrating a process of analyzing the transition of the appearance frequency of a predetermined keyword until it increases. [Figure 6]FIG. 1 illustrates iMap coordinates of the same keyword extracted from two data sources. [Figure 7] 10 is an example of a diagram showing the length of a link between iMap coordinates connecting iMap coordinates of keywords extracted from academic papers and iMap coordinates of keywords extracted from patent documents. [Figure 8] FIG. 1 is a diagram illustrating an example of a system configuration of a data processing system. [Figure 9] FIG. 2 is a diagram illustrating an example of a hardware configuration of an information processing system and a terminal device. [Figure 10] FIG. 1 illustrates an example of a functional configuration of a data processing system. [Figure 11] FIG. 10 is a diagram showing an example of data including keywords used in analyzing changes in appearance frequency. [Figure 12] FIG. 8 is a diagram showing an example of data obtained by converting the number of occurrences of FIG. 7 into a cumulative value. [Figure 13] FIG. 9 is a diagram showing an example of data in which the cumulative value and period of the data in FIG. 8 are normalized. [Figure 14] 10 is an example of a graph showing the transition of the appearance frequency of a certain keyword, with the normalized period as the x-axis and the normalized cumulative value as the y-axis. [Figure 15] FIG. 1 shows an example of a word frequency graph for two different keywords. [Figure 16] 10A and 10B show variations of word frequency graphs with an area A1 of 0.5. [Figure 17] FIG. 10 is a diagram illustrating a pattern analysis of the shape of a word frequency graph based on area A1 and area A2. [Figure 18] FIG. 10 is a diagram showing an example of a plurality of word appearance frequency graphs in which the area A1 gradually changes. [Figure 19] This is an example of a scatter plot of area B2 against area A1. [Figure 20] FIG. 10 is a diagram showing, on an iMap, areas A1 and B2 calculated for keywords that appear more than a certain number of times. [Figure 21]FIG. 10 is a diagram illustrating the correspondence between iMap coordinates created for keywords derived from academic papers and iMap coordinates created for keywords derived from patent documents. [Figure 22] FIG. 1 is a diagram showing an example of an iMap showing the iMap coordinates of keywords that commonly appear in papers and patent documents related to "CFRP (carbon fiber reinforced plastic)." [Figure 23] This is an example of a diagram showing three iMaps with different initial years for the period when creating a word frequency graph. [Figure 24] FIG. 24 is an example of a diagram showing a list of keywords with iMap coordinates shown in FIG. 23. [Figure 25] This is an example of an iMap and word frequency graph for the keyword "high electromagnetic" with the first years of the period being 1969, 2000, and the first year of the paper. [Figure 26] This is an example of an iMap and a word frequency graph for the keyword "structural form" with 1969, 2000, and the first year of the paper as the first year of the period. [Figure 27] This is an example of an iMap and a word frequency graph for the keyword "high speed train" with 1969, 2000, and the first year of the paper as the first year of the period. [Figure 28] This is an example of an iMap and a word frequency graph for the keyword "c / sic composite material" with 1969, 2000, and the first year of the paper as the first year of the period. [Figure 29] 10 is a sequence diagram illustrating an example of a process in which the data processing system extracts an iMap and keywords based on the frequency of appearance of the keywords; [Figure 30] FIG. 2 is a diagram illustrating an example of a data source. DETAILED DESCRIPTION OF THE INVENTION

[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An information processing system and a data processing method performed by the information processing system will be described below as an example of an embodiment of the present invention with reference to the drawings.

[0010] <Regarding time series data trends> There are times when you want to analyze what is currently trending in a certain field, or what is likely to become a hot topic in the future, based on keywords used in that field. For example, if you can detect changes and signs of what technologies are trending in a specific technical field, you can take early action, such as selecting them as future research topics.

[0011] Figure 1 explains the process of extracting multiple keywords from papers in a specific technical field and detecting changes and signs of change. Since papers are generally published in typical technical fields, the keywords contained in these papers are thought to indicate changes in topics and technologies that are likely to become popular in the future. Therefore, as shown in Figure 1, it is effective for the information processing system described below to extract multiple keywords from papers and display how the frequency of each keyword has changed over time. Figure 1 shows the changes in the frequency of appearance of keywords A and B, and it can be seen that keyword A has gradually increased in frequency, while keyword B has recently increased in frequency. For example, if a researcher can identify keyword B, they can consider selecting keyword B as their next research topic.

[0012] Figure 2 is a diagram that illustrates the time cycle of a business that uses technology. A graph showing the maturity, adoption, and societal applicability of a particular technology over time is called a hype cycle (201). As the hype cycle (201) shows, research on a technology generally progresses through experimental studies by many researchers in its early stages. Once research is complete, interest often dies down. Once the market prospects and profitability are established, the technology is commercialized. Once commercialized, research continues for commercialization, and commercialization (products and services) also progresses, as shown in the product lifecycle (202). Before the product lifecycle (202) begins, the hype cycle (201) also shows an upward trend (203). Therefore, the upward trend (203) after the hype cycle (201) temporarily dies down indicates the timing when research on the technology begins again. Companies that can capture this upward trend (203) can commercialize the technology earlier, providing a business advantage. The rising portion 203 can be said to be a sign that the technology in question will become a hot topic in a practical sense, and the data analysis method of this embodiment makes it possible to capture this sign, for example.

[0013] However, as shown in Figure 3, it is difficult to perform a detailed analysis of the target's appearance pattern solely based on keywords showing an increasing trend in the hype cycle 201. Figure 3 shows the cumulative frequency of keywords showing an increasing trend in the hype cycle and in papers. Figure 3(a) shows the cumulative value of the word frequency of keywords (called a word frequency graph). The keywords showing an increasing trend in the papers shown in Figure 3(a) correspond to the rising section 203 of the hype cycle 201 in Figure 3(b). However, the hype cycle 201 often shows rising sections 301 and 203 at least in two places: the dawn (early stage of technology development) and the stage after the disillusionment phase and before commercialization. Thus, simply identifying keywords showing an increasing trend in papers does not allow us to determine whether the signs of the technology being discussed in the papers are the rising section 301 of the hype cycle 201 in the dawn (early stage of technology development) or the rising section 203 after the disillusionment phase and before commercialization. In other words, since the Hype Cycle 201 shows an upward trend both in the early stages of technology development and before the transition to commercialization, it can sometimes be difficult to distinguish between the two.

[0014] <Analysis Method of the Present Embodiment> Therefore, the information processing system of this embodiment creates multiple time-series data for keywords common to multiple data sources and identifies signs of pre-commercialization, making it possible to grasp signs of pre-commercialization with higher accuracy than when identifying signs from a single data source.

[0015] An overview will be provided with reference to Figure 4. Figure 4 illustrates the relationship between a hype cycle 201 and word frequency graphs of keywords in multiple data sources. Figure 4(a) is the same word frequency graph as Figure 3(a), and Figure 4(c) is the same hype cycle 201 as Figure 3(b). Figure 4(b) shows word frequency graphs 305, 306, and 307 of common keywords in two data sources: papers and patent documents. Word frequency graphs 305 and 306 show the cumulative values ​​of the frequency of keywords extracted from papers, and word frequency graph 307 shows the cumulative values ​​of the frequency of keywords (common to either word frequency graph 305 or 306) extracted from patent documents.

[0016] The word frequency graphs 305 and 306 in Figure 4(b) show the word frequency graphs of keywords indicating that technological development has reached a plateau and is about to move toward commercialization. In other words, the word frequency graphs 305 and 306 indicate that the time series pattern of paper data indicating the progress of technological development is "saturated, with almost no increase in recent years." In contrast, the word frequency graph 307 of keywords extracted from patent documents indicates that the time series pattern of patent data indicating the progress toward practical application is "rapidly increasing in recent years, with rapid increase." Therefore, if the word frequency graphs 305 or 306 and the word frequency graph 307 can be identified from the time series data, it is possible to efficiently identify keywords indicating that technological development has reached a plateau and is about to move toward commercialization.

[0017] However, it is not easy to create word frequency graphs 305, 306, and 307 for common keywords in two data sources (papers and patent documents) for each keyword and graph them as shown in Figure 4(b). Even if graphs were created, it would be necessary for a human to compare the word frequency graphs and determine for each keyword whether technological development has reached a plateau and whether the keyword is being commercialized.

[0018] Therefore, in this embodiment, an iMap is introduced as shown in Fig. 5. Fig. 5 is a diagram for explaining the process of analyzing the transition until the appearance frequency of a predetermined keyword increases, and an iMap.

[0019] (1) As shown in FIG. 5(a), the information processing system performs morphological analysis (word segmentation) on the title of a paper, for example, and calculates the frequency of appearance of each keyword included in the title by publication year.

[0020] (2) The information processing system normalizes the publication year and cumulative value of the frequency of appearance for each keyword to between 0 and 1, respectively, so that the trends in frequency of appearance between keywords can be compared. As shown in Figures 5(b-1) and 5(b-2), a word frequency graph is obtained, with the publication year x on the horizontal axis and the cumulative value of frequency of appearance y on the vertical axis. By graphing in this way, the trends in the frequency of appearance of keywords can be visualized. Figure 5(b-1) shows keywords whose frequency of appearance has increased rapidly in recent years, while Figure 5(b-2) shows keywords whose frequency of appearance increased earlier but have rarely been used in recent years.

[0021] (3) The information processing system uses a map to analyze which pattern the shape of the word frequency graph in (2) closely resembles. As will be described in detail later, the information processing system calculates the area formed by the graph and the x-axis, and represents the shape of the word frequency graph for each keyword as a position on the map (Figure 5(c)). Hereinafter, this map will be referred to as iMap 210. iMap 210 indicates the transition pattern of the cumulative value of the occurrence frequency for a single keyword using the positions (iMap coordinates) of data points (e.g., 211-214). The position of the data points determines how the cumulative value of the occurrence frequency has changed over time. For example, the keyword corresponding to data point 212 on the left side of iMap 210 indicates a keyword that has rapidly increased in recent years. The keyword corresponding to data point 214 on the right side indicates a keyword that increased in popularity in the past but was rarely popular. The keyword corresponding to data point 211 on the top side indicates a keyword that suddenly increased in frequency at one time but is now obsolete. The keyword corresponding to the data point 213 at the bottom indicates that it is a keyword that has been around for a long time but fell out of fashion at one point and is now beginning to attract attention again.

[0022] Therefore, the user can understand how the frequency of a keyword has changed by determining where the keyword data point is on iMap 210. For example, it is possible to extract signs of what technologies are trending in a particular technical field and what technologies are likely to become trending in the future.

[0023] Next, we will explain how data points for the keyword "technological development has reached a plateau and is about to move to commercialization" can be analyzed using iMap 210. The information processing system identifies this keyword by the length of the link connecting the iMap coordinates.

[0024] Figure 6 illustrates the iMap coordinates of the same keyword extracted from two data sources. Word frequency graphs 305 and 306 (an example of a first graph) in Figure 6 are created from keywords extracted from academic papers, and word frequency graph 307 (an example of a second graph) is created from keywords extracted from patent documents. As explained in Figure 5, a sudden increase in frequency of occurrence belongs to the leftmost region 311 on the iMap, while a saturation increase belongs to the rightmost region 313 or the topmost region 312 on the iMap. When data points of keywords common to academic papers and patent documents are placed on the same iMap, if the iMap coordinates of the keyword extracted from the patent document are in the leftmost region 311 and the iMap coordinates of the keyword extracted from the academic paper are in the rightmost region 313 or the topmost region 312, it can be determined that the keyword is close to being put into practical use.

[0025] 7 also shows the lengths of links 321, 322 connecting the iMap coordinates of keywords extracted from academic papers and those extracted from patent documents. There is a certain distance between the leftmost region 311 and the rightmost region 313, or between the leftmost region 311 and the topmost region 312. Therefore, for keywords that are close to practical application, the links 321, 322 connecting the iMap coordinates of keywords extracted from academic papers and those extracted from patent documents will be longer.

[0026] In this way, the information processing system of this embodiment can analyze common keywords contained in multiple data sources. As an example of this analysis, it is possible to identify keywords that show signs of being close to practical application by utilizing the length of the links 321, 322 connecting two iMap coordinates.

[0027] <Terminology> Keywords derived from academic papers are keywords extracted from academic papers, and keywords derived from patent documents are keywords extracted from patent documents.

[0028] The data may be anything that includes analysis targets such as keywords and numerical values ​​whose changes over time are analyzed. It is preferable that the data or analysis targets are associated with one or more of the following: year, month, day, hour, minute, and second.

[0029] The analysis target is an object whose change over time is analyzed, and is, for example, a keyword or a numerical value.

[0030] The value related to the analysis target may be any value that can be extracted by processing, such as a value directly included in the analysis target contained in the data, or a value obtained by processing the analysis target. In this embodiment, the value related to the analysis target is explained using the terms "number of occurrences" and "number of transactions."

[0031] <System configuration example> 8 is an example of a system configuration diagram of a data processing system 100. The data processing system 100 includes an information processing system 10 and a terminal device 30. However, the terminal device 30 may be a general-purpose computer and may not be included in the data processing system 100.

[0032] The information processing system 10 and the terminal device 30 are communicatively connected via a wide-area network N1 such as the Internet. The information processing system 10 may be installed in a cloud or a data center, or may be installed on-premise. The information processing system 10 may be a web server for the terminal device 30 that returns processing results in response to requests from the terminal device 30. A server is a computer or software that provides information and processing results in response to requests from clients.

[0033] For example, the information processing system 10 extracts keywords from data (such as titles of multiple documents) specified by the user 9 via the terminal device 30, and displays the transition in the frequency of appearance of each keyword on the iMap 210. Alternatively, the information processing system 10 displays on the iMap 210 how the frequency of appearance of keywords specified by the user via the terminal device 30 has changed within certain data.

[0034] The information processing system 10 may have a data storage device in which data specified by the user for analysis is stored in advance. Alternatively, the information processing system 10 may acquire data from a data server or a NAS (Network Attached Storage). Furthermore, the information processing system 10 may acquire data by performing web scraping over a network. Alternatively, the user may send data to be analyzed from a terminal device 30 to the information processing system 10.

[0035] The information processing system 10 may be compatible with cloud computing. Cloud computing refers to a usage mode in which resources on a network are used without regard to specific hardware resources. Therefore, the information processing system 10 does not need to be housed in a single housing or provided as a single integrated device. The functions of the information processing system 10 may be distributed among multiple information processing devices, or multiple information processing devices may each have all of the functions and the information processing device that processes the information may be switched by load balancing or the like.

[0036] The terminal device 30 is placed in a facility such as a company, an educational institution, or a factory, and is connected to the network N2. The network N2 may be any network that can communicate with the information processing system 10, such as a LAN, Wi-Fi (registered trademark), wide area Ethernet (registered trademark), or a mobile phone network such as 4G, 5G, or 6G.

[0037] The terminal device 30 is a general-purpose computer used by a user. Here, the user refers to a person who uses the information processing system 10. Therefore, a person who uses the information processing system 10 may be any person who wants to analyze the transition of the appearance frequency of keywords, etc. Furthermore, the user may include a person who registers data to be analyzed in the information processing system 10, etc.

[0038] A web browser and a native application dedicated to the information processing system 10 run on the terminal device 30. When the terminal device 30 runs a web browser, the terminal device 30 and the information processing system 10 run a web application. A web application is an application that runs through cooperation between a program written in a programming language (for example, JavaScript (registered trademark)) that runs on the web browser and a program on the web server (information processing system 10) side. When the web application is running, the information processing system 10 may analyze the transition in the frequency of appearance of keywords, or the terminal device 30 that has received the web application may analyze the transition in the frequency of appearance of keywords.

[0039] An application that cannot be executed unless it is installed on the terminal device 30 is called a native application. In this embodiment, the application executed on the terminal device 30 may be a web application or a native application. In this case, too, the analysis of the transition in the appearance frequency of keywords may be performed by the information processing system 10, or may be performed by the terminal device 30 using a native application.

[0040] Furthermore, in this embodiment, the information processing system 10 is described as analyzing the transition of the frequency of appearance of keywords, but as shown in FIG. 8(b), the terminal device 30 may analyze the data by itself. In this case, a native application that analyzes the transition of the frequency of appearance of keywords included in the data runs on the terminal device 30. However, the terminal device 30 may acquire data from a data server on a network. Therefore, even in the configuration of FIG. 8(b), it is preferable that the terminal device 30 be able to connect to a network.

[0041] The terminal device 30 may be, for example, a desktop PC, a notebook PC, a smartphone, a PDA (Personal Digital Assistant), a tablet terminal, or the like used by a user. Alternatively, the terminal device 30 may be any device on which a web browser or a native application runs. The terminal device 30 may also be an electronic whiteboard, a video conference terminal, or the like.

[0042] In this embodiment, unless otherwise specified, the description will be based on the configuration of FIG. 8(a).

[0043] <Hardware configuration example> With reference to FIG. 9, the hardware configuration of the information processing system 10 and the terminal device 30 included in the data processing system 100 according to this embodiment will be described.

[0044] <<Information processing system and terminal device>> Fig. 9 is a diagram showing an example of the hardware configuration of the information processing system 10 and the terminal device 30 according to this embodiment. As shown in Fig. 9, the information processing system 10 and the terminal device 30 are constructed by a computer 500, and include a CPU 501, a ROM 502, a RAM 503, a HD (Hard Disk) 504, an HDD (Hard Disk Drive) controller 505, a display 506, an external device connection I / F (Interface) 508, a network I / F 509, a bus line 510, a keyboard 511, a pointing device 512, an optical drive 514, and a media I / F 516.

[0045] Of these, the CPU 501 controls the overall operation of the information processing system 10 and the terminal device 30. The ROM 502 stores programs, such as an IPL, used to drive the CPU 501. The RAM 503 is used as a work area for the CPU 501. The HD 504 stores various data, such as programs. The HDD controller 505 controls the reading and writing of various data from and to the HD 504 under the control of the CPU 501. The display 506 displays various information, such as a cursor, menus, windows, characters, or images. The external device connection I / F 508 is an interface for connecting various external devices. In this case, external devices include, for example, USB (Universal Serial Bus) memories and printers. The network I / F 509 is an interface for data communication using the network N2. The bus line 510 is an address bus, a data bus, or the like, for electrically connecting the components, such as the CPU 501, shown in FIG. 9.

[0046] The keyboard 511 is a type of input means having multiple keys used to input characters, numbers, various instructions, etc. The pointing device 512 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, etc. The optical drive 514 controls reading and writing of various data from and to CDs, DVDs, and Blu-ray (registered trademark), which are examples of removable optical recording media 513. The media I / F 516 controls reading and writing (storing) of data from and to a recording medium 515, such as a flash memory.

[0047] <About the function> Next, the functional configuration of the data processing system 100 according to this embodiment will be described with reference to Fig. 10. Fig. 10 is a functional configuration diagram of an example of the data processing system 100 according to this embodiment.

[0048] <<Functional configuration of information processing system>> The information processing system 10 includes a communication unit 11, a data acquisition unit 12, a data processing unit 13, an occurrence frequency calculation unit 14, a normalization unit 15, a graph creation unit 16, an area calculation unit 17, a scatter plot creation unit 18, a distance calculation unit 19, and a screen generation unit 20. Each of these units is a function or a means for performing the function that is realized when any of the components shown in Fig. 9 operates in response to an instruction from the CPU 501 in accordance with a program loaded in the RAM 503. The details of each function will be described later with reference to the drawings.

[0049] The communication unit 11 transmits and receives various types of information to and from the terminal device 30. In this embodiment, the communication unit 11 transmits a web application and an iMap, which is an analysis result, to the terminal device 30, and receives various operation contents and instructions from the user.

[0050] The data acquisition unit 12 acquires data including keywords that are the subject of analysis of changes in appearance frequency. In this embodiment, the data acquisition unit 12 acquires data from two data sources. One data source is, for example, a collection of papers, and the other data source is, for example, patent documents. Since the data is analyzed as time-series data, it is preferable that a plurality of papers or patent documents are acquired. The data acquisition unit 12 may receive data from a terminal device 30, or may acquire data from a NAS or a data server. The data acquisition unit 12 may also acquire data by web scraping. In this embodiment, the data acquisition unit 12 acquires two different data.

[0051] The data processing unit 13 extracts keywords by performing morphological analysis on the data as needed and converting it into regular expressions. The data processing unit 13 may extract keywords from a specific column of a table in tabular format, and morphological analysis may not be necessary. For example, morphological analysis may not be necessary if the data format is formatted in a table format (XML, JSON, CSV, etc.).

[0052] The occurrence frequency calculation unit 14 counts the occurrence frequency of each keyword obtained by morphological analysis for each unit period and converts it into a cumulative value for the period. The unit period is the period for counting the number of occurrences, and in the paper analysis described below, it is one year. The unit period may be set appropriately depending on the analysis target and purpose, such as year, month, day, hour, minute, and second. Furthermore, the period is the total period from the beginning to the end of the unit period.

[0053] The normalization unit 15 normalizes the cumulative value and the period for each keyword so that the minimum value is 0 and the maximum value is 1. The normalized period may be the same for all keywords, even if the year of first appearance, etc., differs depending on the keyword.

[0054] The graph creation unit 16 creates a graph (a word occurrence frequency graph, described later) for each keyword, with the period on the x-axis and the cumulative value of the occurrence frequency on the y-axis. In addition, the graph creation unit 16 also creates a graph with the square of the cumulative value on this graph on the y-axis, so that the transition pattern of the occurrence frequency can be distinguished. A graph in which the cumulative value of the original graph is squared is called a square function.

[0055] The area calculation unit 17 calculates the area (this area is called A1) formed between the word appearance frequency graph and the x-axis (area A1 is an example of the first area). The area calculation unit 17 calculates the area (this area is called A2) formed between the square function and the x-axis. By calculating area A2 (area A2 is an example of the second area), it becomes easier to distinguish transition patterns of appearance frequency for keywords that are difficult to distinguish using area A1 alone. Furthermore, the area calculation unit 17 converts area A2 into area B2 for iMap (area B2 is an example of the third area).

[0056] The scatter diagram creation unit 18 creates a scatter diagram with area A1 on the x-axis and area A2 on the y-axis, and arranges data points of area A2 corresponding to area A1 on this scatter diagram. Similarly, the scatter diagram creation unit 18 creates a scatter diagram (iMap) with area A1 on the x-axis and area B2 on the y-axis, and arranges data points of area B2 corresponding to area A1 on this scatter diagram. In either scatter diagram, the position of the data points indicates the trend of how the appearance frequency has increased over time, so the user can understand how the appearance frequency of any keyword has changed over time by the position of the data points.

[0057] The distance calculation unit 19 calculates the distance of the link between the iMap coordinates of the keywords derived from the paper and the iMap coordinates of the keywords derived from the patent document.

[0058] The screen generation unit 20 generates screen information to be displayed by the terminal device 30. When the terminal device 30 executes a web application, the screen information is created using HTML, XML, CSS (Cascade Style Sheet), JavaScript (registered trademark), etc. When the terminal device 30 executes a native application, the screen information is held by the terminal device 30, and the information to be displayed is transmitted in XML, etc.

[0059] <<Terminal Device>> The terminal device 30 is used by a user who wishes to analyze the transition of the frequency of appearance of keywords. The terminal device 30 has a communication unit 31, a display control unit 32, and an operation reception unit 33. Each of these functional units is a function or means realized by the CPU 501 executing instructions contained in one or more programs installed in the terminal device 30. Note that this program may be a web application executed by a web browser, or may be a dedicated native application.

[0060] The communication unit 31 transmits and receives various types of information to and from the information processing system 10. In this embodiment, the communication unit 31 receives screen information such as a Web application or iMap 210 from the information processing system 10, and transmits user operations and instructions to the information processing system 10.

[0061] The display control unit 32 interprets screen information of various screens and displays it on the display 506. The operation accepting unit 33 accepts various operations on the various screens displayed on the display 506 by the user.

[0062] <Normalization of occurrence frequency> The flow of pattern analysis performed by the information processing system 10 will be described in detail below with reference to the drawings. The following describes, as an example, the process in which the information processing system 10 analyzes the transition in the frequency of appearance of keywords contained in technical literature using technical literature as data. In this embodiment, technical literature is, for example, a paper or a patent document. However, the data analysis method of this embodiment can be applied to any data (time-series data) having at least one of the following time periods: year, month, day, hour, minute, and second, acquired from a data source that can be classified into the early stage and the pre-commercialization stage. Furthermore, analysis of the transition in the frequency of appearance of keywords is not limited to analysis of the transition in the frequency of appearance of keywords, but can also be applied to analysis of the transition in any time-series value. Figures 11 to 20 explain the creation of an iMap using keywords derived from papers as an example. Keywords derived from patent documents are processed in the same way.

[0063] FIG. 11 shows an example of the transition in the frequency of appearance of each keyword extracted from a paper. The data acquisition unit 12 acquires files as data from a network or a terminal device 30. The data processing unit 13 extracts the title (text data) and publication year of the paper from, for example, a file specified by the user, and performs morphological analysis on the title as necessary. In FIG. 11, a compound word that has one meaning made up of multiple keywords is acquired, but in this embodiment, it is simply referred to as a keyword. The transition in the frequency of appearance of a single keyword is also possible. The occurrence frequency calculation unit 14 calculates the occurrence frequency for each of these keywords by publication year (an example of a unit period).

[0064] In FIG. 11, for example, the number of times each keyword appears is shown for each year of publication of the paper from 1989 to 2020. In FIG. 11, the number of times the keyword appears for one year is calculated, but the unit period for calculating the number of times the keyword appears can be any period, such as one month, one week, or one day. A single unit period may also be multiple years. Furthermore, if the keywords in the data to be analyzed are associated with time, minutes, or seconds, the unit period for calculating the number of times the keyword appears can also be hours, minutes, or seconds.

[0065] Next, as shown in Fig. 12, the occurrence frequency calculation unit 14 converts the data in Fig. 11 into cumulative values. Fig. 12 shows data in which the number of occurrences in Fig. 11 has been converted into cumulative values. Because it is a cumulative value, the number of occurrences does not decrease over time. By converting it into a cumulative value, the pattern of change in the occurrence frequency does not decrease over time, making pattern analysis easier.

[0066] Next, as shown in FIGS. 13 and 14, the normalization unit 15 normalizes the cumulative value and the period so that the minimum value is 0 and the maximum value is 1, respectively. FIG. 13 shows data in which the cumulative value and the period of the data in FIG. 12 have been normalized. For the period, the normalization unit 15 assigns, for example, 2020-1989=31 years to a value between 0 and 1. For the cumulative value, the normalization unit 15 assigns, for each keyword, a value between 0 and 1 to the difference between the maximum and minimum cumulative values ​​of the keyword.

[0067] FIG. 14 is a graph showing the transition of the frequency of appearance of a certain keyword, with the normalized period on the x-axis and the normalized cumulative value on the y-axis. In FIG. 14, data points 271 are cumulative values ​​for each publication year, and the approximate curve of each data point 271 is shown as graph 272. Hereinafter, the line graph connecting each data point 271 will be referred to as a "word appearance frequency graph." Note that the word appearance frequency graph may also be an approximate curve of the data points 271. The reason for normalizing the cumulative value is to make it easier to compare keywords even if the number of times they appear is different. By normalizing the cumulative value and period, the area formed by the keyword appearance frequency graph and the x-axis can also be normalized, and this area can be used to analyze patterns in changes in appearance frequency.

[0068] <Area calculation> Next, we will explain the area formed by the word frequency graph and the x-axis as one method for quantitatively treating the graph shape of the word frequency graph.

[0069] Figure 15 shows word frequency graphs for two different keywords: Figure 15(a) is a word frequency graph for the keyword (compound) "international scientific conference camstech," and Figure 15(b) is a word frequency graph for the keyword (compound) "peek / carbon composite."

[0070] The graph shape of the word frequency graph in Figure 15(a) is an example of a shape in which the frequency of keyword appearance has increased rapidly in recent years (at the end of a certain period). Figure 15(b) is an example of a graph shape in which the frequency of keyword appearance increased early (at the beginning of a certain period) but has hardly appeared in recent years (at the end of a certain period). The end period can be any period, such as the current time when the analysis is being performed or a year specified by the user. Comparing the two word frequency graphs, it can be seen that the areas formed by the word frequency graphs and the x-axis are significantly different. Therefore, by calculating this area, the area calculation unit 17 can extract keywords that show signs of rapid growth in recent years.

[0071] The method for calculating the area will now be explained. The area calculation unit 17 can simply perform what is called integration on the word frequency graph. Here, a method for calculating the area using trapezoidal approximation will be explained. The area between any two data points using trapezoidal approximation is calculated using formula (1). S is the area of ​​the trapezoid, y is the cumulative value, and x is the period.

[0072]

number

[0073]

number

[0074]

number

[0075]

number

[0076] Figure 16 shows several variations of word frequency graphs with an area A1 of 0.5. The word frequency graphs 221 to 225 in Figures 16(a) to 16(e) all have different shapes, but the area A1 in all of them is 0.5. As such, when the area A1 is close to 0.5, it is not possible to distinguish differences in the transition of the frequency of appearance from the area A1 alone.

[0077] Therefore, in this embodiment, the graph creation unit 16 creates a graph in which the square of the word appearance frequency graph (called a square function) is the value on the y-axis. Figures 16(f) to (j) show square functions 226 to 230 for the word appearance frequency graphs 221 to 225 of Figures 16(a) to (e). The area calculation unit 17 calculates the area formed by the square functions 226 to 230 and the x-axis (this is also called the 2nd moment). Hereinafter, the area formed by the square function and the x-axis will be referred to as area A2.

[0078] The original word frequency graphs 221 to 225 in Figures 16(f) to 16(j) all have an area A1 of 0.5, but the area A2 of the square function is different for each graph. That is, there is a relationship: area A2 (0.475) in Figure 16(f) > area A2 (0.428) in Figure 16(g) > area A2 (0.383) in Figure 16(h) > area A2 (0.333) in Figure 16(i) > area A2 (0.273) in Figure 16(j).

[0079] Therefore, by using the area A2 of the square function, it is possible to determine the transition in the frequency of appearance of the keyword even for the word frequency graphs 221 to 225 whose area is close to 0.5.

[0080] FIG. 17 is a diagram illustrating pattern analysis of the shape of a word frequency graph using area A1 and area A2. FIG. 17(a) shows area A2 of a word frequency graph where area A1 is 0.5 using data points 231 to 235. FIGS. 17(a) to 17(e) were used as word frequency graphs where area A1 is 0.5. Data point 231 is area A2 of word frequency graph 221. Data point 232 is area A2 of word frequency graph 222. Data point 233 is area A2 of word frequency graph 223. Data point 234 is area A2 of word frequency graph 224. Data point 235 is area A2 of word frequency graph 225. For a given keyword, it is possible to analyze which pattern the shape of the word frequency graph is closest to by determining which data point (231 to 235) of the word frequency graph the data point for area A2 is closest to.

[0081] Figure 17(b) shows the range 240 that area A2 can take for all areas A1. As with the case where area A1 is 0.5, this range 240 can be determined by dividing the range of area A1 (0 to 1) into several parts and preparing several word frequency graphs with each having an area A1. For example, by preparing several variations of word frequency graphs with area A1 of 0.1 (variations such as those shown in Figures 16(a) to (e) with area A1 of 0.1) and calculating the area A2 of each word frequency graph, the upper and lower ranges of area A2 when area A1 = 0.1 can be determined. By performing the same process for areas A1 of 0.2 to 1.0, the range 240 can be determined.

[0082] Since the word frequency graph used as a variation in calculating area A2 is known, it is possible to analyze which pattern the shape of the word frequency graph of an arbitrary keyword is closest to by determining where in range 240 the data point of area A2 of the arbitrary keyword is located. By arranging the data points in range 240, scatter plot creation unit 18 can associate the shape of the word frequency graph with an existing pattern. Note that the data point at the bottom left of range 240 corresponds to the graph shape in Figure 15(a), and the data point at the top right of range 240 corresponds to the graph shape in Figure 15(b).

[0083] A supplementary note about graph 250 in Figure 17 is provided. Figure 18 shows multiple word frequency graphs in which the area A1 gradually changes. The word frequency graph in Figure 18 has a standard graph shape showing each area A1. Graph 250 in Figures 17(a) and (b) is a scatter plot of the area A1 and area A2 calculated from each word frequency graph in Figure 18.

[0084] <<iMap>> Although pattern analysis is possible by creating a scatter plot of area A1 and area A2 as shown in Figure 17, more detailed pattern analysis becomes possible by converting area A2 on the vertical axis. First, we will explain how to convert area A2.

[0085] The area calculation unit 17 converts the area A2 into the area B2 using equation (5).

[0086]

number

[0087] Figure 19 shows some data points from a word frequency graph. The data points 211 correspond to a word frequency graph 221 . The data points 216 correspond to a word frequency graph 222 . The data points 218 correspond to a word frequency graph 224 . The data points 219 correspond to a word frequency graph 253 . The data points 213 correspond to a word frequency graph 225 . The data points 212 correspond to a word frequency graph 251 . The data points 217 correspond to the word frequency graph 252 . The data points 220 correspond to a word frequency graph 254 . The data points 214 correspond to a word frequency graph 255 .

[0088] Therefore, the position of the data points (coordinates within iMap 210) allows the transition in keyword frequency to be fitted to a pattern. By arranging the data points on iMap 210, the scatter plot creation unit 18 can associate the shape of the word frequency graph with an existing pattern. For example, the following pattern analysis is possible depending on the position of the data points. The keyword at data point 212 on the left side indicates a keyword that has rapidly increased in recent years (the end of a certain period). The keyword at data point 214 on the right side indicates a keyword that increased in the past (the beginning of a certain period) but was hardly popular at all. The keyword at data point 211 on the top side indicates a keyword that saw a rapid increase in appearance in the middle (near the middle of a certain period) but is now out of fashion. The keyword at data point 213 on the bottom side indicates a keyword that has been around for a long time (the beginning of a certain period), fell out of fashion, and is now starting to attract attention again in recent years (the end of a certain period).

[0089] In addition, since the correspondence between the position of a data point and the shape of the word frequency graph is known, it is possible to analyze which pattern the shape of the word frequency graph of an arbitrary keyword is closest to by determining where the data point of the arbitrary keyword is located on iMap 210. It is also possible to grasp the general trends in the frequency of appearance of unfamiliar keywords (technical themes).

[0090] Figure 20 shows the corresponding points of area A1 and area B2 calculated for keywords with a certain number of occurrences or more on iMap 210. Figure 20 shows the analysis results of keywords extracted from multiple papers in the technical field of CFRP (Carbon Fiber Reinforced Plastics). Keywords with low occurrences are hidden. Figure 20(a) shows iMap 210, with each data point representing one keyword. Users can see the trends in the frequency of occurrence of each keyword. Figure 20(b) shows the number of occurrences of each keyword (the final cumulative value of the number of occurrences) as a bar graph on iMap 210. This allows users to see which pattern the frequency of occurrence of keywords with high and low occurrences closely resembles.

[0091] Furthermore, when the user clicks on any data point with the mouse cursor or the like, the keyword 261, the number of occurrences 262, and the year of first appearance 263 corresponding to the data point are displayed. The user can check the keyword represented by the data point, the number of occurrences, and the year of first appearance.

[0092] <<Another example of a conversion formula for creating an iMap>> The transformation formula for creating an eye map is not limited to formula (5). Various transformation formulas can be considered for creating eye maps for slightly different eye shapes.

[0093] Equation (6) is an equation for calculating an eye map of a different shape using area A1 and area A2. Area calculation unit 17 converts area A1 into area C1 using equation (6), and converts area A1 and area A2 into area C2.

[0094]

number

[0095] <iMaps created based on different types of data sources> So far, we have explained how to arrange the areas A1 and B2 of keywords extracted from the same type of data source (e.g., academic papers) on an iMap. Such iMaps can also be created for the same keywords extracted from different types of data sources (e.g., academic papers and patent documents). Even if iMaps are created for the same keyword, the iMap coordinates on the iMap may differ depending on the data source.

[0096] Figure 21 is a diagram illustrating the correspondence between iMap coordinates created for keywords derived from academic papers and iMap coordinates created for keywords derived from patent documents. Figure 21 shows iMap coordinates 331 and 332 calculated based on the frequency of appearance of keywords that appear in both academic papers and patent documents. Figure 21(a) shows iMap coordinates 331 of keywords derived from academic papers, and Figure 21(b) shows iMap coordinates 332 of keywords derived from patent documents.

[0097] In this embodiment, a line connecting iMap coordinates 331 and 332 of the same keyword is called a "link." If iMap coordinates 331 and 332 of the same keyword extracted from different data sources are different, the difference between iMap coordinates 331 and 332 is reflected in the length of link 330. In this embodiment, attention is focused on the length of this link. In FIG. 21, for the sake of explanation, link 330 is drawn perpendicular to the iMap. However, in FIG. 22 and other figures, the iMaps of FIGS. 21(a) and 21(b) are displayed overlapping one another, and two iMap coordinates (two iMap coordinates of the same keyword extracted from different data sources) are arranged on one iMap. Therefore, link 330 is drawn within a plane within the iMap, and the length of the link is the distance on the iMap plane. Note that the distance of link 330 may be calculated in an iMap created using either equation (5) or equation (6).

[0098] Figure 22 shows an iMap showing the iMap coordinates of keywords that commonly appear in papers and patent documents related to "CFRP (carbon fiber reinforced plastic)." Papers related to "CFRP" are, for example, papers that contain "CREP" in the title, and patent documents related to "CFRP" are, for example, papers that contain "CREP" in the name or abstract of the invention. Keywords may be extracted from the entire paper or patent document (claims or specification).

[0099] In this embodiment, the concept of narrowing down data from a data source, such as "CFRP," is called a "theme." The data acquisition unit 12 acquires, based on the theme, preferably a plurality of papers (an example of first data) and a plurality of patent documents (an example of second data) from each of a data source called papers (an example of a first data source) and a data source called patent documents (an example of a second data source). Note that "CREP" is not a keyword, but the keyword "CREP" may be included in a paper or a patent document. Keywords include nouns extracted by morphological analysis or the like from papers and patent documents on the theme of "CFRP."

[0100] In Figure 22, the iMap coordinates of keywords derived from academic papers (an example of a first data point) and the iMap coordinates of keywords derived from patent documents (an example of a second data point) are shown in different shapes. While Figure 22 is difficult to understand due to the drawing format, it can be seen that there are keywords with the same iMap coordinates even though the data sources are different, and many keywords with similar iMap coordinates. On the other hand, it can be seen that there are keywords with significantly different iMap coordinates even though they are the same keyword, due to the different data sources. In Figure 22, the iMap coordinates of keywords derived from academic papers and the iMap coordinates of the same keywords derived from patent documents are connected by link 330.

[0101] The relationship between iMap coordinates and the shape of the word frequency graph will be explained with reference to Figure 6 above. The iMap in Figure 6 is the same as Figure 22, but it shows three word frequency graphs 305, 306, and 307. Word frequency graphs 305 and 306 are created from keywords derived from academic papers, and word frequency graph 307 is created from keywords derived from patent documents.

[0102] As explained in Figure 19, word frequency graph 307, which shows a sudden increase towards the end of a certain period, is placed in area 311 on the left side of the iMap. Word frequency graph 305, which shows a saturation in frequency from the early stage of a certain period, is placed in area 313 on the right side of the iMap. Word frequency graph 306, which shows a sudden increase in frequency near the middle of a certain period and then saturates, is placed in area 312 on the top side of the iMap.

[0103] Next, we will explain this with reference to Figure 7. Figure 7 is a diagram explaining the link between iMap coordinates of the same keyword extracted from different data sources. The following method can be considered for efficiently identifying keywords related to technology that has reached a plateau in technological development and is about to be commercialized. That is, the information processing system 10 identifies keywords for which the word frequency graph of paper data, which indicates the progress of technological development, indicates "very little increase in recent years = saturated type," and the word frequency graph of patent documents, which indicates the degree of transition to practical application, indicates "rapid increase in recent years = rapid increase type."

[0104] That is, the desired keyword is one whose frequency of appearance increases sharply toward the end of a certain period in patent documents and whose frequency of appearance saturates toward the beginning or middle of a certain period in academic papers. The link of such a keyword forms a straight line connecting the leftmost region 311 with the topmost region 312 or the rightmost region 313, so the desired keyword is considered to have long links 321, 322 on iMap (greater than a threshold). Therefore, in this embodiment, by detecting keywords with long links 321, 322, it is possible to identify keywords whose technological development has reached a plateau and are about to be commercialized.

[0105] <Effects of determining the first year of a period on coordinates> Next, with reference to FIGS. 23 to 28, we will explain how the method for determining the first year of a period when creating an appearance frequency graph affects iMap coordinates. FIG. 23 shows three iMaps with different first years for the period when creating a word appearance frequency graph. In FIG. 23(a), the first year is 1969, in FIG. 23(b), the first year is 2000, and in FIG. 23(c), the first year is the year of first appearance in a paper (hereinafter referred to as the first year of paper). Each figure shows iMap coordinates of keywords extracted from papers and patent documents. 1969, 2000, and the first year of paper are examples of multiple first years. In this embodiment, the "year of first appearance" is used as the start of the period to explain using years as units. However, the start of the period may also be the "month of first appearance" using months as units, the "day of first appearance" using days as units, the "hour of first appearance" using hours as units, the "minute of first appearance" using minutes as units, or the "second of first appearance" using seconds as units.

[0106] Note that 1969 is the year that papers are expected to contain the same keywords as those related to the technology that will be commercialized in the future, and is not limited to 1969. 2000 is the year midway between 1969 and the first year of the paper. This is also not limited to 2000. The first year of the paper is the year the keyword was first used in the paper. Therefore, the first year of the paper may differ depending on the keyword. Also, when calculating the iMap coordinates for keywords derived from patent documents, the year of first appearance should be set to the first year of the paper.

[0107] Even for the same keyword, the shape of the word frequency graph (values ​​on the horizontal and vertical axes) changes depending on how the first year is determined. Therefore, even for the same keyword extracted from both academic papers and patent documents, the iMap coordinates and link lengths change depending on how the first year is determined. This is evident from Figures 23(a)-(c), where the distribution of iMap coordinates differs depending on how the first year is determined (1969, 2000, first year of the paper), and the link lengths also differ significantly depending on how the first year is determined.

[0108] Figure 24 shows a list of keywords with iMap coordinates shown in Figure 23. The keywords in Figure 24 are extracted from keywords (extracted from titles, etc.) common to papers and patent documents related to "CFRP" and have long links. Link lengths were calculated for each of the three iMaps with different initial years. Note that because there were a large number of original keywords extracted from papers and patent documents, the top 20 keywords by link length were extracted for each of the three initial years. Some keywords overlap, while others do not, so more than 20 keywords are extracted in Figure 24. Note that Figure 24 is not sorted in order of link length.

[0109] First, the items in FIG. 24 will be explained. The "Evaluation" item is a human evaluation of the state of the technology related to the keyword. There are three evaluations: "Practical application," "Signs," and "Research stage." In this embodiment, the goal is to extract keywords evaluated as "Signs." The Key Word item is a keyword common to papers and patent documents. The "Total number of papers" item is the total number of papers related to the technology known as CFRP. The "first year of publication" field indicates the year in which the keyword first appeared in the paper. The "Total number of patents" item is the total number of patent documents related to CFRP technology. The "first year of patent" field indicates the year in which the keyword first appeared in the patent document. The "Difference in Year of First Publication" item is the difference between the first year of publication in a paper and the first year of publication in a patent. In this embodiment, keywords where the first year of publication in a paper is less than the first year of publication in a patent are screened. This is because this embodiment assumes that keywords related to early-stage technologies are used in papers, and keywords related to technologies before commercialization are used in patent documents. Also, since it is not possible to determine from iMap coordinates that the first year of publication is greater than or equal to the first year of publication in a patent, this is done to prevent unintended keywords from being extracted based on the length of the link. The "Link Length" item is the length (distance) of the link between iMap coordinates when the first years of the period are 1969, 2000, and the first year of the paper. The link length is the distance within iMap (for example, the longest value on the horizontal axis is 1, and the longest value on the vertical axis is about 0.3).

[0110] In Figure 24, keywords marked with a double circle "◎" are keywords related to technologies that show signs of practical application (including technologies that have been partially put into practical application but are still considered promising despite issues). Keywords marked with a "△" are keywords related to technologies that have already been put into practical application. Keywords marked with a "○" are keywords related to technologies that are still in the research stage. As shown in Figure 24, by extracting keywords based on the length of the link, it is clear that although many keywords have already been put into practical application, keywords that show signs of practical application have also been picked up. Figure 24 shows that eight keywords that show signs of practical application were found.

[0111] The screen generator 20 generates a screen for displaying the list of Fig. 24 and transmits the screen information to the terminal device 30. The screen generator 20 provides the web application to the terminal device 30, so that the terminal device 30 can display the list of Fig. 24 on the screen using a web browser.

[0112] <<Examples of iMap coordinates with different first years of the period>> With reference to FIGS. 25 to 28, the iMap coordinates and link lengths for different initial years will be described for several keywords that show signs of practical use.

[0113] Figure 25 shows an iMap and word frequency graphs for the keyword "high electromagnetic," which is thought to be close to practical application, with the first years being 1969, 2000, and the first year of the paper. Figure 25(a) shows iMap and word frequency graphs 341 and 344 for "high electromagnetic," with 1969 as the first year of the period. Figure 25(b) shows iMap and word frequency graphs 342 and 345 for "high electromagnetic," with 2000 as the first year of the period. Figure 25(c) shows iMap and word frequency graphs 343 and 346 for "high electromagnetic," with the first year of the paper as the first year of the period. Furthermore, word frequency graphs 341-343 on the right side of Figure 25 were created using keywords derived from papers, while word frequency graphs 344-346 were created using keywords derived from patent documents.

[0114] 25(a) to 25(c), there is a link 347 connecting the leftmost region 311 to the top edge, a link 348 connecting the leftmost region 311 to the rightmost region 313, and a link 349 connecting the leftmost region 311 to the rightmost region 313. The start point of each link (the point on the left) is the iMap coordinate of the keyword derived from the patent document, and the end point is the iMap coordinate of the keyword derived from the paper.

[0115] Furthermore, word appearance frequency graphs 344 to 346 of keywords derived from patent documents show a sharp increase in appearance frequency towards the end of a certain period. Word appearance frequency graph 341 of keywords derived from academic papers shows saturation in appearance frequency towards the end of the period. Word appearance frequency graphs 342 and 343 of keywords derived from academic papers show saturation from an early stage of the period. Therefore, these are reflected in the lengths of links 347 to 349, and it can be confirmed from the lengths of links 347 to 349 that keywords showing signs of approaching practical application have been extracted.

[0116] Furthermore, as can be seen by comparing word frequency graphs 341-343 on the right, the area formed by word frequency graphs 341-343 and the x-axis increases as the first year of the period goes from 1969 to 2000 to the first year of the paper. This naturally occurs depending on how the period is chosen. Because the area increases in the order 1969 → 2000 → first year of the paper, the endpoints of links 347-349 in iMap tend to move to the right, resulting in longer links, such as link 347<348<349. Therefore, in order to extract keywords for technologies close to practical application based on link length, it may be effective to use the first year of the paper as the first year of the period.

[0117] Figure 26 shows an iMap and word frequency graphs for the keyword "structural form," which is thought to be close to practical application, with the first years being 1969, 2000, and the first year of the paper. Figure 26(a) shows iMap and word frequency graphs 351 and 354 for "structural form," with 1969 as the first year of the period. Figure 26(b) shows iMap and word frequency graphs 352 and 355 for "structural form," with 2000 as the first year of the period. Figure 26(c) shows iMap and word frequency graphs 353 and 356 for "structural form," with the first year of the paper as the first year of the period. Furthermore, word frequency graphs 351 to 353 on the right side of Figure 26 were created from keywords derived from papers, and word frequency graphs 354 to 356 were created from keywords derived from patent documents.

[0118] 26(a) to 26(c), there is a link 357 connecting the leftmost region 311 to the region roughly in the center, a link 358 connecting the leftmost region 311 to the rightmost region 313, and a link 359 connecting the leftmost region 311 to the rightmost region 313. The starting point of each link (the point on the left) is the iMap coordinate of the keyword derived from the patent document, and the end point is the iMap coordinate of the keyword derived from the paper.

[0119] Furthermore, word appearance frequency graphs 354 to 356 of keywords derived from patent documents show a sharp increase in appearance frequency towards the end of a certain period. Word appearance frequency graph 351 of keywords derived from academic papers saturates towards the end of the period. Word appearance frequency graphs 352 and 353 of keywords derived from academic papers saturate from the middle to the end of the period. Therefore, these are reflected in the lengths of links 357 to 359, and it can be confirmed from the lengths of links 357 to 359 that keywords that show signs of approaching practical application have been extracted.

[0120] Figure 27 shows an iMap and word frequency graphs for the keyword "high speed train," which is thought to be close to practical application, with the first years being 1969, 2000, and the first year of the paper. Figure 27(a) shows iMap and word frequency graphs 361 and 364 for "high speed train," with 1969 as the first year of the period. Figure 27(b) shows iMap and word frequency graphs 362 and 365 for "high speed train," with 2000 as the first year of the period. Figure 27(c) shows iMap and word frequency graphs 363 and 366 for "high speed train," with the first year of the paper as the first year of the period. Furthermore, word frequency graphs 361 to 363 on the right side of Figure 27 are created from keywords derived from papers, and word frequency graphs 364 to 366 are created from keywords derived from patent documents.

[0121] 27(a) to 27(c), there is a link 367 connecting the leftmost region 311 and the topmost region 312, a link 368 connecting the leftmost region 311 and the rightmost region 313, and a link connecting the leftmost region 311 and the rightmost region 313. The start point of each link (the point on the left) is the iMap coordinate of the keyword derived from the patent document, and the end point is the iMap coordinate of the keyword derived from the paper.

[0122] Furthermore, word appearance frequency graphs 364 to 366 of keywords derived from patent documents show a sharp increase in appearance frequency towards the end of a certain period. Word appearance frequency graph 361 of keywords derived from academic papers saturates near the center. Word appearance frequency graphs 362 and 363 of keywords derived from academic papers saturate towards the beginning of the period. Therefore, these are reflected in the lengths of links 367 to 369, and it can be confirmed from the lengths of links 367 to 369 that keywords that show signs of approaching practical application have been extracted.

[0123] Figure 28 shows an iMap and word frequency graphs for the keyword "c / sic composite material," which is considered close to practical application, with the first years being 1969, 2000, and the first year of the paper. Figure 28(a) shows iMap and word frequency graphs 371 and 374 for "c / sic composite material," with 1969 as the first year of the period. Figure 28(b) shows iMap and word frequency graphs 372 and 375 for "c / sic composite material," with 2000 as the first year of the period. Figure 28(c) shows iMap and word frequency graphs 373 and 376 for "c / sic composite material," with the first year of the paper as the first year of the period. Furthermore, word frequency graphs 371-373 on the right side of Figure 28 are created from keywords derived from papers, and word frequency graphs 374-376 are created from keywords derived from patent documents.

[0124] 28(a) to 28(c), there is a link 377 connecting the leftmost region 311 and the topmost region 312, a link 378 connecting the leftmost region 311 and the rightmost region 313, and a link 379 connecting the leftmost region 311 and the rightmost region 313. The start point of each link (the point on the left) is the iMap coordinate of the keyword derived from the patent document, and the end point is the iMap coordinate of the keyword derived from the paper.

[0125] Furthermore, word appearance frequency graphs 374 to 376 of keywords derived from patent documents show a sharp increase in appearance frequency towards the end of a certain period. Word appearance frequency graph 371 of keywords derived from academic papers saturates near the middle of the period. Word appearance frequency graphs 372 and 373 of keywords derived from academic papers saturate from the beginning of the period. Therefore, these are reflected in the lengths of links 377 to 379, and it can be confirmed from the lengths of links 377 to 379 that keywords that show signs of approaching practical application have been extracted.

[0126] <Process or Action> FIG. 29 is a sequence diagram illustrating the process in which the data processing system 100 extracts an iMap and keywords based on the frequency of appearance of the keywords.

[0127] S1: The user connects the terminal device 30 to the information processing system 10 and causes the terminal device 30 to execute a web application. The user specifies a theme and two data sources (papers and patent documents) and instructs the web application executed by the terminal device 30 to analyze the transition in the frequency of appearance of common keywords contained in these data sources for the theme.

[0128] S2: The operation reception unit 33 of the terminal device 30 receives the instruction, and the communication unit 31 transmits an analysis request to the information processing system 10. The communication unit 31 may transmit the data itself.

[0129] S3: The communication unit 11 of the information processing system 10 receives the analysis request, and the data acquisition unit 12 acquires papers and patent documents related to the specified theme. The data acquisition unit 12 may receive the data itself from the terminal device 30, or may acquire it from a network. The data acquisition unit 12 acquires multiple papers and patent documents narrowed down by theme.

[0130] S4: Next, the data processing unit 13 performs morphological analysis on the papers and patent documents to obtain keywords. The data processing unit 13 extracts keywords common to the papers and patent documents. The occurrence frequency calculation unit 14 calculates the occurrence frequency for each common keyword per unit period. The occurrence frequency is calculated for both the papers and patent documents. Depending on the data format, morphological analysis may not be necessary.

[0131] S5: Next, the normalization unit 15 normalizes the period and the cumulative value of the appearance frequency to 0 to 1. The normalization is performed for each of the papers and patent documents.

[0132] S6: Next, the graph creation unit 16 creates a graph of the cumulative value of the frequency of appearance over a period (creates a word frequency graph). The graphing is performed for both the papers and the patent documents.

[0133] S7: Similarly, the graph creation unit 16 creates a square function of the word appearance frequency graph. The creation of the square function is performed for each of the papers and the patent documents.

[0134] S8: The area calculation unit 17 calculates the area A1 formed between the word frequency graph and the x-axis, and the area A2 formed between the square function and the x-axis. The calculation of the areas A1 and A2 is performed for each of the papers and patent documents.

[0135] S9: The area calculation unit 17 converts the area A2 into the area B2. As explained in Fig. 13, if the pattern analysis is performed using the area A2 as is, the processing of step S9 is not necessary. The conversion to the area B2 is performed for both the paper and the patent document.

[0136] S10: The scatter diagram creation unit 18 arranges data points of area B2 corresponding to area A1 for each of the paper and the patent document on the scatter diagram (creates an iMap 210). When analyzing using area A2 as is, the scatter diagram creation unit 18 arranges data points of area A2 corresponding to area A1 on the scatter diagram.

[0137] S11: The distance calculation unit calculates the distance of the link connecting the iMap coordinates of the keyword derived from the paper and the iMap coordinates of the keyword derived from the patent document for each keyword.

[0138] S12: The screen generator 20 extracts some keywords that are ranked high in length.

[0139] S13: The screen generation unit 20 of the information processing system 10 creates a screen for displaying the iMap 210 and a screen for displaying a list of keywords, and the communication unit 11 transmits this screen information to the terminal device 30.

[0140] S14: The communication unit 31 of the terminal device 30 receives the screen information, and the display control unit 32 displays on the display 506 a screen including the iMap 210 and a screen displaying a list of keywords.

[0141] <About the data source> In this embodiment, a case has been described in which papers and patent documents are mainly used as data sources, but the data sources are not limited to these.

[0142] Figure 30 is a diagram explaining an example of a data source. For example, data source 1 for technologies in their infancy (early stages of technological development) includes not only academic papers but also project summaries funded by Kakenhi. Kakenhi is a competitive research grant that aims to significantly advance all academic research across all fields, from the humanities and social sciences to the natural sciences. Kakenhi-funded projects often relate to technologies in their infancy (early stages of technological development).

[0143] Furthermore, data sources 2 relating to technologies before the transition to commercialization, beyond the disillusionment phase, include patent documents, news information, and business summaries of venture or startup companies. News information includes industrial and economic newspapers published for each industry sector (including their websites), new product information, etc., and it is likely that these will contain keywords related to the technology before the transition to commercialization. Furthermore, if it is assumed that a venture or startup company will start a business based on new technology, it is likely that the business summary will contain keywords related to the technology before the transition to commercialization.

[0144] Data sources 1 and 2 can be combined in any way. Furthermore, multiple data sources may be used on the Data Source 1 side, and multiple data sources may be used on the Data Source 2 side.

[0145] <Major Effects> The information processing system of this embodiment can analyze common keywords contained in multiple data sources. As an example of this analysis, it can use the increase in the length of the link connecting two iMap coordinates to identify keywords that indicate signs of practical application.

[0146] <Other application examples> The best mode for carrying out the present invention has been described above using examples, but the present invention is not limited to these examples in any way, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention.

[0147] For example, in the present embodiment, the transition of the frequency of appearance of keywords included in the data is analyzed for patterns, but the transition of the frequency of appearance of a keyword specified by the user may be analyzed for patterns. In this case, the information processing system 10 only needs to analyze the pattern of one keyword.

[0148] Alternatively, the information processing system 10 may extract common keywords from three or more data sources and place three iMap coordinates on the iMap. The information processing system 10 can analyze the relationship between the three iMap coordinates based on the link distance.

[0149] Additionally, keywords common to multiple data sources may be common due to transformations, such as transforming AI into artificial intelligence. For example, keywords with the same meaning may be used in multiple data sources as different terms.

[0150] The information processing system 10 may also detect keywords with short link distances. For example, if there is little difference between the iMap coordinates of a keyword derived from a research paper and the iMap coordinates of a keyword derived from a patent document, it indicates that early research and technological development were conducted in parallel. The locations of the two iMap coordinates from different data sources may vary depending on the pattern of word frequency graphs being searched for.

[0151] Furthermore, the keywords do not have to be converted into text. For example, the information processing system 10 may perform speech recognition on audio data recorded at various conferences, and perform pattern analysis of the transition in frequency of appearance of the included keywords. Users can analyze what topics were discussed and how their patterns changed in various conferences held. The information processing system 10 may also perform speech recognition on audio data of conversations made to a call center, and perform pattern analysis of the transition in frequency of appearance of the included keywords. Users can analyze which keywords are frequently inquired about, and improve their systems and services.

[0152] Furthermore, in this embodiment, the cumulative value of the frequency of appearance is calculated, so the word frequency graph does not decrease over the period, but the word frequency graph may be rotated 180 degrees to perform pattern analysis.

[0153] Furthermore, in this embodiment, the graph creation unit 16 creates a word frequency graph, etc., but the graphs created in this embodiment do not need to be visualized. That is, the graphs shown in this embodiment are for explanation purposes, and as long as the areas A1, A2, and B2 and the link distances can be calculated, they do not need to be displayed. However, by displaying the word frequency graph, etc. together with iMap 210, the user can visually confirm the shape of the word frequency graph.

[0154] In addition, in this embodiment, as an example of the analysis target, the pattern analysis was performed on the transition of the frequency of appearance of keywords included in the titles of papers in the technical field, but the analysis target is not limited to this. For example, the papers do not have to be in the technical field, and may be in the form of papers on medicine, pharmacology, philosophy, art, literature, language, history, geography, anthropology, law, politics, economics, society, education, psychology, mathematics, physics, astronomy, chemistry, energy, biochemistry, agricultural chemistry, civil engineering, sports, or the like.

[0155] The data may also be image data. In this case, the recognition device converts the subject in the image data into keywords. For example, the information processing system 10 extracts keywords by performing optical character recognition processing on the image data.

[0156] Furthermore, the configuration examples in Fig. 10 and the like are divided according to main functions to facilitate understanding of the processing by the terminal device 30 and the information processing system 10. The present invention is not limited by the manner in which the processing units are divided or the names of the processing units. The processing by the terminal device 30 and the information processing system 10 can also be divided into more processing units depending on the processing content. Furthermore, the processing units can also be divided so that one processing unit includes more processes.

[0157] Additionally, the devices described in the examples are merely illustrative of one of several computing environments for implementing the embodiments disclosed herein. In one embodiment, information processing system 10 includes multiple computing devices, such as a server cluster, configured to communicate with each other via any type of communication link, including a network, shared memory, etc., and to perform the processes disclosed herein.

[0158] Furthermore, the information processing system 10 can be configured to share the processing steps disclosed in this embodiment, such as those shown in FIG. 17, in various combinations. For example, a process executed by a specific unit can be executed by multiple information processing devices included in the information processing system 10. Furthermore, the information processing system 10 may be integrated into a single server device, or may be divided into multiple devices.

[0159] Each function of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), and conventional circuit modules designed to perform each of the above-described functions.

[0160] <Aspect> [Aspect 1] Information processing systems, a data acquisition unit that acquires first data and second data, each associated with one or more of year, month, day, hour, minute, and second, from a first data source and a second data source, respectively; a data processing unit that extracts one or more common analysis targets from the first data and the second data acquired by the data acquisition unit; a graph creation unit that creates a first graph and a second graph that are cumulative values ​​over a period of the values ​​related to the analysis target of the first data and the values ​​related to the analysis target of the second data; an area calculation unit that calculates a third area using a first area formed by the graph created by the graph creation unit and an x-axis, and a second area formed by a square function of the graph and the x-axis; a scatter diagram creation unit that arranges, in a scatter diagram with the first area as the x-axis and the third area as the y-axis, first data points that associate the first area and the third area calculated based on the first graph, and second data points that associate the first area and the third area calculated based on the second graph; A program to function as a [Aspect 2] The information processing system, The program according to aspect 1, which causes the program to function as a distance calculation unit that calculates the distance between the first data point and the second data point in the scatter plot. [Aspect 3] The information processing system, The program according to aspect 2, wherein the program functions as a screen generation unit that generates a screen relating to a list of the analysis targets extracted in descending order of the distance calculated by the distance calculation unit. [Aspect 4] the graph creation unit creates the first graph and the second graph using the first appearance year, first appearance month, first appearance day, first appearance hour, first appearance minute, or first appearance second of the analysis target in the first data source as the start of the period; The program according to any one of aspects 1 to 3, wherein the scatter plot creation unit arranges the first data points calculated based on the first graph and the second data points calculated based on the second graph on the scatter plot. [Aspect 5] The graph creation unit The start of the period is changed, and the first graph and the second graph are created for each start of the period; The program according to any one of aspects 1 to 3, wherein the scatter plot creation unit arranges the first data point calculated based on the first graph and the second data point calculated based on the second graph on the scatter plots having different start dates for the periods. [Aspect 6] The program according to aspect 5, wherein the different start dates of the period include the first appearance year, first appearance month, first appearance day, first appearance hour, first appearance minute, or first appearance second of the analysis target in the first data source. [Aspect 7] the area calculation unit calculates a second area A2 formed by the square function of the graph with respect to the x-axis, and converts the second area A2 into a third area B2 using Equation (7) using the first area A1 and the second area A2 formed by the graph with respect to the x-axis; [Number 7] The program according to aspect 2, wherein the scatter plot creation unit arranges the first data points and the second data points on the scatter plot with the first area A1 as the x-axis and the third area B2 as the y-axis. [Aspect 8] the area calculation unit calculates the area converted from the first area A1 by equation (8) as the first area C1; A second area A2 formed by the square function of the graph with the x-axis is calculated, and the area calculated by Equation (8) using the first area A1 and the second area A2 formed by the graph with the x-axis is defined as a third area C2; [Number 8] The program according to aspect 2, wherein the scatter plot creation unit arranges the first data points and the second data points on the scatter plot with the first area C1 as the x-axis and the third area C2 as the y-axis. [Aspect 9] The program according to aspect 7, wherein the scatter plot creation unit determines that the shape of the second graph is a pattern in which the value related to the analysis target increases sharply at the end of the period when a data point is located in the leftmost region in a scatter plot in which the first area A1 is the x-axis and the third area B2 is the y-axis. [Aspect 10] The program according to aspect 9, wherein the scatter plot creation unit determines that when a data point is located in the upper end region in a scatter plot with the first area A1 as the x-axis and the third area B2 as the y-axis, the shape of the first graph is a pattern in which values ​​related to the analysis target increase sharply near the middle of the period but do not increase at the end of the period. [Aspect 11] The program according to aspect 10, wherein the scatter plot creation unit determines that, when a data point is located in the rightmost region in a scatter plot in which the first area A1 is the x-axis and the third area B2 is the y-axis, the shape of the first graph is a pattern in which values ​​related to the analysis target increase sharply at the beginning of the period and do not appear thereafter. [Aspect 12] the first data source is a data source containing data on a technology that is in its infancy or early stages of development; 12. The program according to claim 11, wherein the second data source is a data source including data on a technology that has passed the disillusionment phase and has not yet been commercialized. [Aspect 13] the distance calculation unit is configured to calculate the second data point in the leftmost region of the scatter plot; 13. The program according to any one of aspects 7 to 12, wherein the distance to the first data point located in the upper end region or the right end region is calculated. [Aspect 14] 14. The program according to any one of aspects 1 to 13, wherein the first data source is a paper, and the second data source is a patent document. [Aspect 15] 15. The program according to any one of aspects 1 to 14, wherein the analysis target is a keyword included in the first data and the second data, and the value is the number of times the keyword appears. [Explanation of symbols]

[0161] 10 Information Processing Systems 30 Terminal Equipment 100 Data Processing System [Prior art documents]

Charter Documents

[0162] [Patent Document 1] Patent Gazette No. 5614687

Claims

1. Information processing systems, a data acquisition unit that acquires first data and second data, each associated with one or more of year, month, day, hour, minute, and second, from a first data source and a second data source, respectively; a data processing unit that extracts one or more common analysis targets from the first data and the second data acquired by the data acquisition unit; a graph creation unit that creates a first graph and a second graph that are cumulative values ​​over a period of the values ​​related to the analysis target of the first data and the values ​​related to the analysis target of the second data; an area calculation unit that calculates a third area using a first area formed by the graph created by the graph creation unit and an x-axis, and a second area formed by a square function of the graph and the x-axis; a scatter diagram creation unit that arranges, in a scatter diagram with the first area on the x-axis and the third area on the y-axis, first data points that associate the first area and the third area calculated based on the first graph, and second data points that associate the first area and the third area calculated based on the second graph; A program to function as a

2. The information processing system, The program according to claim 1 , wherein the program functions as a distance calculation unit that calculates a distance between the first data point and the second data point in the scatter diagram.

3. The information processing system, 3. The program according to claim 2, wherein the program functions as a screen generation unit that generates a screen relating to a list of the analysis targets extracted in descending order of the distance calculated by the distance calculation unit.

4. the graph creation unit creates the first graph and the second graph using the first appearance year, first appearance month, first appearance day, first appearance hour, first appearance minute, or first appearance second of the analysis target in the first data source as the start of the period; The program according to any one of claims 1 to 3, wherein the scatter plot creation unit arranges the first data points calculated based on the first graph and the second data points calculated based on the second graph on the scatter plot.

5. The graph creation unit creating the first graph and the second graph for each start point of the period; The program according to any one of claims 1 to 3, wherein the scatter plot creation unit arranges the first data point calculated based on the first graph and the second data point calculated based on the second graph on the scatter plots having different start dates for the periods.

6. The program according to claim 5 , wherein the different start dates of the period include the first appearance year, first appearance month, first appearance day, first appearance hour, first appearance minute, or first appearance second of the analysis target in the first data source.

7. the area calculation unit calculates a second area A2 formed by the square function of the graph with respect to the x-axis, and converts the second area A2 into a third area B2 using Equation (7) using the first area A1 and the second area A2 formed by the graph with respect to the x-axis; [Equation 7] The program according to claim 2 , wherein the scatter diagram creation unit arranges the first data points and the second data points in the scatter diagram with the first area A1 on the x-axis and the third area B2 on the y-axis.

8. The area calculation unit converts the first area A1 into the first area C1 using equation (8), A second area A2 formed by the square function of the graph with respect to the x-axis is calculated, and the area calculated by Equation (8) using the first area A1 formed by the graph with respect to the x-axis and the second area A2 is defined as a third area C2; [Equation 8] The program according to claim 2 , wherein the scatter diagram creation unit arranges the first data points and the second data points in the scatter diagram with the first area C1 on the x-axis and the third area C2 on the y-axis.

9. 8. The program according to claim 7, wherein the scatter plot creation unit determines that the shape of the second graph is a pattern in which values ​​related to the analysis target increase sharply at the end of the period when a data point is located in the leftmost region in a scatter plot in which the first area A1 is the x-axis and the third area B2 is the y-axis.

10. 10. The program according to claim 9, wherein the scatter plot creation unit determines that when a data point is located in the upper end region in a scatter plot in which the first area A1 is the x-axis and the third area B2 is the y-axis, the shape of the first graph is a pattern in which values ​​related to the analysis target increase sharply near the middle of the period but do not increase at the end of the period.

11. 11. The program according to claim 10, wherein the scatter plot creation unit determines that, when a data point is located in the rightmost region in a scatter plot in which the first area A1 is the x-axis and the third area B2 is the y-axis, the shape of the first graph is a pattern in which values ​​related to the analysis target increase sharply at the beginning of the period and do not appear thereafter.

12. the first data source is a data source including data on a technology that is in its infancy or early stages of development; 12. The program according to claim 11, wherein the second data source is a data source containing data on a technology that has passed the disillusionment stage and has not yet been commercialized.

13. the distance calculation unit is configured to calculate the second data point in the leftmost region of the scatter plot; The program according to claim 7 , wherein the distance to the first data point located in the upper end region or the right end region is calculated.

14. 2. The program of claim 1, wherein the first data source is a paper and the second data source is a patent document.

15. 2. The program according to claim 1, wherein the analysis target is a keyword contained in the first data and the second data, and the value is the number of times the keyword appears.

16. a data acquisition unit that acquires first data and second data, each associated with one or more of year, month, day, hour, minute, and second, from a first data source and a second data source, respectively; a data processing unit that extracts one or more common analysis targets from the first data and the second data acquired by the data acquisition unit; a graph creation unit that creates a first graph and a second graph that are cumulative values ​​over a period of the values ​​related to the analysis target of the first data and the values ​​related to the analysis target of the second data; an area calculation unit that calculates a third area using a first area formed by the graph created by the graph creation unit and an x-axis, and a second area formed by a square function of the graph and the x-axis; a scatter diagram creation unit that arranges, in a scatter diagram with the first area on the x-axis and the third area on the y-axis, first data points that associate the first area and the third area calculated based on the first graph, and second data points that associate the first area and the third area calculated based on the second graph; An information processing system comprising:

17. A data processing system in which a terminal device and an information processing system communicate with each other via a network, The information processing system includes: a data acquisition unit that acquires first data and second data, each associated with one or more of year, month, day, hour, minute, and second, from a first data source and a second data source, respectively; a data processing unit that extracts one or more common analysis targets from the first data and the second data acquired by the data acquisition unit; a graph creation unit that creates a first graph and a second graph that are cumulative values ​​over a period of the values ​​related to the analysis target of the first data and the values ​​related to the analysis target of the second data; an area calculation unit that calculates a third area using a first area formed by the graph created by the graph creation unit and an x-axis, and a second area formed by a square function of the graph and the x-axis; a scatter diagram creation unit that arranges, in a scatter diagram with the first area on the x-axis and the third area on the y-axis, first data points that associate the first area calculated based on the first graph with the third area, and second data points that associate the first area calculated based on the second graph with the third area, The terminal device displays the scatter diagram based on screen information transmitted from the information processing system.

18. An information processing system is a data processing method, a process in which a data acquisition unit acquires, from a first data source and a second data source, first data and second data, each associated with one or more of year, month, day, hour, minute, and second; A process in which a data processing unit extracts one or more common analysis targets from the first data and the second data acquired by the data acquisition unit; a process in which a graph creation unit creates a first graph and a second graph which are cumulative values ​​over a period of values ​​related to the analysis target of the first data and values ​​related to the analysis target of the second data; an area calculation unit calculating a third area using a first area formed by the graph created by the graph creation unit and the x-axis, and a second area formed by a square function of the graph and the x-axis; a process in which a scatter diagram creation unit arranges, in a scatter diagram with the first area on the x-axis and the third area on the y-axis, first data points that correspond to the first area and the third area calculated based on the first graph, and second data points that correspond to the first area and the third area calculated based on the second graph; Data processing methods.

Citation Information

Patent Citations

  • Pipe joint and pipe joining method

    JP1981014687A