Information processing system, information processing method, and program
The system addresses the challenge of incomplete researcher data by integrating information from various sources, enabling comprehensive data collection and advanced analysis, including network visualization and collaborative research suggestions.
Patent Information
- Application Number
- JP2024141736
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-10-14
AI Technical Summary
Existing information processing systems struggle to comprehensively collect and integrate researcher information from diverse sources, leading to incomplete and fragmented data on researchers' activities and affiliations.
An information processing system that includes a processor to receive researcher and institution names, acquire corresponding research information from open-source data, extract similar information, and integrate it into a researcher information database, using text analysis and vector representation to identify and match researchers across multiple data sources.
Enables comprehensive collection and integration of researcher information, facilitating advanced search and analysis capabilities, including network visualization and collaborative research partner suggestions.
Smart Images

Figure 2025155548000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]
[0002] Patent Document 1 describes an information provision system that efficiently links various scientific and technological information distributed on the Internet and easily obtains connections between diverse information across fields and industries. Patent Document 1 describes the construction of an information provision system that includes multiple dictionary databases 21, 22, and 5 and a CPU 11 that serves as a computing unit that controls these dictionary databases. Each dictionary database has its own concept group A, B, and C, and stores multiple concepts belonging to each concept group, as well as name and spelling variations for each concept and attributes for identifying the concept, associated with each concept. When a search term is input, the CPU 11 uses the dictionary databases 21, 22, and 5 to distinguish between alternative names for the same concept and alternative concepts with the same spelling, and outputs the alternative names and spelling variations for the search term and the attributes for identifying the concept. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-224952 Summary of the Invention [Problem to be solved by the invention]
[0004] Patent Document 1 proposes that by enhancing the search dictionary, when a search term is entered, it is possible to identify alternative names, spelling variations, and concepts related to the search term. However, there is still room for improvement in the technology of information processing systems that comprehensively collect researcher information. [Means for solving the problem]
[0005] According to one aspect of the present invention, there is provided an information processing system. The information processing system includes at least one processor configured to execute a program to perform the following steps: In the receiving step, the name of a first researcher and the name of the institution to which the first researcher belongs are received. The name and the name of the institution are used as first researcher identification information to identify the first researcher in a researcher information database. In the acquiring step, first research information corresponding to the first researcher identification information is acquired from external open-source data related to research information. In the extracting step, second research information that is similar to the first research information and has a name that matches that of the first researcher is extracted from the open-source data based on the text information of the first research information. In the integrating step, second researcher identification information corresponding to the second research information is integrated with the first researcher identification information to generate or update the researcher information database.
[0006] According to one aspect of the present invention, it is possible to provide a more useful information processing system, etc., as a technology for an information processing system that comprehensively collects researcher information. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a configuration diagram illustrating an information processing system 1. FIG. [Figure 2] FIG. 2 is a block diagram showing the hardware configuration of the server 2. [Figure 3] FIG. 2 is a block diagram showing a hardware configuration of an information processing device 3. [Figure 4] FIG. 10 is a diagram showing content types of open source data. [Figure 5] FIG. 2 is a diagram showing an outline of processing executed by the information processing system 1. [Figure 6] 2 is an activity diagram showing an example of the flow of processing executed by the information processing system 1. FIG. [Figure 7]FIG. 2 is a diagram showing an example of open source data used by the information processing system 1. [Figure 8] FIG. 10 is a diagram showing an example of researcher information 8. [Figure 9] FIG. 10 is a diagram showing an example of a search reception screen 9. [Figure 10] FIG. 10 is a diagram showing an example of a search result 10. DETAILED DESCRIPTION OF THE INVENTION
[0008] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the accompanying drawings. Various features shown in the following embodiments can be combined with each other.
[0009] Incidentally, the program for realizing the software appearing in one embodiment may be provided as a non-transitory computer-readable medium, or may be provided so that it can be downloaded from an external server, or may be provided so that the program is started on an external computer and its functions are realized on a client terminal (so-called cloud computing).
[0010] Furthermore, various information processing according to an embodiment may realize input and output corresponding to the input. Here, the form of information referenced in such information processing (hereinafter referred to as reference information) is not limited as long as an output is obtained as a result of the input. The reference information may be, for example, rule-based information such as a database, a lookup table, or a predetermined function (including a decision formula such as a regression formula constructed using a statistical method), a trained model that has previously trained the correlation between input and output, or a large-scale language model that can output a desired result by inputting a prompt.
[0011] In one embodiment, a "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. In one embodiment, various information is handled, and this information is represented, for example, by physical values of signal values representing voltage and current, high and low signal values as a binary bit set consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculations can be performed on a circuit in the broad sense.
[0012] Furthermore, a circuit in the broad sense is a circuit realized by at least an appropriate combination of a circuit, circuitry, processor, memory, etc. The processor may be a general-purpose processor or a dedicated circuit. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.
[0013] 1. Hardware Configuration This section explains the hardware configuration.
[0014] <Information Processing System 1> FIG. 1 is a configuration diagram illustrating an information processing system 1. The information processing system 1 includes a server 2 and an information processing device 3. The server 2 and the information processing device 3 are configured to be able to communicate with each other via a telecommunications line (network). In an exemplary embodiment, the information processing device 3 can function as a client terminal for the server 2. The server 2 is also configured to be able to communicate with external open source data 4 via the telecommunications line (network). Here, the system exemplified as the information processing system 1 is composed of one or more devices or components. Therefore, it should be noted that the information processing system 1 includes either the server 2 alone or both the server 2 and the information processing device 3. More specifically, the information processing system 1 may include an element selected from the group consisting of the server 2 and the information processing device 3. The unselected element may not be included in the information processing system 1, but may be electrically connected to the selected element as an external element. These components will be described below.
[0015] <Server 2> 2 is a block diagram showing the hardware configuration of server 2. Server 2 includes a communication unit 21, a storage unit 22, and a control unit 23, and these components are electrically connected via a communication bus 20 inside server 2. Each component will be further described below.
[0016] The communication unit 21 is preferably a wired communication means such as USB, IEEE1394, Thunderbolt (registered trademark), wired LAN network communication, etc., but may also include wireless LAN network communication, mobile communication such as 3G / LTE / 5G, BLUETOOTH (registered trademark) communication, etc. as needed. In other words, it is more preferable to implement it as a collection of multiple communication means. In other words, the server 2 may communicate various information from the outside via the communication unit 21 and the network.
[0017] The memory unit 22 stores various pieces of information defined above. This can be implemented, for example, as a storage device such as a solid state drive (SSD) that stores various programs and the like related to the server 2 executed by the control unit 23, or as a memory such as a random access memory (RAM) that stores temporarily required information (arguments, arrays, etc.) related to the program operations. The memory unit 22 stores various programs, variables, etc. related to the server 2 executed by the control unit 23.
[0018] The control unit 23 processes and controls the overall operations related to the server 2. The control unit 23 is, for example, a central processing unit (CPU) not shown. The control unit 23 realizes various functions related to the server 2 by reading out predetermined programs stored in the storage unit 22. In other words, information processing by software stored in the storage unit 22 is specifically realized by the control unit 23, which is an example of hardware, and each step related to each function described below can be executed. This will be described in further detail in the next section. Note that the control unit 23 is not limited to being single, and multiple control units 23 may be provided for each function. A combination of these may also be used.
[0019] <Information processing device 3> 3 is a block diagram showing the hardware configuration of the information processing device 3. The information processing device 3 includes a communication unit 31, a storage unit 32, a control unit 33, a display unit 34, and an input unit 35, and these components are electrically connected via a communication bus 30 inside the information processing device 3. Each component will be further described below. The description of the communication unit 31, the storage unit 32, and the control unit 33 will be omitted as they are the same as the description of each unit in the server 2.
[0020] The display unit 34 may be included in the housing of the information processing device 3 or may be externally attached. The display unit 34 displays a graphical user interface (GUI) screen that can be operated by the user. This is preferably implemented by selectively using display devices such as a CRT display, a liquid crystal display, an organic EL display, and a plasma display depending on the type of the information processing device 3.
[0021] The input unit 35 may be included in the housing of the information processing device 3 or may be externally attached. For example, the input unit 35 may be implemented as a touch panel integrated with the display unit 34. The touch panel allows the user to input tapping, swiping, and the like. Of course, switch buttons, a mouse, a QWERTY keyboard, and the like may be used instead of the touch panel. That is, the input unit 35 accepts an operation input made by the user. The input is transferred as a command signal to the control unit 33 via the communication bus 30, and the control unit 33 can execute predetermined control or calculation as necessary.
[0022] The information processing device 3 may be a smartphone, a tablet terminal, a personal computer, a wearable device, or the like.
[0023] <Open Source Data 4> Figure 4 shows the content types of open source data. Open source data 4 is open source data related to research information. Open source data is data that meets the requirements of being "available to anyone (processing, editing, etc.)," "machine-readable," and "free to use." Open source data 4 includes research information open source data related to multiple research information with different content types. The content types of the multiple research information open source data may include research institution information 41, research budget information 42, academic literature information 43, patent document information 44, and collaborative research information 45. The open source data for research institution information 41 may be published on the homepage of each research institution. The open source data related to research budget information 42, academic literature information 43, patent document information 44, and collaborative research information 45 may be a database containing open source data provided by a public institution. In addition, open source data of content types other than those listed here may be used, or open source data containing data of multiple content types may be used.
[0024] 2. Functional configuration of Server 2 The control unit 23 is configured to execute, for example, the following steps: The following steps can be optionally omitted.
[0025] The control unit 23 is configured to be able to receive information from the information processing device 3 or another device as a receiving step. The control unit 23 is also configured to be able to receive various pieces of information by reading out various pieces of information stored in a storage area that is at least a part of the memory unit 22 and writing the read out information in a working area that is at least a part of the memory unit 22. The storage area is, for example, an area of the memory unit 22 that is implemented as a storage device such as an SSD. The working area is, for example, an area that is implemented as a memory such as a RAM. For example, the receiving step may be a step of receiving the name of the first researcher and the name of the institution to which the first researcher belongs. The receiving step may further be a step of receiving information regarding the position of the first researcher.
[0026] The control unit 23 is configured to be able to acquire information from external open source data via a telecommunications line as an acquisition step. For example, the control unit 23 may acquire first research information corresponding to first researcher identification information from external open source data related to research information as an acquisition step.
[0027] The control unit 23 is configured to be able to extract information from external open source data via a telecommunications line as an extraction step. For example, the control unit 23 may extract second research information from the open source data based on the text information of the first research information, which is similar to the first research information and has a name that matches that of the first researcher as an extraction step. The control unit 23 may extract features from the text information of the first research information as an extraction step, and extract the second research information that is highly similar to the first research information by calculating the similarity between each piece of text information using a vector representation of each piece of text.
[0028] As an integration step, the control unit 23 can integrate the first researcher identification information stored in the researcher information database with information related to the first researcher identification information. For example, as an integration step, the control unit 23 may generate or update the researcher information database by integrating the second researcher identification information with the first researcher identification information.
[0029] The control unit 23 is configured to determine, as a determination step, whether researchers included in multiple pieces of research information are the same researcher. For example, as a determination step, the control unit 23 may determine, based on the second research information and the reference information, that the researcher included in the second research information is the same as the first researcher. Specifically, the control unit 23 may determine, based on the similarity between the information on co-authors in the first research information and the second research information, that the researcher included in the second research information is the same as the first researcher. Furthermore, the control unit 23 may refer to open-source data on external research institution information via an electric communication line and determine that the researcher included in the second research information is the same as the first researcher.
[0030] The control unit 23 is configured to be able to compare information about researchers at a research institution with researcher identification information as a comparison step. For example, the control unit 23 may refer to open source data of external research institution information via a telecommunications line as a comparison step, and compare whether there is only one researcher with the same name or multiple researchers at the research institution.
[0031] As a display control step, the control unit 23 controls the display unit 34 to display visual information such as screens, images including still images or videos, icons, messages, etc. The control unit 23 may generate only rendering information for displaying the visual information on the display unit 34. For example, as a display control step, the control unit 23 may search a researcher information database based on instructions received from a user and display the search results. Furthermore, as a display control step, the control unit 23 may display visual information regarding relationships between researchers.
[0032] The control unit 23 is configured to be able to analyze the relationships between researchers by analyzing the network between multiple pieces of researcher information as an analysis step. As an analysis step, the control unit 23 may analyze the relationships between researchers based on first researcher identification information contained in the researcher information database and research information associated with the first researcher identification information. The control unit 23 may also visualize network structures and patterns based on the analysis results.
[0033] In the presentation step, the control unit 23 presents candidate researchers for collaborative research to the specific researcher based on the first researcher identification information and research information contained in the researcher information database.
[0034] 3. Information processing flow This section describes the flow of an information processing method executed by the information processing system 1. As shown below, the information processing method includes each step executed by the information processing system. The information processing program of this embodiment causes a computer to execute each step of the information processing system. Note that the order of the processes can be changed as appropriate, multiple processes may be executed simultaneously, or some processes may be omitted.
[0035] 3.1 Overview FIG. 5 is a diagram illustrating an overview of the processing executed by the information processing system 1. In this processing, first, the control unit 23 receives the name of a first researcher and the name of the institution to which the first researcher belongs (step S001) as a receiving step. Next, the control unit 23 acquires first research information corresponding to the first researcher identification information from external open source data related to research information (step S002) as an acquiring step. Next, the control unit 23 extracts second research information that is similar to the first research information and has a name that matches the first researcher from the open source data based on the text information of the first research information (step S003) as an extracting step. Next, the control unit 23 generates or updates the researcher information database by integrating second researcher identification information corresponding to the second research information with the first researcher identification information (step S004) as an integrating step.
[0036] In summary, an information processing system according to one embodiment includes at least one processor, and the at least one processor is programmed to perform the following steps: In a receiving step, the control unit 23 receives the name of a first researcher and the name of the institution to which the first researcher belongs. The name and the name of the institution are used as first researcher identification information to identify the first researcher in the researcher information database. In an acquiring step, the control unit 23 acquires first research information corresponding to the first researcher identification information from external open-source data related to research information. In an extracting step, the control unit 23 extracts second research information from the open-source data that is similar to the first research information and has a name that matches the first researcher, based on the text information of the first research information. In an integrating step, the control unit 23 integrates second researcher identification information corresponding to the second research information with the first researcher identification information. This aspect makes it possible to generate or update a researcher information database that integrates researcher information.
[0037] 3.2 Specific examples The above information processing will be described in detail below as an example with reference to FIGS.
[0038] Incidentally, research information about Researcher X is often made public as open source data. However, this research information is posted on multiple websites and databases, and comprehensive information about each researcher is not compiled. Furthermore, researchers often move between multiple research institutions while continuing their research, so it may not be possible to obtain information about past research conducted using only the current researcher's name and the name of the institution to which they belong.
[0039] Therefore, the control unit 23 acquires first research information about Researcher X from open source data using Researcher X's name and the name of the research institution to which he / she belongs, performs text analysis on the text information of the acquired second research information, and collects second research information about Researcher X by utilizing the similarities between the text information contained in the research information. The control unit 23 then collects further research information about Researcher X using Researcher X's name and the name of the institution to which he / she belongs contained in the second research information. The information processing system of this embodiment repeats this process to comprehensively collect research information about Researcher X. Furthermore, the information processing system of this embodiment collects research information on a researcher-by-researcher basis in this way, and can perform searches and analyses using the generated or updated researcher information database.
[0040] FIG. 6 is an activity diagram showing an example of the flow of processing executed by the information processing system 1. This example of the flow may fall within the scope defined in the above-mentioned overview. The following description will be given along with each activity in this activity diagram. Note that the information processing may include any exception handling not shown. Exception handling includes the interruption of the information processing or the omission of each process. Selections or inputs made in the information processing may be based on user operation or may be made automatically without user operation.
[0041] <Registration in the researcher information database> First, the control unit 23 receives the name of the researcher X and the name of the research institution to which the researcher X belongs as the first researcher identification information of the researcher X (activity A101). The name may be the full name written in kanji or the alphabet, or the first letter (initial) of the first name and the last name may be received. The name of the research institution may be only the name of the affiliated institution, or may include the name of the research institution as well as the name of the faculty, department, project name, division name, etc.
[0042] Next, as a comparison step, the control unit 23 may compare the name of the first researcher with information about researchers at the research institution that corresponds to the name of the institution to which Researcher X, an example of a first researcher, belongs (activity A102). If, when compared with open source data related to the research institution, there are multiple researchers whose researcher identification information matches that of Researcher X, the control unit 23 proceeds to activity A103. If it is confirmed that there is only one researcher at the research institution whose name matches, the control unit 23 proceeds to activity A104, since the name and the name of the institution can be used as the first researcher identification information.
[0043] If there are multiple researchers matching the researcher identification information, the control unit 23 then accepts additional information for the first researcher identification information (activity A103). The control unit 23 may also accept information on the position of researcher X as additional information for the first researcher identification information. Positions include professor, associate professor, lecturer, assistant professor, and assistant. In other words, if it is confirmed that there are multiple researchers with the same name at the research institution, the control unit 23 may further accept the researcher's position in the accepting step and use the name, the name of the affiliated institution, and the position as the first researcher identification information. For example, there may be researchers with the same full name written in kanji, and in foreign journals, only initials and last names may be listed as author names, so there may be multiple researchers with the same name at the same research institution. However, since it is rare for researchers with the same research institution name, the same position, and the same name to exist, this configuration makes it highly likely that the first researcher identification information will identify a single researcher.
[0044] Next, the control unit 23 acquires the first research information (activity A104). Specifically, as an acquisition step, the control unit 23 acquires the first research information corresponding to the first researcher identification information from external open source data related to research information. If the control unit 23 acquires the first research information, it proceeds to activity A105. If the control unit 23 cannot acquire the first research information, it proceeds to activity A110.
[0045] Next, the control unit 23 extracts research information similar to the first research information (activity A105). Specifically, as an extraction step, the control unit 23 extracts research information similar to the first research information from the open source data based on the text information of the first research information. The determination of whether the research information is similar to the first research information can be performed by extracting features from the text information of the first research information and calculating the similarity with other research information in the open source data using a vector representation of the text information. If the control unit 23 has extracted research information similar to the first research information, it proceeds to activity A106. If there is no research information similar to the extracted first research information, it proceeds to activity A110.
[0046] In activity A105, the control unit 23 preferably uses the research topic or abstract included in the research information as the text information in the extraction step. The control unit 23 may extract features from the text information of the research topic or abstract included in the research information, and extract second research information that is highly similar to the first research information. This is because research topics and abstracts succinctly express the content of the research, and are often highly similar in a series of research projects by the same researcher.
[0047] Next, the control unit 23 determines whether there is second research information containing the name of researcher X among the research information similar to the extracted first research information (activity A106). If there is second research information containing the name of researcher X, the control unit 23 proceeds to activity A107. If there is no second research information containing the name of researcher X, the control unit 23 proceeds to activity A110.
[0048] Next, as a determination step, the control unit 23 may determine whether the researcher included in the second research information is the same person as the first researcher, researcher X, based on the second research information and the reference information (activity A107). For example, the control unit 23 may refer to research information already associated with researcher X and determine whether the researcher included in the second research information is the same person as researcher X based on the content of the research and information on co-authors. This embodiment makes it possible to avoid mixing research information of another person with the same name as researcher X. If the control unit 23 determines that the second research information belongs to researcher X, it proceeds to activity A108. If it determines that the second research information does not belong to researcher X, it proceeds to activity A110.
[0049] Next, the control unit 23 stores the second research information in association with the first researcher identification information, which is researcher X (activity A108). In this manner, the first research information and the second research information are stored in association with the first researcher identification information. Each piece of stored research information may include information such as the research topic name, the selected topic at the time of initial research adoption, an outline of the research results, a research implementation status report, a research performance report, research keywords, the principal researcher, the research period, the research field, the name of the external funding project, the amount obtained, paper keywords, the name of the journal in which the paper was published, the content of the paper, the paper title, the patent title, patent keywords, the names of co-inventors, and the content of the patent. The control unit 23 may store the full text of reports, academic literature, patent documents, etc.
[0050] Next, the control unit 23 integrates the second researcher identification information into the first researcher identification information (activity A109). In this manner, the control unit 23 can generate or update the researcher information database.
[0051] When second researcher identification information is newly integrated into first researcher identification information, the control unit 23 returns to activity A102 and executes the processing of activities A104-109 based on the newly integrated first researcher identification information. The control unit 23 repeats the information processing of activities A104-109 until it proceeds to activity A110. In other words, when second researcher identification information is integrated into first researcher identification information, the information processing system uses the integrated first researcher identification information to repeat the acquisition step, extraction step, and integration step until there is no more second researcher identification information to be newly integrated.
[0052] The control unit 23 proceeds to activity A110 depending on the results of any of activities A104-107. If there is researcher identification information for the next researcher Y, the control unit 23 returns to activity A101. If there is no researcher identification information for the next researcher Y, the control unit 23 ends registration in the researcher information database.
[0053] FIG. 7 is a diagram showing an example of open source data used by the information processing system 1. In activity A102, it is preferable to use open source data related to research institution information. Furthermore, in activities A104-107, it is preferable to use open source data of multiple content types as external open source data related to research information. As the open source data of multiple specific content types, it is preferable to use, for example, research budget information open source data related to research budgets, academic literature information open source data related to academic literature, patent document information open source data related to patent documents, and collaborative research information open source data related to collaborative research. It is even more preferable to use at least two or more of the multiple specific content type open source data.
[0054] The control unit 23 performs information processing for activity A104-107 for one of the open source data, and then performs information processing for activity A104-107 for another of the open source data. For example, the control unit 23 may perform information processing for activity A104-107 using research budget information open source data related to the research budget, then perform information processing for activity A104-107 using academic literature information open source data related to academic literature, then perform information processing for activity A104-107 using patent document information open source data related to patent documents, and finally perform information processing for activity A104-107 using collaborative research information open source data related to collaborative research. The order in which the open source data are used may be reversed. If second researcher identification information is later integrated, the open source data used once before the second researcher identification information was integrated may be used again to perform information processing for activity A104-107. In other words, the open source data includes multiple external open source data of specific content types, each with a different content type of research information. For each of the plurality of specific content type open source data, it is preferable to repeat the acquisition step, extraction step, and integration step using the first researcher identification information until there is no second researcher identification information to be newly integrated. In this manner, it is possible to collect more comprehensive research information on Researcher X.
[0055] Furthermore, it is preferable that the information processing system first uses the research budget information open source data among the plurality of specific content type open source data to perform the acquisition step, extraction step, and integration step, and then repeats the acquisition step, extraction step, and integration step using the first researcher identification information for each of the plurality of other specific content type open source data until there is no second researcher identification information to be newly integrated. This is because the research budget information open source data often contains accurate descriptions of the researcher's research topic.
[0056] FIG. 8 is a diagram showing an example of researcher information 8. Researcher information 8 includes researcher information 81 and researcher information 82. Researcher information 81 and researcher information 82 include researcher identification information and research information. The researcher identification information includes, as items, for example, the name of the researcher, the name of the research institution to which the researcher belongs, and the researcher's position. The research information includes, as items, for example, research budget information related to the researcher, patent document information related to the researcher, academic document information related to the researcher, and collaborative research information related to the researcher. Researcher information 81 and researcher information 82 each list one item as an example, but multiple pieces of information may be associated with one researcher for each item.
[0057] <Utilizing the researcher information database> Next, using Figures 6 and 9-10, we will explain how to use the researcher information database in which research information is collected and registered on a researcher-by-researcher basis through activities A104-110.
[0058] First, the control unit 23 accepts an instruction from the user (activity A111). FIG. 9 is a diagram showing an example of the search result reception screen 9. The search result reception screen 9 includes areas 90-99 and buttons Bt1-Bt3. Area 90 is an area for accepting input of a researcher's name. Area 91 is an area for accepting the name of the research institution to which the researcher belongs. Area 91 may be configured to accept the name of the user's institution in several stages. For example, area 91 may be configured to accept the university name as a major category, the faculty name as a medium category, and the department name as a minor category. Area 92 is an area for accepting a job title. Area 93 is an area for accepting the name of the research budget representative. Area 94 is an area for accepting the number of papers in the top 10%. Area 95 is an area for accepting the total number of patents. Area 96 is an area for accepting the total number of papers. Area 97 is an area for accepting the international co-authorship rate. Area 98 is an area where the number of hits in the search results is displayed. Area 99 is an area where an instruction to close the refined search is accepted. Bt1 is a button where an instruction to cancel accepted search conditions is accepted. Bt2 is a button where an instruction to execute a search is accepted. Bt3 is a button where an instruction to display the detailed search screen is accepted. Note that the "detailed search" is configured to be able to accept search items other than those on the search acceptance screen 9. The search acceptance screen 9 is merely an example, and the search acceptance screen may be configured with other search items. Other search items include, for example, the research field, keywords of the paper, names of journals where the paper was published, names of externally funded projects, amounts obtained, patent names, patent keywords, names of co-inventors, etc.
[0059] Next, the control unit 23 displays the search results and analysis results based on the received user instructions (activity A112). FIG. 10 is a diagram showing an example of a search result 10. The search result 10 includes areas 101-106. Area 101 is an area where the number of hits in the search result and the number of hits currently being displayed are displayed. Area 102 is an area where the selection of display items is received. Area 103 is an area where the researcher identification information of the researcher who was hit as the first hit is displayed. Area 104 is an area where the researcher identification information of the researcher who was hit as the second hit is displayed. Area 105 is an area where the total number of papers of the researcher who was hit as the first hit is displayed. Area 106 is an area where the total number of papers of the researcher who was hit as the second hit is displayed. In this way, the control unit 23 may search the researcher information database and display the search results based on instructions received from the user as a display control step.
[0060] Furthermore, in activity A111, the control unit 23 may receive an analysis instruction from the user. In the researcher information database, research information is centrally organized by researcher, enabling new researcher-based analyses, such as network analysis between researchers, as well as more detailed analyses by research institution and by research field. The control unit 23 may perform an analysis based on the received analysis instruction and generate visual information related to the analysis results. For example, the control unit 23 may receive an analysis instruction to analyze relationships between researchers based on first researcher identification information and research information contained in the researcher information database as an analysis step. The control unit 23 may display visual information related to the relationships as a display control step. Examples of relationships between researchers that can be analyzed include paper co-author networks, paper citation networks, text similarity networks based on paper text analysis, external funding collaborative researcher networks, patent co-inventor networks, patent citation networks, and affiliation information networks. This configuration allows various networks to be visualized, enabling a bird's-eye view of research trends.
[0061] Furthermore, the analysis instruction from the user may be an analysis instruction to analyze candidate collaborative research partners for a specific researcher. In this case, the instruction received from the user includes at least the name of the specific researcher. In the presentation step, the control unit 23 presents candidate collaborative research partners for the specific researcher based on the first researcher identification information and research information contained in the researcher information database. The control unit 23 may present candidate collaborative research partners based on the content of the paper, the paper title, the name of the external funding project, text information, information on various researcher networks, the patent title, the patent content, etc. This mode can promote collaborative research between researchers.
[0062] In the above embodiment, the researcher's name, research institution name, and position were used as researcher identification information. However, depending on the open source data used, a unique researcher identification code may be assigned. In such cases, the researcher's name and affiliated institution name linked to the researcher identification code may be extracted and compared with the first researcher identification information. If any information differs, it may be integrated into the first researcher identification information. This type of configuration allows for more comprehensive collection of researcher identification information and research information.
[0063] By repeating the processing of activities A101-110 for each researcher in turn, the researcher information database can be updated. By updating it periodically, the latest information can be included.
[0064] When acquiring and extracting data from open source data that provides APIs, it is preferable to use APIs, as this allows for efficient and automated data collection.
[0065] The overall configuration shown in Fig. 1 is an example and is not limited to this. For example, the server 2 may be distributed across two or more devices, or may be replaced by a cloud computing system. Furthermore, all processing may be performed by the server 2, or all processing may be performed by the information processing device 3. An application may be installed on the information processing device 3, and the information processing device 3 and the server 2 may work together to execute the processing described above.
[0066] The server 2 may be an on-premise server or a cloud server. The cloud server 2 may provide the above functions and processes in the form of, for example, SaaS (Software as a Service) or cloud computing.
[0067] In the above embodiment, the server 2 performs various storage and control operations, but multiple external devices may be used instead of the server 2. That is, various information and programs may be distributed and stored in multiple external devices using block chain technology or the like.
[0068] It may be provided in the following manner.
[0069] (1) An information processing system comprising at least one processor, the at least one processor being configured to execute a program to perform the following steps: in a receiving step, a name of a first researcher and a name of the institution to which the first researcher belongs are received, wherein the name and the name of the institution are used as first researcher identification information to identify the first researcher in a researcher information database; in an acquiring step, first research information corresponding to the first researcher identification information is acquired from external open source data related to research information; in an extracting step, second research information that is similar to the first research information and whose name matches that of the first researcher is extracted from the open source data based on text information of the first research information; and in an integrating step, second researcher identification information corresponding to the second research information is integrated with the first researcher identification information, thereby generating or updating the researcher information database.
[0070] According to this embodiment, it is possible to generate or update a researcher information database that integrates information about researchers.
[0071] (2) In the information processing system described in (1) above, the determination step further determines, based on the second research information and reference information, that the researcher included in the second research information is the same person as the first researcher.
[0072] (3) In the information processing system described in (1) or (2) above, when the second researcher identification information is integrated into the first researcher identification information, the system repeats the acquisition step, the extraction step, and the integration step using the integrated first researcher identification information until there is no more second researcher identification information to be integrated.
[0073] (4) In an information processing system described in any one of (1) to (3) above, the open source data includes a plurality of external specific content type open source data, each having a different content type of the research information, and for each of the plurality of specific content type open source data, the acquisition step, the extraction step, and the integration step are repeated using the first researcher identification information until there is no more second researcher identification information to be newly integrated.
[0074] (5) In the information processing system described in (4) above, the plurality of specific content type open source data includes at least two of research budget information open source data related to research budgets, academic literature information open source data related to academic literature, patent document information open source data related to patent documents, and collaborative research information open source data related to collaborative research.
[0075] (6) In the information processing system described in (5) above, the system first uses the research budget information open source data from among the plurality of specific content type open source data to perform the acquisition step, the extraction step, and the integration step.
[0076] (7) In the information processing system described in any one of (1) to (6) above, further, in the matching step, the name of the first researcher is matched with information on researchers at a research institution corresponding to the name of the institution to which the first researcher belongs, and if it is confirmed that there is only one researcher at the research institution whose name matches, the name and the name of the institution are used as the first researcher identification information, and if it is confirmed that there are multiple researchers at the research institution whose name matches, the receiving step further receives the researcher's position, and the name, the name of the institution and the position are used as the first researcher identification information.
[0077] (8) In the information processing system according to any one of (1) to (7) above, in the extraction step, a research topic or an abstract included in the research information is used as the text information.
[0078] (9) In the information processing system described in any one of (1) to (8) above, the display control step further includes searching the researcher information database based on instructions received from the user and displaying the search results.
[0079] (10) In the information processing system described in (9) above, the analysis step further analyzes the relationships between researchers based on the first researcher identification information and the research information contained in the researcher information database, and the display control step displays visual information regarding the relationships.
[0080] (11) In the information processing system described in (9) or (10) above, the instructions received from the user include at least the name of a specific researcher, and further, in the presentation step, potential researchers for collaborative research with the specific researcher are presented based on the first researcher identification information and the research information contained in the researcher information database.
[0081] (12) An information processing method, comprising the steps of the information processing system according to any one of (1) to (11) above.
[0082] (13) A program that causes at least one computer to execute each step in the information processing system according to any one of (1) to (11) above. Of course, this is not the case.
[0083] Finally, while various embodiments of the present invention have been described, these are presented by way of example only and are not intended to limit the scope of the invention. The novel embodiments may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. Such embodiments and modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the inventions and their equivalents as defined in the accompanying claims. [Explanation of symbols]
[0084] 1: Information processing system 2: Server 20: Communication bus 21: Communications Department 22: Storage section 23: Control section 3: Information processing equipment 30: Communication bus 31: Communications Department 32: Storage section 33: Control section 34:Display section 35: Input section 4: Open source data 41: Research institution information 42: Research budget information 43: Academic literature information 44: Patent document information 45: Joint research information 8:Researcher information 81:Researcher information 82:Researcher information 9: Search reception screen 90: area 91: area 92: area 93 :Area 94 :Area 95 :Area 96 :Area 97 :Area 98 :Area 99: area 10: Search results 101 :Area 102 :Area 103 :Area 104 :Area 105 :Area 106 :Area Bt1: Button Bt2: Button Bt3: Button
Claims
1. An information processing system, The method includes at least one processor, the at least one processor being configured to execute a program to perform the following steps: the receiving step receives the name of a first researcher and the name of the institution to which the first researcher belongs, wherein the name and the name of the institution are used as first researcher identification information for identifying the first researcher in a researcher information database; In the acquisition step, first research information corresponding to the first researcher identification information is acquired from external open source data related to research information; In the extraction step, second research information similar to the first research information and having a name that matches that of the first researcher is extracted from the open source data based on text information of the first research information; In the integrating step, the system generates or updates the researcher information database by integrating second researcher identification information corresponding to the second research information into the first researcher identification information.
2. 2. The information processing system according to claim 1, Furthermore, in the determination step, the system determines that the researcher included in the second research information is the same person as the first researcher based on the second research information and the reference information.
3. 2. The information processing system according to claim 1, When the second researcher identification information is integrated with the first researcher identification information, the system repeats the acquisition step, the extraction step, and the integration step using the integrated first researcher identification information until there is no more second researcher identification information to be integrated.
4. 2. The information processing system according to claim 1, The open source data includes a plurality of external open source data of specific content types, each of which has a different content type of the research information; A system that repeats the acquisition step, extraction step, and integration step using the first researcher identification information for each of the plurality of specific content type open source data until there is no new second researcher identification information to be integrated.
5. 5. The information processing system according to claim 4, The system, wherein the plurality of specific content type open source data includes at least two of research budget information open source data related to research budgets, academic literature information open source data related to academic literature, patent document information open source data related to patent documents, and collaborative research information open source data related to collaborative research.
6. 6. The information processing system according to claim 5, The system executes the acquisition step, the extraction step, and the integration step by first using the research budget information open source data from among the plurality of specific content type open source data.
7. 2. The information processing system according to claim 1, Furthermore, in the collation step, information on researchers at a research institution corresponding to the name of the institution to which the first researcher belongs is collated with the name of the first researcher; If it is confirmed that there is only one researcher with the same name at the research institution, the name and the name of the affiliated institution are used as the first researcher identification information; If it is confirmed that there are multiple researchers with the same name at the research institution, the system further accepts the researcher's position in the reception step, and uses the name, the name of the affiliated institution, and the position as the first researcher identification information.
8. 2. The information processing system according to claim 1, In the extraction step, a research topic or an abstract contained in the research information is used as the text information.
9. 2. The information processing system according to claim 1, Furthermore, in the display control step, the system searches the researcher information database based on instructions received from the user and displays the search results.
10. 10. The information processing system according to claim 9, Furthermore, in the analysis step, relationships between researchers are analyzed based on the first researcher identification information and the research information contained in the researcher information database; In the display control step, visual information regarding the relationship is displayed.
11. 10. The information processing system according to claim 9, The instructions received from the user include at least the name of the specific researcher; Furthermore, in the presentation step, the system presents candidate researchers for collaborative research to the specific researcher based on the first researcher identification information and the research information contained in the researcher information database.
12. An information processing method, comprising: A method including the steps in the information processing system according to any one of claims 1 to 11.
13. A program, A program that causes at least one computer to execute each step in the information processing system according to any one of claims 1 to 11.
Citation Information
Patent Citations
Information providing system
JP2010224952A