Research data management infrastructure system, research data management method, and program
The research data management infrastructure system addresses the challenge of accessing lower-level data by using a large-scale language model to extract and link lower-level data from higher-level data, improving data management and search efficiency.
Patent Information
- Application Number
- PCT/JP2025/023086
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2025-06-26
- Publication Date
- 2026-01-08
AI Technical Summary
Existing systems lack the ability to efficiently collect and create lower-level data such as raw data from higher-level data like copyrighted works, making it difficult to access and manage raw data associated with higher-level information.
A research data management infrastructure system that utilizes a large-scale language model to extract relationship information and feature keywords from higher-level data, linking and outputting lower-level data and metadata based on the higher-level data, metadata, and extracted keywords.
Enables efficient collection and management of lower-level data, facilitating access and linkage to raw data from higher-level information, enhancing data utilization and search capabilities.
Smart Images

Figure JP2025023086_08012026_PF_FP_ABST
Abstract
Description
Research data management infrastructure system, research data management method, and program
[0001] The present disclosure relates to a research data management infrastructure system, a research data management method, and a program.
[0002] Patent Document 1 discloses a system for managing and organizing content.
[0003] Special Publication No. 2008-502047
[0004] However, there was no system for collecting and creating lower-level data such as raw data from higher-level data such as copyrighted works. Therefore, the purpose of this disclosure is to provide a research data management infrastructure system that collects and creates lower-level data from higher-level data.
[0005] The research data management infrastructure system disclosed herein is a research data management infrastructure system comprising: an extraction unit that extracts relationship information and feature keywords from a large-scale language model to which higher-order data and metadata associated with the higher-order data have been input; and an output unit that outputs lower-order data and metadata associated with the lower-order data in comparison with the higher-order data based on the higher-order data, the metadata associated with the higher-order data, the relationship information, and the feature keywords.
[0006] The research data management method disclosed herein is a research data management method that extracts relationship information and feature keywords from a large-scale language model to which higher-order data and metadata associated with the higher-order data have been input, and outputs lower-order data and metadata associated with the lower-order data in comparison with the higher-order data based on the higher-order data, the metadata associated with the higher-order data, the relationship information, and the feature keywords.
[0007] The program disclosed herein causes an information processing device to extract relationship information and feature keywords from a large-scale language model to which higher-order data and metadata associated with the higher-order data have been input, and output lower-order data and metadata associated with the lower-order data in comparison with the higher-order data based on the higher-order data, the metadata associated with the higher-order data, the relationship information, and the feature keywords.
[0008] The present disclosure makes it possible to provide a research data management infrastructure system that collects and creates lower-level data from higher-level data.
[0009] FIG. 1 is a diagram showing the relationship between higher-level data and lower-level data according to the present disclosure. FIG. 2 is a block diagram showing the configuration of a research data management infrastructure system according to the present disclosure. FIG. 3 is a flowchart of a research data management method according to the present disclosure. FIG. 4 is a block diagram showing the configuration of a research data management infrastructure system 1 according to the present disclosure. FIG. 5 is a flowchart of feature vector generation for the research data management infrastructure system 1 according to the present disclosure. FIG. 6 is a flowchart of search method 1 for the research data management infrastructure system 1 according to the present disclosure. FIG. 7 is a flowchart of search method 2 for the research data management infrastructure system 1 according to the present disclosure. FIG. 8 is a block diagram showing the configuration of an information processing device according to the present disclosure.
[0010] (Explanation of Higher-Level Data and Lower-Level Data in the Present Disclosure) FIG. 1 is a diagram illustrating the relationship between higher-level data and lower-level data according to the present disclosure. As shown in FIG. 1, information can be hierarchically classified in a pyramidal structure. As shown on the left side of FIG. 1, in English, the following terms are used: DATA 101, INFORMATION 102, KNOWLEDGE 103, and WISDOM 104. As shown on the right side of FIG. 1, in Japanese, the following terms are used: raw data 105, analyzed information 106, papers / works 107, and social implementation 108. Raw data 105 is the data itself, such as data recorded by a seismograph. Analyzed information 106 is seismograph data from a major earthquake on a certain date. Papers / works 107 are papers generated from data from the major earthquake, such as those investigating the cause of the earthquake, such as why it occurred on an active fault in XX. Social implementation 108 could be, for example, creating an ordinance prohibiting the construction of earthquake-vulnerable buildings in XX.
[0011] Low-level data refers to the lower layers of the pyramid in Figure 1, and high-level data is data that is higher than the lower-level data. Therefore, from the perspective of the paper / work 107, the raw data 105 and analyzed information 106 are low-level data. Typically, the lower the level of data, the larger the data volume. Furthermore, raw data that has not been assigned meaning is often left untouched. Furthermore, it is not easy to access the raw data 105 from the paper / work 107. Therefore, the research data management infrastructure system disclosed herein is a system that manages the raw data 105 and analyzed information 106 and can link the raw data 105 and analyzed information 106 from the paper / work 107.
[0012] (Description of Research Data Management Infrastructure System According to an Embodiment) Fig. 2 is a block diagram showing the configuration of a research data management infrastructure system according to the present disclosure. The research data management infrastructure system according to an embodiment will be described with reference to Fig. 2.
[0013] As shown in Figure 2, the research data management infrastructure system 200 of the present disclosure comprises an extraction unit 201 and an output unit 202. The extraction unit 201 extracts relationship information 204 (shown in Figure 4) and feature keywords 205 (shown in Figure 4) from a large-scale language model (LLM) 203 (shown in Figure 4) to which high-level data and metadata associated with the high-level data have been input.
[0014] The higher-level data corresponds to the above-described paper / work 107. Metadata associated with the higher-level data includes, for example, the author of the work, the registrant, the location of an event, or the time of occurrence of the event. The data on the paper / work 107 and the author are input into a large language model (LLM) 203.
[0015] The large-scale language model (LLM) 203 is a language model constructed using an extremely large dataset and deep learning technology. The large-scale language model is capable of fluent conversations close to those of humans and can perform various processes using natural language with high accuracy.
[0016] The LLM 203, to which the high-level data and metadata associated with the high-level data have been input, extracts relationship information 204 and characteristic keywords 205. For example, the characteristic keywords 205 are words characteristic of events such as the Tohoku region, the major earthquake, and the year 2011. The relationship information 204 is information indicating the relationship between keywords including a sentence such as "A major earthquake occurred in the Tohoku region in 2011."
[0017] The output unit 202 outputs the lower-level data and the metadata associated with the lower-level data in comparison with the higher-level data based on the higher-level data, the metadata associated with the higher-level data, the relationship information 204, and the feature keywords 205.
[0018] The research data management infrastructure system stores raw data 105 or has access to the raw data 105. Therefore, raw data, which is lower-level data, is linked and output based on higher-level data, metadata associated with the higher-level data, relationship information 204, and feature keywords 205. For example, seismometer readings from March 11, 2011, are linked to the Great Tohoku Earthquake and output.
[0019] The metadata associated with the low-level data may include, for example, time, location, and registrant information. In addition, if the low-level data is analyzed information such as information about the earthquake swarm in the Tohoku region, the analysis conditions may be included as metadata associated with the low-level data.
[0020] The research data management infrastructure system 100 is constructed using an information processing device 800. As shown in Figure 8, the information processing device 800 physically comprises a memory 801 that stores programs and a processor 802 that executes the programs and performs processing. The information processing device 800 may be composed of a single device or multiple devices. The information processing device 800 may also be a cloud server that distributes and processes some or all of its functions.
[0021] The program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored on a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals.
[0022] The above configuration makes it possible to provide a research data management infrastructure system that collects and creates lower-level data from higher-level data. The extraction unit 201 and output unit 202 may be read as extraction means and output means.
[0023] (Description of Research Data Management Method According to Embodiment) Fig. 3 is a flowchart of the research data management method according to the present disclosure. The research data management method according to the embodiment will be described with reference to Fig. 3. The research data management method is executed, for example, by the above-mentioned research data management infrastructure system.
[0024] As shown in Fig. 3, first, relationship information and feature keywords are extracted (step S301). The extraction unit 201 extracts relationship information 204 and feature keywords 205 from a large-scale language model (LLM) 203 to which high-level data and metadata associated with the high-level data have been input. Next, the low-level data and metadata are output (step S302). The output unit 202 compares the high-level data with the low-level data and the metadata associated with the low-level data based on the high-level data, the metadata associated with the high-level data, the relationship information 204, and the feature keywords 205, and outputs the low-level data and the metadata associated with the low-level data.
[0025] The above configuration provides a research data management infrastructure method for collecting and creating lower-level data from higher-level data.
[0026] (Description of Research Data Management Infrastructure System According to First Embodiment) FIG. 4 is a block diagram showing the configuration of a research data management infrastructure system 1 according to the present disclosure. FIG. 5 is a flowchart of feature vector generation in the research data management infrastructure system 1 according to the present disclosure. FIG. 6 is a flowchart of search method 1 in the research data management infrastructure system 1 according to the present disclosure. FIG. 7 is a flowchart of search method 2 in the research data management infrastructure system 1 according to the present disclosure. The research data management infrastructure system according to the first embodiment will be described with reference to FIGS. 4 to 7.
[0027] The research data management infrastructure system 400 according to the first embodiment differs from the research data management infrastructure system 200 according to the embodiment in that coded feature amount information is extracted.
[0028] The research data management infrastructure system 400 according to the first embodiment inputs papers and works 107, which are papers (evaluation results) that represent secondary data, and metadata 402, which is basic information on registrants, etc., into the LLM 203.
[0029] The LLM 203 extracts relationship information 204 and characteristic keywords 205. The research data management infrastructure system 400 outputs analyzed information 106, metadata 407, and metadata 408 based on the paper / work 107, metadata 402, relationship information 204, and characteristic keywords 205. The analyzed information 106 is secondary data (analysis results) created from the data. The metadata 407 is basic information such as the registrant. The metadata 408 is basic information such as the time and analysis conditions.
[0030] Alternatively, the research data management infrastructure system 400 outputs the raw data 105, which is the registered data, metadata 410 such as basic information on the registrant, and metadata 411, which is essential for data usage, based on the papers / works 107, metadata 402, relationship information 204, and characteristic keywords 205.
[0031] The research data management infrastructure system 400 is equipped with a feature information output unit that extracts encoded feature information 412 from the raw data 105 or analyzed information 106 and generates feature information 413, which is information that vectorizes the features of the data.
[0032] The feature information output unit will be described with reference to Fig. 5. The feature information output unit may be read as feature information output means.
[0033] First, a set of feature digital data and units is acquired (step S501). For raw data such as experimental data, the "type" is converted into a one-hot vector and the neural network is trained using the WordVec method to learn the relationship with related data, thereby enabling data encoding with a small number of dimensions.
[0034] Next, encoding is performed (step S502). The corresponding "value" is used as is since it has been encoded as is. Next, embedding is performed to create a feature vector (step S503). The encoded information is further encoded to create a feature vector so that the "feature distance" can be calculated. Finally, the feature vector is obtained (step S504).
[0035] For example, "queen" is the sum of "king" and "woman," so it has the feature vectors of "king" and "woman." Therefore, if you perform a synonym search for "king," you can find words such as "king," "queen," and "emperor." By giving terms feature vectors in this way, it becomes easier to search for synonyms based on the meaning of the term.
[0036] 6 and 7, the search for lower-level data will be described. It is not easy to identify raw data from higher-level data such as papers. By using the research data management infrastructure system 100 or 400 of the present disclosure, lower-level data can be searched from higher-level data.
[0037] As shown in Figure 6, by using the LLM 203 of the research data management infrastructure system 200, it is possible to extract keywords 604 and Encoder usage data 603 from high-level data such as papers and works 107. Extracting Encoder usage data refers to extracting information from high-level data such as papers and works. Here, the LLM 203 is an LLM basic large-scale language model specialized for academic papers. It is preferable that this LLM 203 be a language model specialized for terminology used only in research fields, such as village slang.
[0038] Feature keywords 205 are extracted from keyword extraction 604. Relationship information 204 is extracted from Encoder-used data extraction 603. Next, a query 608 of the subject to be searched is issued to Decoder-used data extraction 607, which outputs lower-level data. Decoder-used data extraction 607 refers to the extraction of information from raw data or analyzed data.
[0039] Therefore, in other words, a query is issued against the raw data or the analyzed data, and a search result 609 is obtained, which is the raw data found in this way.
[0040] 7, low-level data can be searched for using the characteristic keywords 205 and feature quantity information 413 of the research data management infrastructure system 400. From the characteristic keywords 205, a query 701 can be issued for example data search 703, which selects and specifies a keyword, and for focus point selection 704, which selects closely related words.
[0041] A query 701 can also be issued for a focus point selection 706 that selects information that is highly related from the feature amount information 413. A feature distance search 707 can be performed by issuing a query 701 after weighting the focus points for the feature keywords 205 and the feature amount information 413. Therefore, it is possible to extract 607 data used by the decoder, such as raw data, and obtain search results 709.
[0042] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0043] Each drawing is merely an example for describing one or more embodiments. Each drawing may not relate to only one particular embodiment, but may also relate to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessary to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.
[0044] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes. (Supplementary Note 1) A research data management infrastructure system comprising: extraction means for extracting relationship information and feature keywords from a large-scale language model to which high-level data and metadata associated with the high-level data have been input; and output means for outputting lower-level data and metadata associated with the lower-level data in comparison with the high-level data, based on the high-level data, the metadata associated with the high-level data, the relationship information, and the feature keywords. (Supplementary Note 2) The research data management infrastructure system according to Supplementary Note 1, further comprising feature information output means for outputting feature information obtained by encoding the lower-level data and the metadata associated with the lower-level data. (Supplementary Note 3) The research data management infrastructure system according to Supplementary Note 2, wherein the feature information is encoded so that the encoded information can be converted into a feature vector and feature distance can be calculated. (Supplementary Note 4) The research data management infrastructure system according to Supplementary Note 1, wherein a query is issued to data from which the relationship information and the feature keywords have been extracted, and a search is performed. (Supplementary Note 5) The research data management infrastructure system according to Supplementary Note 2, which issues a query to the feature keywords and the feature quantity information and performs a search by weighting points of interest. (Supplementary Note 6) The research data management infrastructure system according to Supplementary Note 1, wherein the higher-level data is a paper or a work. (Supplementary Note 7) The research data management infrastructure system according to Supplementary Note 1, wherein the metadata accompanying the higher-level data is information on the registrant, the location of an event, or the time of occurrence of an event. (Supplementary Note 8) The research data management infrastructure system according to Supplementary Note 1, wherein the lower-level data is raw data indicating an event or analytical data created from raw data indicating an event. (Supplementary Note 9) The research data management infrastructure system according to Supplementary Note 1, wherein the data accompanying the lower-level data is information on time, location, analysis conditions, or registrant.(Supplementary Note 10) A research data management method comprising: extracting relationship information and feature keywords from a large-scale language model into which higher-order data and metadata associated with the higher-order data have been input; and outputting lower-order data and metadata associated with the lower-order data in comparison with the higher-order data, based on the higher-order data, the metadata associated with the higher-order data, the relationship information, and the feature keywords. (Supplementary Note 11) A program for causing an information processing device to execute the following steps: extracting relationship information and feature keywords from a large-scale language model into which higher-order data and metadata associated with the higher-order data have been input; and outputting lower-order data and metadata associated with the lower-order data in comparison with the higher-order data, based on the higher-order data, the metadata associated with the higher-order data, the relationship information, and the feature keywords.
[0045] Some or all of the elements (e.g., configurations and functions) described in Supplementary Notes 2 to 9 that are dependent on Supplementary Note 1 (e.g., system) may also be dependent on Supplementary Note 10 (e.g., method) and Supplementary Note 11 (e.g., program) in the same dependency relationship as Supplementary Note 2 to Supplementary Note 9. Some or all of the elements described in any Supplementary Note may be applied to various hardware, software, recording means for recording software, systems, and methods.
[0046] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the invention.
[0047] This application claims priority based on Japanese Patent Application No. 2024-106173, filed July 1, 2024, the disclosure of which is incorporated herein in its entirety by reference.
[0048] 101 DATA, 102 INFORMATION, 103 KNOWLEDGE, 104 WISDOM, 105 Raw data, 106 Analyzed information, 107 Papers and works, 108 Social implementation, 200 Research data management infrastructure system, 201 Extraction unit, 202 Output unit, 203 LLM, 204 Relationship information, 205 Feature keywords, 400 Research data management infrastructure system, 402 Metadata, 407 Metadata, 408 Metadata, 410 Metadata, 411 Metadata, 412 Encoded feature information, 413 Feature information, 603 Extraction of data using Encoder, 604 Keyword extraction, 607 Extraction of data using Decoder, 608 Query, 609 Search results, 701 Query, 703 Example data search, 704 Selection of focus point, 706 Selection of focus point, 707 Feature distance search, 709 Search result, 800 Information processing device, 801 Memory, 802 Processor
Claims
1. A research data management infrastructure system comprising: an extraction means for extracting relationship information and characteristic keywords from a large-scale language model to which higher-level data and metadata associated with the higher-level data have been input; and an output means for outputting lower-level data and metadata associated with the lower-level data in comparison with the higher-level data based on the higher-level data, the metadata associated with the higher-level data, the relationship information, and the characteristic keywords.
2. A research data management infrastructure system as described in claim 1, further comprising a feature information output means for outputting feature information that encodes the low-level data and metadata associated with the low-level data.
3. A research data management infrastructure system as described in claim 2, wherein the feature information is encoded so that the coded information can be converted into a feature vector and feature distance can be calculated.
4. A research data management infrastructure system as described in claim 1, which issues a query to the data from which the relationship information and characteristic keywords have been extracted, and performs a search.
5. A research data management infrastructure system as described in claim 2, which issues a query to the characteristic keywords and the characteristic amount information, and performs a search by weighting points of interest.
6. A research data management infrastructure system as described in claim 1, wherein the higher-level data is a paper or a work of authorship.
7. A research data management infrastructure system as described in claim 1, wherein the metadata associated with the higher-level data is information on the registrant, the location of the event, or the time of the event.
8. A research data management infrastructure system as described in claim 1, wherein the lower-level data is raw data representing an event or analytical data created from raw data representing an event.
9. A research data management infrastructure system as described in claim 1, wherein the data accompanying the lower-level data is information on time, location, analysis conditions, or registrant.
10. A research data management method that extracts relationship information and characteristic keywords from a large-scale language model to which higher-level data and metadata associated with the higher-level data are input, and outputs lower-level data and metadata associated with the lower-level data in comparison with the higher-level data based on the higher-level data, the metadata associated with the higher-level data, the relationship information, and the characteristic keywords.
11. A research data management method according to claim 10, further comprising a feature information output means for outputting feature information obtained by encoding the low-level data and metadata associated with the low-level data.
12. A research data management method according to claim 11, wherein the feature information is encoded so that the coded information can be converted into a feature vector and feature distance can be calculated.
13. A research data management method according to claim 10, wherein a query is issued to the data from which the relationship information and the characteristic keywords have been extracted, and a search is performed.
14. A research data management method according to claim 11, wherein a query is issued to the characteristic keywords and the characteristic amount information, and a search is performed by weighting points of interest.
15. A research data management method as described in claim 10, wherein the higher-level data is a paper or a work of authorship.
16. A research data management method as described in claim 10, wherein the metadata associated with the higher-level data is information on the registrant, the location of the event, or the time of the event.
17. A research data management method according to claim 10, wherein the lower-level data is raw data representing an event or analytical data created from raw data representing an event.
18. A research data management method as described in claim 10, wherein the data accompanying the lower-level data is information on time, location, analysis conditions, or registrant.
19. A program that causes an information processing device to extract relationship information and feature keywords from a large-scale language model to which higher-order data and metadata associated with the higher-order data have been input, and output lower-order data and metadata associated with the lower-order data in comparison with the higher-order data based on the higher-order data, the metadata associated with the higher-order data, the relationship information, and the feature keywords.
20. The program according to claim 19, further comprising feature information output means for outputting feature information obtained by encoding the low-level data and metadata associated with the low-level data.
Citation Information
Patent Citations
Method and System for Associating Data with Figures
US20140351678A1