Information provision device, information provision method, and program
Patent Information
- Application Number
- JP2023522338
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-18
- Filing Date
- 2022-04-19
- Publication Date
- 2026-09-17
- Estimated Expiration
- 2042-04-19
Smart Images

Figure 0007923242000001 
Figure 0007923242000002 
Figure 0007923242000003
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an information-providing device that provides information about materials. [Background technology]
[0002] Conventionally, information-providing devices have been proposed that provide information on materials or experiments described in literature such as papers (see, for example, Patent Documents 1 to 3). The numerical search device proposed as an information-providing device in Patent Document 1 extracts numerical data on materials described in each of multiple documents and calculates the similarity between those numerical data. The support device proposed as an information-providing device in Patent Document 2 extracts information on synthesis processes described in papers and provides information showing the synthesis process from starting materials to target materials. The reliability evaluation system proposed as an information-providing device in Patent Document 3 evaluates the reliability of literature related to experiments. In other words, the reliability evaluation system searches for literature in response to the input of keywords such as experimental methods and calculates the reliability of the literature based on a journal usefulness list that maintains the usefulness of the journals to which the hit literature is submitted and the frequency of occurrence of the input keywords. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-80087 [Patent Document 2] International Publication No. 2021 / 039175 [Patent Document 3] Japanese Patent Publication No. 2008-152701 [Overview of the project]
[0004] However, the information-providing devices described in the above patent documents require further improvement to appropriately provide information about materials.
[0005] Therefore, this disclosure provides an information-providing device and the like that can be further improved to appropriately provide information about materials.
[0006] An information providing device according to one aspect of the present disclosure includes: an extraction unit that extracts identification information for identifying each of a plurality of materials and characteristic values of each of the plurality of materials from at least one document information; a derivation unit that derives a reliability level of the characteristic values of a material based on the similarity of the characteristic values between the material and one or more other materials; and an image processing unit that generates a first image that (i) displays the characteristic values of each of the plurality of materials in a display manner corresponding to the reliability level derived for the characteristic values of the material, and (ii) displays them in association with the identification information of the material, and outputs the first image to a display unit.
[0007] This comprehensive or specific embodiment may be implemented as a system, method, integrated circuit, or computer program, computer-readable recording medium, or as any combination of apparatus, system, method, integrated circuit, computer program, and computer-readable recording medium. The recording medium includes, for example, non-volatile recording media such as CD-ROM (Compact Disc-Read Only Memory).
[0008] The information provision device disclosed herein can be further improved to appropriately provide information about materials. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 is a block diagram showing an example configuration of the information provision system in Embodiment 1. [Figure 2A] Figure 2A is a diagram showing an example of a display screen displayed on the display unit by the image processing unit in Embodiment 1. [Figure 2B] Figure 2B shows another example of a display screen displayed on the display unit by the image processing unit in Embodiment 1. [Figure 3]Figure 3 is a diagram showing an example of metadata in the first embodiment. [Figure 4] Figure 4 is a diagram showing an example of text data in the first embodiment. [Figure 5] Figure 5 is a diagram showing an example of material names and characteristic values extracted from sentences of text data in the first embodiment. [Figure 6A] Figure 6A is a diagram showing an example of a figure to be extracted in text data in the first embodiment. [Figure 6B] Figure 6B is a diagram showing an example of characteristic values extracted from the figure to be extracted in the first embodiment. [Figure 7A] Figure 7A is a diagram showing an example of a figure targeted for extraction by an extraction unit in the first embodiment. [Figure 7B] Figure 7B is a diagram showing an example of a figure excluded from extraction targets by the extraction unit in the first embodiment. [Figure 8A] Figure 8A is a diagram showing an example of a table targeted for extraction in text data in the first embodiment. [Figure 8B] Figure 8B is a diagram showing an example of characteristic values extracted from the table to be extracted in the first embodiment. [Figure 9] Figure 9 is a diagram showing an example of normalization processing for material names in the first embodiment. [Figure 10] Figure 10 is an example of conversion processing for characteristic values of materials in the first embodiment. [Figure 11] Figure 11 is a diagram showing an example of an updated extraction information table in the first embodiment. [Figure 12] Figure 12 is a diagram showing an example of a display screen including an updated extraction information table in the first embodiment. [Figure 13] Figure 13 is a diagram showing an example of a graph display screen including a characteristic value graph in the first embodiment. [Figure 14] Figure 14 is a diagram showing an example of a graph display screen indicating reliability in the first embodiment. [Figure 15A]Figure 15A shows another example of the graph display screen showing the reliability in Embodiment 1. [Figure 15B] Figure 15B shows an example of the graph display screen after the bias has been changed in Embodiment 1. [Figure 15C] Figure 15C shows another example of the graph display screen after the bias has been changed in Embodiment 1. [Figure 16] Figure 16 is a flowchart showing an example of the overall processing operation of the information providing device in Embodiment 1. [Figure 17] Figure 17 is a flowchart showing an example of processing by the extraction unit in Embodiment 1. [Figure 18A] Figure 18A is a block diagram showing an example configuration of an information provision system in a modified version of Embodiment 1. [Figure 18B] Figure 18B is a block diagram showing another example of the configuration of the information provision system in a modified example of Embodiment 1. [Figure 19] Figure 19 shows an example of a data display screen in Embodiment 2. [Figure 20] Figure 20 shows another example of the data display screen in Embodiment 2. [Figure 21] Figure 21 is a diagram showing an example of a document display screen in Embodiment 2 where editing materials are displayed. [Figure 22] Figure 22 shows another example of the document display screen in Embodiment 2, where editing materials are displayed. [Figure 23] Figure 23 shows another example of the document display screen in Embodiment 2, where editing materials are displayed. [Figure 24] Figure 24 is a flowchart showing an example of the overall processing operation of the information providing device in Embodiment 2. [Modes for carrying out the invention]
[0010] (Knowledge leading to this disclosure) To synthesize new materials, it is necessary to find materials with appropriate properties from a vast amount of material data. However, discovering the optimal material from such a large amount of data requires a great deal of time and expense, and even experienced researchers find it difficult to discover new materials. For this reason, computer-based material search is being conducted. In other words, a large amount of information on material synthesis is extracted from literature such as papers and patents, the extracted information is stored in a database, and insights into material development are obtained from the stored information. For example, as shown in the above-mentioned Patent Documents 1 to 3, similarity calculations, provision of information indicating synthesis processes, and evaluation of the reliability of literature are performed.
[0011] However, while the information-providing devices described in Patent Documents 1 to 3 above provide various types of information, there is a problem in that it is difficult to determine whether the material characteristic values described in the documents are reliable or not. For example, even if the champion data shows the highest value for a material characteristic, its reliability (i.e., reproducibility) is often low.
[0012] Therefore, the present inventors have discovered that it is possible to extract the characteristic values of multiple materials from literature, derive the reliability of those characteristic values based on the similarity of those extracted characteristic values, and display the multiple characteristic values in a display manner corresponding to that reliability, which led to the present disclosure.
[0013] In other words, an information providing device according to one aspect of the present disclosure includes: an extraction unit that extracts identification information for identifying each of a plurality of materials and characteristic values for each of the plurality of materials from at least one document information; a derivation unit that derives a reliability level for the characteristic values of a material based on the similarity of the characteristic values between the material and one or more other materials; and an image processing unit that generates a first image which displays the characteristic values of each of the plurality of materials in a display manner corresponding to the reliability level derived for the characteristic values of the material, and (ii) in association with the identification information of the material, and outputs the first image to a display unit. For example, the identification information is the name of the material, i.e., the material name, and the material name may be a composition formula. The plurality of materials may be the same material, different materials, or materials that have a common use or elemental species. In a specific example, the plurality of materials are materials used for the positive or negative electrode of a battery. For example, the lower the similarity, the lower the reliability level derived, and the higher the similarity, the higher the reliability level derived.
[0014] As a result, the first image displayed shows the property values of multiple materials in a display format corresponding to the reliability of those property values, and the material identification information is associated with each property value. Therefore, by looking at the first image, users such as materials researchers can easily grasp the reliability of the material property values described in, for example, the vast amount of literature information stored in the database. For example, if the property values of a material are similar to those of many other materials, the user can easily understand that the property values of that material are reliable.
[0015] The extraction unit may extract multiple type characteristic values as the characteristic values. For example, the multiple type characteristic values may be the electrical conductivity and activation energy value of the material.
[0016] This makes it easy to understand the reliability of multiple type characteristics of a material.
[0017] The plurality of type characteristic values include a first type characteristic value and a second type characteristic value. The image processing unit sets a characteristic map in the first image, which is composed of a first coordinate axis for showing the first type characteristic value and a second coordinate axis for showing the second type characteristic value. Marks corresponding to each of the plurality of materials may be superimposed on the characteristic map at positions corresponding to the first type characteristic value and the second type characteristic value of that material, in a display manner corresponding to the reliability derived for the characteristic value of that material. For example, the characteristic map is a characteristic value graph that shows electrical conductivity and activation energy value on the first and second coordinate axes, respectively. Marks corresponding to each material are then placed on the characteristic value graph at positions corresponding to the electrical conductivity and activation energy value of that material, in a display manner corresponding to the reliability of those type characteristic values.
[0018] This makes it easier to understand the overall reliability of the multiple type characteristics of each material.
[0019] The derivation unit may determine the similarity of the characteristic values in the material based on the distance between the mark corresponding to the material and the marks corresponding to each of the other one or more materials. For example, the shorter the average value of these distances, the higher the similarity determined. In other words, the higher the density of the marks, the higher the similarity determined to the characteristic value represented by that mark, and conversely, the lower the density of the marks, the lower the similarity determined to the characteristic value represented by that mark.
[0020] This allows for the appropriate identification of a single similarity for multiple type-specific characteristic values of materials, and as a result, a confidence score based on similarity can be appropriately derived.
[0021] The image processing unit may determine the intensity of a color that is darker the higher the reliability derived from the characteristic value of the material, as the display mode of the characteristic value, and generate the first image that shows the characteristic value with the determined intensity of the color.
[0022] This allows users to easily and visually grasp the reliability of the characteristic values. In the first image, which shows the characteristic values with the determined color intensity, the numerical value of the characteristic value itself may be displayed with the determined color intensity, or the mark indicating the characteristic value may be displayed with the determined color intensity.
[0023] The extraction unit may further extract attribute information indicating the attributes of each of the plurality of materials from at least one document information, and the derivation unit may derive the reliability of the characteristic values in the material based on the similarity of the characteristic values in the material and the attribute information extracted for the material. For example, the derivation unit may derive the reliability of the characteristic values in the material by weighted addition of the similarity of the characteristic values in the material and the attribute values based on the attribute information extracted for the material.
[0024] This allows for the deriving of the reliability of a characteristic value based on the similarity of the characteristic value and the attributes of the material, enabling a multifaceted deriving of reliability and improving the accuracy of that reliability. By adjusting the weights of similarity and attributes as biases, it is possible to derive a reliability that suits the user's purpose.
[0025] The attribute information indicates the publication date of the document containing the identification information and characteristic values of the material corresponding to the attribute information, among the at least one document of the document, and the derivation unit may derive the reliability of the characteristic values using the attribute value indicating the recency of the publication date. For example, the attribute value indicates a larger value the more recent the publication date, and the derivation unit may derive a larger value as the reliability of the characteristic values in the material the higher the similarity of the characteristic values in the material and the larger the attribute value corresponding to the material. For example, the publication date is the publication date.
[0026] This allows for the deriving of relatively high confidence levels for characteristic values described in new literature information. Therefore, it is possible to derive appropriate confidence levels for users who prioritize publication dates.
[0027] The attribute information indicates the number of citations of the document containing the identification information and characteristic values of the material corresponding to the attribute information, among the at least one document, as the attribute, and the derivation unit may derive the reliability of the characteristic value using the attribute value corresponding to the number of citations. For example, the attribute value indicates a larger value the more the number of citations, and the derivation unit may derive a larger value as the reliability of the characteristic value in the material the higher the similarity of the characteristic value in the material and the larger the attribute value corresponding to the material.
[0028] This allows us to derive a relatively high level of confidence in characteristic values described in bibliographic information with a high number of citations. Therefore, we can derive an appropriate level of confidence for users who prioritize the number of citations.
[0029] The attribute information indicates, as the attribute, the author of the document containing the identification information and characteristic values of the material corresponding to the attribute information, among the at least one document of document information. The derivation unit may derive the reliability of the characteristic values using the attribute values, depending on whether the author of the document information is the same as the author of one or more other documents. For example, the attribute value may be larger the more authors among the one or more other documents are different from the author of the document information. The derivation unit may derive a larger reliability value for the characteristic values of the material the higher the similarity of the characteristic values in the material and the larger the attribute value corresponding to the material.
[0030] This allows for a relatively high level of confidence to be derived for characteristic values described in bibliographic information written by authors different from many other sources. In other words, if characteristic values described in many bibliographic sources written by different authors are similar, a high level of confidence can be derived for those characteristic values. On the other hand, even if many bibliographic sources describe similar characteristic values, if those sources were written by the same author, a low level of confidence can be derived for those characteristic values. Therefore, an appropriate level of confidence can be derived for users who place importance on the authorship of the bibliographic information.
[0031] The attribute information indicates the synthesis method of the material corresponding to the attribute information as the attribute, and the derivation unit may derive the reliability of the characteristic value using the attribute value corresponding to the degree of similarity between the synthesis method of the material and the synthesis methods of one or more other materials. For example, the attribute value indicates a larger value the greater the degree of similarity between the synthesis method of the material and the synthesis methods of one or more other materials, and the derivation unit may derive a larger value as the reliability of the characteristic value in the material the higher the similarity of the characteristic value in the material and the larger the attribute value corresponding to the material. For example, the synthesis method of the material may include at least one of the temperature conditions, time conditions, and type of apparatus used in the synthesis of the material.
[0032] Therefore, if the literature describes that the material for that characteristic value is synthesized using the same synthesis method as other literature, a relatively high level of confidence can be derived for that characteristic value. Thus, an appropriate level of confidence can be derived for users who focus on the synthesis method.
[0033] The extraction unit further acquires material conditions and extracts information about each of the one or more materials that meet the material conditions from the at least one literature information as a candidate for display information. The image processing unit further acquires the weights of each of the multiple types of attributes of the material and selects one or more of the candidate display information from the multiple candidates for display information extracted by the extraction unit, each corresponding to a material having a characteristic value for which the confidence level is derived to be above a threshold, as display information. The image processing unit generates a second image in which the display information corresponding to each of the multiple types of attributes is shown in an amount corresponding to the weight of each of the multiple types of attributes, and outputs the second image to the display unit.
[0034] This allows the amount of display information corresponding to each of the multiple types of attributes shown in the second image to be changed by adjusting the bias, which is the weight of each of the multiple types of attributes. Therefore, the user can arbitrarily adjust the amount of display information for each attribute so that more information is displayed for attributes that the user is interested in, and less information is displayed for attributes that the user is not interested in. Furthermore, since this displayed information is information about materials with a confidence level above a threshold, the user can use this displayed information with confidence in tasks such as materials research. Display information for materials that correspond to material conditions is displayed, and these material conditions are, for example, conditions related to the element species contained in the material or the composition of the material. This allows the user to limit the one or more displayed pieces of information to materials that they are interested in.
[0035] The embodiments of this disclosure will be described below with reference to the drawings. The embodiments described below are all specific examples of this disclosure. Therefore, the numerical values, shapes, materials, components, arrangement positions of components, and connection configurations shown in the following embodiments are examples and are not intended to limit this disclosure. Accordingly, among the components in the following embodiments, those components that are not described in the independent claim representing the highest-level concept will be described as optional components.
[0036] Note that each figure is a schematic diagram and not necessarily a strictly accurate representation. In each figure, substantially identical components are denoted by the same reference numerals, and redundant explanations are omitted or simplified.
[0037] (Embodiment 1) [Device configuration] Figure 1 is a block diagram showing an example configuration of the information provision system in Embodiment 1. The information provision system 1000 in this embodiment is a system that provides information about materials and, as shown in Figure 1, comprises an information provision device 100, an input unit 11, a literature database (DB) 12, and a display unit 13.
[0038] The input unit 11 receives user input operations and outputs an input signal corresponding to those operations to the information providing device 100. The input unit 11 can be configured as, for example, a keyboard, touch sensor, touchpad, or mouse. Using this input unit 11 enables more intuitive input operations.
[0039] The display unit 13 acquires an image signal from the information providing device 100 and displays an image corresponding to that image signal. The display unit 13 is, for example, a liquid crystal display, a plasma display, or an organic EL (Electro-Luminescence) display, but is not limited to these.
[0040] The literature database 12 is a recording medium such as a hard disk, and stores multiple pieces of literature information D. Each of the multiple pieces of literature information D is electronic data of a literature. The literature is, for example, a paper on materials, material synthesis, experiments on material synthesis, etc. These pieces of literature information D include text data D1 and metadata D2. Text data D1 is, for example, an electronic paper in PDF or XML format published as an online journal (also called an electronic journal). Metadata D2 is data related to figures, tables, images, authors, research institutions, publication date, title, etc., shown in the text data D1. Such metadata D2 is attached to the text data D1. In this embodiment, metadata D2 is attached to the text data D1, but it does not have to be attached to the text data D1. If metadata D2 is not attached to the main text data D1, the information providing device 100 may download the metadata D2 provided by the publisher of the document corresponding to the main text data D1 from a server or the like and attach it to the main text data D1.
[0041] Such a literature database 12 may be connected to the information providing device 100 via a communication network such as the Internet, or it may be directly connected to the information providing device 100 without going through a communication network. The literature database 12 may be a recording medium other than a hard disk, such as RAM (Random Access Memory), ROM (Read Only Memory), or semiconductor memory. The literature database 12 may be volatile or non-volatile.
[0042] The information providing device 100 provides information about materials based on an input signal output from the input unit 11 and on multiple pieces of literature information D stored in the literature database 12. In other words, the information providing device 100 displays information about materials on the display unit 13. Specifically, the information providing device 100 acquires an input signal from the input unit 11 and searches the literature database 12 for at least one piece of literature information D corresponding to that input signal. Then, the information providing device 100 extracts multiple pieces of information from the at least one piece of literature information D that was retrieved, generates an image based on these multiple pieces of information, and outputs an image signal showing that image to the display unit 13. The information providing device 100 may consist of a processor such as a central processing unit (CPU) and memory. In this case, the processor functions as the information providing device 100 by executing a computer program stored in memory, for example.
[0043] Specifically, as shown in Figure 1, the information providing device 100 includes an extraction unit 101, a first information processing unit 102, a second information processing unit 103, a third information processing unit 104, an output unit 105, and an image processing unit 106.
[0044] The extraction unit 101 acquires the input signal output from the input unit 11 and searches the literature database 12 for at least one literature information D corresponding to that input signal. Furthermore, the extraction unit 101 extracts from these literature information D the first information indicating the name of the material, the second information indicating the characteristic values of the material, and the third information indicating the attributes of the material. The name of the material is also hereafter referred to as the material name and is identification information for identifying the material. If the literature information D describes material synthesis, the material corresponding to the first information, the second information, and the third information is the material that is ultimately produced by that material synthesis and is also called the final material or target material. The extraction unit 101 then outputs the first information to the first information processing unit 102, the second information to the second information processing unit 103, and the third information to the third information processing unit 104.
[0045] The first information processing unit 102 modifies the material name indicated by the first information using a predetermined method, and outputs the modified first information to the derivation unit 105.
[0046] The second information processing unit 103 modifies the material characteristic values indicated by the second information using a predetermined method, and outputs the processed second information to the derivation unit 105.
[0047] The third information processing unit 104 modifies the material characteristic values indicated by the third information using a predetermined method, and outputs the processed third information to the derivation unit 105.
[0048] The derivation unit 105 acquires the processed first information, second information, and third information from the first information processing unit 102, the second information processing unit 103, and the third information processing unit 104. The derivation unit 105 then integrates the processed first information, second information, and third information, and further derives the reliability of the characteristic values of each extracted material based on that information. The derivation unit 105 outputs an output signal containing these derived reliability values to the image processing unit 106.
[0049] The image processing unit 106 acquires an output signal from the output unit 105 and generates a first image by performing image processing according to that output signal. The image processing unit 106 then outputs an image signal representing the first image to the display unit 13.
[0050] In this embodiment, the first information processing unit 102, the second information processing unit 103, and the third information processing unit 104 do not have to modify all of the first information, all of the second information, and all of the third information extracted from the literature database 12. In other words, the first information processing unit 102, the second information processing unit 103, and the third information processing unit 104 may modify some of that information as needed. In this embodiment, the information providing device 100 includes the first information processing unit 102, the second information processing unit 103, and the third information processing unit 104, but it does not have to include at least one of them.
[0051] The details of each component shown in Figure 1 will be explained below.
[0052] [Information extraction from metadata] Figures 2A and 2B show examples of display screens displayed on the display unit 13 by the image processing unit 106.
[0053] As shown in Figure 2A, the image processing unit 106 displays the display screen 20 on the display unit 13. The display screen 20 includes a bibliography window 21, an extraction information table 22, and an extraction start button 23a.
[0054] The bibliography window 21 displays a list of bibliographic information D stored in the bibliographic database 12. Specifically, it displays the icon and bibliographic ID for each bibliographic information D stored in the bibliographic database 12. The bibliographic ID is the identification information used to identify bibliographic information D.
[0055] Extracted Information Table 22 is a table for showing the information extracted from the literature information D. This information includes, for example, the literature ID, publication date, journal name, title, number of citations, final material name, electrical conductivity, and activation energy value. The publication date is the date the paper corresponding to the literature information D was published, and the journal name is the name of the journal in which the paper was published. The title is the title of the paper, and the number of citations is the number of times the paper has been cited in other papers. The final material name is the name of the final material that is ultimately produced by the material synthesis described in the paper. Electrical conductivity is the electrical conductivity or electrical conductivity that represents the ease of electrical conduction in the final material, and will hereafter be simply called "conductivity". The activation energy value is a value that indicates the magnitude of the activation energy of the final material.
[0056] Here, the initial extracted information table 22 does not show the information extracted from those bibliographic information D, but rather the type names of that information are shown. In the examples in Figures 2A and 2B, the type names of that information are shown as "Journal," "Final Material," and "Activation," where "Journal" refers to the journal name, "Final Material" refers to the final material name, and "Activation" refers to the activation energy value.
[0057] The extraction start button 23a is a button used to initiate the extraction of each of the above-mentioned pieces of information from multiple pieces of bibliographic information D stored in the bibliographic database 12.
[0058] The user selects one or more desired bibliographic information icons from all the bibliographic information D displayed in the bibliographic list window 21 by performing an input operation on the input unit 11, and then selects the extraction start button 23a. The input unit 11 outputs an input signal to the extraction unit 101 corresponding to this input operation. In other words, the user selects one or more desired bibliographic information D and instructs the information providing device 100 to extract various types of information from those bibliographic information D.
[0059] When the extraction unit 101 receives the above-mentioned input signal from the input unit 11, it extracts four pieces of bibliographic information from the metadata D2 of one or more selected bibliographic information D, for example, the publication date, journal name, title, and number of citations. The extraction unit 101 then outputs the bibliographic ID of the selected bibliographic information D and the four pieces of bibliographic information extracted from the metadata D2 of the bibliographic information D to the image processing unit 106. The publication date, journal name, title, and number of citations can also be said to be attributes of the material. In other words, the four pieces of bibliographic information can each be said to be third-party information.
[0060] When the image processing unit 106 obtains the aforementioned document ID and four bibliographic information from the extraction unit 101, it updates the display screen 20 displayed on the display unit 13, as shown in Figure 2B. In other words, the image processing unit 106 associates the document ID with the publication date, journal name, title, and number of citations indicated by the four bibliographic information and writes them to the extracted information table 22.
[0061] The extracted information table 22 included in the updated display screen 20 shows the publication date, journal name, title, etc., of each bibliographic information D. Therefore, users can narrow down their analysis to a specific materials field or perform analyses of materials by time period.
[0062] As shown in Figure 2B, the image processing unit 106 includes the replenishment start button 23b in the updated display screen 20. The replenishment start button 23b is a button that initiates the extraction of three pieces of information from the text data D1 of multiple literature information D stored in the literature database 12, indicating the final material name, conductivity, and activation energy value.
[0063] Figure 3 shows an example of metadata.
[0064] Metadata D2 is structured data, such as a bib file. In other words, metadata D2 is a file with the extension ".bib" and is also called a BIB file. Such metadata D2 contains information such as the title of the paper, the names of the authors, the names of the organizations to which the authors belong (affiliations), and the publication date of the paper, as shown in Figure 3. Metadata D2 may also show the number of citations. The extraction unit 101 extracts the four bibliographic information items mentioned above from such metadata D2. Note that the names of the authors are also called author names, and the organizations to which the authors belong are also called research institutions.
[0065] [Information extraction from the main text data] When the user selects the replenishment start button 23b through an input operation to the input unit 11, the extraction unit 101 receives an input signal corresponding to that input operation from the input unit 11 and starts extracting information from the text data D1 of one or more previously selected bibliographic information D.
[0066] Figure 4 shows an example of text data D1.
[0067] As shown in Figure 4, the main text data D1 includes sentences written in natural language, tables, and figures such as graphs. The extraction unit 101 treats these sentences, tables, and figures as extraction targets and extracts the first, second, and third information described above from these extraction targets. In other words, the extraction unit 101 obtains the first, second, and third information described above extracted from the extraction targets. If the main text data D1 contains text data, the extraction unit 101 extracts the first, second, and third information from that text data. In other words, the extraction unit 101 obtains the first, second, and third information extracted from that text data. If the main text data D1 does not contain text data, the extraction unit 101 may convert the sentences that are represented as images in the main text data D1 into text data and extract the first, second, and third information from that text data. In other words, the extraction unit 101 may obtain the first information, second information, and third information extracted from the converted text data.
[0068] When the extraction unit 101 extracts the first, second, and third pieces of information from the text data, it may use natural language processing tools or deep learning tools. Examples of natural language processing tools include CoreNLP or MeCab. Examples of deep learning tools include word2vec, BERT, Tensorflow, or PyTorch. This makes it possible to extract each piece of information with high accuracy. In addition, multiple types of tools may be used in combination to extract a single piece of information. This makes it possible to extract each piece of information with even higher accuracy.
[0069] The names of materials are often written using combinations of element symbols and numbers. Therefore, the extraction unit 101 may use a dictionary in which patterns of element symbol-number combinations are registered, and by comparing the text data with these patterns, extract the words that match the patterns as material names. Alternatively, the extraction unit 101 may represent the combinations of element symbols and numbers contained in the text data using regular expressions and compare these regular expressions with the patterns. Then, the extraction unit 101 may extract the combinations of element symbols and numbers corresponding to the regular expressions that match the patterns as material names. A regular expression is an expression that follows predetermined rules. For example, if the text data shows "The conductivity and activation energy for Li6.25Al0.25La3Zr2O12 with an ion dose of 2.7 x 10⁻¹⁴ cm⁻² are 4.6 x 10⁻³ S cm⁻¹ and 0.11 eV, respectively," the extraction unit 101 describes the combination of element symbols and numbers using a regular expression and extracts the material name "Li6.25Al0.25La3Zr2O12" by matching the regular expression with a pattern. This allows the user to appropriately extract the information they want. In this disclosure, the number to the right of the element symbol indicates the composition ratio or number of atoms of that element, even if the character is not a subscript.
[0070] The extraction unit 101 extracts, for example, "Li6.25Al0.25La3Zr2O12" as the material name from the sentence to be extracted, and then calculates "4.6 × 10 -3 Scm -1 " is extracted as the material's characteristic value, i.e., conductivity, and "0.11eV" is extracted as the material's characteristic value, i.e., activation energy value. In this disclosure, "10" or the number to the right of the unit indicates a multiplier even if it is not a superscript.
[0071] The extraction unit 101 extracts, for example, "3.70 × 10" from the table to be extracted. -4 Scm -1 "1.49 × 10 -4 Scm -1These are extracted as material characteristic values, i.e., conductivity. At this time, the extraction unit 101 extracts "conductivity" or "Scm -1 Since the table contains keywords related to conductivity, such as "[...]", we extract the conductivity from that table.
[0072] The extraction unit 101 extracts, for example, conductivity, which is a characteristic value of the material, from the graph to be extracted. At this time, the extraction unit 101 labels or captions of the graph axes with "conductivity" or "Scm -1 Since the table contains keywords related to conductivity, such as "[...]", we extract the conductivity from the graph.
[0073] Figure 5 shows an example of material names and characteristic values extracted from the text data D1.
[0074] The extraction unit 101 extracts material names from each sentence of the main text data D1. At this time, the extraction unit 101 records the document ID, the extraction line number, and the span in the extraction list 31, associating them with the extracted material name, as shown in Figure 5, for example. The document ID is the identification information of the document information D, which includes the main text data D1 from which the extracted material names were extracted. The extraction line number is the line number in the main text data D1 in which the material name is described. The span indicates the start point where the description of the material name begins and the end point where the description ends in the line of the extraction line number. The start point is indicated by the number of characters from the first character of the line to the first character of the material name, and the end point is indicated by the number of characters from the first character of the line to the last character of the material name.
[0075] If the material name spans multiple lines, the numbers of those multiple lines may be recorded as the extracted line number. The extracted line number may also be the number of the sentence in the main text data D1 in which the material name is described. The sentence number is, for example, a number assigned to all sentences in the main text data D1 to identify that sentence. The extraction unit 101 may record the extracted line number and the page number in the main text data D1 in which the material name is described in the extraction list 31.
[0076] Here, when the extraction unit 101 extracts material names, for example, it may extract a combination of element symbols and numbers as the material name, and if there are parentheses next to the element symbols and numbers, it may also extract the parentheses and the string inside the parentheses as part of the material name. For example, the extraction unit 101 extracts "Li6.25Al0.25La3Zr2O12(LALZ)" as the material name. The extraction unit 101 may recognize the string inside the parentheses as an abbreviation of the material name. For example, the extraction unit 101 recognizes "Li6.25Al0.25La3Zr2O12" in "Li6.25Al0.25La3Zr2O12(LALZ)" as the standard representation of the material name, and recognizes "LALZ" as an abbreviation of the material name "Li6.25Al0.25La3Zr2O12". As a result, the extraction unit 101 also extracts "LALZ" as a material name from the text data D1.
[0077] Furthermore, if the extraction unit 101 determines that a variable is included in the string within the parentheses, it will not recognize the string within the parentheses as an abbreviation of the material name. For example, if the string within the parentheses contains the variable x, such as "x=0.1,0.2,0.3", the extraction unit 101 will not recognize the string within the parentheses as an abbreviation of the material name. The material name extracted by the extraction unit 101 may include the mixing ratio of multiple components. For example, the material name "60Li2SO4*40Li3BO3" may include the mixing ratio "60:40" of the components "Li2SO4" and "Li3BO3". Such variables and mixing ratios may be modified by the processing described later by the first information processing unit 102.
[0078] As described above, the material names extracted by the extraction unit 101 are the names of the target material or the final material. For example, if the text data D1 contains multiple material names, the extraction unit 101 may use a natural language processing tool or a deep learning tool to extract the name of the final material from those multiple material names. For example, if the text data D1 contains the sentence "XXX was synthesized using...", the extraction unit 101 detects "synthesized" in that sentence as a keyword. The extraction unit 101 then determines that "XXX", which is the target of that keyword, is the name of the final material (i.e., the name of the final material), and extracts that name "XXX".
[0079] The extraction unit 101 extracts material names and characteristic values from each sentence of the main text data D1. For example, the characteristic value is conductivity. At this time, the extraction unit 101 records the document ID, the extraction line number, and the span in the extraction list 32, associating them with the extracted characteristic value, as shown in Figure 5, for example. The document ID is the identification information of the document information D, which includes the main text data D1 from which the extracted characteristic value was extracted. The extraction line number is the number of the line in the main text data D1 in which the characteristic value is described. The extraction line number may also be the sentence number, as described above. The span indicates the start point where the description of the material name begins and the end point where the description ends in the line of the extraction line number. The start point is indicated by the number of characters from the first character of the line to the first character of the characteristic value, and the end point is indicated by the number of characters from the first character of the line to the last character of the characteristic value.
[0080] In a specific example, the sentence in text data D1 is "The conductivity and activation energy for Li6.25Al0.25La3Zr2O12 with an ion dose of 2.7 x 10⁻¹⁴ cm⁻² are 4.6 x 10⁻³ S cm⁻¹ and 0.11 eV, respectively." In this case, the extraction unit 101 describes the combination of numbers and units using a regular expression, and by matching that regular expression with a pattern, it recognizes and extracts "4.6 × 10⁻³ S cm⁻¹" as a characteristic value from the sentence.
[0081] Here, the main text data D1 of the paper includes a large variety and a large amount of sentences that do not contain material names and material characteristic values. Examples include sentences containing references, acknowledgments, and the like. That is, when all sentences included in the main text data D1 are extraction targets, there is a possibility that a large amount of noise other than material names and characteristic values will be extracted. In other words, extraction errors are prone to occur. Therefore, the extraction unit 101 in the present embodiment may treat sentences containing words related to materials as extraction targets.
[0082] For example, a sentence containing a material name or a characteristic value may include keywords related thereto. That is, if a sentence contains a keyword, it is highly likely that the sentence contains a material name or a characteristic value. For example, in the sentences "XXX was synthesized using …" and "The conductivity and activation energy for XXX are YYY S cm -1 and ZZZ eV, respectively.", words such as "synthesized", "conductivity", "activation energy", "Scm -1 " are keywords. The extraction unit 101 may treat sentences containing such keywords as extraction targets, and extract material names or characteristic values from the extraction target sentences using pattern matching based on regular expressions or the like. This makes it possible to reliably extract the material name or characteristic value desired by a user.
[0083] When the extraction unit 101 extracts material names and characteristic values, for example, for each processing unit such as a sentence or clause, if one of the material names and characteristic values is extracted from that processing unit, it attempts to extract the other from that processing unit as well. If the extraction unit 101 has extracted one of the material names and characteristic values from the processing unit but cannot extract the other, it may store that processing unit in the database. Such processing units may be used as training data for machine learning. The extraction unit 101 may prompt the user to extract the material name, the characteristic value, or both the material name and characteristic value from that processing unit. In other words, the extraction unit 101 may display an error message on the display unit 13 via the image processing unit 106 to prompt the user to extract the material name, the characteristic value, or both the material name and characteristic value.
[0084] Figure 6A shows an example of a figure to be extracted from the text data D1, and Figure 6B shows an example of characteristic values extracted from the extracted figure.
[0085] The extraction unit 101 extracts material names and characteristic values from graphs such as line graphs and scatter plots in the main text data D1. For example, as shown in Figure 6A, the extraction unit 101 extracts the conductivity and temperature of materials from graphs contained in the main text data D1 of document ID "0001" and document ID "0002". For example, conductivity is a characteristic value. Temperature is one of the conditions used in the synthesis method for synthesizing the material and is an attribute of the material. The extraction unit 101 uses tools such as image processing, image recognition, and deep learning for these extractions.
[0086] In Figure 6A, the vertical axis of the graph represents conductivity, and the horizontal axis represents temperature. The horizontal axis is also called the x-axis, and the vertical axis is also called the y-axis. The graph shows two line graphs: one for the bulk state of the material and another for the total state of the material. Bulk refers to the material in its pure form, while Total refers to the material incorporated into a product. The x-axis is marked at 1000°C, 1200°C, 1400°C, 1600°C, 1800°C, and 2000°C.
[0087] The extraction unit 101 reads the conductivity corresponding to the temperature indicated by each scale on the x-axis of the graph, as shown in Figure 6A, from the Bulk line and the Total line on the graph, respectively. As a result, the conductivity corresponding to the temperatures of 1000°C, 1200°C, 1400°C, 1600°C, 1800°C, and 2000°C is extracted for the Bulk material and the Total material, respectively.
[0088] For example, the extraction unit 101 reads "0.0070 S / cm" as the conductivity on the y-axis, represented by the Bulk line graph, when the temperature on the x-axis is 1000°C in the graph of the main text data D1 of document ID "0001". Note that the temperature on the x-axis is also called the x-value, and the conductivity on the y-axis is also called the y-value.
[0089] Next, as shown in Figure 6B, the extraction unit 101 records the aforementioned document ID, the figure number of the graph included in the text data D1 of that document ID, the x-value which is the temperature on which the x-axis scale is marked, and the y-value which is the conductivity read for that x-value, in the extraction list 33. Furthermore, the extraction unit 101 records the x-axis label and y-axis label of the graph, the material state corresponding to the read conductivity (i.e., Blul or Total), and the graph's caption in the extraction list 33. Since the caption often contains the material name, the caption can be used as a clue when associating the x-value and y-value with the material name.
[0090] Furthermore, as shown in the example in Figure 6A, the extraction unit 101 distinguishes between multiple line graphs and extracts their conductivity and temperature. In the example in Figure 6A, the extraction unit 101 extracts the temperature and the corresponding conductivity for each scale on the x-axis of the graph, but it is not necessary to extract them for each scale. For example, the extraction unit 101 may extract the temperature and the corresponding conductivity for each temperature indicated by an arbitrary temperature interval specified by the user. Specifically, in the example shown in Figure 6A, the temperature and the corresponding conductivity are extracted for each of the temperatures of 1000°C, 1200°C, 1400°C, 1600°C, 1800°C, and 2000°C. In this case, the temperature interval is 200°C. However, the extraction unit 101 may set the temperature interval to 100°C in response to the user's input operation to the input unit 11. In this case, the extraction unit 101 extracts the temperature and the corresponding conductivity for each temperature range indicated by 100°C temperature intervals from 1000°C to 2000°C. That is, the extraction unit 101 extracts the temperatures of 1000°C, 1100°C, 1200°C, 1300°C, 1400°C, 1500°C, 1600°C, 1700°C, 1800°C, 1900°C, and 2000°C, along with the corresponding conductivity. This ensures that the user can reliably extract the conductivity and temperature information they desire. When the extraction unit 101 extracts temperatures other than those marked on a scale, it uses, for example, an image recognition tool to identify the length between two scales and the temperatures at those scales. Next, the extraction unit 101 interpolates the temperature of the midpoint between those scales using the identified length and temperature. This allows for the accurate extraction of temperatures other than those marked on the scale, along with the corresponding conductivity values.
[0091] The extraction list 33 shown in Figure 6B may also record the line number or sentence number of the sentence adjacent to the extracted figure number. This makes it easy to find sentences related to a figure by referring to the extraction list 33.
[0092] Figure 7A shows an example of a figure to be extracted by the extraction unit 101, and Figure 7B shows an example of a figure to be excluded from extraction by the extraction unit 101.
[0093] The main text data D1 contains numerous figures unrelated to material properties, in addition to those relating to property values. These unrelated figures include, for example, diagrams explaining overviews and experimental procedures, or photographs of experimental equipment and materials. Therefore, if all figures included in the main text data D1 are included in the extraction process, errors are likely to occur when extracting property values.
[0094] Therefore, the extraction unit 101 narrows down the figures to be extracted from all the figures included in the main text data D1, as shown in Figure 7A, and excludes the figures shown in Figure 7B from the extraction target. Figures excluded from the extraction target include, for example, figures showing reaction processes and figures showing material structures, as shown in Figure 7B. For example, the extraction unit 101 narrows down the figures to be extracted by using image recognition tools, natural language processing tools, deep learning tools, etc. Image recognition tools include, for example, OpenCV and Pillow. This enables highly accurate narrowing down of the figures.
[0095] Alternatively, the extraction unit 101 may narrow down the figures to be extracted by using regular expression pattern matching. The extraction unit 101 extracts words, units, or strings from the figure's captions, labels, etc., and generates a regular expression from the extracted words, etc. Then, the extraction unit 101 compares the regular expression with a pattern, and treats figures containing the regular expression that matches the pattern as figures to be extracted. The labels mentioned above are, for example, labels attached to the axes of a graph. This allows for highly accurate filtering.
[0096] Specifically, the extraction unit 101 extracts each word from the y-axis label "Conductivity (S / cm)" or the caption "Fig 1. Temperature dependent electrical conductivity of Li6.25Al0.25La3Zr2O12 samples" in the graph shown in Figure 6A, represents these words using regular expressions, and matches these regular expressions against a pattern. For example, the regular expressions for "Conductivity" or "S / cm" in the y-axis label match the pattern, and the regular expression for "conductivity" in the caption also matches the pattern. As a result, the extraction unit 101 treats the graph shown in Figure 6A as the target for extraction. This suppresses the process of attempting to forcibly extract characteristic values from a graph that is not related to the characteristic values.
[0097] Figure 8A shows an example of a table targeted for extraction within the text data D1, and Figure 8B shows an example of characteristic values extracted from the targeted table.
[0098] The extraction unit 101 extracts characteristic values from, for example, Table 1 and Table 2 shown in Figure 8A of the main text data D1. For extracting characteristic values, the extraction unit 101 uses image recognition tools, deep learning tools, or pattern matching using regular expressions, as described above.
[0099] For example, in a table where characteristic values are listed, the column name cells could be "conductivity", "eV", "S / cm", "Scm" -1Keywords related to characteristic values such as "..." are often described. Therefore, the extraction unit 101 first detects these keywords from the table using natural language processing tools, deep learning tools, or regular expression pattern matching. The keywords may be pre-registered in the extraction unit 101. Then, the extraction unit 101 extracts characteristic values from the table in which the keywords were detected. For example, the extraction unit 101 detects keywords from the column name cells of the table, detects material names from the row name cells of the table, and extracts the value stored in the cell where the column of the column name cell and the row of the row name cell intersect as the characteristic value of that material name. This ensures that characteristic values related to pre-registered keywords can be reliably extracted.
[0100] Specifically, in Table 2 of document ID "0001" in Figure 8A, one column name cell contains "Conductivity" and "Scm -1 The column name cell contains the text "Li6.25Al0.25La3Zr2O12(ours)". Therefore, the extraction unit 101 detects the characteristic value keyword from the column name cell, detects the material name "Li6.25Al0.25La3Zr2O12(ours)" from the row name cell, and extracts the value "2.45" stored in the cell where the column of the column name cell and the row of the row name cell intersect as the characteristic value of that material name. The extraction unit 101 then extracts the value "2.45" and the unit "Scm". -1 The string obtained by concatenating " and the other can be extracted as a characteristic value.
[0101] Next, as shown in Figure 8B, the extraction unit 101 records the detected material name, detected keywords, and extracted characteristic values in the extraction list 34. At this time, the extraction unit 101 records the material name in the "row label" column of the extraction list 34, for example, and the keywords in the "column label" column of the extraction list 34. Furthermore, the extraction unit 101 records the figure number of the table from which the characteristic values were extracted, and the document ID of the text data D1 containing that table, in the extraction list 34.
[0102] The extraction unit 101 may further extract values other than conductivity, which is a characteristic value, if they are listed in the table. These values may be, for example, density. If structures such as "Cubic" or "Hexagonal" are shown, as in Table 1 of Figure 8A, the extraction unit 101 may extract information indicating that structure. The density and structure information extracted in this way are recorded in the extraction list 34 along with conductivity. Note that the density and structure information may be treated as third-party information.
[0103] As described above, the extraction unit 101 in this embodiment extracts identification information (i.e., material name) and characteristic values of each of the multiple materials from at least one bibliographic information D. The identification information is the first type of information, and the characteristic values are the second type of information. The extraction unit 101 extracts multiple type characteristic values as the characteristic values of the material. These multiple type characteristic values include a first type characteristic value and a second type characteristic value. For example, the first type characteristic value is conductivity, and the second type characteristic value is activation energy value. Furthermore, for each of the multiple materials, the extraction unit 101 extracts attribute information indicating the attributes of the material from at least one bibliographic information D. The attribute information is the third type of information. For example, attributes include publication date, number of citations, author name, and temperature.
[0104] [Correction process by the first information processing unit] The first information processing unit 102 obtains first information indicating the material name from the extraction unit 101. In other words, the material name is obtained by the first information processing unit 102. In obtaining this first information, the first information processing unit 102 may also obtain extraction lists 31 and 34, shown in Figures 5 and 8B, respectively, which contain the first information, from the extraction unit 101. The first information processing unit 102 then modifies the obtained material name. This modification is also called a correction process or a normalization process.
[0105] Figure 9 shows an example of the normalization process for material names.
[0106] Extracting material names using natural language processing or machine learning such as deep learning is difficult to achieve with 100% accuracy. For example, material names are sometimes written as the full name of the material (e.g., Li6.25Al0.25La3Zr2O12) at the beginning of a paper, but as abbreviations (e.g., LALZO) are sometimes used later in the paper. Furthermore, material names may express the mixing ratio or state of each component. As a result, a single material name may be described in multiple forms. For example, multiple forms such as Sulfonated polyimide, SPI, SPI / poly(vinylidene fluoride)(PVDF)blends, and 50wt% of SPI content may be used. Therefore, as shown in Figures 5 and 9, the extraction list 31 generated by the extraction unit 101 records the material name of the same material in multiple forms.
[0107] Therefore, the first information processing unit 102 converts the extracted list 31 into the modified extracted list 31a by performing a normalization process that modifies the representation of the material names included in the extracted list 31.
[0108] For example, the first information processing unit 102 recognizes the string in parentheses in the material name "Li6.25Al0.25La3Zr2O12(LALZ)" included in the extraction list 31 as an abbreviation of the string outside the parentheses, and recognizes the string outside the parentheses as the standard representation of the material name. The first information processing unit 102 then determines that the abbreviation is equivalent to the standard representation and removes the parentheses and the abbreviation inside the parentheses from the material name. Furthermore, the first information processing unit 102 replaces the abbreviation included in the extraction list 31 with the standard representation. Specifically, the first information processing unit 102 modifies the material name "LALZ" included in the extraction list 31 to "Li6.25Al0.25La3Zr2O12". Alternatively, the first information processing unit 102 removes the abbreviation from the extraction list 31.
[0109] The first information processing unit 102 recognizes that x is a variable in the material name "Li6.25AlxLa(1-x)Zr2O12(x=0.1,0.2,0.3)" included in the extraction list 31. This variable x represents the mixing ratio of Al and La. The first information processing unit 102 then removes (x=0.1,0.2,0.3) from "Li6.25AlxLa(1-x)Zr2O12(x=0.1,0.2,0.3)" and assigns the value "x=0.1,0.2,0.3" to the variable x of "Li6.25AlxLa(1-x)Zr2O12". As a result, the first information processing unit 102 decomposes "Li6.25AlxLa(1-x)Zr2O12(x=0.1,0.2,0.3)" into "Li6.25Al0.1La0.9Zr2O12", "Li6.25Al0.2La0.8Zr2O12", and "Li6.25Al0.3La0.7Zr2O12". Then, the first information processing unit 102 modifies the original "Li6.25AlxLa(1-x)Zr2O12" included in the extraction list 31 into the three representations obtained by the decomposition.
[0110] The first information processing unit 102 compares the material name "60Li2SO4*40Li3BO3" and the material name "Li2SO4-Li3BO3" included in the extraction list 31 and determines that the material name "60Li2SO4*40Li3BO3" contains a mixing ratio. As a result, the first information processing unit 102 deletes "60Li2SO4*40Li3BO3" from the extraction list 31. Alternatively, the first information processing unit 102 modifies "60Li2SO4*40Li3BO3" from the extraction list 31 to "Li2SO4-Li3BO3".
[0111] The first information processing unit 102 modifies the representation of the material names included in this extraction list 31, thereby generating the modified extraction list 31a.
[0112] Furthermore, the first information processing unit 102 may search the extraction list 31 for two material names whose start or end points of the span are close to each other and whose representation forms are similar. If such two material names are found, it may determine that those material names are used for the same material. The first information processing unit 102 may also determine that multiple representation forms of material names are used for the same material by performing dependency parsing or other natural language processing on the sentences in the main text data D1.
[0113] Extraction of material names and other information by the extraction unit 101 may result in character corruption, extraction errors, etc. Character corruption can occur, for example, due to errors when converting image data contained in a PDF or XML file to text data. Extraction errors can occur, for example, due to character distortion. The first information processing unit 102 may detect such character corruption and extraction errors and correct the material names included in the extraction list 31.
[0114] [Correction process by the second information processing unit] The second information processing unit 103 acquires second information, which indicates the material's characteristic values, from the extraction unit 101. In other words, the characteristic values are acquired by the second information processing unit 103. In acquiring this second information, the first information processing unit 102 may acquire extraction lists 32, 33, and 34, which include the second information, from the extraction unit 101, as shown in Figures 5, 6B, and 8B, respectively. The second information processing unit 103 modifies these acquired characteristic values. This modification is also called a modification process or a conversion process.
[0115] Figure 10 shows an example of the process for converting material property values.
[0116] The second information processing unit 103 converts the combination of numbers and units included in the characteristic value extracted by the extraction unit 101 into numbers. The second information processing unit 103 uses, for example, a dictionary in which units frequently appearing in materials science papers are pre-registered to convert the numbers in the characteristic value extracted by the extraction unit 101 into numbers expressed using standard units, and removes the units included in the characteristic value. The second information processing unit 103 may also convert the numerical representation of the characteristic value that uses powers of 10 into a numerical representation that does not use powers of 10.
[0117] For example, as shown in Figure 10, the second information processing unit 103 processes the extracted characteristic value "4.6 × 10 -3 Scm -1 By replacing "10-3" with 0.001 in the expression and performing the calculation "4.6 × 0.001", the characteristic value "4.6 × 10 -3 Scm -1 The second information processing unit 103 converts the extracted characteristic value "5.4 mScm" to "0.0046". -1 By replacing "m" with 0.001 in the expression and performing the calculation "5.4 × 0.001", the characteristic value "5.4 mScm" is obtained. -1 Convert " to "0.0054".
[0118] For example, "2.0-2.2 mScm" -1 In some cases, the extracted characteristic value may indicate a numerical range, as shown above. In this case, the second information processing unit 103 recognizes that the characteristic value indicates a numerical range by detecting the hyphenation symbol. As a result, the second information processing unit 103 interpolates the characteristic value between those numerical ranges to obtain the characteristic value "2.0-2.2mScm". -1 This converts "" to "0.0020, 0.0021, 0.0022". In this example, the second information processing unit 103 interpolates the characteristic values at intervals of 0.0001, but it may change the interval automatically, or it may be set according to the user's input operation to the input unit 11.
[0119] For example, "4.38(6) × 10 -3 Scm -1As shown above, the extracted characteristic value may contain an error. The error is indicated by the number in parentheses. In this case, the third information processing unit 104 performs the calculation, for example, "0.00438±0.00006", to obtain the characteristic value "4.38(6)×10 -3 Scm -1 This converts "" to "0.00444, 0.00432". In other words, characteristic values containing errors are converted to the maximum and minimum values of the characteristic value within that error range.
[0120] The characteristic values extracted from the figures in the text data D1 by the extraction unit 101 may be incorrect. For example, the image of the figure to be extracted may be unclear, multiple lines in the figure may overlap too much, the lines may be unclear, or the figure may be too small. In such cases, the extracted characteristic values may be incorrect. Specifically, as shown in Figure 6A, the graph in the text data D1 of document ID "0002" includes a magnified portion. This magnified portion may cause the characteristic values extracted by the extraction unit 101 to be incorrect.
[0121] The second information processing unit 103 may correct or delete such erroneous characteristic values. For example, the second information processing unit 103 may use an image recognition tool to estimate areas in the figure where errors are likely to occur and delete characteristic values extracted from those areas. Alternatively, the second information processing unit 103 may compare characteristic values extracted from the figure with characteristic values extracted from other figures in the same text data D1, detect characteristic values that are significantly different from each other, and delete those characteristic values. If the characteristic value is conductivity, the second information processing unit 103 may, for example, use a dictionary in which the numerical range that conductivity can take is pre-registered, and delete the extracted conductivity if it falls outside that registered numerical range. Alternatively, the second information processing unit 103 may clip the conductivity to the upper or lower limit of the numerical range.
[0122] The extraction list 34, which contains characteristic values extracted from the table of text data D1 by the extraction unit 101, may be incorrect. For example, as shown in Figure 8A, Table 2 for document ID "0001" has two rows of column labels. That is, there is a column label that describes "Experimental1" and a column label that describes "Conductivity(mScm-1)". In this case, as shown in Figure 8B, the extraction unit 101 may mistakenly record "Experimental1" in the column of the column label in the extraction list 34, where it should record "Conductivity(mScm-1)".
[0123] In this case, the second information processing unit 103 compares the characteristic value "2.45" shown in the extraction list 34 with the numerical range that conductivity can take, which is registered in the dictionary mentioned above. The second information processing unit 103 then determines that if the characteristic value falls within the numerical range, it associates that characteristic value with "Conductivity(mScm-1)", even if "Experimental1" is associated with that characteristic value in the extraction list 34. As a result, the second information processing unit 103 replaces "Experimental1" in the extraction list 34 with "Conductivity(mScm-1)" and treats that characteristic value as conductivity.
[0124] The characteristic values extracted from the table of text data D1 by the extraction unit 101 may be incorrect. For example, if the table is unclear, the extracted characteristic values may be incorrect. In this case as well, the second information processing unit 103 may detect the error in the characteristic value by comparing it with the numerical range that the characteristic value can take in the dictionary.
[0125] [Correction process by the 3rd Information Processing Unit] The third information processing unit 104 acquires third information indicating the attributes of the material from the extraction unit 101. In other words, the attributes are acquired by the third information processing unit 104. In acquiring this third information, the third information processing unit 104 may acquire an extraction list 33, shown in Figure 6B, which includes the third information, from the extraction unit 101. The extraction list 33 shows temperature, which is the x value, as the third information. The third information processing unit 104 modifies this acquired third information, i.e., the attributes, in the same way as the first information processing unit 102 and the second information processing unit 103.
[0126] [Integration processing of each piece of information] The derivation unit 105 generates integrated information by acquiring and integrating the information output from the first information processing unit 102, the second information processing unit 103, and the third information processing unit 104 for each text data D1. In other words, the derivation unit 105 associates the processed first information output from the first information processing unit 102, the processed second information output from the second information processing unit 103, and the processed third information output from the third information processing unit 104 with the document ID of the text data D1. The processed first information, processed second information, and processed third information indicate the corrected material name, corrected characteristic value, and corrected attribute, respectively. This generates integrated information in which the document ID is associated with the material name, characteristic value, and attribute.
[0127] For example, the derivation unit 105 obtains a modified extraction list (e.g., extraction list 31a) from the first information processing unit 102, the second information processing unit 103, and the third information processing unit 104 for each text data D1. Then, the derivation unit 105 identifies the distance between material names and characteristic values, and the distance between material names and attributes, based on the figure numbers, extraction line numbers, spans, etc., shown in these extraction lists, and associates the material names, characteristic values, and attributes. The derivation unit 105 associates material names, characteristic values, and attributes extracted from the same sentence or figure. Furthermore, it can also associate material names, characteristic values, and attributes extracted from multiple adjacent sentences, etc. Note that the distance may be the difference in extraction line numbers in which the two pieces of information are described, or it may be a distance based on the extraction line number and span. For example, the derivation unit 105 may identify the distance between a material name and a characteristic value by calculating the number of characters from the end of the material name to the start of the characteristic value. The derivation unit 105 may create association rules according to the distance, or it may associate material names, characteristic values, and attributes using a natural language processing dependency tool. An example of a natural language processing dependency tool is ChemDataExtractor.
[0128] The derivation unit 105 then outputs the integrated information of each of the multiple text data D1 to the image processing unit 106. The image processing unit 106 retrieves this integrated information from the derivation unit 105 and updates the extracted information table 22 by including this integrated information in the extracted information table 22.
[0129] Figure 11 shows an example of the updated extracted information table 22.
[0130] The updated extracted information table 22 contains integrated information extracted from each text data D1, namely the first information, second information, and third information. The first information indicates the final material name, the two second information entries indicate conductivity and activation energy values, respectively, and the two third information entries indicate the author name and research institution name, respectively. In the extracted information table 22, "author" and "research institution" are listed as information type names and refer to the author name and research institution name, respectively.
[0131] In the example shown in Figure 11, the author's name and the research institution's name are extracted from the main text data D1 as third-party information, i.e., as attributes. However, these attributes may also be extracted from the metadata D2 of the bibliographic information D. In the example shown in Figure 11, "LiSo..." is shown as the final material name, but specific examples of this final material name include "Li7La3Zr2O12", "LiCoO2", "Li3P-LiCl", "Li1.33Ti1.67O4", and "Li2CO3".
[0132] Figure 12 shows an example of a display screen 20 that includes the updated extracted information table 22.
[0133] As shown in Figure 12, the image processing unit 106 displays a display screen 20, including the updated extracted information table 22, on the display unit 13. At this time, the image processing unit 106 includes a graph display start button 23c in the display screen 20. The graph display start button 23c is a button for starting the display of the characteristic value graph. In the example in Figure 12, conductivity and activation energy value are characteristic values, respectively. The characteristic value graph is a graph showing the characteristic values shown in each bibliographic information D. Note that the extracted information table 22 shown in Figure 12 does not show the author's name and the name of the research institution, but these names may be shown.
[0134] By referring to the extracted information table 22 on such a display screen 20, the user can check which publication dates describe materials with what characteristic values. The display screen 20 shown in Figure 12 does not include the literature list window 21 shown in Figures 2A and 2B, but it may include that literature list window 21. In this case, the user may deselect or select new icons for the literature information D within the literature list window 21 by performing input operations on the input unit 11. The image processing unit 106 may update the extracted information table 22 in response to such selections and deselections. Such an extracted information table 22 allows for analysis by date, analysis focused on specific material fields, and so on.
[0135] Figure 13 shows an example of a graph display screen that includes a characteristic value graph. In Figures 13, 14, 15A, 15B, and 15C, "activation" refers to "activation energy".
[0136] For example, the user selects the graph display start button 23c on the display screen 20 shown in Figure 12 by performing an input operation on the input unit 11. As a result, the input unit 11 outputs an input signal to the image processing unit 106 prompting the display of the characteristic value graph. Upon receiving this input signal, the image processing unit 106 switches the display screen 20 shown on the display unit 13 to the graph display screen 20a.
[0137] As shown in Figure 13, the graph display screen 20a includes a characteristic value graph 24, a confidence level display button 23d, and a return button 23e. The confidence level display button 23d is a button for displaying the confidence level of the characteristic value corresponding to the marks plotted on the characteristic value graph 24. The return button 23e is a button for returning the graph display screen 20a to the display screen 20 in Figure 12, and this return button 23e is labeled with the text, for example, "Return to Table".
[0138] The characteristic value graph 24 is a graph in which the horizontal axis shows the activation energy value and the vertical axis shows the conductivity. For each material name (i.e., final material name) shown in the extracted information table 22, the image processing unit 106 plots a mark on the characteristic value graph 24 at a position corresponding to two types of characteristic values shown in association with that material name. These two types of characteristic values are the first type characteristic value and the second type characteristic value, specifically conductivity and activation energy value. In the example shown in Figure 13, the marks are x marks. The image processing unit 106 assigns the material name corresponding to the mark near the mark. As a result, the characteristic value graph 24 showing the conductivity and activation energy values of each of the multiple materials extracted from multiple literature information D selected by the user is displayed on the graph display screen 20a. The characteristic value graph 24 shows the relationship between conductivity, activation energy value and material name.
[0139] Note that the materials or material names corresponding to each mark plotted in the characteristic value graph 24 may be the same or different from each other. For example, these materials may be materials used for the same purpose, such as the negative or positive electrode of a battery, and may each contain the element Li. These materials may each contain the same multiple elements, but the composition ratios of those elements may differ from each other. In the examples shown in Figures 13 to 15C, the material name is shown as "LiSo··", but specific examples of this material name include "Li7La3Zr2O12", "LiCoO2", "Li3P-LiCl", "Li1.33Ti1.67O4", and "Li2CO3".
[0140] For example, users such as materials researchers can refer to this graph display screen 20a to find materials that are suitable for their experimental environment or subject, taking into account the balance of characteristic values.
[0141] The image processing unit 106 may also automatically adjust the scales of the vertical and horizontal axes of the characteristic value graph 24. Such an automatic scaling function is expected to help users accurately grasp the full picture of the characteristic values. For example, the image processing unit 106 determines the scale according to the maximum and minimum values of the characteristic values shown in the extracted information table 22. In a specific example, the image processing unit 106 sets the minimum value of the characteristic value as the origin of the characteristic value graph 24 and the maximum value of the characteristic value as the end of the axis of the characteristic value graph 24. However, if the extracted information table 22 includes outliers as characteristic values, the presence of these outliers may make it difficult to observe characteristic values other than the outliers on the characteristic value graph 24. In a specific example, 99% of all characteristic values included in the extracted information table 22 are within the range of 0 to 0.001, and the remaining 1% are within the range of 1.0 to 2.0. In this case, those 1% of characteristic values are outliers. If the scale is determined using these characteristic values, a 1% outlier will cause the axis corresponding to that characteristic value in the characteristic value graph 24 to be marked in increments of 0.1 within the range of 0 to 2.0. As a result, it becomes difficult to properly observe the 99% of characteristic values that are originally intended to be observed. Therefore, the image processing unit 106 may omit the outliers included in the extracted information table 22 and generate the characteristic value graph 24 based on the characteristic values other than the outliers. This allows for proper observation of the characteristic values that are originally intended to be observed.
[0142] Alternatively, the image processing unit 106 may adjust the scales of the vertical and horizontal axes of the characteristic value graph 24 according to the type of material. For example, the image processing unit 106 identifies the type of material from its material name and sets the scale of the characteristic value associated with that type to the scales of the vertical and horizontal axes of the characteristic value graph 24. In a specific example, if a solid electrolyte is identified as the type of material, the conductivity of a solid electrolyte is 0 to 0.001 S / cm, so the image processing unit 106 sets the scale of the axis corresponding to conductivity in the characteristic value graph 24 to 0 to 0.001. Such a range of conductivity for solid electrolytes may be registered in memory in advance. The image processing unit 106 may also adjust the scales of the vertical and horizontal axes of the characteristic value graph 24 in response to user input operations to the input unit 11. If the image processing unit 106 determines that the characteristic value is conductivity, it may set a predetermined logarithmic scale for conductivity to the axis of the characteristic value graph 24.
[0143] In response to an input operation by the user to the input unit 11, the image processing unit 106 may, when a mark is selected, acquire literature information D containing the material name and two types of characteristic values corresponding to that mark via the extraction unit 101 and display it on the display unit 13.
[0144] [Reliability Derivation Process] For example, the user selects the reliability display button 23d on the graph display screen 20a shown in Figure 13 by performing an input operation on the input unit 11. This causes the input unit 11 to output an input signal to the derivation unit 105 prompting the display of reliability. Upon receiving this input signal, the derivation unit 105 derives the reliability values for the conductivity and activation energy of each material shown in the characteristic value graph 24. The derivation unit 105 then outputs an output signal containing these derived reliability values to the image processing unit 106. Upon receiving the output signal from the derivation unit 105, the image processing unit 106 changes the display mode of each material mark plotted on the characteristic value graph 24 according to the derived reliability value for the conductivity and activation energy of that material.
[0145] In a specific example, the derivation unit 105 calculates the distance between the mark of the material to be derived and the respective marks of several other materials, and then calculates the average value of these distances. This distance may be Euclidean distance, Manhattan distance, or any other distance. Next, the derivation unit 105 calculates the similarity of the characteristic values of the material to be derived. A larger average value of the distances between the mark of the material to be derived and the respective marks of several other materials indicates a smaller similarity, while a smaller average value indicates a larger similarity. For example, the similarity is expressed as a numerical value within the range of 0 to 100. These characteristic values are conductivity and activation energy value. Next, the derivation unit 105 determines this similarity as a confidence score. In this embodiment, one confidence score is calculated for two types of characteristic values of the material to be derived.
[0146] The above process is illustrated in the case where eight different coordinates are displayed in the characteristic value graph 24 of Figure 13 (the updated extracted information table 22 in Figure 12 contains data for IDs 001 to 008). The coordinates corresponding to the first pair of conductivity and activation energy values identified by ID 001 (0.41, 0.048) are called the first coordinate, and the coordinates corresponding to the eighth pair of conductivity and activation energy values identified by ID 008 (0.15, 0.077) are called the eighth coordinate. The number of different coordinates may be other than eight.
[0147] The derivation unit 105 calculates the distance L12, ~ between the first and second coordinates, L18 between the first and eighth coordinates, L23, ~ between the second and third coordinates, L28, ~ between the second and eighth coordinates, and L78 between the seventh and eighth coordinates. The number of distances calculated is (8 × 7) / 2 = 28.
[0148] The derivation unit 105 calculates L(avg1)=(L12+···+L18) / 7, which is the average with respect to the first coordinate, L(avg2)=(L12+L23+···+L28) / 7, ~, and L(avg8)=(L18+L28+···+L78) / 7, which is the average with respect to the eighth coordinate.
[0149] The derivation unit 105 sets a smaller similarity between the conductivity and activation energy values for the nth set when L(avgn) is large, and a larger similarity between the conductivity and activation energy values for the nth set when L(avgn) is small. n is a natural number between 1 and 8.
[0150] Figure 14 shows an example of a graph display screen 20a showing the reliability level.
[0151] As shown in Figure 14, the image processing unit 106 displays each mark plotted on the characteristic value graph 24 in a manner corresponding to the reliability of the two characteristic values of the material indicated by that mark, namely conductivity and activation energy value. In a specific example, these marks are circular, and the image processing unit 106 darkens the color of the mark corresponding to the two characteristic values as the reliability of those two characteristic values increases, and lightens the color of the mark corresponding to the two characteristic values as the reliability of those two characteristic values decreases. In other words, the image processing unit 106 darkens the color of marks that are densely packed together on the characteristic value graph 24, and lightens the color of marks that are less densely packed. Therefore, marks that are far from many other marks are displayed lightly, allowing the user to determine that the characteristic value of the material corresponding to that mark is relatively unreliable. Conversely, marks that are clustered close together are displayed darkly, allowing the user to determine that the characteristic value of the material corresponding to those marks is relatively reliable.
[0152] The image processing unit 106 includes a confidence level hide button 23f on its graph display screen 20a instead of the confidence level display button 23d described above. This confidence level hide button 23f is a button for hiding the confidence level, and is used, for example, to return the graph display screen 20a shown in Figure 14 back to the graph display screen 20a shown in Figure 13.
[0153] In this embodiment, the user is presented with a confidence level for the characteristic values extracted from the literature information D. If the multiple literature information sources D selected by the user are electronic data of papers on material synthesis experiments, the confidence level is an index that represents the reproducibility of a single experiment described in a paper, estimated based on the information extracted from that paper. Reproducibility is an index that shows the ratio of successes to the number of trials when an experiment is actually performed multiple times. If one author performs similar experiments, obtains materials with similar characteristic values, and publishes them in multiple papers, the more publications there are, the higher the probability that the synthesis experiment of that material will be successful. Therefore, reproducibility can also be calculated from the number of papers published by a single author. If similar characteristic values are described in papers published by different authors or different research institutions, a high confidence level can be presented, and the confidence level can be visualized considering the development skills of the authors or research institutions. Even if multiple characteristic values are measured with different devices, if those characteristic values are close, the reliability of those characteristic values can be judged to be high, and the effect of increasing the accuracy of the confidence level can be expected. Therefore, this embodiment is expected to be useful for materials researchers who place importance on the number of previously reported papers on material synthesis when searching for desired materials.
[0154] In this embodiment, confidence levels are presented for two characteristic values, conductivity and activation energy value. However, these may be other characteristic values, and confidence levels may be presented for only one characteristic value. In this embodiment, the characteristic value graph 24 is composed of two axes showing conductivity and activation energy value, but it may be composed of three axes showing three characteristic values. In other words, the characteristic value graph 24 may be composed of a three-dimensional graph. One of the three characteristic values may be the conditions used in material synthesis, such as temperature or time. Furthermore, the characteristic value graph 24 may be composed of four or more axes. In other words, this characteristic value graph 24 shows four or more characteristic values. Since it is difficult to visualize a characteristic value graph 24 with four or more axes, the characteristic value graph 24 may be constructed by reducing the number of axes to two or three using dimensionality reduction techniques such as PCA (Principal Component Analysis) or t-SNE (t-Distributed Stochastic Neighbor Embedding). The image processing unit 106 may select the type and number of axes in response to the user's input operation to the input unit 11. By selecting different types of axes and increasing the number of axes, the expressive power of characteristic values can be enhanced, allowing for more detailed analysis.
[0155] The function that changes the intensity of the mark corresponding to a material according to the level of confidence in its property values is essential for materials researchers to quickly check a large amount of property values, and effectively highlights areas of interest within the property value graph 24. For example, when a user prioritizes high reproducibility when searching for materials, they can focus on materials with darker marks. Conversely, when a user prioritizes rarity when searching for materials, they can focus on materials with lighter marks. This enables materials research and other users to search for materials according to their objectives.
[0156] The color of the material marks plotted on the characteristic value graph 24 may be the same as the color used for the main component element of that material in the periodic table that the user normally uses. For example, if the main component element of the material is Li, and the Li region is shown in yellow in the periodic table, the image processing unit 106 may display the material mark in yellow on the display unit 13. In a specific example, the image processing unit 106 pre-stores data showing the periodic table that the user normally uses, identifies the main component element of the material shown in the extracted information table 22, and determines the color corresponding to that element based on the aforementioned data that it has pre-stored. Then, the image processing unit 106 sets the color of the mark corresponding to that material to the color determined as described above, and plots the mark of that color on the characteristic value graph 24. This allows the user to easily grasp the main component of the material corresponding to each mark shown on the characteristic value graph 24 from its color, thereby further improving the efficiency of data analysis.
[0157] As described above, the derivation unit 105 in this embodiment derives the reliability of the characteristic value of each of the multiple materials based on the similarity of the characteristic values of that material with one or more other materials. The image processing unit 106 then generates a first image which (i) displays the characteristic values of each of the multiple materials in a display manner corresponding to the reliability derived for the characteristic value of that material, and (ii) displays them in association with the identification information of that material, and outputs the first image to the display unit 13. The first image is, for example, the graph display screen 20a in Figure 14. As a result, in the displayed first image, the characteristic values of each of the multiple materials are displayed in a display manner corresponding to the reliability of the characteristic value, and the identification information of the material is associated with the characteristic value. Therefore, by looking at the first image, users such as materials researchers can easily grasp the reliability of the characteristic values of the materials described in the vast amount of literature information D stored in the literature database 12 from the first image.
[0158] The image processing unit 106 sets a characteristic map in the first image, which consists of a first coordinate axis for showing the first type characteristic value, which is conductivity, and a second coordinate axis for showing the second type characteristic value, which is activation energy value. This characteristic map is, for example, the characteristic value graph 24 in Figure 14. The image processing unit 106 then superimposes marks corresponding to each of the multiple materials onto the characteristic map at positions corresponding to the first type characteristic value and the second type characteristic value of that material, in a display manner that corresponds to the reliability derived for the characteristic value of that material. This makes it easier to grasp the overall reliability of the multiple type characteristic values that each material possesses.
[0159] The derivation unit 105 identifies the similarity of characteristic values in a material based on the distance between a mark corresponding to that material and marks corresponding to each of one or more other materials. This allows for the appropriate identification of a single similarity for multiple type characteristic values of a material, and as a result, a confidence score based on the similarity can be appropriately derived.
[0160] The image processing unit 106 determines the display mode for the material's characteristic value, assigning a darker color intensity to each characteristic value based on its derived reliability, and generates a first image showing the characteristic value with the determined color intensity. This allows the user to easily grasp the reliability of the characteristic value visually. In this embodiment, a characteristic value graph 24 is used, and the marks representing the characteristic value are displayed with the determined color intensity. However, if the characteristic value graph 24 is not used, the numerical value of the characteristic value itself may be displayed with the determined color intensity.
[0161] Figure 15A shows another example of the graph display screen 20a that shows the reliability.
[0162] In the example illustrated in FIG. 14, the deriving unit 105 derives the similarity based on the above-mentioned distance as the reliability. That is, when the similarity is R0 and the reliability is K, the deriving unit 105 calculates the reliability K by K=1×R0. However, for the calculation of the reliability K, the deriving unit 105 may reflect the similarity R0 and an attribute value corresponding to the attribute of a material in the reliability K. For example, when the four attribute values are R1, R2, R3, and R4, the deriving unit 105 calculates the reliability K by K=(p×R0)+(a×R1)+(b×R2)+(c×R3)+(d×R4), that is, by weighted addition of the four attribute values R1 to R4 and the similarity R0. p is a weight for the similarity R0, a is a weight for the attribute value R1, and b is a weight for the attribute value R2. c is a weight for the attribute value R3, and d is a weight for the attribute value R4. That is, the deriving unit 105 applies a plurality of types of biases to the derivation of the reliability K. The weights p and a to d satisfy p+a+b+c+d=1, and are set to values according to an input operation to an input unit 11 by a user. Note that the weight p for the similarity R0 satisfies 0<p≦1, and each of the weights a to d for the attribute values R1 to R4 satisfies 0≦[a,b,c,d]<1. The weight is also referred to as a bias.
[0163] The image processing unit 106 displays slider bars 25a to 25d for adjusting the weights a to d of the attribute values R1 to R4 on a graph display screen 20a, for example as illustrated in FIG. 15A.
[0164] The slider bar 25a is a display element for setting the weight a of the attribute value R1 based on the publication date, according to the position of the slider 1sa. The publication date is the publication date of the literature information D from which the material name, conductivity, and activation energy value of the material are extracted, and is an attribute of that material. The derivation unit 105 determines the attribute value R1 based on the publication date, with a value closer to 100 the more recent the publication date and a value closer to 0 the older the publication date. The derivation unit 105 sets the weight a to a value that is 0 when the slider 1sa is at the left end and increases as the slider 1sa approaches the right end. For example, if the slider 1sa is in the center of the slider bar 25a, the weight a may be 0.5.
[0165] The slider bar 25b is a display element for setting the weight b of the attribute value R2 based on the number of citations, according to the position of the slider 1sb. The number of citations is the number of citations of the literature information D from which the material name, conductivity, and activation energy value of the material have been extracted, and is an attribute of that material. The derivation unit 105 determines the attribute value R2 based on the number of citations, and the value that is closer to 100 the more citations there are, and the value that is closer to 0 the fewer citations there are. The derivation unit 105 sets the weight b to a value that is 0 when the slider 1sb is at the left end and increases as the slider 1sb approaches the right end. For example, if the slider 1sb is in the center of the slider bar 25b, the weight b may be 0.5.
[0166] The slider bar 25c is a display element for setting the weight c of the attribute value R3 based on the author's name, according to the position of the slider 1sc. The author's name is the author's name from the literature information D from which the material name, conductivity, and activation energy value of the material have been extracted, and is an attribute of that material. The derivation unit 105 determines the attribute value R3 based on the author's name, with a value closer to 100 the more well-known the author's name is, and a value closer to 0 the less well-known the author's name is. The information providing device 100 may also hold data indicating the well-knownness of each author's name. The derivation unit 105 sets the weight c to a value that is 0 when the slider 1sc is at the left end, and increases as the slider 1sc approaches the right end. For example, if the slider 1sc is in the center of the slider bar 25c, the weight c may be 0.5.
[0167] The slider bar 25d is a display element for setting the weight d of the temperature-based attribute value R4 according to the position of the slider 1sd. This temperature is the temperature used to synthesize the material of the extracted material name, and is an attribute of that material. This temperature can also be said to be one of the conditions included in the material synthesis method. The derivation unit 105 calculates the difference between the temperature of that material and the respective temperatures of one or more other materials, and determines the attribute value R4 based on the temperature of that material, with a value closer to 100 the smaller the average of the differences, and closer to 0 the larger the average. In other words, this attribute value R4 indicates the similarity of temperatures. The derivation unit 105 sets the weight d to a value that is 0 when the slider 1sd is at the left end, and increases as the slider 1sd approaches the right end. For example, if the slider 1sd is in the center of the slider bar 25d, the weight d may be 0.5.
[0168] Furthermore, the publication date, number of citations, author's name, and temperature are each extracted as third-party information by the extraction unit 101.
[0169] The user adjusts the positions of sliders 1sa to 1sd on slider bars 25a to 25d by performing input operations on the input unit 11. For example, the user moves each of sliders 1sa to 1sd to the leftmost position. In this case, the derivation unit 105 sets the weights a to d for publication date, number of citations, author name, and temperature to 0, and sets the weight p for similarity R0 to 1, in response to the input signal output from the input unit 11 by the user's input operation. As a result, the derivation unit 105 calculates the confidence score K using K = (1 × R0) + (0 × R1) + (0 × R2) + (0 × R3) + (0 × R4), i.e., by the similarity score R0, similar to the example shown in Figure 14. As a result, the image processing unit 106 displays the characteristic value graph 24, similar to the example in Figure 14, on the display unit 13, including it on the graph display screen 20a, as shown in Figure 15A.
[0170] Thus, in this embodiment, the derivation unit 105 derives the confidence level K of the characteristic values in a material based on the similarity R0 of the characteristic values in the material and the attribute information extracted for that material. In other words, the derivation unit 105 derives the confidence level K of the characteristic values in the material by weighted addition of the similarity R0 of the characteristic values in the material and the attribute values R1 to R4 based on the attribute information extracted for that material. As a result, the confidence level K of the characteristic values is derived based on the similarity R0 of the characteristic values and the attributes of the material, so the confidence level K can be derived from multiple perspectives, and the accuracy of the confidence level K can be improved. By adjusting the weights of the similarity R0 and the attributes as biases, a confidence level K that suits the user's purpose can be derived.
[0171] Figures 15B and 15C show an example of the graph display screen 20a after the bias has been changed.
[0172] As shown in Figure 15B, the user performs an input operation on the input unit 11, moving slider 1sa from the position shown in Figure 15A to the rightmost end of the sliders 1sa to 1sd of slider bars 25a to 25d. In this case, the derivation unit 105 sets the weight a of the publication date to, for example, 0.9, the weights b to d of the number of citations, author name, and temperature to 0, and the weight p of the similarity R0 to 0.1, according to the input signal output from the input unit 11 by the input operation. As a result, the derivation unit 105 calculates the confidence score K using K = (0.1 × R0) + (0.9 × R1) + (0 × R2) + (0 × R3) + (0 × R4), that is, by heavily biasing it towards the publication date. As a result, the image processing unit 106 displays a characteristic value graph 24, different from the examples in Figures 14 and 15A, on the display unit 13, including it on the graph display screen 20a, as shown in Figure 15B. In other words, the publication date has a greater impact on the confidence score K than the similarity score R0, and the earlier the publication date of the literature information D, the more likely it is that the marks corresponding to the material name, conductivity, and activation energy value extracted from that literature information D will be displayed in a darker color.
[0173] In other words, simply having a high density of plotted marks does not make those marks appear darker; rather, the more densely plotted the marks are, and the more recent the publication date corresponding to those marks, the darker the color of those marks will appear. Thus, with the slider bars 25a to 25d shown in Figure 15B, the newer the literature information D from which characteristic values are extracted, the higher the reliability K of those characteristic values, and therefore the darker the mark corresponding to those characteristic values will appear. Consequently, this embodiment is expected to be useful for materials researchers who prioritize the recency of literature information D, such as papers, when searching for desired materials.
[0174] Thus, in this embodiment, the attribute information indicates the publication date of at least one bibliographic information D that contains the identification information and characteristic values of the material corresponding to that attribute information. The publication date is the publication date as described above, or it may be the publication date. The derivation unit 105 then derives the confidence level K of the characteristic value using the attribute value R1 which indicates the recency of the publication date. Specifically, the attribute value R1 is larger the more recent the publication date. The derivation unit 105 then derives a larger confidence level K for the characteristic value of the material the higher the similarity R0 of the characteristic values in the material and the larger the attribute value R1 corresponding to that material. As a result, a relatively high confidence level K can be derived for characteristic values described in newer bibliographic information D. Therefore, an appropriate confidence level K can be derived for users who place importance on the publication date.
[0175] Next, as shown in Figure 15C, the user moves sliders 1sa to 1sd of slider bars 25a to 25d to the right by performing an input operation on the input unit 11. The derivation unit 105 sets weights p and a to d according to the input signal output from the input unit 11 by the input operation, that is, according to the distance from the left end of each of the sliders 1sa to 1sd. The ratio of each weight a to d is set to the ratio of the distance from the left end of each of the sliders 1sa to 1sd. For example, in the example in Figure 15C, weights a to d are set as follows: (weight d for temperature) > (weight c for author name) > (weight b for citation count) = (weight a for publication date), and the weight p for similarity R0 is set by p = 1 - (a + b + c + d).
[0176] As a result, the derivation unit 105 calculates the confidence score K using the formula K=(p×R0)+(a×R1)+(b×R2)+(c×R3)+(d×R4), that is, by applying a bias of a value greater than 0 to each of the similarity R0, publication date, number of citations, author name, and temperature. As a result, the image processing unit 106 displays a characteristic value graph 24 on the graph display screen 20a and the display unit 13, as shown in Figure 15C, which is different from the examples in Figures 14, 15A, and 15B. In other words, the more similar the temperature of material synthesis is to the temperature of other materials, the more well-known the author's name is, the more citations there are, and the earlier the publication date, the more likely the mark color of the corresponding material will be displayed more intensely. Such slider bar settings 25a~25d are expected to be effective for materials researchers who want to search for materials from multiple perspectives or at a high level.
[0177] Thus, in this embodiment, the attribute information indicates, as an attribute, the number of citations of at least one bibliographic information D that contains the identification information and characteristic values of the material corresponding to that attribute information. The derivation unit 105 then derives the confidence level K of the characteristic value using the attribute value R2 corresponding to the number of citations. Specifically, the attribute value R2 is larger the more citations a material receives. The derivation unit 105 derives a larger confidence level K for the characteristic value of a material the higher the similarity R0 of the characteristic values in the material and the larger the attribute value R2 corresponding to that material. As a result, a relatively high confidence level K can be derived for characteristic values described in bibliographic information D with a large number of citations. Therefore, an appropriate confidence level K can be derived for users who place importance on the number of citations.
[0178] In this embodiment, attribute information indicates the synthesis method of the material corresponding to that attribute information as an attribute. The derivation unit 105 then derives the confidence level K of the characteristic value using attribute values R4 corresponding to the degree of similarity between the synthesis method of that material and the synthesis methods of one or more other materials. The synthesis method of that material is the temperature conditions used in material synthesis, as shown in the example in Figure 15C. Note that the synthesis method may be the time conditions or the type of equipment used in material synthesis, rather than temperature conditions. In other words, the synthesis method of the material may include at least one of the temperature conditions, time conditions, and the type of equipment used in material synthesis. Specifically, the attribute value R4 is larger the greater the degree of similarity between the synthesis method of that material and the synthesis methods of one or more other materials. The derivation unit 105 derives a larger confidence level K of the characteristic value of that material the higher the similarity R0 of the characteristic value of that material and the larger the attribute value R4 corresponding to that material. Therefore, if the material for that characteristic value is described in the literature information D as being synthesized using the same synthesis method as other literature information D, a relatively high confidence level K can be derived for that characteristic value. Thus, an appropriate confidence level K can be derived for users who focus on the synthesis method.
[0179] In this embodiment, the attribute information indicates the author of at least one bibliographic information D that contains the identification information and characteristic values of the material corresponding to that attribute information. The derivation unit 105 then derives the confidence level K of the characteristic value using an attribute value R3 corresponding to the fame of the author of the bibliographic information D. Specifically, the attribute value is larger the author is more famous. The derivation unit 105 derives a larger confidence level K for the characteristic value of the material the higher the similarity R0 of the characteristic value in that material and the larger the attribute value R3 corresponding to that material. As a result, a relatively high confidence level K can be derived for characteristic values described in bibliographic information D of authors who are famous. Therefore, an appropriate confidence level K can be derived for users who place importance on the fame of the authors of bibliographic information D.
[0180] Furthermore, the derivation unit 105 may derive the confidence level K of the characteristic value using an attribute value R3a that corresponds to whether the author of the document information D is the same as the author of one or more other documents information D. In other words, the attribute value R3a may be used instead of the attribute value R3 mentioned above. Specifically, the attribute value R3a is larger the more authors there are among the one or more other documents information D whose authors are different from the author of document information D. The derivation unit 105 derives a larger confidence level K of the characteristic value for that material the higher the similarity R0 of the characteristic value for that material is, and the larger the attribute value R3a corresponding to that material is. As a result, a relatively high confidence level K can be derived for characteristic values described in documents D written by authors different from the authors of many other documents D. In other words, if the characteristic values described in many documents D written by different authors are similar, a high confidence level K can be derived for those characteristic values. On the other hand, even if many bibliographic sources D describe similar characteristic values, if those bibliographic sources D are written by the same author, a low confidence level K can be derived for those characteristic values. Therefore, an appropriate confidence level can be derived for users who place importance on the author identity of the bibliographic sources D.
[0181] In the examples in Figures 15A to 15C, bias is adjusted for each of the four attributes: publication date, number of citations, author name, and temperature. However, bias may also be adjusted for other attributes (i.e., third information) included in the extracted information table 22 in Figure 11 or Figure 12. Depending on the user's input operation to the input unit 11, the image processing unit 106 may increase or decrease the number of slider bars for bias adjustment.
[0182] If the information source D is an article, the publication date, submission date, and acceptance date may be shown in metadata D2. In this case, the extraction unit 101 may extract the submission date and acceptance date from the metadata D2 as attributes of the material, and the derivation unit 105 may include these attributes in the extracted information table 22. The derivation unit 105 may then use an attribute value R5 corresponding to the period from the submission date to the acceptance date instead of the attribute value R1 based on the publication date, or it may use it together with the attribute value R1 to calculate the confidence score K. In this case, the derivation unit 105 determines the attribute value R5 to be, for example, a value closer to 100 the shorter the period from the submission date to the acceptance date, and a value closer to 0 the longer the period. The derivation unit 105 then calculates the confidence score K by weighted addition including the attribute value R5.
[0183] This allows for the presentation of a confidence score K that takes into account the period from submission to acceptance, providing significant information for materials researchers who consider such periods important. In other words, if the period from submission to acceptance is long, the paper may have undergone repeated revisions based on criticisms from reviewers during that time. Such repeated revisions can lower the reliability of the paper, i.e., the reliability of the material properties presented in that paper. Therefore, as described above, by presenting a confidence score K that takes into account the period from submission to acceptance, the accuracy of the confidence score K can be improved.
[0184] If the journal name and impact factor (also known as IF) are included as attributes in the extracted information table 22, the derivation unit 105 may calculate the confidence level K by weighted addition using attribute values based on those attributes. Attribute values based on the journal name are closer to 100 the more well-known the journal is and closer to 0 the less well-known it is. Attribute values based on the IF are closer to 100 the larger the IF is and closer to 0 the smaller the IF is.
[0185] The popularity of such journal names and impact factor (IF) can be said to indicate the quality of the articles, which are the bibliographic information D. Therefore, by using attribute values based on journal names and IF to present a confidence level K that takes the quality of the articles into account, the accuracy of the confidence level K can be improved.
[0186] If the journal name and IF are not extracted from the bibliographic information D and are not included as attributes in the extracted information table 22, the derivation unit 105 may estimate these attributes. For example, the derivation unit 105 compares bibliographic information D with known journal names and IFs with bibliographic information D with unknown journal names and IFs, for example, using a natural language processing tool, and determines whether the similarity of the bibliographic information D is above a threshold. If the similarity is above a threshold, the derivation unit 105 may estimate the known journal names and IFs as the journal names and IFs of the unknown bibliographic information D. This makes it possible to calculate a confidence score K with high accuracy even for unknown bibliographic information D.
[0187] In the examples shown in Figures 15A to 15C, the derivation unit 105 determines the similarity of the material temperatures as the attribute value R4. However, the derivation unit 105 may also determine the attribute value R4 based on the material temperature, such that the higher the material temperature, the closer it is to 100, and the lower the temperature, the closer it is to 0.
[0188] [Processing operation] Figure 16 is a flowchart showing an example of the overall processing operation of the information providing device 100 in this embodiment.
[0189] (Step S110) First, the extraction unit 101 of the information providing device 100 receives an input signal output from the input unit 11 in response to an input operation by the user to the input unit 11. This input signal is, for example, a signal for identifying one or more bibliographic information D selected by the user's input operation from among multiple bibliographic information D stored in the bibliographic database 12. For example, the input signal may indicate the bibliographic ID of the bibliographic information D.
[0190] (Step S120) Next, the extraction unit 101 identifies one or more pieces of literature information D identified by the input signal received in step S110, and extracts various information for each of the multiple materials from those pieces of literature information D. This various information includes the first information, second information, and third information described above, namely the material name, characteristic value, and attribute. The material name is identification information for identifying the material. In other words, for each of the multiple materials, the extraction unit 101 extracts identification information for identifying that material and its characteristic value from at least one piece of literature information D.
[0191] (Step S130) Next, the derivation unit 105 derives the confidence level for the characteristic value of each material based on the material name, characteristic value, and attributes of the multiple materials extracted in step S120. For example, for each of the multiple materials, the derivation unit 105 derives the confidence level for the characteristic value of that material based on the similarity of the characteristic values between that material and one or more other materials.
[0192] (Step S140) Next, the image processing unit 106 generates a first image that displays the characteristic values of each material in a manner corresponding to the confidence level derived in step S130. This first image is, for example, the characteristic value graph 24 shown in Figures 14 to 15C. That is, the image processing unit 106 generates a first image that (i) displays the characteristic values of multiple materials in a manner corresponding to the confidence level derived for the characteristic values of those materials, and (ii) displays them in association with the identification information of those materials.
[0193] (Step S150) Next, the image processing unit 106 displays the first image generated in step S140 on the display unit 13. That is, the image processing unit 106 outputs the first image to the display unit 13.
[0194] Figure 17 is a flowchart illustrating an example of processing by the extraction unit 101. In other words, Figure 17 is a flowchart that shows in detail the processing operation of step S120 in Figure 16.
[0195] (Step S121) First, the extraction unit 101 obtains the literature information D identified by the input signal from the literature database 12.
[0196] (Step S122) Next, the extraction unit 101 performs a process to search for and extract the first information, namely the material name, from the literature information D.
[0197] (Step S123) Next, the extraction unit 101 determines whether or not the material name was extracted from the literature information D by the process in step S122. If the extraction unit 101 determines that it was not possible to extract the material name (no in step S123), it repeats the process from step S121. If the process in step S121 is repeated, the extraction unit 101 obtains literature information D that has not yet been acquired from the literature database 12.
[0198] (Step S124) Then, if the extraction unit 101 determines in step S123 that a material name has been extracted from the literature information D (yes in step S123), it extracts the second and third pieces of information from that literature information D. The second and third pieces of information are the characteristic values and attributes of the material corresponding to the material name that has already been extracted.
[0199] (Modified version of Embodiment 1) Figure 18A is a block diagram showing an example configuration of the information provision system in a modified version of Embodiment 1. Note that, for components in this modified version that are the same as those in Embodiment 1, the same reference numerals are used, and detailed descriptions are omitted.
[0200] In this modified example, the information provision system 1001 comprises an information provision device 100, an input unit 11, a literature database 12, a display unit 13, and a material properties database (DB) 14, as shown in Figure 18A.
[0201] The material properties database 14 is a recording medium that stores, for each of several materials, the material name, its characteristic values, and its attributes in a pre-associated manner. Such a recording medium may be a hard disk drive, RAM, ROM, or semiconductor memory. The recording medium may be volatile or non-volatile.
[0202] In this modified example, the derivation unit 105 of the information providing device 100 uses its material properties database 14. For example, the derivation unit 105 determines whether characteristic values and attributes are associated with each of the material names of the multiple materials shown in the extracted information table 22. If the derivation unit 105 determines that characteristic values and attributes are not associated with a material name, it treats that material name as the material name to be processed and refers to the material properties database 14. The derivation unit 105 then searches the material properties database 14 for the material name to be processed and extracts the characteristic values and attributes associated with that material name in the material properties database 14. The derivation unit 105 records the extracted characteristic values and attributes in the extracted information table 22, associating them with the material name to be processed.
[0203] This means that even if it is not possible to extract material characteristic values and attributes from multiple bibliographic information D stored in the bibliographic database 12, those characteristic values and attributes can be supplemented by using the material characteristic database 14.
[0204] Figure 18B is a block diagram showing another example of the configuration of the information provision system 1001 in a modified example of Embodiment 1.
[0205] The information provision system 1001 may include a material database 15 instead of the material properties database 14 shown in Figure 18A. The material database 15 is an open-access database and a recording medium that stores various information about multiple materials. Such a recording medium may be a hard disk drive, RAM, ROM, or semiconductor memory. The recording medium may be volatile or non-volatile.
[0206] The derivation unit 105 extracts characteristic values and attributes corresponding to the material names to be processed from the material database 15, similar to the extraction unit 101. The derivation unit 105 may also estimate the characteristic values and attributes corresponding to the material names to be processed by referring to the material database 15.
[0207] (Embodiment 2) The information provision system 1000 and information provision device 100 in this embodiment have the same configuration as the information provision system 1000 and information provision device 100 in Embodiment 1. In addition to the functions of the information provision device 100 in Embodiment 1, namely the reliability derivation function, the information provision device 100 in this embodiment has a data display function. The data display function is a function that displays some of the information from a plurality of bibliographic information D stored in the bibliographic database 12 as data on the display unit 13. Of the components in this embodiment, components that are the same as those in Embodiment 1 are denoted by the same reference numerals as in Embodiment 1, and detailed descriptions are omitted.
[0208] Figure 19 shows an example of a data display screen in this embodiment.
[0209] After the reliability level is derived by the reliability level derivation function, the user instructs the information providing device 100 to display the document display screen 40 by performing an input operation on the input unit 11. When the image processing unit 106 receives an input signal corresponding to the input operation from the input unit 11, it displays the document display screen 40, for example, as shown in Figure 19, on the display unit 13.
[0210] The data display screen 40 includes a bibliography window 21, a materials window 41, a data window 42, a quantity adjustment window 46, and a return button 43b.
[0211] The bibliography window 21 displays a list of bibliographic information D stored in the bibliographic database 12. In this list, icons for bibliographic information D selected by the user for confidence level derivation are displayed in a manner that distinguishes them from the icons for bibliographic information D that have not been selected.
[0212] The material window 41 is a window for displaying material conditions entered by the user. Material conditions include, for example, conditions related to the elemental species or composition that make up the material.
[0213] The quantity adjustment window 46 is a window for adjusting the amount of data displayed, and includes slider bars 46a to 46f. Slider bar 46a is a display element for adjusting the amount of data displayed regarding publication date, depending on the position of slider 2sa. Slider bar 46b is a display element for adjusting the amount of data displayed regarding conductivity, depending on the position of slider 2sb. Slider bar 46c is a display element for adjusting the amount of data displayed regarding temperature, depending on the position of slider 2sc. Slider bar 46d is a display element for adjusting the amount of data displayed regarding structure, depending on the position of slider 2sd. Slider bar 46e is a display element for adjusting the amount of data displayed regarding the name of the research institution, depending on the position of slider 2se. Slider bar 46f is a display element for adjusting the amount of data displayed regarding the name of the author, depending on the position of slider 2sf.
[0214] The sliders 2sa to 2sf indicate quantities that are smaller when the slider is closer to the left end and larger when it is closer to the right end. Publication date, conductivity, temperature, structure, name of research institution, and author name are all attributes of the material. In other words, these sliders 2sa to 2sf are used to adjust the bias of the amount of data displayed for each attribute, such as the publication date. This bias in the amount of data is also called the weight of the amount of data.
[0215] The data window 42 is a window for displaying data related to materials that match the material conditions entered by the user. This data window 42 displays data on publication date, conductivity, temperature, structure, name of research institution, and name of author, in quantities adjusted by the quantity adjustment window 46.
[0216] The return button 43b is used to return the data display screen 40 displayed on the display unit 13 back to the display screen 20 or graph display screen 20a of Embodiment 1.
[0217] Figure 20 shows another example of the data display screen 40 in this embodiment.
[0218] The user inputs material conditions by performing an input operation on the input unit 11 and operates the slider 2sb on the slider bar 46b. The extraction unit 101 then acquires the material conditions corresponding to the input operation. The image processing unit 106 then displays "Li-Al-La-Zr-O" as the material conditions in the material window 41. Furthermore, the image processing unit 106 moves the slider 2sb on the slider bar 46b to the rightmost position. The sliders 2sa and 2sc-2sf on slider bars 46a and 46c-46f are at the leftmost position. This state of the quantity adjustment window 46 means that the conductivity-related data will be displayed across the entire data window 42. Additionally, the image processing unit 106 displays the data display button 43a within the data display screen 40. This data display button 43a is used to display the data in the data window 42.
[0219] Next, the extraction unit 101 extracts information regarding the conductivity of materials whose material names correspond to the material condition "Li-Al-La-Zr-O" from each document information D in the document database 12 selected by the user. Specifically, the extraction unit 101 first searches for material names containing the element species Li, Al, La, Zr, and O from the extraction lists 31-34 and 31a of Embodiment 1. Next, the extraction unit 101 identifies the document ID and conductivity associated with that material name from those extraction lists 31-34 and 31a. Then, the extraction unit 101 extracts graphs, figures, tables, or text describing the conductivity from the document information D that has that document ID as extracted images.
[0220] The extraction unit 101 outputs information including the document ID, conductivity, and extracted image to the derivation unit 105 as display information candidates. Here, if the extraction unit 101 finds multiple material names through the material name search described above, it performs the same processing as described above for each of those multiple material names and outputs display information candidates to the derivation unit 105.
[0221] When the derivation unit 105 acquires multiple display information candidates, it narrows down the display information from those candidates based on the reliability of the conductivity included in those candidates. In other words, the derivation unit 105 selects a display information candidate from the multiple display information candidates that has a reliability of a threshold or higher as the display information. The derivation unit 105 then outputs the selected one or more display information to the image processing unit 106.
[0222] When the image processing unit 106 obtains one or more display information from the derivation unit 105, it combines that one or more display information to generate an editable document, and displays the editable document in the document window 42 of the document display screen 40.
[0223] Figure 21 shows an example of a document display screen 40 where editing materials are displayed.
[0224] The image processing unit 106 displays the edited material as a "Report on Li-AL-La-Zr-O Material" on the material display screen 40. This edited material is, for example, a document showing a conductivity ranking, as shown in Figure 21. In the example in Figure 21, since slider 2sb of the slider bar 46b corresponding to conductivity is at the far right, the data on conductivity is displayed across the entire material display screen 40. In the conductivity ranking, the conductivity and the document ID of the document information D in which that conductivity is described are shown in descending order of conductivity. For the 1st and 2nd ranked conductivitys, extracted images such as graphs describing that conductivity are also displayed.
[0225] The image processing unit 106 may also display a save button 43c on the document display screen 40. This save button 43c is for saving the edited document. When the user selects this save button 43c through an input operation to the input unit 11, the image processing unit 106 saves the edited document to a recording medium such as memory provided in the information providing device 100.
[0226] Figure 22 shows another example of the document display screen 40 where editing materials are displayed.
[0227] Next, the user operates the sliders 2sa and 2sc of slider bars 46a and 46c by performing input operations on the input unit 11. As a result, the image processing unit 106 moves these sliders 2sa and 2sc to the right. For example, the amount of movement to the right of slider 2sb, which corresponds to conductivity, is the largest, while the amount of movement to the right of slider 2sa, which corresponds to publication date, and slider 2sc, which corresponds to temperature, is smaller than that of slider 2sb.
[0228] In this case, the extraction unit 101 searches for a material name corresponding to the material condition "Li-Al-La-Zr-O" from the extraction lists 31-34 and 31a, etc., and identifies the document ID, publication date, conductivity, and temperature of the document information D associated with that material name. Then, the extraction unit 101 extracts the graph, figure, table, or text describing the conductivity from the document information D that has that document ID as an extracted image.
[0229] Next, the extraction unit 101 outputs information including the document ID, material name, publication date, conductivity, temperature, and extracted image to the derivation unit 105 as display information candidates. Here, if the extraction unit 101 finds multiple material names through the material name search described above, it performs the same processing as described above for each of those multiple material names and outputs display information candidates to the derivation unit 105.
[0230] When the derivation unit 105 acquires multiple display information candidates, it narrows down the display information from those candidates based on the reliability of the conductivity included in those candidates. In other words, the derivation unit 105 selects a display information candidate from the multiple display information candidates that has a reliability of a threshold or higher as the display information. The derivation unit 105 then outputs the selected one or more display information to the image processing unit 106.
[0231] When the image processing unit 106 obtains one or more display information from the derivation unit 105, it combines that information to generate an edited document and displays the edited document in the document window 42 of the document display screen 40. The edited document consists of the first document, the second document, and the third document, which will be described later. Specifically, based on the state of the slider bars 46a to 46f of the quantity adjustment window 46, the image processing unit 106 allocates half of the total area of the document window 42 to conductivity, one-quarter to the publication date, and the remaining one-quarter to temperature. Then, the image processing unit 106 displays the first document, which shows a conductivity ranking similar to the example in Figure 21, in the left half of the area of the document window 42. Furthermore, the image processing unit 106 displays a graph relating conductivity and temperature as a second source in the upper right quarter of the source window 42, and a graph relating conductivity and publication date as a third source in the lower right quarter of the source window 42. The second source is a graph showing temperature on the horizontal axis and conductivity on the vertical axis, with points plotted at positions corresponding to the temperature and conductivity of each material on the graph. The material name may also be labeled at these positions. The third source is a graph showing publication date on the horizontal axis and conductivity on the vertical axis, with multiple points plotted on the graph. Each of these points is plotted at a position corresponding to the publication date and conductivity of the bibliographic information D in which the material is described. The material name may also be labeled at these positions.
[0232] Figure 23 shows another example of the document display screen 40 where editing materials are displayed.
[0233] The user operates the sliders 2sa and 2sd to 2sf of slider bars 46a and 46d to 46f by performing input operations on the input unit 11, thereby changing the state of the volume adjustment window 46 in Figure 19. As a result, the image processing unit 106 moves these sliders 2sa and 2sd to 2sf to the right. For example, the amount of movement to the right of slider 2sd, which corresponds to structure, is the largest, while the amount of movement to the right of slider 2sa, which corresponds to publication date, slider 2se, which corresponds to research institution name, and slider 2sf, which corresponds to author name, is smaller than that of slider 2sd.
[0234] In this case, the extraction unit 101 searches for a material name corresponding to the material condition "Li-Al-La-Zr-O" from the extraction lists 31-34 and 31a, etc., and identifies the literature ID, publication date, conductivity, structure, name of research institution, and author name of the literature information D associated with that material name. Then, the extraction unit 101 extracts graphs, figures, or tables describing the structure from the literature information D that has that literature ID as extracted images.
[0235] The extraction unit 101 outputs information including the document ID, publication date, conductivity, structure, name of research institution, author name, and extracted image to the derivation unit 105 as display information candidates. Here, if the extraction unit 101 finds multiple material names through the material name search described above, it performs the same processing as described above for each of those multiple material names and outputs display information candidates to the derivation unit 105.
[0236] When the derivation unit 105 acquires multiple display information candidates, it narrows down the display information from those candidates based on the reliability of the conductivity included in those candidates. In other words, the derivation unit 105 selects a display information candidate from the multiple display information candidates that has a reliability of a threshold or higher as the display information. The derivation unit 105 then outputs the selected one or more display information to the image processing unit 106.
[0237] When the image processing unit 106 obtains one or more display information from the derivation unit 105, it combines that information to generate an edited document and displays the edited document in the document window 42 of the document display screen 40. The edited document consists of the fourth document, the fifth document, and the sixth document, which will be described later. Specifically, based on the state of the slider bars 46a to 46f of the volume adjustment window 46, the image processing unit 106 allocates half of the total area of the document window 42 to the structure and publication date, one-quarter to the author's name, and the remaining one-quarter to the research institution's name. The image processing unit 106 then displays the fourth document, which shows extracted images of the structure of each material in order of publication date, in the left half of the area of the document window 42. Furthermore, the image processing unit 106 displays information about the author's name as a fifth document in the upper right quarter of the document window 42, and displays information about the research institution's name as a sixth document in the lower right quarter of the document window 42.
[0238] In Document 4, the earlier the publication date of the bibliographic information D, the closer the extracted image of the structure described in that bibliographic information D is placed to the beginning. The extracted image may also be labeled with the bibliographic ID and publication year.
[0239] The fifth document displays the names of three authors who are ranked 1st, 2nd, and 3rd in terms of fame among the one or more author names included in one or more display information acquired by the image processing unit 106, along with information related to those three authors. The image processing unit 106 may generate the fifth document by referring to author data that shows the fame of multiple authors and information about those authors, for example. The author data may be stored on a server connected to the information providing device 100 via a communication network such as the Internet, or it may be stored in the image processing unit 106 or the information providing device 100.
[0240] The sixth document displays the names of the research institutions ranked first and second in terms of recognition among the one or more research institution names included in the one or more display information acquired by the image processing unit 106, along with information related to those two research institutions. The image processing unit 106 may generate the sixth document by referring to research institution data that shows the recognition of multiple research institutions and information about those research institutions, for example. The research institution data may be stored on a server connected to the information providing device 100 via a communication network such as the Internet, or it may be stored in the image processing unit 106 or the information providing device 100.
[0241] In this embodiment, the extraction unit 101 acquires material conditions and extracts information about each of the one or more materials that meet those material conditions from at least one bibliographic information D as a candidate for display information. The image processing unit 106 then acquires the weights of each of the multiple types of attributes of the material. The weights of each of the multiple types of attributes are, for example, the weights or biases of publication date, conductivity, temperature, structure, name of research institution, and author name. Next, the image processing unit 106 selects one or more display information candidates from the multiple display information candidates extracted by the extraction unit 101, each corresponding to a material having characteristic values for which a confidence level above a threshold has been derived, as display information. The image processing unit 106 then generates a second image that shows the display information corresponding to each of the multiple types of attributes from the one or more display information, in an amount corresponding to the weight of each of the multiple types of attributes, and outputs the second image to the display unit 13. The second image is, for example, the edited material in the material display screen 40 or material window 42 shown in Figures 21-23.
[0242] This allows the amount of display information corresponding to each of the multiple types of attributes shown in the second image to be changed by adjusting the bias, which is the weight of each of the multiple types of attributes. Therefore, the user can arbitrarily adjust the amount of display information for each attribute so that more display information is shown for attributes that the user is interested in, and less display information is shown for attributes that the user is not interested in. Furthermore, since this displayed information is information about materials with a confidence level above a threshold, the user can use this displayed information with confidence in tasks such as materials research. Display information for materials that correspond to material conditions is displayed, and these material conditions are, for example, conditions related to the element species contained in the material or the composition of the material. This allows the user to limit the one or more displayed information to materials of interest. In this embodiment, edited materials according to the user's wishes can be automatically generated as technical documents, thus reducing the man-hours required to create technical documents.
[0243] Figure 24 is a flowchart showing an example of the overall processing operation of the information providing device 100 in this embodiment.
[0244] (Step S210) First, the information providing device 100 performs a reliability derivation process. This reliability derivation process, as in Embodiment 1, is a process for deriving the reliability for each characteristic value of a plurality of materials, and may include the processes of steps S110 to S130 in Figure 16, or it may include the processes of steps S110 to S150.
[0245] (Step S220) The extraction unit 101 acquires material conditions in response to user input operations to the input unit 11. Furthermore, the image processing unit 106 acquires the bias of each of several types of attributes in response to user input operations to the input unit 11.
[0246] (Step S230) Next, the extraction unit 101 searches for a material name that matches the material conditions obtained in step S220. For example, the extraction unit 101 searches for the material name from the extraction lists 31-34 and 31a, etc., which are generated in the reliability derivation process in step S210.
[0247] (Step S240) Next, the extraction unit 101 generates candidate display information for each of the one or more material names found in step S230. That is, the extraction unit 101 identifies various information, including the literature ID associated with the material name, from the extraction lists 31-34 and 31a, etc. This various information includes the characteristic values of the material having that material name. Furthermore, the extraction unit 101 extracts graphs, figures, tables, or text describing the various information from the literature information D having that literature ID as extracted images. Then, the extraction unit 101 generates information including the literature ID, various information, and extracted images as candidate display information.
[0248] (Step S250) Next, the derivation unit 105 acquires multiple display information candidates and then narrows down the display information from those candidates based on the reliability of the conductivity included in those candidates. In other words, the derivation unit 105 selects a display information candidate from the multiple display information candidates that has a reliability of a threshold or higher as the display information.
[0249] (Step S260) Next, the image processing unit 106 obtains one or more display information from the derivation unit 105 and combines that one or more display information to generate a second image. In other words, the image processing unit 106 generates a second image in which, from the one or more display information, the display information corresponding to each of the multiple types of attributes is shown in an amount corresponding to the bias of each of the multiple types of attributes obtained in step S220.
[0250] (Step S270) The image processing unit 106 then displays the generated second image on the display unit 13.
[0251] The information processing apparatus according to one or more aspects has been described above based on the above embodiments and modifications. However, the present disclosure is not limited to these embodiments and modifications. Unless it deviates from the spirit of the present disclosure, various modifications conceivable by those skilled in the art made to the above embodiments and modifications, and configurations constructed by combining components from different embodiments or different modifications may also be included within the scope of the present disclosure.
[0252] For example, in each of the above embodiments and modifications, the document information D is electronic data of papers, but it may also be electronic data of textbooks, magazines, patent documents, or the like.
[0253] In each of the above embodiments and modifications, the information providing apparatus 100 is configured as a single apparatus, such as a personal computer, for example, but may be configured from a plurality of apparatuses. In this case, the extracting unit 101, the first information processing unit 102, the second information processing unit 103, the third information processing unit 104, the deriving unit 105, and the image processing unit 106 are not provided in a single apparatus, but are distributed among a plurality of apparatuses.
[0254] Each of the screens shown in FIGS. 13 to 15C and FIGS. 21 to 23 in the above embodiments does not include the document list window 21 shown in FIGS. 2A and 2B, but the document list window 21 may be included therein. In this case, the user may deselect an icon of document information D in the document list window 21 or select a new icon by performing an input operation on the input unit 11. In response to the selection or deselection, the image processing unit 106 may update the characteristic value graph 24 or the material window 42.
[0255] In each of the above embodiments and modifications, the types of characteristic values and attributes of the material extracted by the extracting unit 101 may be specified in accordance with an input operation performed by a user on the input unit 11. That is, the extracting unit 101 may acquire an input signal from the input unit 11, and extract the characteristic values and attributes of the types indicated by the input signal from the document information D. As a result, the extracting unit 101 may, for example, extract density or the like as a characteristic value instead of the conductivity of the material. The types of the extracted characteristic values and attributes may be predetermined.
[0256] In the above-described second embodiment, display information is narrowed down from a plurality of display information candidates based on the reliability of the characteristic values, but the narrowing down does not have to be performed. In this case, the image processing unit 106 may treat all the display information candidates generated by the extracting unit 101 as respective pieces of display information, and generate the second image by combining these pieces of display information.
[0257] In each of the above embodiments and modifications, each component may be configured by dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory. Here, the software program that realizes the information providing apparatus 100 of each of the above embodiments and modifications causes a computer to execute each step included in at least one flowchart shown in FIG. 16, FIG. 17, and FIG. 24.
[0258] It goes without saying that the present disclosure is not limited to the above embodiments and modifications. The following cases are also included in the present disclosure.
[0259] (1) Specifically, each of the above devices is a computer system consisting of a microprocessor, ROM, RAM, hard disk unit, display unit, keyboard, mouse, etc. A computer program is stored in the RAM or hard disk unit. Each device achieves its function by operating according to the computer program through the microprocessor. Here, a computer program is composed of a combination of multiple instruction codes that indicate commands to the computer in order to achieve a predetermined function.
[0260] (2) Some or all of the components constituting each of the above devices may be made up of a single system LSI (Large Scale Integration). A system LSI is a highly functional LSI manufactured by integrating multiple components onto a single chip, and specifically, it is a computer system that includes a microprocessor, ROM, RAM, etc. The RAM stores computer programs. The system LSI achieves its function by operating the microprocessor in accordance with its computer programs.
[0261] (3) Some or all of the components constituting each of the above devices may consist of a removable IC card or a standalone module. The IC card or module is a computer system consisting of a microprocessor, ROM, RAM, etc. The IC card or module may include the above-mentioned multi-functional LSI. The microprocessor operates according to a computer program, thereby enabling the IC card or module to achieve its function. The IC card or module may be tamper-resistant.
[0262] (4) The disclosure may also be the methods described above. These methods may be computer programs that implement them using a computer, or they may be digital signals consisting of computer programs.
[0263] This disclosure may also refer to a computer program or digital signal recorded on a computer-readable recording medium, such as a flexible disk, hard disk, CD-ROM, MO, DVD, DVD-ROM, DVD-RAM, BD (Blu-ray® Disc), semiconductor memory, etc. It may also refer to a digital signal recorded on one of these recording media.
[0264] This disclosure may also include the transmission of the computer program or digital signal via telecommunications lines, wireless or wired communication lines, networks such as the Internet, data broadcasting, etc.
[0265] This disclosure relates to a computer system comprising a microprocessor and memory, wherein the memory stores the aforementioned computer program, and the microprocessor may operate in accordance with the computer program.
[0266] The program or digital signal may be carried out by another independent computer system by recording it on a recording medium and transferring it, or by transferring the program or digital signal via a network or the like.
[0267] (others) An apparatus relating to one aspect of this disclosure may be an apparatus as shown below.
[0268] An extraction unit that extracts multiple material names, multiple values for a first property, and multiple values for a second property from at least one document, An introduction unit for calculating multiple distances, where each of the multiple distances is the distance between two different coordinates included in n different coordinates, and n is a natural number greater than or equal to 2. The aforementioned n distinct coordinates are the first coordinate corresponding to the first material name, ~, and the nth coordinate corresponding to the nth material name. The first coordinate is a pair of the first value of the first characteristic and the first value of the second characteristic, and the nth coordinate is a pair of the nth value of the first characteristic and the nth value of the second characteristic, The aforementioned multiple material names include the first material name, ~, the nth material name, The multiple values of the first characteristic include the first value, ~, and the nth value of the first characteristic. The multiple values of the second characteristic include the first value, ~, and the nth value of the second characteristic. An apparatus including an image processing unit that determines the display configuration of the first coordinate, ~, and the n coordinate on a two-dimensional plane based on the calculated plurality of distances, and outputs the determined display correspondence for the first coordinate and the determined display correspondence for ~ and the n coordinate to a display unit.
[0269] The first material name, ~, and the nth material name may be any of the multiple final material names listed in the column for final material names in Figure 11.
[0270] The first value of the first characteristic, ~, and the nth value of the first characteristic may be any of the multiple values listed in the column of conductivity values shown in Figure 11.
[0271] The first value of the second characteristic, ~, and the nth value of the second characteristic may be any of the values listed in the column of activation energy values shown in Figure 11.
[0272] The first coordinate, ~nth coordinate may be a coordinate where the activation energy value corresponding to ID=001 in Figure 11 is the x-axis coordinate and the conductivity value is the y-axis coordinate, or a coordinate where the activation energy value corresponding to ~, ID=n (not shown) is the x-axis coordinate and the conductivity value is the y-axis coordinate.
[0273] The first coordinate, ~, and the nth coordinate may be multiple coordinates indicated by × in Figure 13.
[0274] The representation of the first coordinate, ~, and the nth coordinate may be the intensity (number of points inside the circle) shown inside each of the multiple circles in Figure 14.
[0275] The plurality of distances may be: the distance between two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=001 and two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=002 in FIG. 11, ..., the distance between two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=001 and two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=n (not shown), the distance between two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=002 and two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=003, ..., the distance between two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=002 and two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=n (not shown), ..., and the distance between two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=(n-1) and two-dimensional data (activation energy value, conductivity value) corresponding to Document ID=n (not shown).
[0276] The number of the plurality of distances for n different coordinates is n(n-1) / 2. [Industrial Applicability]
[0277] The information providing apparatus of the present disclosure can appropriately provide information related to materials, and is useful for an apparatus or system for conducting material research, material development, or new material synthesis. [Description of Reference Numerals]
[0278] 1sa to 1sd Slider 2sa to 2sf Slider 11 Input unit 12 Document database 13 Display unit 20 Display screen 20a Graph display screen 21 Document list window 22 Extracted information table 23a Extraction start button 23b Supplement start button 23c Graph display start button 23d Confidence level display button 23e Return button 23f Hide Trust Level button 24. Characteristic Value Graph 25a~25d Slider bar 31 Extraction List 31a Extraction List 32 Extraction List 33 Extraction List 34 Extraction List 40 Material display screen 41 Material Window 42 Document Window 43a Document Display Button 43b Return button 43c Save button 46 Quantity adjustment window 46a~46f Slider bar 100 Information provision device 101 Extraction part 102 First Information Processing Unit 103 Second Information Processing Unit 104 Third Information Processing Unit 105 Derivation part 106 Image Processing Unit D Literature information D1 Main text data D2 Metadata
Claims
1. An extraction unit that extracts identification information for identifying each of multiple materials from at least one piece of literature information, characteristic values for each of the multiple materials, and attribute information indicating the attributes of each of the multiple materials, A derivation unit that derives the reliability of the characteristic values of the material based on the similarity of characteristic values between the material and one or more other materials and the attribute information of the material, An image processing unit that generates a first image which displays the characteristic values of each of the plurality of materials in a display manner corresponding to the reliability derived for the characteristic values of the materials, and in association with the identification information of the materials, and outputs the first image to a display unit. An information-providing device equipped with the following features.
2. An extraction unit that extracts identification information for identifying each of multiple materials and characteristic values for each of the multiple materials from at least one piece of literature information, A derivation unit that derives the reliability of the characteristic values of the material based on the similarity of characteristic values between the material and one or more other materials, An image processing unit that generates a first image which displays the characteristic values of each of the plurality of materials in a display manner corresponding to the reliability derived for the characteristic values of the materials, and in association with the identification information of the materials, and outputs the first image to a display unit. Equipped with, The extraction unit further, For each of the aforementioned multiple materials, attribute information indicating the attributes of the material is extracted from at least one of the aforementioned literature sources. The aforementioned derivation section is, Based on the similarity of the characteristic values in the material and the attribute information extracted for the material, the reliability of the characteristic values in the material is derived. Information provision device.
3. The aforementioned derivation section is, The reliability of the characteristic values in the material is derived by weighting and adding the similarity of the characteristic values in the material and the attribute values based on the attribute information extracted for the material. ru, The information providing device according to claim 2.
4. The attribute information indicates, as the attribute, the publication date of the document containing the identification information and characteristic values of the material corresponding to the attribute information, among the at least one document of the document. The aforementioned derivation section is, The reliability of the characteristic value is derived using the attribute value indicating the recency of the publication date. The information providing device according to claim 3.
5. The attribute value indicates that the more recent the publication date, the larger the value. The aforementioned derivation section is, The higher the similarity of the characteristic values in the material, and the larger the attribute value corresponding to the material, the greater the reliability value of the characteristic value in the material. The information providing device according to claim 4.
6. The attribute information indicates, as the attribute, the number of citations of the document containing the identification information and characteristic values of the material corresponding to the attribute information, among the at least one document information. The aforementioned derivation section is, The reliability of the characteristic value is derived using the attribute value corresponding to the number of citations. The information providing device according to claim 3.
7. The attribute value indicates that the value increases as the number of citations increases. The aforementioned derivation section is, The higher the similarity of the characteristic values in the material, and the larger the attribute value corresponding to the material, the greater the reliability value of the characteristic value in the material. The information providing device according to claim 6.
8. The attribute information indicates, as the attribute, the author of the document containing the identification information and characteristic values of the material corresponding to the attribute information, among the at least one document. The aforementioned derivation section is, The reliability of the characteristic value is derived using the attribute value, which is determined by whether or not the author of the aforementioned document information is the same as the author of one or more other documents. The information providing device according to claim 3.
9. The attribute value indicates a larger value the more authors among the other one or more sources of bibliographic information are different from the authors of the aforementioned bibliographic information. The aforementioned derivation section is, The higher the similarity of the characteristic values in the material, and the larger the attribute value corresponding to the material, the greater the reliability value of the characteristic value in the material. The information providing device according to claim 8.
10. The attribute information indicates the synthesis method of the material corresponding to the attribute information as the attribute, The aforementioned derivation section is, The reliability of the characteristic value is derived using the attribute value corresponding to the degree of similarity between the synthesis method of the aforementioned material and the synthesis method of one or more other materials. The information providing device according to claim 3.
11. The attribute value indicates a larger value the greater the degree of similarity between the synthesis method of the material and the synthesis methods of each of the other one or more materials. The aforementioned derivation section is, The higher the similarity of the characteristic values in the material, and the larger the attribute value corresponding to the material, the greater the reliability value of the characteristic value in the material. The information providing device according to claim 10.
12. The method for synthesizing the material includes at least one of the temperature conditions, time conditions, and type of apparatus used in the synthesis of the material. The information providing device according to claim 10.
13. An extraction unit that extracts identification information for identifying each of multiple materials and characteristic values for each of the multiple materials from at least one piece of literature information, A derivation unit that derives the reliability of the characteristic values of the material based on the similarity of characteristic values between the material and one or more other materials, An image processing unit that generates a first image which displays the characteristic values of each of the plurality of materials in a display manner corresponding to the reliability derived for the characteristic values of the materials, and in association with the identification information of the materials, and outputs the first image to a display unit. Equipped with, The extraction unit further acquires material conditions and, for each of the one or more materials that meet the material conditions, extracts information about the material as a candidate for display information from the at least one document information. The aforementioned image processing unit further, Obtain the weights of each of the multiple attributes of the material. From the plurality of display information candidates extracted by the extraction unit, one or more display information candidates corresponding to materials having characteristic values for which the reliability of the above-the-threshold value is derived are selected as display information. A second image is generated in which, from the one or more of the aforementioned display information, the display information corresponding to each of the multiple types of attributes is shown in an amount corresponding to the weight of each of the multiple types of attributes, and the second image is output to the display unit. Information provision device.
14. A method of providing information that is performed by one or more computers, Extract identification information for identifying each of multiple materials, characteristic values for each of the multiple materials, and attribute information indicating the attributes of each of the multiple materials from at least one piece of literature information. Based on the similarity of characteristic values between the material and one or more other materials and the attribute information of the material, the reliability of the characteristic values of the material is derived. (i) The characteristic values of each of the plurality of materials are displayed in a display manner corresponding to the reliability derived for the characteristic values of the material, and (ii) a first image is generated that is displayed in association with the identification information of the material, and the first image is output to the display unit. Information provision method.
15. Extract identification information for identifying each of multiple materials from at least one piece of literature information, characteristic values for each of the multiple materials, and attribute information indicating the attributes of each of the multiple materials. Based on the similarity of characteristic values between the material and one or more other materials and the attribute information of the material, the reliability of the characteristic values of the material is derived. (i) The characteristic values of each of the plurality of materials are displayed in a display manner corresponding to the reliability derived for the characteristic values of the material, and (ii) a first image is generated that is displayed in association with the identification information of the material, and the first image is output to the display unit. A program that causes a computer to perform a task.
Citation Information
Patent Citations
System for evaluating reliability of document
JP2008152701A
Numerical-value retrieving device, numerical-value retrieving method, and numerical-value retrieving program
JP2020080087A
Material characteristics prediction device and material characteristics prediction method
JP2020128962A
Similar material search system, test apparatus and computer program
JP2020190863A
Assistance device, generation device, analysis device, assistance method, generation method, analysis method,and program
WO2021039175A1