Learning device, estimation device, program, learning method, and estimation method
The learning device and method address the issue of varying numerical value distributions by normalizing and refining units based on object and attribute combinations, enhancing the accuracy of quantitative expression analysis.
Patent Information
- Application Number
- JP2024576947
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-04-21
AI Technical Summary
Conventional models fail to account for distribution differences in numerical values across different objects and attributes, leading to inaccurate training and estimation when comparing product specifications.
A learning device and method that identifies quantitative expressions, normalizes numerical values and units, and refines units to correspond to specific object and attribute combinations, using a quantitative expression language model for accurate estimation.
Enables learning and estimation to consider distribution differences for each object and attribute, improving the accuracy of quantitative expression analysis.
Smart Images

Figure 0007752793000001 
Figure 0007752793000002 
Figure 0007752793000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a learning device, an estimation device, a program, a learning method, and an estimation method. [Background technology]
[0002] There has long been a need to search for quantitative expressions contained in data such as documents, when comparing product specifications, such as when consumers are considering a purchase, when manufacturing companies are procuring parts, or when automatically determining whether a product complies with a standard or regulation; when customers or sales representatives input required specifications and search for similar products from the past; or when inputting the specifications of one's own company's products and creating a comparison table with similar products from other companies, and when searching for similar specifications.
[0003] In response to such needs, for example, Patent Document 1 describes a device that extracts and learns numerical values as quantities rather than unknown words in a method of learning embedding vectors from text using a neural network. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2021-149935 Summary of the Invention [Problem to be solved by the invention]
[0005] Here, when a model is pre-trained, it is often trained using a large amount of document data, but there are cases where a certain physical quantity is used in the same unit for different attributes of different objects. In such cases, the distribution of numerical values varies greatly depending on the object and attribute, so in order to train an accurate model, it is necessary to perform training taking these distribution differences into account. However, the conventional techniques do not take into consideration the differences in distribution for each object and attribute.
[0006] Therefore, one or more aspects of the present disclosure aim to enable learning and estimation to be performed taking into account differences in distribution for each object and attribute. [Means for solving the problem]
[0007] A learning device according to one aspect of the present disclosure is characterized by comprising: a quantitative expression identification unit that identifies, from original learning data, quantitative expressions that represent quantities using numerical values and units; a numerical normalization unit that normalizes the numerical values; a unit normalization unit that normalizes the units; a unit refinement unit that identifies, in the original learning data, objects that are physical entities and attributes that are properties of the objects indicated by the numerical values and the units, and converts the normalized units into refinement units that uniquely correspond to the combination of the identified objects, the identified attributes, and the normalized units; and a quantitative expression learning unit that uses data including the refinement units and the normalized numerical values to learn a quantitative expression language model for estimating the normalized numerical values.
[0008] An estimation device according to one aspect of the present disclosure is characterized by comprising: an estimation target data acquisition unit that acquires estimation target data, using a predetermined expression format and unit, requesting estimation of a numerical value at a position where the predetermined expression format is located; a unit normalization unit that normalizes the unit; a unit refinement unit that identifies, in the estimation target data, an object that is a physical entity and an attribute that is a property of the object indicated by the predetermined expression format and the unit, and converts the normalized unit into a refinement unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; and a quantitative expression estimation unit that estimates a numerical value at a position where the predetermined expression format is located by inputting data including the predetermined expression format and the converted refinement unit into a quantitative expression language model trained using data including a refinement unit and a normalized numerical value.
[0009] A program according to a first aspect of the present disclosure causes a computer to function as a quantitative expression identification unit that identifies, from original training data, quantitative expressions that represent quantities using numerical values and units; a numerical normalization unit that normalizes the numerical values; a unit normalization unit that normalizes the units; a unit refinement unit that identifies, in the original training data, objects that are physical entities and attributes that are properties of the objects indicated by the numerical values and the units, and converts the normalized units into refinement units that uniquely correspond to the combination of the identified objects, the identified attributes, and the normalized units; and a quantitative expression learning unit that uses data including the refinement units and the normalized numerical values to learn a quantitative expression language model for estimating the normalized numerical values.
[0010] A program according to a second aspect of the present disclosure causes a computer to function as an estimation target data acquisition unit that acquires estimation target data, using a predetermined expression format and unit, requesting estimation of a numerical value at a position where the predetermined expression format is placed; a unit normalization unit that normalizes the unit; a unit refinement unit that identifies, in the estimation target data, an object that is a physical entity and an attribute that is a property of the object indicated by the predetermined expression format and the unit, and converts the normalized unit into a refinement unit that uniquely corresponds to the combination of the identified object, the identified attribute, and the normalized unit; and a quantitative expression estimation unit that estimates a numerical value at a position where the predetermined expression format is placed by inputting data including the predetermined expression format and the converted refinement unit into a quantitative expression language model trained using data including the refinement unit and a normalized numerical value.
[0011] A learning method according to one aspect of the present disclosure includes: The quantifier specifying part is Identifying quantitative expressions that represent quantities using numerical values and units from the original learning data; The numerical normalization part is normalizing said values; The unit normalization part is normalizing the units; The unit refinement sectionidentifying, in the original learning data, an object that is a physical entity and an attribute that is a property of the object, which is indicated by the numerical value and the unit, and converting the normalized unit into a refined unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; The Quantitative Expression Learning Department The method is characterized in that a quantitative expression language model for estimating the normalized numerical value is trained using data including the refinement unit and the normalized numerical value.
[0012] An estimation method according to one aspect of the present disclosure includes: The estimation target data acquisition unit Using a predetermined expression format and unit, obtain estimation target data that requests estimation of a numerical value at a position where the predetermined expression format is arranged; The unit normalization part is normalizing the units; The unit refinement section In the estimation target data, an object that is a physical entity and an attribute that is a property of the object, which is indicated by the predetermined expression format and the unit, are identified, and the normalized unit is converted into a refinement unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; The quantitative expression estimation unit The method is characterized in that data including the predetermined expression format and the converted elaboration unit is input into a quantitative expression language model trained using data including the elaboration unit and a normalized numerical value, thereby estimating a numerical value at a position where the predetermined expression format is located. [Effects of the Invention]
[0013] According to one or more aspects of the present disclosure, learning and estimation can be performed taking into account differences in distribution for each object and attribute. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a block diagram schematically illustrating a configuration of a learning estimation system according to first to third embodiments. [Figure 2] FIG. 1 is a block diagram schematically illustrating a configuration of a learning device according to first and third embodiments. [Figure 3]10(A) to 10(C) are schematic diagrams for explaining the processing in the numerical expression specification unit. [Figure 4] FIG. 2 is a schematic diagram showing a first example of sentence structure data. [Figure 5] FIG. 10 is a schematic diagram illustrating an example of a numerical normalization result. [Figure 6] FIG. 10 is a schematic diagram illustrating an example of unit normalization information. [Figure 7] FIG. 10 is a schematic diagram illustrating an example of a unit normalization result. [Figure 8] FIG. 10 is a schematic diagram illustrating an example of target attribute-specific detailed unit information. [Figure 9] FIG. 2 is a block diagram illustrating a schematic configuration of a unit refining unit. [Figure 10] FIG. 2 is a schematic diagram illustrating an example of attribute synonym dictionary information. [Figure 11] FIG. 10 is a schematic diagram illustrating an example of a neighborhood attribute detection result. [Figure 12] 10A and 10B are schematic diagrams for explaining the processing in the dependency attribute detection unit. [Figure 13] 10A and 10B are schematic diagrams for explaining the processing in the table attribute detection unit. [Figure 14] FIG. 10 is a schematic diagram illustrating an example of a result of detecting attributes in a table. [Figure 15] FIG. 10 is a schematic diagram showing a second example of document structure data. [Figure 16] FIG. 10 is a schematic diagram illustrating an example of target synonym dictionary information. [Figure 17] FIG. 10 is a schematic diagram illustrating an example of a result of detecting an object within a document. [Figure 18] FIG. 10 is a schematic diagram illustrating an example of a result of intra-sentence object detection. [Figure 19] FIG. 10 is a schematic diagram illustrating an example of word hierarchical relationship information. [Figure 20] FIG. 10 is a schematic diagram illustrating an example of a result of detecting an object in a table. [Figure 21] 10A and 10B are schematic diagrams for explaining the processing in the part-whole relationship determination unit. [Figure 22] FIG. 10 is a schematic diagram illustrating an example of partial-whole information. [Figure 23] 10A and 10B are schematic diagrams for explaining the processing in the differential expression determination unit. [Figure 24] FIG. 2 is a schematic diagram illustrating an example of differential expression pattern dictionary information. [Figure 25] 10A and 10B are schematic diagrams for explaining the processing in the unit conversion unit. [Figure 26] FIG. 2 is a block diagram illustrating an example of a hardware configuration. [Figure 27] 4 is a flowchart showing the operation of the learning device in the first embodiment. [Figure 28] 4 is a flowchart showing the operation of the quantitative expression learning unit in the first embodiment. [Figure 29] 1 is a block diagram schematically illustrating a configuration of an estimation device according to a first embodiment. [Figure 30] 4 is a flowchart showing the operation of the estimation device in the first embodiment. [Figure 31] FIG. 10 is a block diagram schematically illustrating the configuration of a learning device according to a second embodiment. [Figure 32] 10 is a flowchart showing the operation of the learning device in the second embodiment. [Figure 33] FIG. 10 is a block diagram schematically illustrating the configuration of an estimation device according to a second embodiment. [Figure 34] 10 is a flowchart showing the operation of the estimation device in the second embodiment. [Figure 35] 11 is a flowchart showing the operation of the quantitative expression learning unit in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] Embodiment 1 FIG. 1 is a block diagram schematically illustrating a configuration of a learning estimation system 100 according to the first embodiment. The learning and estimation system 100 includes a learning device 110 and an estimation device 150 . The learning device 110 and the estimation device 150 are connected to a network 101 such as the Internet.
[0016] FIG. 2 is a block diagram schematically showing the configuration of learning device 110 according to the first embodiment. The learning device 110 includes a learning source data acquisition unit 111, a quantitative expression identification unit 112, a numerical normalization unit 113, a unit normalization information storage unit 114, a unit normalization unit 115, a target attribute-specific refinement unit storage unit 116, a unit refinement unit 117, a quantitative expression learning unit 118, a quantitative expression language model storage unit 119, and a communication unit 120.
[0017] The learning original data acquisition unit 111 acquires learning original data, which is data that serves as the basis for learning in the learning device 110. The learning original data acquisition unit 111 may acquire the learning original data, for example, via an input unit (not shown), or may acquire the learning original data from another device via the communication unit 120. The acquired learning original data is provided to the quantification expression identification unit 112.
[0018] The quantitative expression identification unit 112 identifies, from the original learning data, quantitative expressions that represent quantities using numerical values and units.
[0019] 3A to 3C are schematic diagrams for explaining the processing in the numerical expression specification unit 112. In FIG. For example, when the original learning data is learning text data as shown in Figure 3(A), the quantitative expression identification unit 112 extracts a sentence of text shown in the learning text data as input text. Here, the quantitative expression identification unit 112 extracts a sentence of text using paragraphs and periods. In the example shown in Figure 3(A), the input texts are "New River" and "A tributary of the Kanawha River, which belongs to the Mississippi River system, with a total length of approximately 515 km."
[0020] The quantitative expression identification unit 112 divides the input text into words using a known morphological analysis technique, identifies the parts of speech of the divided words, and then identifies, from the identified parts of speech, a portion where a numerical value and a unit are consecutive, as a quantitative expression.
[0021] FIG. 3(B) is a schematic diagram showing a numerical expression identification result 112a#1, which is a result of identifying a numerical expression from the input text "New River." The quantitative expression identification result 112a#1 is data in a table format including a notation row 112b#1, a part of speech row 112c#1, and a division row 112d#1.
[0022] The notation line 112b#1 shows the words divided by the morphological analysis technique in the order in which they appear in the input text. Part-of-speech row 112c#1 indicates the part of speech of the word shown in the same column of notation row 112b#1. The division row 112d#1 indicates the division of the words shown in the same column of the notation row 112b#1. The division here is a division of quantitative expressions, which are "numbers" or "units." "-" indicates that there is no division. As shown in FIG. 3(B), the input text "New River" does not contain any quantitative expressions, so the division row 112d#1 is all "-."
[0023] FIG. 3(C) is a schematic diagram showing a numerical expression identification result 112a#2, which is the result of identifying a numerical expression from the input text "It is a tributary of the Kanawha River, which belongs to the Mississippi River system, and has a total length of about 515 km."
[0024] In the input text "A tributary of the Kanawha River, belonging to the Mississippi River system, with a total length of approximately 515 km," the numerical value "515" and the unit "km" are consecutive, so the quantitative expression identification unit 112 identifies these as a quantitative expression. Therefore, in the quantitative expression identification result 112a#2, "numerical value" is stored in the section row 112d#2 of the column of the word "515," and "unit" is stored in the section row 112d#2 of the column of the word "km."
[0025] 3A to 3C show an example in which the original training data is text data for training, but the original training data may be text data or document structure data. For example, if an explanatory article available on the Web is written in a markup language such as HTML (HyperText Markup Language), the tag information can be used to extract the text data for training and the document structure data from the original training data using known techniques.
[0026] For example, as shown in Figure 4, quantitative expression identification results 112a#1 and 112a#2 as shown in Figures 3(B) and (C) can be generated from sentence structure data containing items and content. In the sentence structure data, the quantitative expression specification unit 112 only needs to treat the content of one item as one sentence. However, if the content of one item includes a table, one line included in the table is treated as one sentence.
[0027] Then, the numerical expression specification unit 112 provides the numerical expression specification result indicating the numerical expression specified as above to the numerical value normalization unit 113 .
[0028] The numeric normalization unit 113 normalizes the numeric values. For example, the numeric normalization unit 113 normalizes the numeric values included in the quantitative expression identified by the quantitative expression identification unit 112 by converting them into exponential notation. The numeric normalization result obtained by the numeric normalization unit 113 is provided to the unit normalization unit 115.
[0029] For example, if the unit included in the quantitative expression identified by the quantitative expression identification unit 112 is a predetermined unit, the numerical value normalization unit 113 converts the numerical value that is continuous with that unit into an exponential expression. The predetermined units here are units that are to be refined by the unit refinement section 117, which will be described later.
[0030] Conversion to exponential notation is performed to ensure that the number of digits in the integer part of a number is the same. Here, conversion to exponential notation is performed so that the number is represented as a single integer part and a decimal part. For example, the number "515" is converted to exponential notation consisting of a mantissa of "5.15", the string "[EXP]" indicating that conversion to exponential notation has been performed, and "+02" indicating the exponent of 10 to be multiplied by the mantissa.
[0031] FIG. 5 is a schematic diagram showing a numerical value normalization result 113a#2, which is the result of converting the numerical value included in the quantitative expression identification result 112a#2 shown in FIG. 3(C) into an exponential expression by the numerical value normalization unit 113. As shown in Figure 5, in the numeric normalization result 113a#2, the numeric value "515" contained in the quantitative expression identification result 112a#2 shown in Figure 3(C) is converted into an exponential expression consisting of a mantissa of "5.15", a string of characters "[EXP]", and an exponent of "+02".
[0032] The unit normalization information storage unit 114 stores unit normalization information, which is information for normalizing units.
[0033] FIG. 6 is a schematic diagram illustrating an example of unit normalization information. The unit normalization information 114a shown in FIG. 6 includes a unit row 114b, a standard unit row 114c, and an exponent adjustment row 114d.
[0034] Units row 114b stores units that are to be normalized to standard units. The standard unit row 114c stores standard units that are converted to normalize the units stored in the same column of the unit row 114b. For example, as shown in Figure 6, "km", "cm", and "mm" are converted to "m" as the standard unit.
[0035] The exponent adjustment row 114d stores the value of the exponent to be adjusted in the exponential representation of a numerical value adjacent to a unit to be converted in the input text when the unit stored in the same column of the unit row 114b is converted to the standard unit stored in the same column of the standard unit row 114c. As shown in FIG. 6, when converting "km" to "m," "1 km" is "1000 m," so the number of exponents in the exponential representation must be increased by three. Therefore, the exponent adjustment row 114d in the same column as "km" and "m" stores "+03," which is a string formed by concatenating "+," which indicates addition, with "03," which is the value to be added. Similarly, the exponent adjustment row 114d in the same column as "mm" and "m" stores "-03," which is a string formed by concatenating "-," which indicates subtraction, with "03," which is the value to be subtracted.
[0036] As described above, the unit normalization information associates the unit to be converted, the standard unit to be converted, and the adjustment method of the exponent to be adjusted during conversion.
[0037] The unit normalization unit 115 normalizes the units. For example, the unit normalization unit 115 normalizes the units included in the numerical value normalization result from the numerical value normalization unit 113 by referring to the unit normalization information stored in the unit normalization information storage unit 114. The unit normalization result, which is the result of unit normalization by the unit normalization unit 115, is provided to the unit refinement unit 117.
[0038] FIG. 7 is a schematic diagram showing a unit normalization result 115a#2 obtained by normalizing the units included in the numeric value normalization result 113a#2 shown in FIG. As shown in Figure 7, in the unit normalization result 115a#2, the unit "km" included in the numerical value normalization result 113a#2 shown in Figure 5 is converted to the standard unit "m", and the exponent "+02" is converted to the exponent "+05".
[0039] The object attribute-specific detailed unit storage unit 116 stores object attribute-specific detailed unit information, which is information for detailing a standard unit according to an object and an attribute.
[0040] FIG. 8 is a schematic diagram showing an example of target attribute-specific detailed unit information. The target attribute-specific detailed unit information 116a shown in FIG. 8 is information in a table format including a target column 116b, an attribute column 116c, a standard unit column 116d, and a detailed unit column 116e.
[0041] The object column 116b stores objects, which are physical entities whose attributes can be expressed quantitatively. The attribute column 116c stores attributes that are properties of an object that can be expressed quantitatively. The standard unit column 116d stores the standard unit used to express the target attribute in terms of quantity. The refinement unit column 116e stores refinement units for converting standard units corresponding to a combination of an object and an attribute. The refinement unit may be any character string as long as it is uniquely determined by the combination of the object, the attribute, and the standard unit. Here, the refinement unit is refinement unit identification information that can identify the combination of the object, the attribute, and the standard unit.
[0042] The unit refinement unit 117 analyzes the unit normalization result from the unit normalization unit 115, detects combinations of objects and attributes, and converts standard units of exponential expressions representing the quantities of the detected objects and attributes into refinement units corresponding to the objects, attributes, and standard units, while suppressing erroneous conversion to refinement units.
[0043] For example, the unit refinement unit 117 identifies, in the original learning data, objects that are physical entities and attributes that are properties of the objects and are indicated by numerical values and units, and converts the normalized units into refinement units that uniquely correspond to the combination of the identified objects, the identified attributes, and the normalized units. Specifically, the unit refinement unit 117 identifies refinement units that uniquely correspond to the combination of the identified objects, the identified attributes, and the normalized units by referring to object attribute-specific refinement unit information that associates multiple objects, multiple attributes, multiple units, and multiple refinement units, respectively.
[0044] Here, when a predetermined first word is included in the vicinity of the quantitative expression, the unit refinement section 117 identifies the first word as an attribute. Furthermore, when a predetermined first word is included in a plurality of words that have a modification relationship with a quantitative expression, the unit refining section 117 identifies the first word as an attribute. Furthermore, when a quantitative expression is included in a table in the original data for learning, if a predetermined first word is included in at least one of the row heading name and the column heading name of the table, the unit detailing unit 117 identifies the first word as an attribute. The unit refining section 117 can resolve spelling variations by replacing synonyms of the first word with the first word and then identifying the attribute.
[0045] The unit detailing unit 117 identifies a predetermined second word as a target when the original data for learning is document structure data having items and contents corresponding to the items, and the contents include quantitative expressions, and the title of the item includes the predetermined second word. Furthermore, when a predetermined second word is included in the vicinity of the quantitative expression, the unit refinement section 117 identifies the second word as the target. Furthermore, when a quantitative expression is included in a table in the original data for learning, if a predetermined second word is included in at least one of the row heading name and the column heading name of the table, the unit detailing unit 117 identifies the second word as a target. The unit detailing section 117 can resolve spelling variations by replacing synonyms of the second word with the second word and then identifying the target.
[0046] However, the unit detailing unit 117 does not convert the normalized unit into a detailed unit if a predetermined word indicating a part of the identified object, or a predetermined word indicating the whole including the identified object, is included near the quantitative expression or among multiple words that have a dependency relationship with the quantitative expression. In addition, the unit detailing unit 117 does not convert a normalized unit into a detailed unit if a predetermined word indicating a difference is included in the vicinity of the quantitative expression or among multiple words that have a dependency relationship with the quantitative expression. In this way, the unit refining section 117 can prevent erroneous conversion into refinement units.
[0047] FIG. 9 is a block diagram showing a schematic configuration of the unit refining section 117. As shown in FIG. The unit refining unit 117 includes an attribute detecting unit 130 , an object detecting unit 134 , an erroneous conversion suppressing unit 138 , and a unit converting unit 141 .
[0048] The attribute detection unit 130 detects attributes from words included in the unit normalization result from the unit normalization unit 115 . The attribute detection unit 130 includes a neighborhood attribute detection unit 131 , a dependency attribute detection unit 132 , and an in-table attribute detection unit 133 .
[0049] The neighboring attribute detection unit 131 detects attributes from words included in a predetermined range (here, the number of words) from the exponential expression in the unit normalization result. Here, words that match attributes stored in the target attribute-specific detailing unit information stored in the target attribute-specific detailing unit storage unit 116 are detected. If there are multiple attributes within the predetermined range, the neighboring attribute detection unit 131 detects the attribute that is closer to the exponential representation.
[0050] Furthermore, the neighboring attribute detection unit 131 can correct spelling variations of attributes by using attribute synonym dictionary information 131a as shown in FIG. The attribute synonym dictionary information 131a is information in a table format that includes an attribute column 131b and an attribute synonym column 131c. The attribute column 131b stores the attribute stored in the target attribute-specific detailed unit information. The attribute synonym column 131c stores synonyms of the attributes stored in the attribute column 131b.
[0051] By referring to the attribute synonym dictionary information 131a, if the unit normalization result contains a synonym stored in the attribute synonym column 131c, the neighboring attribute detection unit 131 can correct spelling variations of the attribute by converting the synonym into an attribute stored in the attribute column 131b.
[0052] Then, the neighborhood attribute detection unit 131 stores the "attribute" in the row of the same column as the detected attribute in the unit normalization result. If the unit normalization result includes multiple exponential expressions, identification information for identifying which exponential expression it is is added to the detected attribute's category.
[0053] FIG. 11 is a schematic diagram showing a neighborhood attribute detection result 131d#2, which is the result of the neighborhood attribute detection unit 131 detecting neighborhood attributes for the unit normalization result 115a#2 shown in FIG. As shown in FIG. 11, in the neighborhood attribute detection result 131d#2, "attribute" is stored in the division row of the detected attribute "total length". Then, the neighboring attribute detection unit 131 provides the neighboring attribute detection result to the dependency attribute detection unit 132 .
[0054] The dependency attribute detection unit 132 detects attributes from words that have a dependency relationship with the exponential expression included in the neighborhood attribute detection result from the neighborhood attribute detection unit 131. When an attribute is detected from a word having a dependency relationship with a certain exponential expression, if another attribute has already been detected for that exponential expression by the neighboring attribute detection unit 131, the attribute detected from the word having the dependency relationship is used, and the word detected from the neighborhood is excluded from the classification as an attribute. In other words, the attribute detected from the dependency relationship has priority over the attribute detected from the neighborhood relationship.
[0055] 12A and 12B are schematic diagrams for explaining the processing in the dependency attribute detection unit 132. FIG. Figure 12(A) is a schematic diagram that explains the process of detecting dependency attributes for words contained in a line of notation in the neighboring attribute detection results, in which quantitative expressions have been identified, numerical values have been normalized, units have been normalized, and neighboring attributes have been detected, using the input text "The total width of Maenami Tunnel is 5.5 m and its height is 3.7 m."
[0056] 12(A), the exponential expressions "5.5", "[EXP]", "+00", and "m" have dependency relationships with the words "total width", "is", "at", and "aru". The dependency attribute detection unit 132 detects "width" as an attribute from these dependency-related words by referring to the target attribute-specific refinement unit information stored in the target attribute-specific refinement unit storage unit 116.
[0057] Similarly, the dependency attribute detection unit 132 detects the attribute "height" from the words "height", "is", "de", and "aru" which are in a dependency relationship with the exponential expressions "3.7", "[EXP]", "+00", and "m". The dependency attribute detection unit 132 detects "height" as an attribute from these dependency words by referring to the target attribute-specific refinement unit information stored in the target attribute-specific refinement unit storage unit 116.
[0058] The modification attribute detection unit 132 may also correct spelling variations of attributes by referring to the attribute synonym dictionary information 131a.
[0059] FIG. 12(B) is a schematic diagram showing a dependency attribute detection result 132a#3, which is the result of the dependency attribute detection process described with reference to FIG. 12(A). In the dependency attribute detection result 132a#3, "Attribute" is stored in the dividing line between "Total Width", which is an attribute detected from the dependency relationships of the exponential expressions "5.5", "[EXP]", "+00", and "m", and "Height", which is an attribute detected from the dependency relationships of the exponential expressions "3.7", "[EXP]", "+00", and "m". In the modification attribute detection result 132a#3, since the input text contains two numerical expressions (exponential expressions), in order to distinguish between them, a number indicating the order in which they appear in the input text is added to the "attribute." In other words, "attribute 1" indicates the attribute of the numerical expression (exponential expression) that appears first in the input text, and "attribute 2" indicates the attribute of the numerical expression (exponential expression) that appears second in the input text.
[0060] Then, the dependency attribute detection unit 132 provides the dependency attribute detection result to the in-table attribute detection unit 133 .
[0061] When a table is included in the document structure data as the original data for learning, the table attribute detection unit 133 detects the attributes of the input text extracted from the table from the row header names and column header names of the table.
[0062] 13A and 13B are schematic diagrams for explaining the processing in the table attribute detection unit 133. FIG. If the document structure data includes a table such as that shown in FIG. 13(A), the input text "X switch 250V1A" is extracted from the first line of the table. In such input text, even if quantitative expressions are identified, numerical values are normalized, units are normalized, neighbor attributes are detected, and dependency attributes are detected, no attributes are detected.
[0063] Here, the table attribute detection unit 133 checks the row item names and column item names of the table from which the input text was extracted, and detects "voltage" as an attribute in the same column for "2.5," "[EXP]," "+02," and "V," which are exponential expressions of "250V," and detects "current" in the same column for "1.0," "[EXP]," "+00," and "A," which are exponential expressions of "1A." Then, the in-table attribute detection unit 133 adds the detected attribute as an extended attribute to the dependency attribute detection result.
[0064] 13(B), there are cases where units are written in the row or column heading names of a table, and only numerical values are shown in the table. In such cases, the in-table attribute detection unit 133 adds the units included in the row or column heading names to the numerical values in the table, performs preprocessing to make the format equivalent to that of FIG. 13(A), and then performs the same in-table attribute keyword detection process as above.
[0065] Figure 14 is a schematic diagram showing the in-table attribute detection result 133a#4 in which extended attributes are added to the dependency attribute detection result in which quantitative expressions are identified, numerical values are normalized, units are normalized, neighboring attributes are detected, and dependency attributes are detected using the input text ``X switch 250V1A.''
[0066] In the table attribute detection result 133a#4, in addition to the notation row 133b#4, part of speech row 133c#4, and division row 133d#4, an extended notation row 133e#4, an extended part of speech 133f#4, and an extended division 133g#4 are added, and the attribute words, parts of speech, and divisions detected from the row and column item names of the table are stored. The expanded notation row 133e#4, the expanded part of speech 133f#4, and the expanded section 133g#4 store information about the detected attributes in the order in which they were detected, in other words, in the order in which they appear in the input text.
[0067] Then, the in-table attribute detection unit 133 provides the in-table attribute detection result, which is the detection result of the in-table attribute, to the object detection unit 134.
[0068] The object detection unit 134 receives the detection result from the attribute detection unit 130 and detects objects. The object detection unit 134 includes an intra-document object detection unit 135 , an intra-sentence object detection unit 136 , and an intra-table object detection unit 137 .
[0069] The document object detection unit 135 detects objects of the input text extracted from the training source data from the title of the training source data. Here, words that match objects stored in the target attribute-specific detail unit information stored in the target attribute-specific detail unit storage unit 116 are detected.
[0070] For example, if the original data for learning is document structure data as shown in Figure 4, the document object detection unit 135 uses "river" extracted from the content of the article title as the target of the table attribute detection result from the table attribute detection unit 133, which is generated based on the input text extracted from the document structure data.
[0071] 15, the document object detection unit 135 uses "Device" extracted from the contents of the document title as the target of the in-table attribute detection result from the in-table attribute detection unit 133, which is generated based on the input text extracted from the document structure data. Furthermore, the document object detection unit 135 uses "Model" extracted from the contents of the chapter / section title as the target of the in-table attribute detection result from the in-table attribute detection unit 133, which is generated based on the input text extracted from the chapter / section. Furthermore, the document object detection unit 135 uses "Switch" extracted from the contents of the table title as the target of the in-table attribute detection result from the in-table attribute detection unit 133, which is generated based on the input text extracted from the table. In other words, when the original data for learning is document structure data, the document object detection unit 135 regards the object extracted from the title of the document structure data as the object in the input text extracted from the range covered by the title of the document structure data.
[0072] Furthermore, the document object detection unit 135 can correct spelling variations of the object by using the object synonym dictionary information 135a as shown in FIG. The target synonym dictionary information 135a is information in a table format that includes a target column 135b and a target synonym column 135c. The target column 135b stores the target stored in the target attribute-specific detailed unit information. The target synonym column 135c stores the target synonym stored in the target column 135b.
[0073] By referring to the target synonym dictionary information 135a, if the title of the original data for learning contains a synonym stored in the target synonym column 135c, the document target detection unit 135 can correct the spelling variation of the target by converting the synonym to the target stored in the target column 135b.
[0074] Then, the document object detection unit 135 adds the detected object as an extension object to the table attribute detection result from the table attribute detection unit 133.
[0075] FIG. 17 is a schematic diagram showing an intra-document object detection result 135d#2 in which an extended attribute is added to the intra-table attribute detection result from the intra-table attribute detection unit 133. In FIG. An intra-document target detection result 135#2 shows an example in which the intra-table attribute detection result from the intra-table attribute detection unit 133 is the same as the neighboring attribute detection result 131d#2 shown in FIG.
[0076] In the document object detection result 135d#2, in addition to the notation row 135e#2, part of speech row 135f#2, and division row 135g#2, an extended notation row 135h#2, an extended part of speech 135i#2, and an extended division 135j#2 are added, and the target words, parts of speech, and divisions detected by the table attribute detection unit 133 are stored. The expanded notation row 135h#2, the expanded part of speech 135i#2, and the expanded section 135j#2 store information on the detected objects in the order in which they were detected.
[0077] Then, the intra-document object detection unit 135 provides the intra-document object detection result, which is the detection result of the intra-document object, to the intra-sentence object detection unit 136 .
[0078] The intra-sentence object detection unit 136 detects objects from words included in a predetermined range (here, the number of words) from the exponential representation of the intra-document object detection result. Here, words that match objects stored in the object attribute-specific detail unit information stored in the object attribute-specific detail unit storage unit 116 are detected. If there are multiple objects within the predetermined range, the intra-sentence object detection unit 136 detects the object that is closer to an exponential expression.
[0079] Furthermore, the intra-sentence object detection unit 136 can correct spelling variations of the object by using the object synonym dictionary information 135a as shown in FIG.
[0080] Specifically, by referring to the target synonym dictionary information 135a, if the in-document target detection result contains a synonym stored in the target synonym column 135c, the in-sentence target detection unit 136 can correct the spelling variation of the target by converting the synonym into the target stored in the target column 135b.
[0081] Then, in the document object detection result, the intra-sentence object detection unit 136 stores "object" in the same column and row of the detected object. If the document object detection result includes multiple exponential expressions, identification information for identifying which exponential expression the detected object belongs to is added to the detected object's classification.
[0082] FIG. 18 is a schematic diagram showing an intra-sentence object detection result 136a#3, which is a result of intra-sentence object detection by intra-sentence object detection unit 136. As shown in FIG. The intra-sentence object detection result 136a#3 is an example of a case where the intra-document object detection result from the intra-document object detection unit 135 is the same as the modification attribute detection result 132a#3 shown in FIG.
[0083] As shown in FIG. 18, in the in-text target detection result 136a#3, "Target 1, Target 2" is stored in the classification row of the target "tunnel". In FIG. 18, a state is shown where a word that was originally written as "tunnel" in the original sentence is converted to "tunnel" by referring to the target synonym dictionary information 135a. Here, "Target 1" indicates that it is the target of the first quantity expression (exponential expression) that appears in the input text, and "Target 2" indicates that it is the target of the second quantity expression (exponential expression) that appears in the input text. In the example shown in FIG. 18, the first quantity expression (exponential expression) that appears in the input text is "5.5", "[EXP]", "+00", and "m", and the second quantity expression (exponential expression) that appears in the input text is "3.7", "[EXP]", "+00", and "m".
[0084] Then, the in-text target detection unit 136 provides the in-text target detection result to the in-table target detection unit 137.
[0085] When a table is included in the document structure data as learning source data, the in-table target detection unit 137 detects the target of the input text extracted from the table from the row item names and column item names of the table.
[0086] At this time, as shown in FIG. 19, the in-table target detection unit 137 can detect a common target for many words with a common specification by referring to the word hierarchy relationship information indicating the hierarchy relationship between words. In the hierarchy relationship indicated by the word hierarchy relationship information, it is assumed that one hierarchy includes words that match the targets stored in the target attribute-by-detail unit information stored in the target attribute-by-detail unit storage unit 116.
[0087] If the document structure data includes a table such as that shown in Fig. 13(A), the row item name of the first row of the table includes "X switch." This "X switch" can be determined to be a subordinate word of "electrical component" based on the word hierarchical relationship information shown in Fig. 19, and therefore the in-table object detection unit 137 can extract the object "electrical component" from "X switch." Then, the in-table object detection unit 137 adds the detected object as an extension object to the in-sentence object detection result.
[0088] FIG. 20 is a schematic diagram showing an in-table object detection result 137a#4, which is a result of in-table object detection by in-table object detection unit 137. The in-table object detection result 137a#4 is an example of a case where the in-sentence object detection result from the in-sentence object detection unit 136 is the same as the in-table attribute detection result 133a#4 shown in FIG.
[0089] In the table object detection result 137a#4, in addition to the notation row 137b#4, part of speech row 137c#4, and division row 137d#4, an extended notation row 137e#4, extended part of speech 137f#4, and extended division 137g#4 are added, and the target words, parts of speech, and divisions detected from the row item names and column item names of the table are stored. The expanded notation row 137e#4, the expanded part of speech 137f#4, and the expanded section 137g#4 store information about the detected objects in the order in which they were detected, in other words, in the order in which they appear in the input text.
[0090] Then, the in-table object detection unit 137 provides the in-table object detection result, which is the detection result of the in-table object, to the erroneous conversion suppression unit 138 .
[0091] The erroneous conversion suppression unit 138 suppresses erroneous conversion of a unit included in the in-table object detection result from the in-table object detection unit 137 into a refined unit. The erroneous conversion suppression unit 138 includes a part-whole relationship determination unit 139 and a differential expression determination unit 140 .
[0092] If an attribute detected from the proximity of an exponential expression or from a dependency relationship in the in-table object detection result from the in-table object detection unit 137 is an attribute of part or the whole of the object detected by the object detection unit 134, the part-whole relationship determination unit 139 prevents the unit associated with the attribute and object from being converted into a refinement unit. Specifically, when a word having a predetermined part-whole relationship is detected in the vicinity or dependency relationship of an exponential expression associated with an attribute detected from the vicinity or dependency relationship in the in-table object detection results, the part-whole relationship determination unit 139 prevents the unit associated with that attribute from being converted into a refined unit.
[0093] For example, a case will be described in which the in-table object detection result from in-table object detection section 137 is in-table object detection result 137a#5 as shown in FIG. 21(A). The in-table object detection result 137a#5 includes the attributes detected by the nearby attribute detection unit 131 and the objects detected by the in-document object detection unit 135. Specifically, "width" near the exponential expression is detected as the attribute, and "train" from the article title is detected as the object.
[0094] However, "width" is not the "width" of the "train" itself, but the "width" of the "door," which is one of the components of the "train." It is possible to define the "width of the train door" by further refining the target attribute-specific refinement unit information, but there is a trade-off in that increasing the number of refinement units reduces the amount of training data for each refinement unit, so excessive refinement leads to a decrease in accuracy and has the opposite effect.
[0095] In such a case, the part-whole relationship determination unit 139 determines whether the attributes and objects included in the in-table object detection results are in a part-whole relationship by referring to part-whole information, for example, as shown in Figure 22. For example, in the in-table object detection result 137a#5, the word "door" detected near the exponential expression has a part-whole relationship with the object "train," so the part-whole relationship determination unit 139 determines that the attribute "width" is an attribute of "door." The part-whole relationship determination unit 139 then changes the classification of the object "train" and the attribute "width" to "part-whole," which is a classification other than "object" and "attribute," thereby preventing erroneous refinement in the unit refinement process described below. For example, in the part-whole relationship determination result 139a#5 shown in FIG. 21(B), the classification information "object" of the expanded notation "train" and the classification information "attribute" of the notation "width" are each rewritten to "part-whole." Note that words that are part or whole with respect to the object may be detected from words that have a dependency relationship with the exponential expression. Also, here we have shown an example of suppressing the attribute of "door" which is on the lower side of the part-whole relationship, but depending on the application, it may be possible to suppress the attribute of "train" which is on the higher side and refine the unit of the attribute of "door".
[0096] Then, the part-whole relationship determination unit 139 provides the part-whole relationship determination result to the differential expression determination unit 140 .
[0097] If the quantity of an attribute detected from a neighborhood or dependency relationship in the part-whole relationship determination result from the part-whole relationship determination unit 139 is expressed by a differential expression, the differential expression determination unit 140 prevents the unit associated with that attribute from being converted into a refined unit. Specifically, when a predetermined differential expression is detected in the vicinity or dependency relationship of an exponential expression associated with an attribute detected from the vicinity or dependency relationship in the part-whole relationship determination result, the differential expression determination unit 140 prevents the unit associated with the attribute from being converted into a refined unit.
[0098] For example, a case will be described in which the part-whole relationship determination result from part-whole relationship determination section 139 is part-whole relationship determination result 139a#6 as shown in FIG. 23(A). The part-whole relationship determination result 139a#6 includes the attribute detected by the nearby attribute detection unit 131 and the object detected by the intra-document object detection unit 135. Specifically, "long" near the exponential expression is detected as the attribute, and "train" from the article title is detected as the object.
[0099] However, the "long" does not refer to the "length" of the "train" in question, but rather to the increase from another model, the "105 series."
[0100] In such a case, the differential expression determining unit 140 determines the differential expression included in the part-whole relationship determination result by referring to the differential expression pattern dictionary information 140a shown in FIG.
[0101] The differential expression pattern dictionary information 140a is information in a table format having a part of speech column 140b and a notation column 140c. The part of speech column 140b stores the part of speech of the word. The notation column 140c stores words. The differential expression pattern dictionary information 140a allows a differential expression to be identified by a combination of a word and a part of speech. In other words, if a word stored in the notation column 140c is a part of speech stored in the part of speech column 140b, that word becomes a differential expression.
[0102] For example, the "noun" "enhancement" in the differential expression pattern dictionary information 140a is detected from the part-whole relation determination result 139a#6, so the differential expression determination unit 140 changes the classification of the attribute and object "long" and "train" associated with the nearby exponential expression to "difference," a classification other than "attribute" and "object," thereby preventing erroneous refining in the unit refining process described below. For example, in the differential expression determination result 140d#6 shown in Figure 23(B), the classification information "object" of the extended notation "train" and the classification information "attribute" of the notation "long" are each rewritten to "difference."
[0103] Then, the differential expression determination unit 140 provides the differential expression determination result to the unit conversion unit 141 .
[0104] The unit conversion unit 141 refers to the target attribute-specific detail unit information stored in the target attribute-specific detail unit storage unit 116, and converts the unit included in the exponential expression in the differential expression determination result from the differential expression determination unit 140 into the detail unit associated with that unit and the target and attribute associated with that exponential expression.
[0105] For example, if the differential expression determination result from the differential expression determination unit 140 is differential expression determination result 140e#2 shown in Figure 25(A), the unit "m" associated with the object "river" and the attribute "total length" is associated with "m_river_length" in the object attribute-specific detailed unit information 116a shown in Figure 8, and therefore the unit "m" in the differential expression determination result 140e#2 is converted to "m_river_length", for example, as in the unit conversion result 141a#2 shown in Figure 25(B).
[0106] When a differential expression determination result generated from one sentence includes multiple exponential expressions, the unit conversion unit 141 repeatedly performs the above process for each of the multiple units included in the multiple exponential expressions.
[0107] Then, the unit conversion section 141 provides the unit conversion result, which is the processing result of the unit conversion process, to the quantitative expression learning section 118.
[0108] The numerical expression learning unit 118 uses data including refinement units and normalized numerical values to learn a numerical expression language model for estimating the normalized numerical values. For example, the numerical expression learning unit 118 uses the unit conversion result from the unit conversion unit 141 to learn a numerical expression language model, which is a model for estimating numerical expressions.
[0109] The quantitative expression language model storage unit 119 stores the quantitative expression language model learned by the quantitative expression learning unit 118 .
[0110] The communication unit 120 transmits the quantitative expression language model stored in the quantitative expression language model storage unit 119 to the estimation device 150.
[0111] Some or all of the above-described learning source data acquisition unit 111, quantitative expression identification unit 112, numerical normalization unit 113, unit normalization unit 115, unit refinement unit 117, and quantitative expression learning unit 118 can be configured, for example, as shown in FIG. 26 , by a memory 10 and a processor 11 such as a CPU (Central Processing Unit) that executes a program stored in the memory 10. Such a program may be provided via a network or may be provided by being recorded on a recording medium. That is, such a program may be provided, for example, as a program product.
[0112] In other words, the learning device 110 can be realized by a so-called computer. The unit normalization information storage unit 114, the target attribute-specific detailed unit storage unit 116, and the quantitative expression language model storage unit 119 can be realized by an auxiliary storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). Furthermore, the communication unit 120 can be realized by a communication I / F such as a network interface card (NIC).
[0113] FIG. 27 is a flowchart showing the operation of learning device 110 in the first embodiment. First, the learning original data acquisition unit 111 acquires learning original data (S10). The acquired learning original data is provided to the quantification expression identification unit 112.
[0114] Next, the numerical expression identification unit 112 analyzes the received learning source data and identifies the numerical expressions (S11). The numerical expression identification result indicating the identified numerical expressions is provided to the numerical value normalization unit 113.
[0115] The numeric normalization unit 113 normalizes the received quantitative expression identification result by converting the numeric part of the quantitative expression into an exponential expression (S12). The numeric normalization result indicating the normalized numeric value is provided to the unit normalization unit 115.
[0116] The unit normalization unit 115 normalizes the units in the received numeric value normalization result by referring to the unit normalization information stored in the unit normalization information storage unit 114, and adjusts the exponent part of the exponential expression in accordance with the change in unit (S13). The unit normalization result indicating the normalized units is provided to the unit refinement unit 117.
[0117] The unit refinement unit 117 refines the units in the received unit normalization result by referring to the target attribute-specific refinement unit information stored in the target attribute-specific refinement unit storage unit 116 (S14). The unit conversion result indicating the refined units is provided to the quantitative expression learning unit 118.
[0118] The quantitative expression learning unit 118 performs a learning process using the received unit conversion result to generate a quantitative expression language model (S15). The generated quantitative expression language model is stored in the quantitative expression language model storage unit 119 and is sent to the estimation device 150 via the communication unit 120.
[0119] FIG. 28 is a flowchart showing the operation of the numerical expression learning unit 118 in the first embodiment. First, the quantitative expression learning unit 118 extracts the text stored in the notation line as the target text from the unit conversion result (S20).
[0120] Next, the numerical expression learning unit 118 performs a masking process to mask the part of the target text where prediction is to be made, thereby generating a masked text (S21).
[0121] Next, the numerical expression learning unit 118 performs a process of predicting the masked portion of the masked text using the numerical expression language model (S22).
[0122] Next, the numerical expression learning unit 118 detects the difference between the prediction result in step S22 and the correct answer obtained from the target text before masking, and performs parameter update processing for the numerical expression language model (S23).
[0123] The quantitative expression learning unit 118 performs the above process on all unit conversion results provided by the unit refining unit 117, thereby learning the quantitative expression language model.
[0124] For the processes in steps S21 to S23, for example, the technique described in the following document 1 can be used.
[0125] FIG. 29 is a block diagram schematically illustrating a configuration of an estimation device 150 according to the first embodiment. The estimation device 150 includes a communication unit 151, a quantitative expression language model storage unit 152, an estimation target data acquisition unit 153, a quantitative expression identification unit 154, a numerical normalization unit 155, a unit normalization information storage unit 156, a unit normalization unit 157, a target attribute-specific refinement unit storage unit 158, a unit refinement unit 159, and a quantitative expression estimation unit 160.
[0126] The communication unit 151 receives the quantitative expression language model from the learning device 110. The received quantitative expression language model is provided to the quantitative expression language model storage unit 152.
[0127] The quantitative expression language model storage unit 152 stores the quantitative expression language model received by the communication unit 151 .
[0128] The estimation target data acquisition unit 153 acquires estimation target data that requests estimation of a numerical value at a position where the predetermined expression format is arranged, using a predetermined expression format and unit. The acquired estimation target data is provided to the quantitative expression specification unit 154. The predetermined expression format is, for example, "*^*", but is not limited to this example.
[0129] The inference target data acquisition unit 153 may receive input of the inference target data via, for example, an input unit (not shown). Alternatively, the inference target data acquisition unit 153 may acquire the inference target data from another device via the communication unit 151. For example, the inference target data acquisition unit 153 may acquire text data such as "a car with an overall width of 1.2 m and an overall length of ** m" as the inference target data.
[0130] If the estimation target data contains a numerical expression, the numerical expression identification unit 154 identifies the numerical expression. The method for identifying the numerical expression is the same as the method for identifying the numerical expression by the numerical expression identification unit 112 of the learning device 110. However, if the numerical part is in the predetermined expression format described above, it is also identified as a numerical expression.
[0131] For example, when the estimation target data is "a car with an overall width of 1.2 m and an overall length of *^* m," the quantitative expression specification unit 154 specifies "1.2 m" and "an overall length of *^* m" as the quantitative expressions. Then, the numerical expression specification unit 154 provides the numerical expression specification result indicating the numerical expression specified as described above to the numerical value normalization unit 155 .
[0132] The numeric normalization unit 155 normalizes the numeric values included in the quantitative expression identified by the quantitative expression identification unit 154 by converting them into exponential notation. The method of normalizing the numeric values is the same as the method of normalizing the numeric values by the numeric normalization unit 113 of the learning device 110. However, if the numeric portion is in the predetermined representation format described above, there is no need to convert it into exponential notation. The numeric normalization result, which is the result of the normalization of the numeric values performed by the numeric normalization unit 155, is provided to the unit normalization unit 157.
[0133] Unit normalization information storage unit 156 stores unit normalization information, which is information for normalizing units. This unit normalization information may be the same as the unit normalization information stored in unit normalization information storage unit 114 of learning device 110.
[0134] Unit normalization unit 157 normalizes the units included in the numerical value normalization result from numerical value normalization unit 155 by referring to the unit normalization information stored in unit normalization information storage unit 156. The unit normalization method used by unit normalization unit 157 is the same as the unit normalization method used by unit normalization unit 115 in learning device 110. However, if the numerical value portion is in the predetermined representation format described above, there is no need to adjust the exponent. The unit normalization result, which is the result of unit normalization by unit normalization unit 157, is provided to unit refinement unit 159.
[0135] The unit refinement unit 159 analyzes the unit normalization result from the unit normalization unit 157 to detect combinations of objects and attributes, and converts the standard units of the exponential representation representing the detected objects and attributes into refined units corresponding to the objects, attributes, and standard units while suppressing erroneous conversion to refined units. The unit refinement method of the unit refinement unit 159 is the same as the unit refinement method of the unit refinement unit 117 in the learning device 110. The unit refinement unit 159 then provides the unit conversion result to the quantitative expression estimation unit 160.
[0136] The quantitative expression estimation unit 160 estimates the numerical value at the position where the predetermined expression form is located by inputting data including a predetermined expression form and a converted elaboration unit into a quantitative expression language model trained using data including a elaboration unit and a normalized numerical value. Here, the numerical expression estimation unit 160 estimates the numerical expression using the unit conversion result from the unit refinement unit 159. For example, the quantitative expression estimation unit 160 performs estimation by extracting estimated data from the notation line of the unit conversion result and inputting the estimated data into the quantitative expression language model stored in the quantitative expression language model storage unit 152. Here, the part "*^*", which is a predetermined expression format, is estimated.
[0137] Part or all of the estimation target data acquisition unit 153, the quantitative expression identification unit 154, the numerical normalization unit 155, the unit normalization unit 157, the unit refinement unit 159, and the quantitative expression estimation unit 160 described above can be configured, for example, as shown in Fig. 26, by a memory 10 and a processor 11 that executes a program stored in the memory 10. Such a program may be provided via a network or may be provided by being recorded on a recording medium. That is, such a program may be provided, for example, as a program product.
[0138] In other words, the estimation device 150 can be realized by a so-called computer. The unit normalization information storage unit 156, the target attribute-specific refinement unit storage unit 158, and the quantitative expression language model storage unit 152 can be realized by an auxiliary storage device. Furthermore, the communication unit 151 can be realized by a communication I / F.
[0139] FIG. 30 is a flowchart showing the operation of the estimating device 150 according to the first embodiment. First, the estimation target data acquisition unit 153 acquires estimation target data (S30). The acquired estimation target data is provided to the quantitative expression identification unit 154.
[0140] Next, the numerical expression identification unit 154 analyzes the received estimation target data and identifies the numerical expressions (S31). The numerical expression identification result indicating the identified numerical expressions is provided to the numerical value normalization unit 155.
[0141] The numeric normalization unit 155 normalizes the received result of identifying the quantitative expression by converting the numerical part of the quantitative expression into an exponential expression (S32). The result of the numeric normalization indicating the normalized numerical value is provided to the unit normalization unit 157.
[0142] The unit normalization unit 157 normalizes the units in the received numeric value normalization result by referring to the unit normalization information stored in the unit normalization information storage unit 156, and adjusts the exponent part of the exponential expression in accordance with the change in unit (S33). The unit normalization result indicating the normalized units is provided to the unit refinement unit 159.
[0143] The unit refinement unit 159 refines the units in the received unit normalization result by referring to the target attribute-specific refinement unit information stored in the target attribute-specific refinement unit storage unit 158 (S34). The unit conversion result indicating the refined units is provided to the quantitative expression estimation unit 160.
[0144] The quantitative expression estimation unit 160 performs estimation using the received unit conversion result (S35).
[0145] According to embodiment 1, the distribution of numerical values, which varies greatly depending on the object and attribute, can be adjusted to a specific digit, thereby enabling quantitative expressions to be reliably learned and the learned model to be used to reliably estimate quantitative expressions.
[0146] Embodiment 2 As shown in FIG. 1, a learning estimation system 200 according to the second embodiment includes a learning device 210 and an estimation device 250.
[0147] FIG. 31 is a block diagram showing a schematic configuration of learning device 210. As shown in FIG. The learning device 210 includes a learning source data acquisition unit 111, a quantitative expression identification unit 112, a numerical normalization unit 113, a unit normalization information storage unit 114, a unit normalization unit 115, a target attribute-specific refinement unit storage unit 116, a unit refinement unit 117, a quantitative expression learning unit 118, a quantitative expression language model storage unit 119, a communication unit 120, and a numerical rounding unit 221.
[0148] The learning original data acquisition unit 111, the quantitative expression identification unit 112, the numerical normalization unit 113, the unit normalization information storage unit 114, the unit normalization unit 115, the target attribute-based detailed unit storage unit 116, the unit detailed unit 117, the quantitative expression learning unit 118, the quantitative expression language model storage unit 119, and the communication unit 120 of the learning device 210 in embodiment 2 are similar to the learning original data acquisition unit 111, the quantitative expression identification unit 112, the numerical normalization unit 113, the unit normalization information storage unit 114, the unit normalization unit 115, the target attribute-based detailed unit storage unit 116, the unit detailed unit 117, the quantitative expression learning unit 118, the quantitative expression language model storage unit 119, and the communication unit 120 of the learning device 110 in embodiment 1. However, numeric value normalization unit 113 of learning device 210 in embodiment 2 provides the numeric value normalization result to numeric value rounding unit 221. Furthermore, unit normalization unit 115 of learning device 210 in embodiment 2 performs numeric value normalization on the numeric value rounding result from numeric value rounding unit 221 instead of the numeric value normalization result.
[0149] The numeric rounding unit 221 performs a numeric rounding process to round the numeric value included in the numeric normalization result from the numeric normalization unit 113 to a predetermined number of digits. In the second embodiment, the numeric rounding unit 221 rounds the numeric value to a predetermined number of digits to get the predetermined number of digits. For example, if the numeric value included in the numeric normalization result is "3.5569", the numeric rounding unit 221 rounds the numeric value to one decimal place to get the numeric value included in the numeric normalization result to "3.6".
[0150] In the second embodiment, the numeric rounding unit 221 rounds off a numeric value to a predetermined number of digits to obtain a predetermined number of digits, but the present invention is not limited to this example. For example, the numeric rounding unit 221 may round down a numeric value to a predetermined number of digits. Then, the numeric rounding unit 221 provides the unit normalization unit 115 with the numeric rounding result including the numeric value rounded to a predetermined number of digits. As described above, the numerical expression learning unit 118 learns the numerical expression language model using the normalized numerical values after the rounding process.
[0151] The above-described numerical value rounding unit 221 can also be configured, for example, as shown in FIG. 26, by a memory 10 and a processor 11 that executes a program stored in the memory 10.
[0152] FIG. 32 is a flowchart showing the operation of learning device 210 in the second embodiment. First, the learning original data acquisition unit 111 acquires learning original data (S40). The acquired learning original data is provided to the quantification expression identification unit 112.
[0153] Next, the numerical expression identification unit 112 analyzes the received original learning data and identifies the numerical expressions (S41). The numerical expression identification result indicating the identified numerical expressions is provided to the numerical value normalization unit 113.
[0154] The numeric normalization unit 113 normalizes the received result of identifying the numeric expression by converting the numeric part of the numeric expression into an exponential expression (S42). The numeric normalization result indicating the normalized numeric value is provided to the numeric rounding unit 221.
[0155] The numeric rounding unit 221 rounds off a predetermined number of digits of the numeric value included in the numeric normalization result to make the numeric value have a predetermined number of digits (S43). The numeric rounding result indicating the numeric value with the limited number of digits is provided to the unit normalization unit 115.
[0156] The unit normalization unit 115 normalizes the units in the received rounded numeric value result by referring to the unit normalization information stored in the unit normalization information storage unit 114, and adjusts the exponent part of the exponential representation in accordance with the change in unit (S44). The unit normalization result indicating the normalized unit is provided to the unit refinement unit 117.
[0157] The unit refinement unit 117 refines the units in the received unit normalization result by referring to the target attribute-specific refinement unit information stored in the target attribute-specific refinement unit storage unit 116 (S45). The unit conversion result indicating the refined units is provided to the quantitative expression learning unit 118.
[0158] The quantitative expression learning unit 118 performs a learning process using the received unit conversion result to generate a quantitative expression language model (S46). The generated quantitative expression language model is stored in the quantitative expression language model storage unit 119 and is sent to the estimation device 250 via the communication unit 120.
[0159] FIG. 33 is a block diagram schematically showing the configuration of an estimation device 250 according to the second embodiment. The estimation device 250 includes a communication unit 151, a quantitative expression language model storage unit 152, an estimation target data acquisition unit 153, a quantitative expression identification unit 154, a numerical normalization unit 155, a unit normalization information storage unit 156, a unit normalization unit 157, a target attribute-specific refinement unit storage unit 158, a unit refinement unit 159, a quantitative expression estimation unit 160, and a numerical rounding unit 261.
[0160] The communication unit 151, the quantitative expression language model storage unit 152, the estimation target data acquisition unit 153, the quantitative expression identification unit 154, the numerical normalization unit 155, the unit normalization information storage unit 156, the unit normalization unit 157, the target attribute-specific detailing unit storage unit 158, the unit detailing unit 159, and the quantitative expression estimation unit 160 of the estimation device 250 in embodiment 2 are similar to the communication unit 151, the quantitative expression language model storage unit 152, the estimation target data acquisition unit 153, the quantitative expression identification unit 154, the numerical normalization unit 155, the unit normalization information storage unit 156, the unit normalization unit 157, the target attribute-specific detailing unit storage unit 158, the unit detailing unit 159, and the quantitative expression estimation unit 160 of the estimation device 150 in embodiment 1.
[0161] However, the numeric value normalization unit 155 of the estimating device 250 in the second embodiment provides the numeric value normalization result to the numeric value rounding unit 261. Furthermore, the unit normalization unit 157 of the estimating device 250 in the second embodiment normalizes the numeric value using the numeric value rounding result from the numeric value rounding unit 261 instead of the numeric value normalization result.
[0162] The numeric rounding unit 261 rounds the numeric value included in the numeric normalization result from the numeric normalization unit 155 to a predetermined number of digits. In the second embodiment, the numeric rounding unit 261 rounds off the numeric value to the predetermined number of digits to get the predetermined number of digits.
[0163] In the second embodiment, the numeric rounding unit 261 rounds off a numeric value to a predetermined number of digits to obtain a predetermined number of digits, but the present invention is not limited to this example. For example, the numeric rounding unit 261 may round down a numeric value to a predetermined number of digits. Then, the numeric rounding unit 261 provides the unit normalization unit 157 with the numeric rounding result including the numeric value rounded to a predetermined number of digits.
[0164] The above-described numerical value rounding unit 261 can also be configured, for example, as shown in FIG. 26, by a memory 10 and a processor 11 that executes a program stored in the memory 10.
[0165] FIG. 34 is a flowchart showing the operation of the estimating device 250 according to the second embodiment. First, the estimation target data acquisition unit 153 acquires estimation target data (S50). The acquired estimation target data is provided to the quantification expression identification unit 154.
[0166] Next, the numerical expression identification unit 154 analyzes the received estimation target data and identifies the numerical expressions (S51). The numerical expression identification result indicating the identified numerical expressions is provided to the numerical value normalization unit 155.
[0167] The numeric normalization unit 155 normalizes the received result of identifying the numerical expression by converting the numerical part of the numerical expression into an exponential notation (S52). The result of the numeric normalization indicating the normalized numerical value is provided to the numeric rounding unit 261.
[0168] The numeric rounding unit 261 rounds off a predetermined number of digits of the numeric value included in the numeric normalization result to make the numeric value have a predetermined number of digits (S53). The numeric rounding result indicating the numeric value with the limited number of digits is provided to the unit normalization unit 157.
[0169] The unit normalization unit 157 normalizes the units in the received rounded numeric value result by referring to the unit normalization information stored in the unit normalization information storage unit 156, and adjusts the exponent part of the exponential representation in accordance with the change in unit (S54). The unit normalization result indicating the normalized unit is provided to the unit refinement unit 159.
[0170] The unit refinement unit 159 refines the units in the received unit normalization result by referring to the target attribute-specific refinement unit information stored in the target attribute-specific refinement unit storage unit 158 (S55). The unit conversion result indicating the refined units is provided to the quantitative expression estimation unit 160.
[0171] The quantitative expression estimation unit 160 performs estimation using the received unit conversion result (S56).
[0172] According to the second embodiment, since the number of numerical values to be learned increases, a more appropriate quantitative expression language model can be learned. Furthermore, by using such a quantitative language model, the estimation accuracy of the quantitative expressions can be improved.
[0173] Embodiment 3 As shown in FIG. 1, the learning estimation system 300 according to the third embodiment includes a learning device 310 and an estimation device 150. The estimation device 150 of the learning estimation system 300 according to the third embodiment is similar to the estimation device 150 of the learning estimation system 100 according to the first embodiment.
[0174] As shown in Figure 2, the learning device 310 in embodiment 3 includes a learning source data acquisition unit 111, a quantitative expression identification unit 112, a numerical normalization unit 113, a unit normalization information storage unit 114, a unit normalization unit 115, a target attribute-specific detailed unit storage unit 116, a unit detailed unit 117, a quantitative expression learning unit 318, a quantitative expression language model storage unit 119, and a communication unit 120.
[0175] The learning original data acquisition unit 111, the quantitative expression identification unit 112, the numerical normalization unit 113, the unit normalization information storage unit 114, the unit normalization unit 115, the target attribute-specific detail unit storage unit 116, the unit detailing unit 117, the quantitative expression language model storage unit 119, and the communication unit 120 of the learning device 310 in embodiment 3 are similar to the learning original data acquisition unit 111, the quantitative expression identification unit 112, the numerical normalization unit 113, the unit normalization information storage unit 114, the unit normalization unit 115, the target attribute-specific detail unit storage unit 116, the unit detailing unit 117, the quantitative expression language model storage unit 119, and the communication unit 120 of the learning device 110 in embodiment 1.
[0176] The numerical expression learning unit 318 uses the unit conversion result from the unit conversion unit 141 to learn a numerical expression language model, which is a model for estimating numerical expressions. In the third embodiment, the numerical expression learning unit 318 performs a masking process that emphasizes the learning of numerical expressions. For example, the numerical expression learning unit 318 increases the probability of masking numerical values compared to the probability of masking other parts. Specifically, in the following document 1, 80% of the tokens in the input text are masked, 10% are replaced with random tokens, and 10% are used as the original tokens for learning. However, in the third embodiment, the numerical expression learning unit 318 can emphasize the learning of numerical expressions by setting the probability of masking numerical values in the input text to 90% and the probability of masking parts other than numerical values to 75%. Note that the ratio of the probability of masking numerical values to parts other than numerical values can also be dynamically set according to the proportion of tokens corresponding to numerical values in the input text. For example, the probability of masking numerical values can be increased as the proportion of tokens corresponding to numerical values in the input text increases.
[0177] Furthermore, the numerical expression learning unit 318 learns the numerical expression language model so that when an error occurs in the estimation of a normalized numerical value, the penalty is greater than when an error occurs in the estimation of other parts. For example, the numerical expression learning unit 318 performs parameter updating processing that emphasizes numerical values in the parameter updating processing of the numerical expression language model. For example, the numerical expression learning unit 318 updates the parameters by giving a larger penalty when the exponents in the exponential expressions are different. Furthermore, the numerical expression learning unit 318 can also perform parameter update processing using the ratio between the predicted numerical value and the correct numerical value before masking. For example, the numerical expression learning unit 318 can update the parameters by giving a larger penalty as the ratio between the predicted numerical value and the correct numerical value before masking increases.
[0178] FIG. 35 is a flowchart showing the operation of the numerical expression learning unit 318 in the third embodiment. First, the quantitative expression learning unit 318 extracts the text stored in the notation line as the target text from the unit conversion result (S60).
[0179] Next, the numerical expression learning unit 318 performs a masking process to mask the part of the target text to be predicted, thereby generating a masked text (S61). In the third embodiment, the numerical expression learning unit 318 masks numerical values more frequently than other parts in the masking process.
[0180] Next, the numerical expression learning unit 318 performs a process of predicting the masked portion of the masked text using the numerical expression language model (S62).
[0181] Next, the numerical expression learning unit 318 detects the difference between the prediction result in step S62 and the correct answer obtained from the target text before masking, and performs parameter update processing for the numerical expression language model (S63). Here, in the third embodiment, the numerical expression learning unit 318 updates the parameters so that a larger penalty is imposed when the predicted exponents are different.
[0182] The numerical expression learning unit 318 performs the above process on all unit conversion results provided by the unit refining unit 117, thereby learning the numerical expression language model.
[0183] As described above, according to the third embodiment, it is expected that the quantitative expression language model can be optimized by placing emphasis on the quantitative expressions contained in the original training data.
[0184] In the above-described first to third embodiments, the processing in the erroneous conversion suppression unit 138 does not necessarily have to be performed.
[0185] Furthermore, in the above-described first to third embodiments, each process has been described using numerical expressions in Japanese, but the first to third embodiments are not limited to such examples. For example, numerical expressions in other languages, such as English, can also be applied to the first to third embodiments. In this case, the same effects as those described in the first to third embodiments can be achieved.
[0186] Reference 1: Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019 [Explanation of symbols]
[0187] 100,200,300 Learning estimation system, 110,210,310 Learning device, 111 Learning original data acquisition unit, 112 Quantitative expression identification unit, 113 Numeric normalization unit, 114 Unit normalization information storage unit, 115 Unit normalization unit, 116 Object attribute-specific refinement unit storage unit, 117 Unit refinement unit, 118,318 Quantitative expression learning unit, 119 Quantitative expression language model storage unit, 120 Communication unit, 221 Numeric rounding unit, 130 Attribute detection unit, 131 Neighborhood attribute detection unit, 132 Dependency attribute detection unit, 133 Table attribute detection unit, 134 Object detection unit, 135 Document object detection unit, 136 Sentence object detection unit, 137 Table object detection unit, 138 Misconversion suppression unit, 139 part-whole relationship determination unit, 140 differential expression determination unit, 141 unit conversion unit, 150,250 estimation device, 151 communication unit, 152 quantitative expression language model storage unit, 153 estimation target data acquisition unit, 154 quantitative expression identification unit, 155 numeric normalization unit, 156 unit normalization information storage unit, 157 unit normalization unit, 158 target attribute-specific refinement unit storage unit, 159 unit refinement unit, 160 quantitative expression estimation unit, 261 numeric rounding unit.
Claims
1. a quantitative expression identification unit that identifies, from the original learning data, a quantitative expression that represents a quantity using a numerical value and a unit; a numerical value normalization unit that normalizes the numerical value; a unit normalization unit that normalizes the units; a unit refinement unit that identifies, in the original learning data, an object that is a physical entity and an attribute that is a property of the object, which is indicated by the numerical value and the unit, and converts the normalized unit into a refinement unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; a numerical expression learning unit that uses data including the refinement unit and the normalized numerical value to learn a numerical expression language model for estimating the normalized numerical value. A learning device characterized by:
2. The unit detailing unit refers to object attribute-specific detailing unit information that associates a plurality of objects, a plurality of attributes, a plurality of units, and a plurality of detailing units, and thereby identifies the detailing unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit.
2. The learning device according to claim 1, wherein:
3. The unit refinement unit specifies a predetermined first word as the attribute when the predetermined first word is included in the vicinity of the quantitative expression.
2. The learning device according to claim 1, wherein:
4. The unit refinement unit specifies a predetermined first word as the attribute when the predetermined first word is included in the vicinity of the quantitative expression.
3. The learning device according to claim 2, wherein:
5. the unit refinement unit specifies a predetermined first word as the attribute when the first word is included in a plurality of words having a modification relationship with the quantitative expression; 2. The learning device according to claim 1, wherein:
6. the unit refinement unit specifies a predetermined first word as the attribute when the first word is included in a plurality of words having a modification relationship with the quantitative expression; 3. The learning device according to claim 2, wherein:
7. When the quantitative expression is included in a table in the original data for learning, the unit refining unit specifies a predetermined first word as the attribute if the predetermined first word is included in at least one of a row heading name and a column heading name of the table.
2. The learning device according to claim 1, wherein:
8. When the quantitative expression is included in a table in the original data for learning, the unit refining unit specifies a predetermined first word as the attribute when the predetermined first word is included in at least one of a row heading name and a column heading name of the table.
3. The learning device according to claim 2, wherein:
9. The unit refinement unit replaces a synonym of the first word with the first word, and then identifies the attribute.
4. The learning device according to claim 3, wherein:
10. The unit refinement unit replaces a synonym of the first word with the first word, and then identifies the attribute.
6. The learning device according to claim 5,
11. The unit refinement unit replaces a synonym of the first word with the first word, and then identifies the attribute.
8. The learning device according to claim 7,
12. the unit refinement unit specifies, when the original data for learning is document structure data having items and contents corresponding to the items, and when the contents include the quantitative expression, a predetermined second word is included in the title of the item, the second word is specified as the target.
2. The learning device according to claim 1, wherein:
13. the unit refinement unit specifies, when the original data for learning is document structure data having items and contents corresponding to the items, and when the contents include the quantitative expression, a predetermined second word is included in the title of the item, the second word is specified as the target.
3. The learning device according to claim 2, wherein:
14. the unit refinement unit specifies, when the original data for learning is document structure data having items and contents corresponding to the items, and when the contents include the quantitative expression, a predetermined second word is included in the title of the item, the second word is specified as the target.
4. The learning device according to claim 3, wherein:
15. the unit refinement unit specifies, when the original data for learning is document structure data having items and contents corresponding to the items, and when the contents include the quantitative expression, a predetermined second word is included in the title of the item, the second word is specified as the target.
6. The learning device according to claim 5,
16. the unit refinement unit specifies, when the original data for learning is document structure data having items and contents corresponding to the items, and when the contents include the quantitative expression, a predetermined second word is included in the title of the item, the second word is specified as the target.
8. The learning device according to claim 7,
17. the unit refinement unit specifies, when the original data for learning is document structure data having items and contents corresponding to the items, and when the contents include the quantitative expression, a predetermined second word is included in the title of the item, the second word is specified as the target. The learning device according to claim 9,
18. The unit refinement unit specifies a predetermined second word as the target when the predetermined second word is included in the vicinity of the quantitative expression.
2. The learning device according to claim 1, wherein:
19. The unit refinement unit specifies a predetermined second word as the target when the predetermined second word is included in the vicinity of the quantitative expression.
3. The learning device according to claim 2, wherein:
20. The unit refinement unit specifies a predetermined second word as the target when the predetermined second word is included in the vicinity of the quantitative expression.
4. The learning device according to claim 3, wherein:
21. The unit refinement unit specifies a predetermined second word as the target when the predetermined second word is included in the vicinity of the quantitative expression.
6. The learning device according to claim 5,
22. The unit refinement unit specifies a predetermined second word as the target when the predetermined second word is included in the vicinity of the quantitative expression.
8. The learning device according to claim 7,
23. The unit refinement unit specifies a predetermined second word as the target when the predetermined second word is included in the vicinity of the quantitative expression. The learning device according to claim 9,
24. When the quantitative expression is included in a table in the original data for learning, the unit refinement unit specifies a predetermined second word as the target when at least one of a row heading name and a column heading name of the table includes the second word.
2. The learning device according to claim 1, wherein:
25. When the quantitative expression is included in a table in the original data for learning, the unit refinement unit specifies a predetermined second word as the target when at least one of a row heading name and a column heading name of the table includes the second word.
3. The learning device according to claim 2, wherein:
26. When the quantitative expression is included in a table in the original data for learning, the unit refinement unit specifies a predetermined second word as the target when at least one of a row heading name and a column heading name of the table includes the second word.
4. The learning device according to claim 3, wherein:
27. When the quantitative expression is included in a table in the original data for learning, the unit refinement unit specifies a predetermined second word as the target when at least one of a row heading name and a column heading name of the table includes the second word.
6. The learning device according to claim 5,
28. When the quantitative expression is included in a table in the original data for learning, the unit refinement unit specifies a predetermined second word as the target when at least one of a row heading name and a column heading name of the table includes the second word.
8. The learning device according to claim 7,
29. When the quantitative expression is included in a table in the original data for learning, the unit refinement unit specifies a predetermined second word as the target when at least one of a row heading name and a column heading name of the table includes the second word. The learning device according to claim 9,
30. The unit refinement unit replaces a synonym of the second word with the second word, and then identifies the target. The learning device according to claim 12,
31. The unit refinement unit replaces a synonym of the second word with the second word, and then identifies the target.
19. The learning device according to claim 18,
32. The unit refinement unit replaces a synonym of the second word with the second word, and then identifies the target.
25. The learning device according to claim 24,
33. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a part of the specified object or a predetermined word indicating the whole including the specified object is included in the vicinity of the quantitative expression or among a plurality of words having a dependency relationship with the quantitative expression. The learning device according to claim 12,
34. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a part of the specified object or a predetermined word indicating the whole including the specified object is included in the vicinity of the quantitative expression or among a plurality of words having a dependency relationship with the quantitative expression.
19. The learning device according to claim 18,
35. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a part of the specified object or a predetermined word indicating the whole including the specified object is included in the vicinity of the quantitative expression or among a plurality of words having a dependency relationship with the quantitative expression.
25. The learning device according to claim 24,
36. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a part of the specified object or a predetermined word indicating the whole including the specified object is included in the vicinity of the quantitative expression or among a plurality of words having a dependency relationship with the quantitative expression.
31. The learning device according to claim 30,
37. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a difference is included in the vicinity of the quantitative expression or among a plurality of words in a dependency relationship with the quantitative expression. The learning device according to claim 12,
38. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a difference is included in the vicinity of the quantitative expression or among a plurality of words in a dependency relationship with the quantitative expression.
19. The learning device according to claim 18,
39. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a difference is included in the vicinity of the quantitative expression or among a plurality of words in a dependency relationship with the quantitative expression.
25. The learning device according to claim 24,
40. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a difference is included in the vicinity of the quantitative expression or among a plurality of words in a dependency relationship with the quantitative expression.
31. The learning device according to claim 30,
41. The unit refinement unit does not convert the normalized unit into the refined unit when a predetermined word indicating a difference is included in the vicinity of the quantitative expression or among a plurality of words in a dependency relationship with the quantitative expression.
34. The learning device according to claim 33,
42. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
2. The learning device according to claim 1, wherein:
43. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
3. The learning device according to claim 2, wherein:
44. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
4. The learning device according to claim 3, wherein:
45. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
6. The learning device according to claim 5,
46. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
8. The learning device according to claim 7,
47. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process. The learning device according to claim 9,
48. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process. The learning device according to claim 12,
49. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
19. The learning device according to claim 18,
50. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
25. The learning device according to claim 24,
51. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
31. The learning device according to claim 30,
52. a numerical value rounding unit that performs a numerical value rounding process to make the number of digits of the normalized numerical value a predetermined number of digits, the numerical expression learning unit learns the numerical expression language model using the normalized numerical values after the numerical rounding process.
34. The learning device according to claim 33,
53. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
2. The learning device according to claim 1, wherein:
54. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
3. The learning device according to claim 2, wherein:
55. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
4. The learning device according to claim 3, wherein:
56. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
6. The learning device according to claim 5,
57. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
8. The learning device according to claim 7,
58. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value. The learning device according to claim 9,
59. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value. The learning device according to claim 12,
60. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
19. The learning device according to claim 18,
61. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
25. The learning device according to claim 24,
62. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
31. The learning device according to claim 30,
63. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
34. The learning device according to claim 33,
64. the numerical expression learning unit learns the numerical expression language model by setting a probability of masking the normalized numerical value to be higher than a probability of masking a portion other than the normalized numerical value.
43. The learning device according to claim 42,
65. The numerical expression learning unit learns the numerical expression language model so that when an error occurs in the estimation of the normalized numerical value, a larger penalty is imposed than when an error occurs in the estimation of other parts.
65. A learning device according to any one of claims 1 to 64.
66. an estimation target data acquisition unit that acquires estimation target data that requests estimation of a numerical value at a position where the predetermined expression format is arranged, using a predetermined expression format and unit; a unit normalization unit that normalizes the units; a unit refinement unit that identifies, in the estimation target data, an object that is a physical existence and an attribute that is a property of the object, which is indicated by the predetermined expression format and the unit, and converts the normalized unit into a refinement unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; a quantitative expression estimation unit that estimates a numerical value at a position where the predetermined expression format is located by inputting data including the predetermined expression format and the converted detailed unit into a quantitative expression language model that has been trained using data including the detailed unit and a normalized numerical value. An estimation device comprising:
67. The unit detailing unit refers to object attribute-specific detailing unit information that associates a plurality of objects, a plurality of attributes, a plurality of units, and a plurality of detailing units, and thereby identifies the detailing unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit.
67. The estimation device of claim 66.
68. the unit refinement unit specifies a predetermined first word as the attribute when the predetermined first word is included in the predetermined expression format and in the vicinity of the unit; 67. The estimation device of claim 66.
69. the unit refinement unit specifies a predetermined first word as the attribute when the predetermined first word is included in the predetermined expression format and in the vicinity of the unit; 68. The estimation device of claim 67.
70. the unit refinement unit specifies a predetermined first word as the attribute when the predetermined first word is included in a plurality of words having a dependency relationship with the predetermined expression format and the unit; 67. The estimation device of claim 66.
71. the unit refinement unit specifies a predetermined first word as the attribute when the predetermined first word is included in a plurality of words having a dependency relationship with the predetermined expression format and the unit; 68. The estimation device of claim 67.
72. The unit refinement unit replaces a synonym of the first word with the first word, and then identifies the attribute.
69. The estimation device of claim 68.
73. The unit refinement unit replaces a synonym of the first word with the first word, and then identifies the attribute.
71. The estimation device according to claim 70,
74. the unit refinement unit specifies a predetermined second word as the target when the predetermined expression format and the unit include the predetermined second word in the vicinity of the unit.
74. An estimation device according to any one of claims 66 to 73, characterized in that
75. The unit refinement unit replaces a synonym of the second word with the second word, and then identifies the target.
75. The estimation device of claim 74.
76. Computer, a quantitative expression identification unit that identifies, from the original learning data, a quantitative expression that represents a quantity using a numerical value and a unit; a numerical value normalization unit that normalizes the numerical value; a unit normalization unit that normalizes the units; a unit refinement unit that identifies, in the original learning data, an object that is a physical entity and an attribute that is a property of the object, which is indicated by the numerical value and the unit, and converts the normalized unit into a refinement unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; and and functioning as a numerical expression learning unit that uses data including the refinement unit and the normalized numerical value to learn a numerical expression language model for estimating the normalized numerical value. A program characterized by.
77. Computer, an estimation target data acquisition unit that acquires estimation target data that requests estimation of a numerical value at a position where the predetermined expression format is arranged, using a predetermined expression format and unit; a unit normalization unit that normalizes the units; a unit refinement unit that identifies an object, which is a physical entity, and an attribute, which is a property of the object, indicated by the predetermined expression format and the unit in the estimation object data, and converts the normalized unit into a refinement unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; and and inputting data including the predetermined expression format and the converted refinement unit into a quantitative expression language model trained using data including a refinement unit and a normalized numerical value, the quantitative expression language model functions as a quantitative expression estimation unit that estimates a numerical value at a position where the predetermined expression format is located. A program characterized by.
78. A quantitative expression identification unit identifies, from the original learning data, a quantitative expression that represents a quantity using a numerical value and a unit; a numerical value normalization unit normalizing the numerical value; a unit normalization unit normalizing the units; a unit refinement unit that identifies, in the original learning data, an object that is a physical entity and an attribute that is a property of the object, which is indicated by the numerical value and the unit, and converts the normalized unit into a refinement unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; a quantitative expression learning unit that uses data including the refinement unit and the normalized numerical value to learn a quantitative expression language model for estimating the normalized numerical value; A learning method characterized by:
79. An estimation target data acquisition unit acquires estimation target data, which requires estimation of a numerical value at a position where the predetermined expression format is arranged, using a predetermined expression format and unit; a unit normalization unit normalizing the units; a unit refinement unit that identifies, in the estimation target data, an object that is a physical existence and an attribute that is a property of the object, which is indicated by the predetermined expression format and the unit, and converts the normalized unit into a refinement unit that uniquely corresponds to a combination of the identified object, the identified attribute, and the normalized unit; a quantitative expression estimation unit inputting data including the predetermined expression format and the converted refinement unit into a quantitative expression language model trained using data including the refinement unit and a normalized numerical value, thereby estimating a numerical value at a position where the predetermined expression format is arranged; An estimation method characterized by:
Citation Information
Patent Citations
Information processing apparatus and method
JP2021149935A
Machine-learned desking vehicle recommendation
US20230080589A1