Information processor, information processing method and program

A large-scale language model is used to identify and compare monetary expressions in documents, enhancing accuracy and reducing human error in data processing by automatically determining and displaying identical monetary expressions.

JP2025180839APending Publication Date: 2025-12-11BLENDING TECH CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024088448
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Conventional methods for detecting inconsistencies in textual data require storing a vast amount of factual data in a database, making it difficult to cover all necessary information for accurate inconsistency detection.

Method used

Utilizing a large-scale language model to extract and identify monetary expressions in documents, assign attribute information, and determine identical expressions, with a display control unit to overlay the results on the document.

Benefits of technology

Enables accurate determination of monetary expressions in documents, reducing the need for manual matching and minimizing human error in data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025180839000001_ABST
    Figure 2025180839000001_ABST
Patent Text Reader

Abstract

To appropriately determine the identity of monetary expressions included in a disclosure document disclosed by a company using a large-scale language model.SOLUTION: An information processor 100 includes: an assignment unit 120 which extracts money amount data contained in an object document 10 disclosed by a company and attribute information for identifying the money amount data using an LLM and assigns the extracted attribute information to the extracted money amount data; determination part 130 which identifies a plurality of money amount data of identical type from the money amount data contained in the object document 10 according to the attribute information and compares the plurality of money amount data to determine whether the plurality of money amount data are identical or not; and an output control part 160 which displays the determination result from the determination part 130 over the object document 10 in which the attribute information is added.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program capable of handling linguistic information. [Background technology]

[0002] Conventionally, there are techniques for handling linguistic information. For example, a technique has been proposed in which inconsistent data and corresponding expressions in text are corrected using a factual data database (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 11-167576 Summary of the Invention [Problem to be solved by the invention]

[0004] In the above-mentioned conventional technology, in order to detect inconsistencies from factual data contained in text, it is necessary to store factual data indicating a large number of facts in a factual data database. However, it is considered difficult to store all of the factual data necessary for appropriate inconsistency detection in the factual data database.

[0005] The present invention aims to utilize a large-scale language model to appropriately determine the identity of monetary expressions contained in disclosure documents disclosed by companies. [Means for solving the problem]

[0006] One aspect of the present invention is an information processing device that includes an assignment unit that uses a large-scale language model to extract monetary expressions and attribute information that can identify the monetary expressions, which are included in disclosure documents disclosed by a company, and assigns the extracted attribute information to the extracted monetary expressions; a determination unit that identifies multiple monetary expressions of the same type from among the monetary expressions included in the disclosure documents based on the attribute information, and compares the multiple monetary expressions to determine whether the multiple monetary expressions are identical; and a display control unit that overlays and displays the determination result by the determination unit on the disclosure document to which the attribute information has been added. [Effects of the Invention]

[0007] According to the present invention, it is possible to appropriately determine the identity of monetary expressions contained in disclosure documents disclosed by companies by utilizing a large-scale language model. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram illustrating an example of a functional configuration of an information processing device. [Figure 2] FIG. 2 is a diagram showing an example of a target document to be processed by the information processing device. [Figure 3] FIG. 3 is a diagram illustrating an example of attribute information assigned by the assigning unit. [Figure 4] FIG. 4 is a diagram showing a schematic flow of an extraction process when extracting amount data and attribute information of the amount data contained in a target document using LLM. [Figure 5] FIG. 5 is a diagram showing an example of the relationship between input data input to the LLM and output data output from the LLM. [Figure 6] FIG. 6 is a simplified diagram showing a tagged target document output from the LLM in response to an input of the target document. [Figure 7] FIG. 7 is a diagram showing a comparative example of price data and its attribute information included in a tagged target document. [Figure 8] FIG. 8 is a diagram showing an example of a notification when a named entity in which an error has been detected in the determination process is notified to the user. [Figure 9] FIG. 9 is a flowchart showing an example of the determination process. [Figure 10] FIG. 10 is a diagram illustrating an example of use of the information processing device. [Figure 11] FIG. 11 is a diagram illustrating an example of use of the information processing device. [Figure 12] FIG. 12 is a block diagram illustrating an example of a functional configuration of an information processing device. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings.

[0010] [Configuration example of information processing device] 1 is a block diagram showing an example of the functional configuration of an information processing device 100. The information processing device 100 can be realized by an information processing device or electronic device such as a server, a personal computer, a smartphone, or a tablet terminal.

[0011] FIG. 2 is a diagram illustrating an example of a target document 10 to be processed by the information processing device 100. The target document 10 is an example of a document including text information 11 and table information 12. While FIG. 2 illustrates a target document 10 including both text information 11 and table information 12, the present embodiment is not limited to this example and can also be applied to a target document including either text information or table information. The target document 10 is an example of a financial statement disclosed by a company. This financial statement is, for example, a securities report. For example, numbers in a securities report are often related between tables and between tables and text. Furthermore, when matching monetary data in a securities report, manual matching is likely to result in human error. Therefore, in this embodiment, a large-scale language model is used to automatically determine the identity of monetary data included in the target document 10, such as a securities report. This improves the accuracy of matching monetary data.

[0012] Here, the text information 11 is information in which various information such as characters, numbers, symbols, etc. are written in a sentence format, and the table information 12 is information in which information such as characters, numbers, symbols, etc. are written in a table format (for example, a table format or a graph format).

[0013] 1, the information processing device 100 includes an acquisition unit 110, an attachment unit 120, a determination unit 130, a recording control unit 140, a storage unit 150, an output control unit 160, and an output unit 170. Each of the acquisition unit 110, the attachment unit 120, the determination unit 130, the recording control unit 140, and the output control unit 160 is realized by, for example, one or more processing circuits such as a central processing unit (CPU) or a graphics processing unit (GPU).

[0014] The acquiring unit 110 accepts target information (for example, target document 10) input by the user, and outputs the target information to the tagging unit 120. This target information is information to be tagged.

[0015] For example, the acquisition unit 110 can be an input device (e.g., a keyboard, a mouse, a recording medium reader, an imaging device, a voice input device, or a scanner) that can input target information including characters, numbers, etc. For example, the voice input device can be a microphone that can input information such as characters, numbers, etc. by voice, a voice input device such as an input device dedicated to voice recognition, etc. For example, the imaging device can be an image acquisition device such as a camera that can capture information such as characters, numbers, etc. and acquire the image information. Note that when image information such as characters, numbers, etc. is acquired by an imaging device, a scanner, etc., it is possible to acquire the information such as characters, numbers, etc. contained in the image information using known character recognition technology. For example, the target information can be read and acquired from a recording medium (e.g., a memory card, a Universal Serial Bus (USB) memory, a Hard Disk Drive (HDD), a Compact Disc (CD), a Digital Versatile Disc (DVD), a Blu-ray (registered trademark) Disc (BD), etc.) that stores a file in which the target information is stored (e.g., a file of the target document 10). Furthermore, for example, a recording medium on which the target information is stored can be connected to the information processing device 100 via wireless or wired communication, and the acquisition unit 110 can read and acquire the target information from the recording medium. Furthermore, if the recording medium is built into the information processing device 100, the acquisition unit 110 can read and acquire the target information from the recording medium. In this way, the acquisition unit 110 is realized by an input interface, a file input device, an imaging device, etc. Note that the file format of the information acquired by the acquisition unit 110 is not particularly limited. For example, any format such as PDF or HTML may be used.

[0016] The assigning unit 120 extracts named entities included in the target information output from the acquiring unit 110 and attribute information that can identify the named entities, and assigns the extracted attribute information to the extracted named entities. That is, the assigning unit 120 tags the named entities included in the target information. The assigning unit 120 then outputs the tagged target information (tagged target document 20 (see FIG. 6 )) to the determining unit 130, the recording control unit 140, and the output control unit 160.

[0017] Here, a named entity refers to a word or phrase with a unique name, and often refers to a proper noun that is specifically limited to nouns. Examples of proper nouns include monetary amounts such as "100 yen," "198 million yen," or "987 US dollars," place names such as "Tokyo," "Osaka," or "Sapporo," or names such as "Yamada Ichiro" or "Tanaka Goro."

[0018] For example, the Message Understanding Conference (MUC) defines the following seven types as named entities: Organization name (ORGANIZATION), person name (PERSON), place name (LOCATION), date expression (DATE), time expression (TIME), amount expression (MONEY), percentage expression (PERCENT)

[0019] Named Entity Recognition (NER), which extracts named entities contained in a target text, is also known as a natural language processing technology. For example, various extraction methods can be used for named entity extraction, such as extraction using dictionary information, rule-based extraction, and machine learning extraction.

[0020] In this embodiment, an example will be described in which a securities report is the target document 10 and monetary expressions (also referred to as monetary data) are extracted as named entities. However, as will be described later, this embodiment can also be applied to cases in which other named entities are extracted from other target documents.

[0021] In addition, in this embodiment, an example of extracting one or more pieces of attribute information related to a named entity will be described. Also, associating attribute information extracted for a named entity with the named entity will be described as tagging. For example, it is possible to tag attribute information (1) to (4) described below as attribute information related to monetary amount data. Note that the attribute information can also be referred to as tag information, metadata, accompanying information, additional information, etc. Note that the method of extracting attribute information will be described in detail with reference to Figs. 4 to 6, etc.

[0022] The determination unit 130 determines whether the named entities tagged with attribute information by the assignment unit 120 are identical, and outputs the determination result to the recording control unit 140 and the output control unit 160. For example, the determination unit 130 identifies multiple named entities of the same type from among the named entities included in the target information based on the attribute information tagged to the named entities. The determination unit 130 then compares the multiple named entities and determines whether the multiple named entities are identical. This determination process will be described in detail with reference to FIG. 7 etc.

[0023] The recording control unit 140 executes recording control to record the target information tagged by the tagging unit 120 and the determination result by the determination unit 130 in the storage unit 150 .

[0024] The storage unit 150 is a storage medium that stores various types of information. For example, the storage unit 150 stores various types of information (e.g., control programs) required for the acquisition unit 110, the assignment unit 120, the determination unit 130, the recording control unit 140, and the output control unit 160 to perform various processes. As the storage unit 150, various storage media such as a read-only memory (ROM), a random access memory (RAM), a static random access memory (SRAM), a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof can be used.

[0025] The output control unit 160 executes output control to cause the output unit 170 to output the target information tagged by the tagging unit 120 and the determination result by the determination unit 130.

[0026] The output unit 170 outputs various types of information based on the control of the output control unit 160. For example, the output unit 170 can be configured with a display unit 171 (see FIG. 8 ) and a sound output unit. The display unit 171 displays various images based on instructions from the output control unit 160. For example, a display panel such as an organic EL (Electro Luminescence) panel or an LCD (Liquid Crystal Display) panel can be used as the display unit 171. The sound output unit outputs various sounds based on instructions from the output control unit 160. For example, one or more speakers can be used as the sound output unit. Note that the output unit 170 is an example of a user interface, and other user interfaces may also be used. For example, the information processing device 100 may output the target information tagged by the tagging unit 120 and the determination result by the determination unit 130 to an external output device (for example, a display device or a sound output device), and cause the output device to output the tagged target information and the determination result.

[0027] 1 shows an example in which the acquisition unit 110, the assignment unit 120, the determination unit 130, the recording control unit 140, the storage unit 150, the output control unit 160, and the output unit 170 are provided in the information processing device 100, but at least one of these may be used as a separate device different from the information processing device 100. For example, by registering at least one of the acquisition device, the storage device, and the output device in advance in the information processing device 100 (for example, by pairing or connecting using a wired or wireless line), it is possible to make the registered device function as the acquisition unit, the storage unit, and the output unit of the information processing device 100. Furthermore, as shown in FIG. 10, the information processing device 100 may be used as a user terminal, or as shown in FIG. 11, the information processing device 100 may be used as a server.

[0028] [Tag format example] 3 is a diagram showing an example of attribute information assigned by the assigning unit 120. In FIG. 3, the data of the amount expression (amount data) included in the target document 10 is <money>Specifically, as shown in a rectangle 200, the following attribute information (1) to (4) is given as attributes of the MONEY tag. (1) type (2) fiscal (3) unit (4) amount

[0029] (1) type is attribute information that indicates the type of monetary data. In other words, it is attribute information that indicates what type of monetary data it is. For example, it can be assigned "sales," "operating profit," "capital," "capital reserve," etc.

[0030] (2) Fiscal is attribute information that indicates the fiscal year of the monetary data. For example, it is possible to assign "previous period" or "current period."

[0031] (3) unit is attribute information indicating the unit of the monetary amount data. For example, "yen" or "million yen" can be assigned. Note that instead of the unit of the monetary amount data, attribute information indicating the type of currency of the monetary amount data (for example, "yen" or "US dollar") may be used.

[0032] (4) amount is attribute information that indicates the integer value of the amount data. For example, if the amount data is "7,118 million yen," it is possible to assign "7118000000."

[0033] The character strings extracted for (1) to (4) above are stored in the "string" within dotted rectangles 201 to 204 shown in rectangle 200. The method for extracting the amount data and attribute information will be described later.

[0034] Note that the attribute information of the MONEY tag shown in Figure 3 is an example, and some of it may be omitted or other attribute information may be added as needed. For example, the above-mentioned (2) may be omitted, and only the above-mentioned attribute information (1), (3), and (4) may be added and used. For example, if the target document contains information on capital, capital reserves, etc., it is possible to add and use only the above-mentioned attribute information (1), (3), and (4).

[0035] Also, for example, identification information "id" to be assigned to the amount data can be assigned as attribute information. For example, a serial number (1, 2, 3, ..., etc.) can be used as this identification information "id."

[0036] Additionally, for example, department-related information "segment" can be added as attribute information. In other words, the attribute information "segment" indicates which business segment within the reporting segment it is. For example, attribute information "segment" can be "consulting business," "business solutions business," "education business," etc.

[0037] Also, for example, information "title" indicating a heading corresponding to a parent concept can be assigned as attribute information. For example, "Financial status and business performance status" can be assigned as attribute information "title."

[0038] Also, for example, information "increase_decrease" (or "up_down") that indicates in detail what type of amount the amount data is can be assigned as attribute information. For example, information such as "total" or "increase ratio" can be assigned as attribute information "increase_decrease" to further classify amount data classified as "type" into more detailed types.

[0039] When extracting each of these pieces of attribute information using LLM, it is important to determine attribute names that are consistent with the information that LLM can handle. Therefore, in this embodiment, an example will be shown in which the names of each of the above-mentioned attribute information (1) to (4) are used.

[0040] [Example of attribute information extraction] Next, a description will be given of an extraction method for extracting amount data included in the target document 10 and an extraction method for extracting attribute information of the amount data. These extraction processes are executed by the attachment unit 120 shown in FIG.

[0041] For example, the attachment unit 120 can extract amount data and attribute information of the amount data contained in the target document 10 using an AI (Artificial Intelligence) model (e.g., a machine learning model generated by machine learning). The term "learning" used in this embodiment refers to discovering patterns behind a large amount of data based on the data. The AI ​​model generated by learning used in this embodiment is generated using various learning algorithms. For example, various types of content (e.g., text, images) and text describing each of these contents can be read as training data to learn in advance, and this AI model can then be used in the extraction process.

[0042] Examples of AI models that can be used include large language models (LLMs) and multimodal LLMs. Examples of LLMs that can be used include various natural language processing models (e.g., Bidirectional Encoder Representations from Transformers (BERT)), Generative Pre-trained Transformer (ChatGPT), GPT-4, GPT-4o, Bard, and Large Language Model Meta AI (Llama). These are merely examples, and other AI models may also be used. Figure 4 shows an example of an extraction process that uses an LLM 121 to extract amount data and attribute information for the amount data contained in a target document 10.

[0043] [Example of extraction processing using LLM] FIG. 4 is a diagram showing a flow of extraction processing when the LLM 121 is used to extract amount data and attribute information of the amount data contained in the target document 10. In FIG.

[0044] The LLM 121 is realized by an information processing device 100 (see FIGS. 10 and 11) that can execute various extraction processes using the LLM. For example, the attachment unit 120 transmits the target document 10 acquired by the acquisition unit 110 to a server, and the server executes an extraction process of the amount data contained in the target document 10 and the attribute information of the amount data. The attachment unit 120 can receive and use the processing result (tagged target document 20) from the server. The tagged target document 20 is information in which the amount data extracted from the target document 10 and the attribute information of the amount data are associated with the target document 10. Note that these can be executed by installing a predetermined application in the information processing device 100. Furthermore, for example, if the information processing device 100 is a device with a high computing speed (e.g., a server), the LLM process may be executed in the information processing device 100.

[0045] 5 is a diagram showing an example of the relationship between input data 210 input to the LLM 121 and output data 220 output from the LLM 121. The input data 210 corresponds to a part of a sentence included in the target document 10.

[0046] Here, when ChatGPT or GPT-4 is used as the LLM 121, instruction information (prompts) to be input to these LLMs will be described. For example, when assigning tags using an LLM, as described above, it is important to set attribute names that the LLM 121 can accept in the attribute information. In other words, it is important to set attribute names that are consistent with the information in the LLM 121. Therefore, in this embodiment, the attribute names (1) to (4) described above are set as attribute names that satisfy this condition.

[0047] For example, a user can create sample data that includes one (or more) sentences before tagging and sample sentences with the sentences tagged, and input instructions (prompts) to tag each sentence in the same way as this sample data into the LLM, and the output data (target document 10, amount data, attribute information for the amount data) can be used as tagged target information (e.g., tagged target document 20).

[0048] For example, the sentences contained in the input data 210 shown in Figure 5(A) and the sentences and attribute information contained in the output data 220 shown in Figure 5(B) are used as sample data, and instructions (prompts) to tag each sentence in the same way as this sample data are input to LLM121, and the corresponding output data (target document 10, amount data, attribute information of the amount data) can be used as tagged target information (for example, tagged target document 20).

[0049] Also, for example, monetary data can be extracted as a named entity, and instruction information (prompts) to extract (1) type, (2) fiscal, (3) unit, and (4) amount as attribute information of the extracted monetary data can be input to the LLM, and the corresponding output data (target document 10, monetary data, and attribute information of the monetary data) can be used as tagged target information (e.g., tagged target document 20). Note that these prompts are merely examples and are not limiting. Other prompts that can extract monetary data contained in target document 10 and attribute information of the monetary data can also be used.

[0050] For example, in the case of LLM, it is possible to process information such as characters, numbers, and symbols in table information 12 written in tabular form. In this case, for example, it is possible to identify and process information such as characters, numbers, and symbols in table information 12 based on each piece of information (e.g., item, cost) written above or to the side of the table in table information 12. For example, because "unit (million yen)" is written in the upper right corner of table information 12 (see FIG. 2), LLM can determine that each number written in table information 12 is related to monetary data related to the currency "yen."

[0051] Note that other extraction methods using machine learning may also be employed. For example, machine learning may be performed using a large amount of text and training data labeled with named entities and their corresponding attribute information to generate a trained model, and the trained model may be used to extract named entities and their corresponding attribute information.

[0052] Alternatively, an extraction method may be used that extracts named entities using dictionary information. For example, named entities to be extracted may be registered in advance as dictionary information (database), and named entities may be extracted by referencing this dictionary information. However, when an extraction method using dictionary information is used, there is a possibility that duplicate personal names and place names may be extracted. For example, the word "Tamachi" may be a personal name "Tamachi" and a place name "Tamachi." Furthermore, since it is difficult to register all numerical values ​​in dictionary information, it is preferable to extract numerical information using other character recognition techniques.

[0053] Alternatively, an extraction method using a rule base to extract named entities may be used. For example, certain rules for extracting named entities may be defined in advance, and named entities may be extracted based on those rules. For example, in the case of the word "Tamachi," a rule may be defined in advance that if "Tamachi" is followed by "san," "kun," or "chan," it is a person's name, and otherwise it is a place name. In other words, it is important to assign a part of speech to each word to be extracted and to define rules based on the surrounding words and parts of speech. In this way, extraction methods that extract named entities using dictionary information or a rule base involve humans devising certain rules and extracting information to be extracted based on whether or not the rules apply. Therefore, it may be difficult to determine all the rules in advance when there is a large amount of information to be extracted, or when the information to be extracted is numerical information. For this reason, the following describes an example of extracting named entities and attribute information using a large-scale language model.

[0054] [Example of output result] Fig. 6 is a simplified diagram showing a tagged target document 20 output from the LLM 121 in response to input of the target document 10. As shown in Fig. 6, a tagged target document 20 is generated in which the amount data extracted by the LLM 121 is associated with corresponding attribute information. Note that in Fig. 6, information corresponding to the text information 11 (see Fig. 2) is shown as text information 21, and information after conversion of the table information 12 (see Fig. 2) into the character information and numeric information contained therein is shown as text information 22.

[0055] Fig. 7 is a diagram showing the amount data and its attribute information included in the tagged target document 20 shown in Fig. 6, enclosed in dotted-line rectangles 31 to 39. The type, fiscal, unit, and amount included in each rectangle 31 to 39 correspond to the attribute information (1) to (4) described above.

[0056] 7, among the monetary data shown enclosed in dotted rectangles 31-39, monetary data identified as multiple monetary data of the same type based on their attribute information are shown connected by arrows C1-C3. That is, arrow C1 indicates dotted rectangles 31 and 34, arrow C2 indicates dotted rectangles 32 and 35, and arrow C3 indicates dotted rectangles 33 and 36. For these multiple monetary data of the same type, the monetary data is compared to determine whether the respective monetary data are the same.

[0057] Specifically, the determination unit 130 identifies the amount data included in the tagged target document 20 based on the attribute information of the amount data included in the tagged target document 20. For example, the determination unit 130 determines the amount data included in the tagged target document 20 based on the attribute information of the amount data included in the tagged target document 20. <money type="…"> …circle< / money> The part containing " " is identified. A known character recognition technique can be used to identify the character string within " ". Next, the determination unit 130 determines whether the identified character string " <money type="…"> …circle< / money> ' to obtain the attribute information of the amount data included in the specified string ' <money type="…"> …circle< / money> The attribute information of the amount data included in " " can be easily identified because each position is uniform.

[0058] Next, the determination unit 130 determines whether the specified character string " <money type="…"> …circle< / money> ', the determination unit 130 identifies the same type of amount data from among the amount data included in the identified character string ' <money type="…"> …circle< / money> Among the attribute information (type, fiscal, unit) included in the attribute information, the determination unit 130 identifies multiple pieces of monetary data whose attribute information matches. That is, the determination unit 130 identifies monetary data of the same type. For example, as shown in FIG. 7, the monetary data within the dotted rectangles 31 and 34 indicated by the arrow C1, the monetary data within the dotted rectangles 32 and 35 indicated by the arrow C2, and the monetary data within the dotted rectangles 33 and 36 indicated by the arrow C3 are identified as monetary data of the same type. Note that in the example shown in FIG. 7, the unit of the attribute information includes both "yen" and "million yen." Therefore, when identifying monetary data of the same type, it may be determined whether the portion of the unit of the attribute information corresponding to the currency (e.g., "yen") matches, or the determination of whether the unit matches may be omitted. Furthermore, if it is possible to identify monetary data of the same type using only one of the attribute information's type and fiscal, only one of them may be used. That is, it is possible to identify monetary data of the same type using at least one piece of attribute information. The value of the "amount" in the attribute information can also be calculated using the "unit" in the attribute information and the tagged amount data. For example, if the "unit" is "yen" and the tagged amount data is "277," then the amount can be calculated as 277. If the "unit" is "million yen" and the tagged amount data is "277," then the amount can be calculated as 277,000,000 (=277×1,000,000). For this reason, extraction of the amount can be omitted in the tagging process using LLM, and the determination unit 130 can calculate the amount using the "unit" in the attribute information and the tagged amount data.

[0059] Next, the determination unit 130 compares the multiple amount data identified as the same type and determines whether these amount data are the same. Specifically, the determination unit 130 compares the amount values ​​of the amount data identified as the same type and determines whether they are the same. For example, the values ​​of the amount data within dotted rectangles 31 and 34 indicated by arrow C1 are the same. The values ​​of the amount data within dotted rectangles 32 and 35 indicated by arrow C2 are the same. On the other hand, the values ​​of the amount data within dotted rectangles 33 and 36 indicated by arrow C3 are different. Therefore, although the values ​​of the amount data indicated by arrows C1 and C2 are the same, it is possible to determine that at least one of the values ​​of the amount data indicated by arrow C3 is incorrect. In this way, it is also possible that the attribute information (type, fiscal, unit) matches, but the amount information does not match. In this case, it is possible to determine that one of the amount information (i.e., amount data) being compared is incorrect. If the attribute information (type, fiscal, unit) matches but the amount information does not match, it is possible that an error has occurred in the tagging process by the tagging unit 120. In this case, the error must be resolved by user verification.

[0060] In this way, by using the attribute information of the amount data, it is possible to automatically perform the matching work of each amount data contained in the target document 10. This eliminates the need for an operator to visually perform the matching work of each amount data contained in the target document 10. It is also possible to improve the accuracy of matching each amount data contained in the target document 10.

[0061] Furthermore, since it is possible to use information (amount) that quantifies the total amount of the amount data as attribute information of the amount data, the determination unit 130 can easily grasp the numerical values ​​to be compared. Therefore, it is possible to omit the process of quantifying the total amount of the amount data to be compared in the comparison process by the determination unit 130. Therefore, it is possible to reduce the calculation process related to the comparison process.

[0062] Furthermore, as the attribute information of the amount data, it is possible to use at least one of the type of amount data, fiscal year information of the amount data, and unit type of the amount data. Therefore, it is possible to perform appropriate comparison processing using a relatively small amount of attribute information. This makes it possible to reduce the amount of calculation processing involved in the comparison processing.

[0063] The recording control unit 140 stores the tagged target document 20 in the storage unit 150. In this case, various information related to the tagged target document 20 can be stored as appropriate. For example, the tagged target document 20 may be stored in the storage unit 150 in its original format (see FIG. 7 ), or named entities and attribute information may be extracted from the tagged target document 20, and the named entities and attribute information may be associated with the tagged target document 20 (or target document 10) and stored in the storage unit 150. Alternatively, only the attribute information of named entities included in the tagged target document 20 may be stored in the storage unit 150. In this way, the various information related to the tagged target document 20 stored in the storage unit 150 can be used as appropriate when the user needs it.

[0064] The output control unit 160 outputs the tagged target document 20 from the output unit 170. In this case, various information related to the tagged target document 20 can be output as appropriate. For example, the tagged target document 20 may be displayed on the display unit 171 of the output unit 170 in its original format (see FIG. 7 ), or named entities and attribute information may be extracted from the tagged target document 20, and the named entities and attribute information may be associated with the tagged target document 20 (or target document 10) and displayed on the display unit 171 of the output unit 170. Alternatively, only the attribute information of the named entities included in the tagged target document 20 may be displayed on the display unit 171 of the output unit 170. Furthermore, various information related to the tagged target document 20 may be output as audio, or may be transmitted to another device and output from that device. In this way, various information related to the tagged target document 20 can be provided as needed by the user by displaying, outputting as audio, transmitting, etc.

[0065] [Example of notification of a named entity spelling error] Fig. 8 is a diagram showing an example of notification when a named entity in which an error has been detected in the determination process by the determination unit 130 is notified to the user. Fig. 8 shows an example in which the tagged target document 20 is displayed on the display unit 171 of the output unit 170, and the portion of the named entity in which an error has been detected in the tagged target document 20 is indicated by being surrounded by triangles E1 and E2, and an example in which audio information S1 indicating that an error has been detected is output from the audio output unit of the output unit 170. Fig. 8 also shows an example in which the portion of the named entity that is the determination target in the tagged target document 20 is surrounded by ellipses J1 to J4. Note that the determination process by the determination unit 130 is the same as the determination process described with reference to Fig. 7. Note that the error notification method shown in Fig. 8 is just an example, and other notification methods may be used.

[0066] Here, in the case of documents with a fixed format, such as securities reports, the position of each item in the text and the flow of each content are almost fixed. Therefore, workers who review target documents such as securities reports often know the location of the parts of the target document where the content to be reviewed is described. Therefore, even if tagged target document 20 is displayed as a notification screen that notifies the worker reviewing a target document such as a securities report of the review results, the worker can easily understand the content to be reviewed. Note that, if page information needs to be displayed, tagged target document 20 with page information of target document 10 added may be displayed. Similarly, if other information that needs to be displayed exists, tagged target document 20 with the necessary information added can be displayed.

[0067] It is also possible to display attribute information in a table format as the output result of the LLM. For example, it is possible to assign an ID (e.g., identification information such as a serial number) to each piece of attribute information, and display a list of attribute information associated with this ID in a table format (e.g., in serial number order). Here, because the LLM processes information incorporating surrounding information, when checking the output result of the LLM, it is important to verify the accuracy by checking the content before and after the original location. For this reason, for example, if only attribute information is displayed as the output result of the LLM, it is expected that it will be difficult to verify why the attribute information was extracted by the LLM. In contrast, in this embodiment, the tagged target document 20 containing the amount data and the attribute information to be evaluated is displayed on the display unit 171, making it possible to verify the accuracy by checking the content before and after the part to be evaluated. This makes it easy to verify why the attribute information was extracted by the LLM.

[0068] [Example of operation of information processing device] 9 is a flowchart showing an example of a determination process in information processing device 100. This determination process is executed based on a program stored in storage unit 150. This determination process is executed when target document 10 is acquired by acquisition unit 110. This determination process will be explained with appropriate reference to FIGS. 1 to 8.

[0069] FIG. 9 shows an example in which, if an error is detected in the determination process, the error is notified to the user.

[0070] In step S501 , the acquisition unit 110 acquires the target document 10 input by the user, and outputs the acquired target document 10 to the attachment unit 120 .

[0071] In step S502, the attachment unit 120 extracts named entities contained in the target document 10 output from the acquisition unit 110, and extracts attribute information related to the extracted named entities. The attachment unit 120 also performs an attachment process to tag the target document 10 with the extracted named entities and the attribute information related to them. The attachment unit 120 then outputs the tagged target document 20 that has been subjected to the attachment process to the determination unit 130, the recording control unit 140, and the output control unit 160. This attachment process is the same as the attachment process described above.

[0072] In step S503, the determination unit 130 executes a determination process to determine whether or not the named entities included in the tagged target document 20 to which attribute information was assigned in step S502 are identical. Then, the determination unit 130 outputs the result of the determination process to the recording control unit 140 and the output control unit 160. This determination process is the same as the determination process shown in FIG.

[0073] In step S504, the output control unit 160 determines whether or not a named entity error was detected by the determination process in step S503. If a named entity error is detected, the process proceeds to step S506. On the other hand, if a named entity error is not detected, the process proceeds to step S505. For example, in the example shown in FIG. 7, an error was detected in at least one of the amount data within the dotted rectangles 33 and 36 indicated by arrow C3, so the process proceeds to step S506.

[0074] In step S505, the output control unit 160 causes the output unit 170 to output information that the named entities included in the target document 10 are the same. For example, the output control unit 160 can cause the display unit 171 of the output unit 170 to display the target document 10 (or the tagged target document 20) and display information that the named entities of the target document 10 are the same. The output control unit 160 may also cause the audio output unit to output audio information that the named entities of the target document 10 are the same.

[0075] In step S506, the output control unit 160 causes the output unit 170 to output, in an identifiable manner, the named entities in which an error has been detected among the named entities contained in the target document 10. For example, as shown in Fig. 8, the output control unit 160 can cause the display unit 171 of the output unit 170 to display the tagged target document 20 and add triangles E1 and E2 to the named entities in which an error has been detected to indicate that an error has been detected. The output control unit 160 may also cause the audio output unit to output audio information S1 indicating that the named entity in the tagged target document 20 is incorrect.

[0076] In step S507, the recording control unit 140 stores the tagged target document 20 to which the attribute information has been assigned in step S502 in the storage unit 150. In this case, the recording control unit 140 may store in the storage unit 150 the tagged target document 20, the target document 10 acquired in step S501, and the attribute information (see FIG. 7) assigned to the corresponding tagged target document 20 in association with each other.

[0077] [Examples of using information processing equipment] 1 shows an example in which the acquisition process, assignment process, determination process, etc. are executed in the information processing device 100. In this case, the user terminal used by the user U1 can be used as the information processing device 100. An example of use in this case is shown in FIG.

[0078] 10 is a diagram showing an example of how the information processing device 100 is used. For example, information processing devices and electronic devices such as a personal computer, a smartphone, a tablet terminal, etc. can be used as the information processing device 100. For example, when a user U1 inputs a target document 10 into the information processing device 100, a tagged target document 20 corresponding to the target document 10 and a determination result for the tagged target document 20 are displayed on the display unit 171.

[0079] Furthermore, all or part of each process, such as the acquisition process, the assignment process, and the determination process, may be executed by another device. In this case, an information processing system is configured by the devices that execute part of each process. For example, at least part of each process can be executed by a device that user U1 can use (e.g., a smartphone, a tablet terminal, a personal computer), various information processing devices such as a server that can be connected via a predetermined network such as the Internet, and various electronic devices. For example, a device different from the user terminal used by user U1 can be used as information processing device 100. An example of use in this case is shown in FIG. 11.

[0080] Furthermore, a part (or all) of the information processing system capable of executing the functions of the information processing device 100 may be provided by an application that can be provided via a predetermined network such as the Internet. This application is, for example, SaaS (Software as a Service).

[0081] FIG. 11 is a diagram showing an example of use of the information processing device 100. For example, the information processing device 100 can be an information processing device such as a server or an electronic device. The user terminal 180 is, for example, an information processing device or an electronic device such as a personal computer, a smartphone, or a tablet terminal. The network NW1 is a network such as a public line network or the Internet. The user terminal 180 and the information processing device 100 are connected to the network NW1 by a communication method using wireless communication or a communication method using wired communication, or by both methods.

[0082] For example, when user U1 inputs target document 10 into user terminal 180, user terminal 180 transmits the target document 10 to information processing device 100. When information processing device 100 receives the target document 10, it performs an assignment process, a determination process, etc. on the target document 10 and transmits the processing results to user terminal 180. User terminal 180 displays the processing results, that is, a tagged target document 20 corresponding to the target document 10 and the determination results for the tagged target document 20, on display unit 171.

[0083] [Example of pre-processing and post-processing] The above has shown an example of performing the assignment process, determination process, etc. using the input target document 10. Here, when the target document 10 contains a large amount of text information, a large amount of table information, etc., it is possible to reduce the computational load related to the calculation processes such as the assignment process, determination process, etc. by performing a predetermined preprocessing. Therefore, Fig. 12 shows an example of performing the assignment process, determination process, etc. after performing the preprocessing on the target information to be processed, and then performing the postprocessing after performing each of these processes.

[0084] [Configuration example of information processing device] Fig. 12 is a block diagram showing an example of the functional configuration of the information processing device 400. Note that the information processing device 400 is a partial modification of the information processing device 100 shown in Fig. 1, with the addition of a pre-processing unit 410 and a post-processing unit 420, and apart from these additions, the information processing device 400 is common to the information processing device 100. For this reason, parts common to the information processing device 100 are assigned the same reference numerals as those in the information processing device 100, and descriptions thereof will be omitted.

[0085] The preprocessing unit 410 performs predetermined preprocessing on the target information output from the acquisition unit 110, and outputs the target information after preprocessing to the annotation unit 120. For example, if the target information includes both text information and table information, the text information and table information can be separated, and the combination of the separated text information and table information can be used as the target information for annotation processing, determination processing, etc. For example, in the case of an HTML file, a table tag is attached to the beginning of the table information, making it possible to recognize the table information. In this way, known segmentation processing can be used to segment the text information and table information.

[0086] Furthermore, for example, when dividing text information, it is possible to divide it based on semantic chunks. In other words, if a sentence contains a meaningful structure, it is preferable to divide it without destroying that structure. For example, it is possible to divide it based on paragraphs, chapters, or the like of the sentence. For dividing these sentences, known language processing can be used.

[0087] Furthermore, for example, in the case of an HTML file, attribute information included in the HTML data (for example, attribute information related to font type, size, color, etc.) is considered to be very long and difficult to handle. Therefore, in the case of an HTML file, the preprocessing unit 410 may perform preprocessing to temporarily delete unnecessary parts (attribute information included in the HTML data).

[0088] The post-processing unit 420 performs predetermined post-processing on the tagged target information output from the tagging unit 120, and outputs the tagged target information after the post-processing (e.g., tagged target document 20 (see FIG. 4)) to the determination unit 130, the recording control unit 140, and the output control unit 160. The post-processing unit 420 performs post-processing corresponding to the pre-processing performed by the pre-processing unit 410. For example, if the pre-processing unit 410 performs a division process to divide text information and table information, the post-processing unit 420 performs a combination process to combine the divided text information and table information. For example, if the pre-processing unit 410 performs a division process to divide text information, the post-processing unit 420 performs a combination process to combine the divided text information. For example, if the pre-processing unit 410 performs a pre-processing to delete unnecessary parts (e.g., attribute information included in HTML data), the post-processing unit 420 performs a process to restore the deleted parts. For these restoration processes, known processing methods can be used.

[0089] [Examples of using other named entities] The above example shows how to extract and use monetary expressions (monetary data) as named entities. However, as mentioned above, it is also possible to extract other named entities and determine errors. Therefore, the following example shows how to extract other named entities and perform the determination process.

[0090] [Example using organization name] First, an example in which an organization name (ORGANIZATION) is extracted and used as a named entity will be described.

[0091] For example, the assigning unit 120 assigns the following to the data of the organization (organization data) included in the target document 10: <orgnization>It is possible to add a tag called "ORGNIZATION tag" to the <orgnization type="String" name="String" en_name="String">It is possible to use "[organization data]" as the tag format (i.e., the expression in quotation marks " "). In this tag format, [organization data] is the extracted organization name. For example, in the case of ABC Co., Ltd., "ABC Co., Ltd." or "(ABC Co., Ltd.)" is stored in [organization data]. In addition, the following attribute information (O1) to (O3) can be assigned as attributes of the ORGNIZATION tag. (O1)type (O2) name (O3)en_name

[0092] (O1) type is attribute information that indicates the type of organizational data. In other words, it is attribute information that indicates what type of organization the organizational data is. For example, in the case of ABC Co., Ltd., it is possible to assign "KK". Note that KK is information that indicates a corporation.

[0093] (O2) name is attribute information indicating the name of the organization data. For example, in the case of ABC Co., Ltd., it is possible to assign "ABC."

[0094] (O3) en_name is attribute information indicating the English notation of the organization data. For example, in the case of ABC Co., Ltd., it is possible to assign "ABC".

[0095] The attribute information of the ORGNIZATION tag shown here is an example, and some of it may be omitted as necessary, or other attribute information may be added.

[0096] Furthermore, when performing error determination processing, the determination unit 130 identifies organizational data of the same type based on the attribute information of the organizational data. Specifically, the determination unit 130 identifies multiple organizational data whose attribute information (e.g., (O1) type, (O2) name, (O3) en_name) matches. The determination unit 130 then compares the multiple organizational data identified as the same type and determines whether these organizational data are the same. For example, the determination unit 130 compares the [organization data] of the organizational data identified as the same type and determines whether they are the same.

[0097] [Example of using a person's name] First, an example in which a person's name (PERSON) is extracted and used as a named entity will be described.

[0098] For example, the assigning unit 120 assigns the following to the data of a person's name (personal name data) included in the target document 10: <person>It is possible to add a tag called "PERSON tag" to a person. For example, <person family_name="String" given_name="String" role="String" en_name="String"> [Personal name data]< / person> " can be used as the tag format (i.e., the expression within " "). In this tag format, [person's name data] is the extracted person's name. In addition, the following attribute information (P1) to (P4) can be assigned as attributes of the PERSON tag. (P1)family_name (P2)given_name (P3)role (P4)en_name

[0099] (P1) family_name is attribute information that indicates the gender and surname of a person's name. For example, in the case of Yamada Taro, it is possible to assign "Yamada."

[0100] (P2) given_name is attribute information indicating the name of a person. For example, in the case of Yamada Taro, it is possible to give "Taro."

[0101] (P3) role is attribute information that indicates the role of a person's name. For example, in the case of the CEO of ABC Co., Ltd., it is possible to assign "CEO." Also, in the case of a section manager of ABC Co., Ltd., it is possible to assign "Section Manager."

[0102] (P4) en_name is attribute information indicating the English spelling of personal name data. For example, in the case of Yamada Taro, it is possible to assign "TARO YAMADA." It is preferable to separate en_name into first and last names.

[0103] The attribute information of the PERSON tag shown here is an example, and some of it may be omitted as necessary, or other attribute information may be added.

[0104] Furthermore, when performing error determination processing, the determination unit 130 identifies personal name data of the same type based on the attribute information of the personal name data. Specifically, the determination unit 130 identifies multiple personal name data whose attribute information (e.g., (O1) type, (P2) given_name, (P3) role, (P4) en_name) matches. Then, the determination unit 130 compares the multiple personal name data identified as the same type and determines whether these personal name data are identical. For example, the determination unit 130 compares the [personal name data] of the personal name data identified as the same type and determines whether they are identical.

[0105] [Configuration example and effects of this embodiment] The information processing devices 100 and 400 utilize an LLM (an example of a large-scale language model) to extract monetary data (an example of a named entity) and attribute information (e.g., type, fiscal, unit, amount) that can identify the monetary data from a target document 10 (an example of target information), and include an assignment unit 120 that assigns the extracted attribute information to the monetary data. The information processing devices 100 and 400 also include a determination unit 130 that identifies multiple monetary data of the same type from the monetary data contained in the target document 10 based on the attribute information, and compares the multiple monetary data to determine whether the multiple monetary data are identical. The information processing method according to this embodiment includes these processes. The program according to this embodiment is a program that causes a computer to execute these processes. In other words, the program according to this embodiment is a program that causes a computer to realize each function executable by each information processing device. The information processing devices 100 and 400 shown here may be configured as a single device or multiple devices. Furthermore, instead of the information processing devices 100 and 400, an information processing system may be configured with a plurality of devices capable of executing the processes realized by the information processing devices 100 and 400. The target document 10 is, for example, a disclosure document (e.g., a securities report) disclosed by a company.

[0106] According to this configuration, it is possible to use LLM to extract attribute information of the amount data included in the target document 10, and to use this attribute information to appropriately determine the identity of the amount data. Furthermore, since it is possible to use LLM to perform a determination process using the attribute information assigned to the target document 10, it is possible to reduce the computational load involved in the determination process. Furthermore, since it is possible to perform a determination process using the attribute information assigned to the target document 10 using LLM, it is possible to improve the determination accuracy of the determination process.

[0107] The assignment unit 120 extracts the target document 10 (an example of target information), amount data (an example of a named entity), and attribute information (e.g., type, fiscal, unit, amount) related to the amount data from the target document 10, inputs a prompt (an example of instruction information) that instructs the assignment of the extracted attribute information to the extracted amount data into an LLM (an example of a large-scale language model), and obtains the output result from the LLM, a tagged target document 20 in which attribute information has been assigned to the amount data.

[0108] According to this configuration, by inputting the target document 10 and a predetermined prompt into the LLM, it is possible to obtain from the LLM a tagged target document 20 in which attribute information is added to the amount data. This makes it possible to appropriately extract the amount data and the attribute information of that amount data using the LLM.

[0109] Target document 10 (an example of target information) includes both text information 11 and table information 12. Addition unit 120 adds attribute information related to amount data (an example of a named entity) to the position of the amount data in text information 11 to generate text information 21, and adds attribute information related to the amount data to the position of the amount data in text information 22 (character information and numerical information) related to converted table information 12, which is obtained by converting table information 12 into the character information and numerical information contained in that table information 12.

[0110] For example, as described above, workers who check target documents such as securities reports often know the location of the part of the target document where the content to be checked is described. Therefore, by displaying tagged target document 20 to workers who check target documents such as securities reports, the content to be checked can be easily understood.

[0111] The information processing devices 100 and 400 further include an output control unit 160 that displays, on the display unit 171, a determination result obtained by using attribute information added to text information 11 and attribute information added to text information 22 (character information and numerical information) related to table information 12, superimposed on the text information 21 to which the attribute information has been added and the text information 22 (character information and numerical information) to which the attribute information has been added. For example, as shown in Fig. 8, the determination results (ovals J1 to J4, triangles E1 and E2) are superimposed and displayed on the text information 21 to which the attribute information has been added and the text information 22 to which the attribute information has been added.

[0112] For example, as described above, workers reviewing target documents such as securities reports often know the location of the portion of the target document where the content to be reviewed is described. Therefore, by displaying the tagged target document 20 with the judgment results (ovals J1-J4, triangles E1 and E2) superimposed, workers reviewing target documents such as securities reports can easily understand the content to be reviewed. Furthermore, by displaying the tagged target document 20, it is possible to verify the accuracy by checking the content before and after the portion to be reviewed. Therefore, it is possible to easily verify why the extracted attribute information was output by the LLM.

[0113] The assignment unit 120 extracts the type (type) of amount data (amount expression) in the target document 10 (an example of target information), the fiscal year information (fiscal) of the amount data, the unit type (unit) of the amount data, and the total numerical value (amount) of the amount data as attribute information.

[0114] According to this configuration, it is possible to appropriately extract attribute information that can improve the accuracy of the determination process by the determination unit 130.

[0115] The judgment unit 130 uses at least one of the type of monetary data (monetary expression), fiscal information of the monetary data, and unit type of the monetary data to identify multiple monetary data of the same type from the monetary data, compares the total amount of the numerical values ​​of the multiple monetary data, and judges whether the multiple monetary data are the same.

[0116] According to this configuration, it is possible to identify multiple amount data of the same type using one or multiple attribute information, and the results of this identification can be used to appropriately determine whether the multiple amount data are the same.

[0117] Note that each processing procedure shown in this embodiment is an example for realizing this embodiment, and the order of some of the processing procedures may be changed within the scope that makes it possible to realize this embodiment, and some of the processing procedures may be omitted or other processing procedures may be added.

[0118] Each process described in this embodiment is executed based on a program that causes a computer to execute each processing procedure. Therefore, this embodiment can also be understood as an embodiment of a program that realizes the function of executing each process and a recording medium that stores the program. For example, an update process for adding a new function to an information processing device can store the program in the storage device of the information processing device. This makes it possible to cause the updated information processing device to execute each process described in this embodiment.

[0119] Although the embodiments of the present invention have been described above, the above embodiments merely illustrate some of the application examples of the present invention, and it is not intended that the technical scope of the present invention be limited to the specific configurations of the above embodiments. [Explanation of symbols]

[0120] 100, 400 information processing device, 110 acquisition unit, 120 assignment unit, 130 determination unit, 140 recording control unit, 150 storage unit, 160 output control unit, 170 output unit, 171 display unit, 180 user terminal, 410 pre-processing unit, 420 post-processing unit, NW1 network< / person> < / orgnization> < / orgnization> < / money>

Claims

1. an attribute unit that uses a large-scale language model to extract monetary expressions and attribute information that can identify the monetary expressions included in disclosure documents disclosed by companies, and assigns the extracted attribute information to the extracted monetary expressions; a determination unit that identifies a plurality of monetary expressions of the same type from among the monetary expressions included in the disclosure document based on the attribute information, and compares the plurality of monetary expressions to determine whether the plurality of monetary expressions are the same; a display control unit that displays the determination result by the determination unit superimposed on the disclosure document to which the attribute information has been added; An information processing device comprising:

2. The assignment unit inputs the disclosure document and instruction information instructing the large-scale language model to extract the monetary expressions and the attribute information related to the monetary expressions from the disclosure document and assign the extracted attribute information to the extracted monetary expressions, and obtains the disclosure document in which the attribute information is assigned to the monetary expressions, which is an output result from the large-scale language model. The information processing device according to claim 1 .

3. the disclosure document includes both textual and tabular information; The adding unit adds the attribute information related to the monetary amount expression to the position of the monetary amount expression in the text information, and adds the attribute information related to the monetary amount expression to the position of the monetary amount expression in converted table information obtained by converting the table information into character information and numeric information included in the table information. The information processing device according to claim 1 .

4. The assigning unit extracts, as the attribute information, the type of monetary expression in the disclosure document, fiscal year information of the monetary expression, the type of unit of the monetary expression, and the total amount of the numerical value of the monetary expression. The information processing device according to claim 1 .

5. The determination unit uses at least one of the type of monetary expression, the fiscal year information of the monetary expression, and the type of unit of the monetary expression to identify a plurality of monetary expressions of the same type from among the monetary expressions, and compares the total amounts of the numerical values ​​of the monetary expressions related to the plurality of monetary expressions to determine whether the plurality of monetary expressions are the same. The information processing device according to claim 4 .

6. an assignment process that uses a large-scale language model to extract monetary expressions and attribute information that can identify the monetary expressions included in the disclosure documents disclosed by the company, and assigns the extracted attribute information to the extracted monetary expressions; a determination process for identifying a plurality of monetary expressions of the same type from among the monetary expressions included in the disclosure document based on the attribute information, and comparing the plurality of monetary expressions to determine whether the plurality of monetary expressions are the same; a display control process for displaying the determination result in the determination process superimposed on the disclosure document to which the attribute information has been added; An information processing method including:

7. an assignment procedure that uses a large-scale language model to extract monetary expressions and attribute information that can identify the monetary expressions included in disclosure documents disclosed by companies, and assigns the extracted attribute information to the extracted monetary expressions; a determination procedure for identifying a plurality of monetary expressions of the same type from among the monetary expressions included in the disclosure document based on the attribute information, and comparing the plurality of monetary expressions to determine whether the plurality of monetary expressions are the same; a display control step of superimposing and displaying the determination result in the determination step on the disclosure document to which the attribute information has been added; A program that causes a computer to execute the following.

Citation Information

Patent Citations

  • Link setting device and method using information extracted from text

    JP2005157853A

  • Data processing system, data processing method, and data structure

    JP2018185716A

  • Information processing device, control method, and program

    JP2020201965A

  • Data obtainment device, data obtainment method, and data obtainment program

    JP2021009591A

  • Information processing device, information processing method and information processing program

    JP2022180714A