Method for recommending xbrl taxonomy element and computing device performing the same
Patent Information
- Application Number
- KR1020250169236
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-09-23
- Estimated Expiration
- 2045-11-11
Smart Images

Figure 112025125753217-PAT00002_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a method for recommending XBRL taxonomy items and a computing device for performing the same. More specifically, it relates to a system and method for efficiently searching for and recommending the most suitable eXtensible Business Reporting Language (XBRL) taxonomy items by analyzing the structural characteristics of a table in table-type accounting data. Background Technology
[0002] As the Extensible Business Reporting Language (XBRL) establishes itself as the standard for digital financial reporting, corporate finance professionals face the challenge of accurately mapping vast amounts of tabular data to XBRL taxonomy items. Given that XBRL taxonomy is a vast and complex system containing thousands or tens of thousands of items, manually locating the appropriate item for specific data is time-consuming and labor-intensive, and carries a high risk of errors resulting from the personnel's subjective judgment.
[0003] Existing simple keyword-based search methods have limitations in that they fail to understand the context of table data. For example, searching simply for the word 'profit' may return dozens of irrelevant taxonomy entries. This is because they fail to utilize structural information regarding where in the table the entry is located (e.g., 'Operating Profit' in the 'Consolidated Statement of Comprehensive Income').
[0004] In addition, existing simple matching methods are vulnerable to changes in table formats or various representation styles, resulting in a problem where they fail to provide consistent mapping even for identical or similar financial information. Therefore, there is a need for new technology that can dramatically improve search accuracy by understanding the structural context of tables. The problem to be solved
[0005] The problem that the present invention aims to solve is to provide a system that maximizes the accuracy and efficiency of searching by automatically recognizing structural information, such as header hierarchy and data membership relationships, in table-type accounting data, and searching for appropriate XBRL taxonomy items based thereon.
[0006] The problem that the present invention aims to solve is to implement a method for providing optimal XBRL taxonomy item recommendation results by splitting queries to search for complex user queries effectively and merging each search result through a Reciprocal Rank Fusion (RRF) algorithm.
[0007] The problem that the present invention aims to solve is to provide a method that enables rapid response solely through metadata updates, without modifying the core logic of the system, even when the XBRL taxonomy is periodically updated or changed.
[0008] The problem that the present invention aims to solve is to drastically reduce work time compared to the existing method of manually searching for XBRL taxonomy items, and to provide consistent mapping results by reducing variation among workers.
[0009] The problems to be solved through the various embodiments of the present invention are not limited to those mentioned above, and other unmentioned problems will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0010] To solve the above problem, the technical concept of the present invention provides an XBRL taxonomy item recommendation method performed by a computing device, comprising: a step of obtaining accounting data in a table form while the XBRL taxonomy is stored in a vector database; a step of analyzing the hierarchical structure of a first cell of the accounting data; a step of generating at least one subquery including a header to which the first cell belongs based on the hierarchical structure of the first cell; and a step of performing a similarity search in the vector database based on the at least one subquery of the first cell to derive an XBRL taxonomy item recommendation list corresponding to the first cell.
[0011] To solve the above problem, the technical concept of the present invention provides a computing device comprising a memory in which an instruction is stored and at least one processor that executes said instruction, wherein the processor performs the XBRL taxonomy item recommendation method by executing said instruction.
[0012] To solve the above problem, the technical concept of the present invention is to provide a recording medium readable by a computing device, wherein a program for executing an XBRL taxonomy item recommendation method on a computing device is recorded thereon, and wherein the XBRL taxonomy item recommendation method is the XBRL taxonomy item recommendation method of claim 1.
[0013] Specific details of other embodiments are included in the detailed description and drawings. Effects of the invention
[0014] The method for recommending eXtensible Business Reporting Language (XBRL) taxonomy items according to the technical concept of the present invention can maximize work efficiency by automating table structure analysis and relatively reducing work time compared to the existing method of manually searching for XBRL taxonomy items.
[0015] The XBRL taxonomy item recommendation method according to the technical concept of the present invention utilizes contextual information such as the location of data and the header hierarchy, and thus can improve search accuracy by recommending relevant XBRL taxonomy items with relatively higher accuracy compared to simple keyword search.
[0016] The XBRL taxonomy item recommendation method according to the technical concept of the present invention can reduce errors caused by interpretation deviations between operators and provide consistent mapping results by recommending consistent XBRL taxonomy items in the same context.
[0017] The XBRL taxonomy item recommendation method according to the technical concept of the present invention can immediately reflect changes in the XBRL taxonomy to the system by simply updating the metadata, thereby enabling rapid response to changes and increasing flexibility and maintenance convenience, which reduces system operating costs.
[0018] The XBRL taxonomy item recommendation method according to the technical concept of the present invention can recommend optimal XBRL taxonomy items even in tables with complex hierarchical structures through query splitting and reciprocal rank fusion (RRF) algorithms.
[0019] The effects of the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description in the claims. Brief explanation of the drawing
[0020] FIG. 1 is a block diagram schematically showing a computing device (100) according to one embodiment of the present invention. FIG. 2 is a flowchart schematically illustrating an XBRL taxonomy item recommendation method according to one embodiment of the present invention. FIG. 3 is a flowchart schematically illustrating the hierarchical structure analysis step of a first cell according to one embodiment of the present invention. FIG. 4 is a flowchart schematically illustrating the subquery generation step according to one embodiment of the present invention. FIG. 5 is a flowchart schematically illustrating the steps for deriving an XBRL taxonomy recommendation list according to one embodiment of the present invention. Specific details for implementing the invention
[0021] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Throughout the specification, the same reference numerals refer to the same components.
[0022] The embodiments described herein will be explained with reference to cross-sectional and / or plan views, which are exemplary illustrations of the invention. In the drawings, the thicknesses of the films and regions are exaggerated for effective explanation of the technical content. Accordingly, the regions illustrated in the drawings are schematic in nature, and the shapes of the regions illustrated in the drawings are intended to illustrate specific forms of regions of the device and are not intended to limit the scope of the invention.
[0023] In the various embodiments of this specification, terms such as first, second, third, etc., have been used to describe various components, but these components should not be limited by such terms. These terms are used merely to distinguish one component from another. The embodiments described and illustrated herein also include their complementary embodiments.
[0024] The terms used herein are for describing the embodiments and are not intended to limit the invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. As used herein, "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components, steps, actions, and / or elements to the mentioned components, steps, actions, and / or elements.
[0025] Unless otherwise defined, all terms used in this specification (including technical and scientific terms) may be used in a meaning that is commonly understood by those skilled in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. Hereinafter, the concept of the present invention and embodiments thereof will be described in detail with reference to the drawings.
[0026] FIG. 1 is a block diagram schematically illustrating a computing device (100) according to an embodiment of the present invention. FIG. 2 is a flowchart schematically illustrating an XBRL taxonomy recommendation method (S100) according to an embodiment of the present invention. FIG. 3 is a flowchart schematically illustrating a step of analyzing the hierarchical structure of a first cell (S120) according to an embodiment of the present invention. FIG. 4 is a flowchart schematically illustrating a step of generating a sub-query (S130) according to an embodiment of the present invention. FIG. 5 is a flowchart schematically illustrating a step of deriving an XBRL taxonomy recommendation list (S140) according to an embodiment of the present invention.
[0027] Referring to FIG. 1, a computing device (100) according to one embodiment of the present invention may include a processor (101), a communication unit (102), a memory (103), and a vector database (104).
[0028] The computing device (100) is a system that performs recommendations for eXtensible Business Reporting Language (XBRL) taxonomy items and can transmit and receive information with a user terminal (200) and other external terminals. The computing device (100) can analyze accounting data in a table format and automatically recommend the most suitable XBRL taxonomy item.
[0029] In some embodiments, the computing device (100) may be a server that provides an XBRL taxonomy item recommendation platform. The computing device (100) may communicate with a user terminal (200) via a network and may be implemented in the form of a web server, an application server, a cloud server, or a database server.
[0030] In some embodiments, the computing device (100) and the user terminal (200) may be physically separated separate devices. For example, the computing device (100) is implemented as a server, and the user terminal (200) can access the server via a network to use the XBRL taxonomy item recommendation service.
[0031] In some embodiments, the user terminal (200) itself may operate as the computing device (100) of the present invention. In this case, the processor (101), communication unit (102), memory (103), and vector database (104) are all installed within the user terminal (200), such as a smartphone, tablet PC, laptop, or desktop computer, so that all functions, such as table structure analysis, query generation, vector search, and taxonomy item recommendation, can be performed directly on the user terminal (200) without a separate external server.
[0032] In some embodiments, some functions of the computing device (100) may be performed at a user terminal (200). For example, a processor (101), a communication unit (102), a memory (103), and a vector database (104) may be mounted within a single device such as a smartphone, a tablet PC, or a laptop, so that the terminal itself can operate as the computing device of the present invention.
[0033] The memory (103) stores instructions for performing an XBRL taxonomy item recommendation method and may temporarily store data generated at each step. For example, the memory (103) may store instructions such as a table structure analysis algorithm, natural language query generation logic, and a vector search algorithm. In some embodiments, the memory (103) may include at least one of ROM (read only memory), RAM (random access memory), and flash memory.
[0034] In some embodiments, the memory (103) may temporarily store table-type accounting data entered by the user. For example, table-type accounting data refers to a grid-type data structure consisting of rows and columns, and each cell may contain text, numbers, or formulas.
[0035] For example, a header may be located in the first row or first column of a table to indicate the category or item of each data, while specific data values can be stored in the remaining cells. For instance, in the case of a balance sheet, the structure may feature account titles such as "Assets," "Liabilities," and "Equity" in the first column, periods such as "Current Period" and "Previous Period" as headers in the first row, and corresponding monetary values entered in the intersecting cells.
[0036] Accounting data may be in Excel file format (.xlsx, .xls), CSV (Comma-Separated Values) file, or other spreadsheet formats. This accounting data may include various forms of financial statements, such as the balance sheet, income statement, comprehensive income statement, cash flow statement, and statement of changes in equity.
[0037] The vector database (104) can convert XBRL taxonomy into a vector and store it. XBRL taxonomy refers to an entire classification system for standardizing and representing financial information, and can serve as a standard dictionary or library for defining and classifying individual financial items.
[0038] For example, an XBRL taxonomy can contain thousands to tens of thousands of taxonomy elements, such as "Current Assets," "Trade Receivables," and "Profit for the period." Each taxonomy element can have metadata, such as a Korean label, English name, and definition, along with a unique identifier (e.g., ifrs-full_CurrentAssets).
[0039] In some embodiments, what is stored in the vector database (104) may be metadata of these XBRL taxonomy entries. The metadata may include a Korean label, an English name, and a definition of each taxonomy entry.
[0040] For example, for a taxonomy item named "ifrs-full_CurrentAssets", metadata such as the Korean label "Current Assets", the English name "Current Assets", and the definition "assets expected to be converted into cash or consumed within 12 months after the end of the reporting period" may be stored.
[0041] In some embodiments, the vector database (104) may convert and store this metadata into vectors using a natural language processing model. Through vector conversion, semantically similar taxonomy items are placed in close positions in the vector space, thereby enabling efficient similarity search.
[0042] In some embodiments, the vector database (104) may update and store only the changed metadata when the XBRL taxonomy is updated. This allows for a rapid response to taxonomy changes without complex system modifications.
[0043] In some embodiments, the vector database (104) may be implemented as a non-volatile storage device in the form of at least one of a hard disk drive (HDD), a solid-state drive (SSD), and cloud storage.
[0044] The processor (101) can perform operations to recommend XBRL taxonomy items by reading and executing instructions stored in memory (103), analyzing table-shaped accounting data, generating natural language queries, and performing similarity searches in a vector database (104).
[0045] In some embodiments, the processor (101) may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, an application processor, or a digital signal processor (DSP).
[0046] In some embodiments, the processor (101) may perform operations of each step, such as table structure analysis, query generation, vector search, and result merging, sequentially or in parallel. The XBRL taxonomy item recommendation method performed by the processor (101) will be described later with reference to FIGS. 2 to 5.
[0047] The communication unit (102) can transmit and receive information from a user terminal (200) and an external terminal. For example, the communication unit (102) may include a communication module that supports one of various wired and wireless communication methods. For example, the communication unit (102) may support at least one of Wireless LAN, Wi-Fi (Wireless Fidelity), Bluetooth, Wired LAN, 3G, 4G, and 5G.
[0048] The communication unit (102) can receive accounting data in a table format from the user terminal (200). Additionally, the communication unit (102) can transmit a recommended XBRL taxonomy item list to the user terminal (200).
[0049] In some embodiments, the communication unit (102) can update the vector database (104) by receiving the latest XBRL taxonomy information from an external server. For example, it can automatically download XBRL taxonomy update information provided by the Financial Supervisory Service or the Korea Accounting Standards Board and reflect it in the system.
[0050] The user terminal (200) can communicate with the computing device (100) to transmit accounting data and receive XBRL taxonomy item recommendation results. The user terminal (200) may be various types of electronic devices, such as a smartphone, tablet PC, desktop computer, or laptop.
[0051] In some embodiments, the user terminal (200) may access the computing device (100) through a web browser or a dedicated application. The user may upload an Excel file or select a specific cell to request an XBRL taxonomy item recommendation for that cell.
[0052] Referring to FIGS. 2 to 5, an XBRL taxonomy item recommendation method (S100) performed by a computing device (100) will be described.
[0053] Referring to FIG. 2, the XBRL taxonomy item recommendation method (S100) may include the step of obtaining accounting data in a table form (S110), the step of analyzing the hierarchical structure of a first cell of the accounting data (S120), the step of generating at least one sub-query based on the hierarchical structure of the first cell (S130), and the step of deriving an XBRL taxonomy item recommendation list based on the sub-query (S140).
[0054] First, the computing device (100) can acquire accounting data in the form of a table that requires XBRL taxonomy item mapping. After transmitting the accounting data to the computing device (100), the user can select a specific cell or range of cells and request XBRL taxonomy item recommendations for that part.
[0055] In some embodiments, when the computing device (100) is implemented as a server, the computing device (100) can receive accounting data from a user terminal (200) through a communication unit (102). In some embodiments, when the user terminal (200) itself operates as the computing device (100), the user can directly retrieve an accounting data file stored on the user terminal (200) or retrieve accounting data from an external storage device or cloud storage.
[0056] In some embodiments, accounting data may include various financial statements such as a balance sheet, income statement, comprehensive income statement, cash flow statement, and statement of changes in equity. Each financial statement has a table structure consisting of rows and columns and includes headers and data values. For example, accounting data may be in an Excel file format and may have extensions such as .xlsx, .xls, and .csv.
[0057] In some embodiments, tabular accounting data may have a hierarchical structure. The top-level header represents the overall title of the financial statement and may be displayed, for example, as "Statement of Consolidated Financial Position." The upper-level header below it represents major classification items and may be broad categories such as "Assets," "Liabilities," and "Equity." The lower-level header may represent items that further subdivide the upper-level header.
[0058] For example, accounting data may have a hierarchical structure where "current assets" and "non-current assets" are located under "assets," and "cash and cash equivalents," "accounts receivable," and "inventory" are located under "current assets."
[0059] Next, the computing device (100) can analyze the hierarchical structure of the first cell selected by the user in the acquired accounting data. The first cell refers to any cell where the user wants to recommend an XBRL taxonomy item.
[0060] In some embodiments, the user may select a single cell or select an area composed of multiple cells. If multiple cells are selected, hierarchical analysis and taxonomy item recommendations may be performed individually for each cell. For convenience of explanation below, the case where a single cell is selected will be referred to as the first cell.
[0061] Referring to FIG. 3, the step of analyzing the hierarchical structure of the first cell (S120) may include the step of analyzing the hierarchical structure of headers included in accounting data (S121), and the step of determining the membership relationship of the header to which the first cell belongs based on the hierarchical structure of the analyzed headers (S122).
[0062] The computing device (100) can analyze the hierarchical structure of headers included in accounting data. In some embodiments, the computing device (100) can analyze the hierarchical structure of headers using a rule-based algorithm with at least one of cell formatting, merge information, and indentation of the accounting data.
[0063] In some embodiments, the computing device (100) can determine the hierarchy of the header by analyzing formatting information such as the font size, weight, color, and background color of the cell. For example, headers displayed in a larger font or bold text can be classified as upper hierarchy, and headers displayed in a smaller font or regular text can be classified as lower hierarchy.
[0064] In some embodiments, the computing device (100) can determine the hierarchy of the headers by analyzing cell merging information. For example, a header merged across multiple columns can be classified as a top-level header, and a header displayed as a single cell can be classified as a lower-level header. In the example above, "Statement of Consolidated Financial Position" is merged across three columns so it is recognized as a top-level header, and "Current Period" and "Previous Period" each occupy a single column so they can be recognized as lower-level headers.
[0065] In some embodiments, the computing device (100) can determine the hierarchy of headers by analyzing the degree of indentation of the text. For example, items with no indentation can be classified as Level 1 headers, items indented by one space as Level 2 headers, and items indented by two spaces as Level 3 headers. For example, "Assets" can be classified as Level 1 headers because it has no indentation, "Current Assets" can be classified as Level 2 headers because it is indented by one space, and "Cash and Cash Equivalents" can be classified as Level 3 headers because it is indented by two spaces.
[0066] In some embodiments, the computing device (100) can determine the hierarchical structure between headers by comprehensively analyzing these cell formatting, merge information, and indentation information. The hierarchical structure can be represented as a tree-shaped data structure.
[0067] The computing device (100) can determine the membership relationship of the header to which the first cell belongs based on the hierarchical structure of the analyzed headers. This membership relationship represents the context of the first cell and can subsequently be used to generate a natural language query.
[0068] In some embodiments, the computing device (100) can identify the headers to which the cell belongs by analyzing the row and column positions of the first cell. After determining which row and column the first cell is located in within the table and identifying the headers corresponding to the row and column, the entire membership relationship of the first cell can be determined by tracking the parent headers of each header using the previously analyzed hierarchical structure information.
[0069] In some embodiments, the computing device (100) may search the previously generated tree-shaped hierarchical data to sequentially track parent headers from the member header of the first cell. Alternatively, based on indentation information, it may search the rows located above the first cell in reverse order and identify headers with progressively decreasing indentation levels as parent headers. Through this method, not only the direct header of the first cell but all parent headers can be systematically identified.
[0070] Next, the computing device (100) can generate at least one sub-query based on the hierarchical structure of the first cell.
[0071] Referring to FIG. 4, the step of generating a subquery (S130) may include the step of generating a hierarchical query (S131) in which the headers to which the first cell belongs are listed based on the hierarchical structure of the first cell, and the step of generating at least one subquery by dividing the hierarchical query into header units (S132).
[0072] First, the computing device (100) can generate a hierarchical query listing the headers to which the first cell belongs, based on the hierarchical structure of the first cell. In some embodiments, the hierarchical query may be in the form of a natural language sentence connecting all headers included in the membership relationships of the first cell with spaces or specific delimiters.
[0073] For example, if the previously identified membership relationship of the first cell is "Statement of Financial Position, Assets, Current Assets, Cash and Cash Equivalents, Current Period," the hierarchical query corresponding to the first cell can be generated as "Statement of Financial Position Assets Current Assets Cash and Cash Equivalents Current Period."
[0074] In some embodiments, hierarchical queries may be listed sequentially from the top-level header to the bottom-level header. This allows for the generation of a query that encompasses the entire context of the first cell. Such a hierarchical query expresses the context in which the first cell is located within the table in natural language form, and can indicate which major, medium, or minor category of which financial statement the cell belongs to, and which period the data belongs to.
[0075] In some embodiments, redundant or unnecessary headers may be excluded when generating a hierarchical query. For example, headers with the same meaning that are repeated, or auxiliary expressions such as "sum" or "total," may be excluded from the query.
[0076] Next, the computing device (100) can generate at least one sub-query by splitting the generated hierarchical query into header units. In some embodiments, query splitting is a process of dividing a complex composite query into simple queries of semantic units. Through this, individual searches for each header information can be performed, and the search results can then be combined to derive an optimal recommendation result.
[0077] For example, the hierarchical query in cell 1, "Consolidated Statement of Financial Position Assets Current Assets Cash and Cash Equivalents Current Period", can be divided into 5 sub-queries as follows. Sub-query 1 may be "Consolidated Statement of Financial Position", sub-query 2 may be "Assets", sub-query 3 may be "Current Assets", sub-query 4 may be "Cash and Cash Equivalents", and sub-query 5 may be "Current Period".
[0078] In some embodiments, the computing device (100) may generate subqueries while maintaining hierarchical information of the headers. For example, hierarchical level information or header types (row headers, column headers, etc.) may be assigned as metadata to each subquery. This metadata may be used to apply different weights in a subsequent search result merging step.
[0079] In some embodiments, the number of subqueries may vary depending on the number of headers to which the first cell belongs. For tables with a complex hierarchical structure, more subqueries may be generated, while for simple tables, fewer subqueries may be generated. For example, in financial statements with a deep hierarchy, more than 7 subqueries may be generated, while in simple tables, only 2 to 3 subqueries may be generated.
[0080] In some embodiments, semantic units may be considered when splitting subqueries. For example, a header consisting of a single compound term, such as "cash and cash equivalents," may be maintained as a single subquery without splitting. On the other hand, a header consisting of multiple independent words may be further split as needed.
[0081] Next, the computing device (100) can perform a similarity search in a vector database based on subqueries of the first cell to derive an XBRL taxonomy item recommendation list.
[0082] Referring to FIG. 5, the step of deriving an XBRL taxonomy item recommendation list (S140) may include the step of deriving a search result list by performing a similarity search on each of at least one subquery of a first cell in a vector database (S141), and the step of generating an XBRL taxonomy item recommendation list by merging the search result lists corresponding to each of at least one subquery of the first cell (S142).
[0083] First, the computing device (100) can derive a list of search results by performing a similarity search on each of at least one subqueries of the first cell in the vector database (104). In some embodiments, the similarity search may be performed in the vector database (104). Metadata of the XBRL taxonomy is converted into a vector and stored in the vector database (104).
[0084] In some embodiments, the similarity search process may be performed as follows.
[0085] First, the subqueries are converted into vectors. Each subquery text is converted into a vector using a natural language processing model. The natural language processing model used at this time may be the same model as the one used to store XBRL taxonomy items in the vector database (104), or a compatible model. Through this, the query vector and the taxonomy item vector are mapped to the same vector space, allowing the semantic similarity to be calculated accurately.
[0086] Second, similarity is calculated. Similarity is calculated between the transformed query vector and all XBRL taxonomy item vectors stored in the vector database (104). Similarity can be calculated using methods such as cosine similarity, Euclidean distance, and dot product. In some embodiments, cosine similarity may be generally used, which can effectively reflect semantic similarity by measuring directional similarity between vectors.
[0087] Third, the top results are extracted. The top N XBRL taxonomy items with high similarity are extracted as search results. For example, the top 5, 10, or 20 taxonomy items can be extracted. The number of items extracted can be adjusted according to system settings and can be determined by considering the balance between search quality and processing speed.
[0088] In some embodiments, since a vector search is performed individually for each subquery, a list of search results as many times as there are subqueries may be generated. For example, if there are 5 subqueries, a list of search results 1 for subquery 1 "balance sheet" may be generated, a list of search results 2 for subquery 2 "assets" may be generated, a list of search results 3 for subquery 3 "current assets" may be generated, a list of search results 4 for subquery 4 "cash and cash equivalents" may be generated, and a list of search results 5 for subquery 5 "current period" may be generated.
[0089] Each list of search results contains XBRL taxonomy items that are highly similar to the corresponding subquery, and each item is ranked along with its similarity score. For example, the search results for subquery 4 "Cash and Cash Equivalents" may be ranked as follows: "ifrs-full_CashAndCashEquivalents" as 1st, "ifrs-full_Cash" as 2nd, and "ifrs-full_CashEquivalentsAtCarryingValue" as 3rd.
[0090] In some embodiments, the search results may also return metadata for each taxonomy item. The metadata may include the unique identifier, Korean label, English name, definition, etc., of the taxonomy item, which can be used later when presenting recommendation results to the user.
[0091] Next, the computing device (100) can merge the search result lists corresponding to each of at least one subquery of the first cell to generate an XBRL taxonomy item recommendation list corresponding to the first cell. In some embodiments, the merging of the search result lists may be performed using a Reciprocal Rank Fusion (RRF) algorithm.
[0092] The RRF algorithm is a technique that merges multiple search result lists to generate a single integrated ranking list. The RRF algorithm can calculate a comprehensive score by considering the ranking of specific items in each search result list.
[0093] In some embodiments, the RRF algorithm may assign high scores to XBRL taxonomy items that rank highly in common across multiple search results. XBRL taxonomy items that appear evenly across multiple subqueries receive higher scores than items that rank highly in only specific subqueries, thereby enabling a balanced reflection of information from all subqueries.
[0094] In some embodiments, the computing device (100) can calculate RRF scores for XBRL taxonomy items appearing in all search result lists and sort them in order of highest RRF scores to generate a final XBRL taxonomy item recommendation list.
[0095] In some embodiments, the final XBRL taxonomy item recommendation list may include only the top N items. For example, the top 5, 10, or 20 taxonomy items may be recommended to the user. The number of recommended items may be adjusted according to user settings or system policies.
[0096] In some embodiments, each item in the XBRL taxonomy item recommendation list may be provided along with information such as an RRF score, Korean label, English name, definition, and unique identifier. Users can select XBRL taxonomy items by referring to this information. Additionally, detailed information regarding each item, such as which subquery the item was retrieved from and its ranking in each search result list, may also be provided to help users understand the basis of the recommendation results.
[0097] In some embodiments, the computing device (100) may transmit the generated recommendation list to a user terminal (200) to present it to the user. The user may review the recommendation list, select the taxonomy item deemed most suitable, and map it to a first cell.
[0098] In some embodiments, the XBRL taxonomy may be updated periodically. The Financial Supervisory Service or the Korea Accounting Standards Board may add, modify, or delete the XBRL taxonomy in accordance with new accounting standards or reporting requirements.
[0099] In some embodiments, metadata updates may be automated. A computing device (100) may periodically check an external server to automatically check for updates to the XBRL taxonomy and, if there are changes, automatically download and update the vector database (104).
[0100] In some embodiments, the computing device (100) of the present invention can flexibly respond to such taxonomy changes. Since the vector database (104) stores metadata of XBRL taxonomy items, there is no need to modify the table structure analysis logic, query generation logic, vector search algorithm, etc., which are core algorithms of the XBRL taxonomy item recommendation method (S100), even if the XBRL taxonomy changes.
[0101] In some embodiments, the computing device (100) can recommend XBRL taxonomy entries collectively for the entire table. Instead of the user selecting a specific cell, if the user uploads the entire table, the system can automatically analyze the hierarchy for all data cells and recommend taxonomy entries.
[0102] For example, if the entire balance sheet is uploaded, the system can automatically recommend suitable XBRL taxonomy items for each data cell—that is, for every cell containing numerical data, such as current period assets or prior period liabilities—and present the results to the user in a visualized table format. Users can view all recommended results at a glance and, if necessary, modify the recommendations for specific cells.
[0103] In some embodiments, the batch processing method can be efficiently performed through parallel processing. Since the processes of hierarchical structure analysis, query generation, vector search, and result merging for each cell can be performed independently, the total processing time can be reduced by utilizing multiple processing units to process multiple cells simultaneously.
[0104] In some embodiments, the computing device (100) may recommend XBRL taxonomy items while interacting interactively with the user. When the user selects a cell, the system provides a recommendation list, and when the user selects a specific item or requests additional information, the system may provide additional more detailed descriptions or related taxonomy items.
[0105] For example, if a user requests to display only foreign currency cash separately after receiving a recommendation for "cash and cash equivalents," the system can recommend additional foreign currency-related attributes or detailed taxonomy items. Furthermore, if the user selects one of the recommended items, the system can provide detailed information such as the definition, usage examples, and relevant accounting standards for that item.
[0106] In some embodiments, the computing device (100) may provide a confidence score for each recommended taxonomy item. The confidence score may be an RRF score, a similarity score, or a combination thereof, and may be displayed in the form of a percentage or a rating. Users may prioritize items with high confidence.
[0107] For example, a recommendation list may be provided as follows. The first item is "ifrs-full_CashAndCashEquivalents" with a confidence level of 95 percent, the second item is "ifrs-full_Cash" with a confidence level of 78 percent, and the third item is "ifrs-full_CurrentAssets" with a confidence level of 65 percent.
[0108] In some embodiments, the confidence score may be visualized using color coding. For example, high confidence of 90 percent or more may be displayed in green, medium confidence of 70 percent to 90 percent in yellow, and low confidence of less than 70 percent in red. This allows the user to intuitively understand the reliability of the recommendation results.
[0109] In some embodiments, various factors may be considered when calculating the confidence score. In addition to the RRF score, information such as how many subqueries the item was retrieved from, the average rank in each subquery, and whether the item has a history of being previously selected in a similar context may be comprehensively reflected.
[0110] In some embodiments, the computing device (100) can improve XBRL taxonomy item recommendation performance by storing and analyzing the user's selection history. When the user selects a specific item from the recommendation list, that information is recorded, and the ranking of that item can be increased in similar situations.
[0111] For example, if a specific company always selects "ifrs-full_CashAndCashEquivalents" for "Cash and Cash Equivalents," that item can be placed at the top of the recommendation list the next time. These personalized recommendations can be applied on a per-company or per-user basis, and more accurate recommendations can be provided by learning the past selection patterns of the company or user.
[0112] In some embodiments, the collected feedback data may be periodically analyzed and utilized to improve the system's recommendation algorithm. For example, if a pattern is discovered in which a specific taxonomy item is frequently selected within a specific type of table structure or a specific header combination, this can be reflected in the recommendation algorithm to increase the ranking of that item in similar situations.
[0113] In some embodiments, the computing device (100) can handle various exceptional situations that may occur during the table analysis process. For example, it may be designed to derive meaningful analysis results as much as possible even in situations where the structure of the table is irregular, the header is unclear, or the cell merging is complex.
[0114] In some embodiments, the computing device (100) may request confirmation from the user when the analysis of the table structure is uncertain. For example, if it is unclear whether a particular row is a header or a data row, the computing device (100) may present an estimated result to the user and obtain confirmation. The user's feedback may be immediately reflected and utilized in subsequent analysis.
[0115] In some embodiments, the computing device (100) can verify consistency between taxonomy items selected by the user. For example, it may display a warning if the relationship between a parent item and a child item is not logically consistent, or if conflicting taxonomy items are used within the same financial statement.
[0116] The above-described XBRL taxonomy item recommendation method can be implemented as code readable by a computing device on a recording medium readable by a computing device. A recording medium readable by a computing device may include all types of recording devices in which data readable by a computing device is stored.
[0117] For example, a recording medium readable by a computing device may include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc. Additionally, the recording medium readable by a computing device may be distributed across networked computer systems, and code readable by a computing device may be stored and executed in a distributed manner.
[0118] For example, a program for implementing an XBRL taxonomy item recommendation method according to an embodiment of the present invention may be implemented in the form of program instructions that can be executed through various computing devices and may be recorded on a recording medium readable by a computing device.
[0119] Although preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above. Various modifications are possible by those skilled in the art without departing from the essence of the invention as claimed in the patent claims, and such modifications should not be understood individually from the technical spirit or perspective of the present invention. Explanation of the symbols
[0120] 100: Computing device 101: Processor 102: Communication Unit 103: Memory 104: Vector Database 200: User Terminal
Claims
Claim 1 An XBRL taxonomy item recommendation method performed by a computing device, comprising: a step of obtaining accounting data in a table form while the XBRL taxonomy is stored in a vector database; a step of analyzing the hierarchical structure of a first cell of the accounting data; a step of generating at least one subquery including a header to which the first cell belongs based on the hierarchical structure of the first cell; and a step of deriving a search result list by performing a similarity search on each of the at least one subquery in the vector database, and merging the search result lists corresponding to each of the at least one subquery through a Reciprocal Rank Fusion (RRF) algorithm to derive an XBRL taxonomy item recommendation list corresponding to the first cell. Claim 2 In claim 1, the step of analyzing the hierarchical structure of the first cell includes the step of analyzing the hierarchical structure of headers included in the accounting data and the step of identifying the membership relationship of the header to which the first cell belongs based on the analyzed hierarchical structure of the headers, and in the step of analyzing the hierarchical structure of the headers included in the accounting data, the method of recommending an XBRL taxonomy item by analyzing using a rule-based algorithm utilizing at least one of the cell format, merge information, and indentation of the accounting data. Claim 3 In claim 1, the step of generating at least one sub-query comprises: generating a hierarchical query in which the headers to which the first cell belongs are listed based on the hierarchical structure of the first cell, and dividing the hierarchical query into header units to generate at least one sub-query, thereby providing an XBRL taxonomy item recommendation method. Claim 4 delete Claim 5 delete Claim 6 A method for recommending XBRL taxonomy items according to claim 1, wherein metadata of the XBRL taxonomy is converted into a vector and stored in the vector database, and the metadata includes a Korean label, English name, and definition of the XBRL taxonomy item. Claim 7 A computing device comprising a memory in which an instruction is stored and at least one processor that executes said instruction, wherein the processor performs the XBRL taxonomy item recommendation method of claim 1 by executing said instruction. Claim 8 A recording medium readable by a computing device, wherein a program for executing an XBRL taxonomy item recommendation method on a computing device is recorded thereon, said XBRL taxonomy item recommendation method is the XBRL taxonomy item recommendation method of claim 1, a recording medium readable by a computing device.
Citation Information
Patent Citations
Method, system, and computer program to dynamically provide sub-item recommendation list for each item included in search results based on search query
KR1020230032811A
A method and apparatus for question-answering using a paraphraser model
KR1020240049526A
Systems and methods for XBRL tag suggestion and validation
US20220138403A1
Method for constructing XBRL taxonomy with multidimensional attributes
KR100925725B1
Data Mapping Method for Extensible Business ReportingLanguage and Editing Method of Document using the DataMapping Method
KR1020070040734A