System and method for XBRL tag suggestion and verification

A machine learning-based system for suggesting and validating XBRL tags addresses the challenge of tag selection, achieving high accuracy and improving reporting efficiency.

JP2026031944APending Publication Date: 2026-02-25WORKIVA INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025178511
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-11-04
Filing Date
2025-10-23
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Selecting the correct XBRL tag from thousands of available tags is a challenge for preparers of financial statements or business reports, leading to inefficiencies in data reporting.

Method used

A system and method utilizing a trained machine learning model to suggest and validate XBRL tags, incorporating natural language processing and neural networks to predict tags with high accuracy and associate them with confidence values, and a graphical user interface for user interaction.

Benefits of technology

The system achieves greater than 95% accuracy in predicting XBRL tags and provides a user-friendly interface for validating and selecting tags, enhancing the efficiency of data reporting processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026031944000001
    Figure 2026031944000001
  • Figure 2026031944000002
    Figure 2026031944000002
  • Figure 2026031944000003
    Figure 2026031944000003
Patent Text Reader

Abstract

To provide a method and a system for proposing and verifying business report language (XBRL) tags.SOLUTION: A method according to an XBRL tag suggestion system includes receiving an XBRL document associated with one or more assigned XBRL tags, parsing the XBRL document using a trained machine learning model to generate one or more suggested XBRL tags and determine one or more corresponding confidence values, comparing the one or more assigned XBRL tags with the one or more suggested XBRL tags to generate a comparison result, and determining a tag confidence value associated with each assigned XBRL tag of the one or more assigned XBRL tags based on the comparison result.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Technical Field The present disclosure relates to reporting business data through documents using XBRL (Extensible Business Reporting Language), and more particularly to systems and methods for automated XBRL document tag suggestion and validation. [Background technology]

[0002] background

[0002] XBRL is a standardized computer language that enables companies to efficiently and accurately communicate business data with each other and with regulatory agencies. Extensible Business Reporting Language (XBRL) 2.1 is available at http: / / www.xbrl.org / Specification / XBRL-2.1 / REC-2003-12-31 / XBRL-2.1-REC-2003-12-31+corrected-errata-2013-02-20.html. XBRL is a markup language not very different from XML (Extensible Markup Language) and HTML (Hypertext Markup Language). HTML was designed to represent general-purpose data in a standardized way, XML was designed to transport and store general-purpose data in a standardized way, and XBRL was designed to transport and store business data in a standardized way.

[0003]

[0003] A taxonomy is a reporting and subject-specific dictionary used by the XBRL community. A taxonomy contains specific tags, called XBRL tags, used for individual data items (e.g., "revenue," "operating expenses"), their attributes, and their interrelationships. Different business reporting objectives often require different taxonomies.

[0004]

[0004] XBRL is bringing about a dramatic change in the way people think about exchanging business information. Financial disclosure is a prime example of an industry built around paper-based processes being pushed into the technological age. This transition involves a paradigm shift from a pixel-perfect world of building unstructured reports to a digital world where structured data dominates. Summary of the Invention

[0005] overview At least some aspects of the present disclosure are directed to a method for validating XBRL tags. The method is implemented on a computer system having one or more processors and a memory. The method includes receiving an XBRL document associated with one or more assigned XBRL tags, analyzing the XBRL document using a trained machine learning model to generate one or more suggested XBRL tags and determining one or more corresponding trust values, each suggested XBRL tag of the one or more suggested XBRL tags associated with a trust value, comparing the one or more assigned XBRL tags with the one or more suggested XBRL tags to generate a comparison result, and determining a tag trust value associated with each assigned XBRL tag of the one or more assigned XBRL tags based on the comparison result.

[0006]

[0006] At least some aspects of the present disclosure are directed to a method for suggesting XBRL tags. The method is implemented on a computer system having one or more processors and a memory. The method includes receiving a plurality of XBRL datasets, training a machine learning model using the plurality of XBRL datasets, receiving a document, and predicting one or more XBRL tags associated with the document using the trained machine learning model, wherein at least some of the plurality of XBRL datasets are received. Each contains a row header and at least one of a table type. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The accompanying drawings, which are incorporated in and constitute a part of this specification, serve to explain the features and principles of the disclosed embodiments. [Brief explanation of the drawings]

[0007] [Figure 1]

[0008] 1 illustrates an exemplary system diagram of an XBRL tag proposal and validation system, according to certain embodiments of the present disclosure. [Figure 2A]

[0009] 1 illustrates an example flow diagram of an XBRL tag suggestion system, according to certain embodiments of the present disclosure. [Figure 2B]

[0010] 1 illustrates another example flow diagram of an XBRL tag validation system, in accordance with certain embodiments of the present disclosure. [Figure 2C]

[0011] 1 illustrates an example flow diagram of a training process for a neural network used within an XBRL tag proposal and validation system, according to certain embodiments of the present disclosure. [Figure 3]

[0012] 1 shows an illustrative example of a document with forecast / proposal XBRL tags. [Figure 4]

[0013] 1 shows an illustrative example of a neural network architecture, according to certain embodiments of the present disclosure. [Figure 5]

[0014] 1 shows an illustrative example of a graphical user interface for an XBRL tag suggestion system. [Figure 6]

[0015] 1 shows an illustrative example of a representation of an XBRL document with assigned XBRL tags and tag confidence values. DETAILED DESCRIPTION OF THE INVENTION

[0008] Detailed Description

[0001] Unless otherwise specified, all numbers expressing size, quantity, and physical properties of features used in the specification and claims are to be understood as being modified in all instances by the term "about." Accordingly, unless indicated to the contrary, the numerical parameters set forth in the above specification and appended claims are approximations that may vary depending upon the desired properties sought to be obtained by those of ordinary skill in the art utilizing the teachings disclosed herein. The use of numerical ranges by endpoints includes all numbers within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5) and any range within that range.

[0009]

[0016] Although an example method may be represented by one or more diagrams (e.g., flow charts, communication flows, etc.), the diagrams should not be construed as implying any requirement of the various steps disclosed herein or a particular order among or between such steps. However, as may be expressly described herein and / or understood from the nature of the steps themselves, some specific embodiments may require certain steps and / or a certain order among certain steps (e.g., performance of some steps may depend on the results of previous steps). Additionally, a "set," "subset," or "group" of items (e.g., inputs, algorithms, data values, etc.) can include one or more items, and similarly, a subset or subgroup of items can include one or more items. "Plurality" means more than one.

[0010]

[0017] As used herein, the term "based on" is not meant to be limiting, but rather indicates that a determination, identification, prediction, calculation, etc. is made by using at least the term following "based on" as an input. For example, predicting an outcome based on a particular piece of information may additionally or alternatively base the same determination on another piece of information.

[0011]

[0018] One of the current challenges faced by preparers of XBRL financial statements or other business reports is selecting the correct XBRL tag from thousands of XBRL tags. At least some embodiments of the present disclosure are directed to systems and methods for suggesting XBRL tags using a trained machine learning model. In some cases, the trained machine learning model can predict XBRL tags with greater than 95% accuracy. In some embodiments, each suggested XBRL tag is associated with a confidence value. As used herein, confidence value refers to the probability that a given tag is correct based on inputs provided to the trained machine learning model and the data on which the machine learning model was trained. At least some embodiments of the present disclosure are directed to systems and methods for validating XBRL tags by comparing tags suggested by a trained machine learning model with tags in an XBRL document.

[0012]

[0019] 1 illustrates an exemplary system diagram of an XBRL tag proposal and validation system 100 in accordance with certain embodiments of the present disclosure. As shown, system 100 includes a document processor 120, a machine learning processor 130, an interface engine 140, a presentation engine 145, and an XBRL data repository 150. One or more components of system 100 are optional. In some cases, system 100 may include additional components. In some cases, system 100 interfaces with one or more other systems 160, such as filing systems, financial systems, vendor systems, etc.

[0013]

[0020] In some embodiments, the document processor 120 includes natural language processing capabilities. In some cases, the document processor 120 parses the incoming document into n-grams and generates multiple terms based on the n-grams. As used herein, an n-gram refers to a continuous sequence of n words, including numbers and symbols, from a data stream, typically a meaningful phrase or word sequence. An n-gram may include numbers and symbols such as commas, periods, and dollar signs. In some cases, the document processor 120 normalizes the parsed n-grams. In some cases, the document processor 120 generates multiple normalized sections with normalized terms based on the n-grams. In one example, multiple incoming terms include normalized n-grams. In one example, the n-grams are dates, and the normalized terms are dates in a predetermined format (e.g., year-month-day). In some cases, the document processor 120 determines the context of the normalized terms. In one example, the context is part of the same sentence as the normalized terms. In one example, the natural language processor 120 parses the n-grams and labels them based on context, e.g., time period, expenses, revenue, etc. In some embodiments, the document processor 120 uses a natural language model to process the documents and parsed n-grams. For example, the natural language model can be a statistical language model, a neural network language model, etc.

[0014]

[0021] In some embodiments, the machine learning processor 130 is configured to train a machine learning model using the XBRL dataset and use the trained machine learning model to predict XBRL tags associated with an input document, including an XBRL document. As used herein, an XBRL document refers to a document that has been tagged with one or more XBRL tags. The machine learning model may include any suitable machine learning model, deep learning model, etc. In some cases, the trained machine learning model includes at least one of a decision tree, a random forest, a support vector machine, a neural network, a convolutional neural network, a recurrent neural network, etc. In some embodiments, the XBRL tag suggestion system receives multiple XBRL datasets for training and / or testing. In some cases, the system may select a subset of the XBRL datasets for training. In some cases, the subset of the XBRL datasets is selected based on the completeness of the XBRL datasets. In some cases, the first subset of the XBRL datasets is selected based on the taxonomy of the XBRL datasets. In some cases ... XBRL A subset of the dataset is used for testing (e.g., one-third of the XBRL dataset). In some cases, the XBRL dataset or subset selected as training data is compiled to include raw header data along with corresponding XBRL data and metadata. Metadata may include SIC ("Standard Industrial Classification") codes, table types, etc.

[0015]

[0022] In some embodiments, the machine learning model includes a neural network. In some cases, the neural network includes multiple layers. In one embodiment, the neural network includes at least an input layer, a concatenation layer, and a dropout layer, an example of which is shown in FIG. 4. In some embodiments, the machine learning processor 130 receives documents, such as business documents, report documents, matched documents, etc. In some cases, the machine learning processor 130 may receive documents by interfacing with other systems 160. In some cases, the machine learning processor 130 may receive documents by retrieving the documents through XBRL data repository 150.

[0016]

[0023] The machine learning processor 130 is configured to predict one or more XBRL tags associated with a document using the trained machine learning model. Figure 3 is an illustrative example of a document 300 having predicted XBRL tags 301. In this example, data associated with the XBRL tags 301 is stored in a data structure for the XBRL tags 301. In some embodiments, the system predicts one or more XBRL tags associated with a section of the document.

[0017]

[0024] In some embodiments, the machine learning processor 130 is configured to determine one or more confidence values, with each of the one or more tags associated with a respective confidence value. In some embodiments, the machine learning processor 130 is configured to apply a multivariate function to the document to determine the confidence values. One example of a multivariate function is a softmax function. In one example, the softmax function is implemented as a layer of a neural network that the machine learning processor 130 uses to analyze the document.

[0018]

[0025] In some embodiments, the machine learning processor 130 can determine the set of selected XBRL tags and generate an output of selected XBRL tags. The selected XBRL tags can be stored in the XBRL data repository 150. In one embodiment, the system can determine the set of selected XBRL tags based at least in part on confidence values. In one example, each of the set of selected XBRL tags has the highest confidence value of the confidence values ​​corresponding to a group of proposed XBRL tags associated with the same document section. In some embodiments, the system can determine the set of selected XBRL tags based at least in part on user input. The selected XBRL tags can be used to tag a document.

[0019]

[0026] In some embodiments, system 100 performs a review / validation process on an XBRL document associated with one or more assigned tags. In some embodiments, the XBRL document includes multiple sections, and each assigned tag is associated with a respective section of the XBRL document. In some cases, machine learning processor 130 analyzes each section of the multiple sections and determines one or more suggested XBRL tags for each section. In some embodiments, machine learning processor 130 determines a confidence value for each suggested tag of the one or more suggested tags.

[0020]

[0027] In some embodiments, the machine learning processor 130 can compare one or more assigned tags with one or more suggested tags to generate a comparison result. The comparison result may include one or more assigned tags that match one or more suggested tags. The comparison result may include one or more assigned tags that do not match any suggested tags. In one embodiment, the comparison results may include, for each tagging section associated with each assigned tag and one or more suggested tags, whether the each assigned tag matches any one of the one or more suggested tags.

[0021]

[0028] In some embodiments, the machine learning processor 130 may determine a tag confidence value for each assigned tag based on the comparison results. In some embodiments, the tag confidence value of an assigned tag is set equal to the matching suggested tag (e.g., 95%). In some embodiments, the tag confidence value of an assigned tag is set equal to the matching suggested tag of the one or more suggested tags for the corresponding tagged section. In one embodiment, if an assigned tag does not match any one of the one or more suggested section tags for the corresponding tagged section, the tag confidence value of the assigned tag is set to logical low. In one embodiment, logical low is represented by 0%. In some cases, a tagged section may be a table (e.g., an income statement table, etc.). In some cases, a tagged section may represent a hierarchy within an individual XBRL taxonomy. For example, a document may include one or more tagged sections, and each tagged section may include one or more subsections.

[0022]

[0029] Optionally, the machine learning processor 130 may determine a category for each tag confidence value. In one embodiment, the tag confidence categories may include a high confidence category, a medium confidence category, and a low confidence category. In one embodiment, each tag confidence category is associated with a predetermined range. For example, the high confidence category corresponds to confidence values ​​in the range of 40% to 100%, the medium confidence category corresponds to confidence values ​​in the range of 10% to 40%, and the low confidence category corresponds to confidence values ​​in the range of 0% to 10%. In another embodiment, the tag confidence categories may include more than three categories. In yet another embodiment, the tag confidence categories may include two categories.

[0023]

[0030] In some embodiments, interface engine 140 is configured to interface with other systems 160. In some embodiments, interface engine 140 is configured to connect to electronic filing systems or financial systems 160 via a software interface. In some cases, interface engine 140 is configured to use a set of predetermined protocols via the software interface. In some cases, the software interface includes at least one of an application programming interface and a web services interface.

[0024]

[0031] In some embodiments, the presentation processor 145 is configured to generate a representation of the XBRL document, the proposed / predicted XBRL tags, the labels for the proposed / predicted XBRL tags, and / or the tag confidence values. In some cases, the representation representing the XBRL tags includes a representation of the labels associated with the tags. One illustrative example of a graphical user interface is shown in FIG. 5 , which shows, for an XBRL document or document section, recommended / suggested XBRL tags 510 represented by associated labels and associated confidence values ​​520. In some cases, the graphical user interface is received by a client application and rendered on a user's computing device. In some cases, the presentation processor 145 is configured to receive input or requests submitted by a user. In the example of FIG. 5 , the user can click button 530 to submit user input. In this example, the received input is a user selection. Another illustrative example of a representation is shown in FIG. 6 , which shows a review of an XBRL document with XBRL tags 610 and corresponding tag confidence values ​​620 determined by a tag validation system. In some cases, the presentation engine may be configured to render the representation to a user. In some cases, the presentation engine 145 is configured to receive the user's computing device type (e.g., laptop, smartphone, tablet computer, etc.) and generate a graphical presentation adapted to the computing device type.

[0025]

[0032] In some embodiments, the representation of the tag trust values ​​includes an indication of the tag trust value's corresponding category. In one example, the tag trust categories are represented by colors. For example, high trust categories are represented by green, medium trust categories are represented by yellow, and low trust categories are represented by red. Figure 6 shows one illustrative example of a representation of an XBRL document with assigned XBRL tag 610 labels and tag trust values ​​620. In one embodiment, an indication of the tag trust category for each tag trust value is included in the representation.

[0026]

[0033] In some embodiments, XBRL data repository 150 may include taxonomy data, XBRL datasets, proposed XBRL tags, selected XBRL tags, documents (including XBRL documents) received for analysis, etc. XBRL data repository 150 may be implemented using any one of the configurations described below. The data repository may include random access memory, flat files, XML files, and / or one or more database management systems (DBMS) running on one or more database servers or data centers. The database management systems may be relational (RDBMS), hierarchical (HDBMS), multidimensional (MDBMS), object-oriented (ODBMS or OODBMS), or object-relational (ORDBMS) database management systems, etc. The data repository may be, for example, a single relational database. In some cases, the data repository may include multiple databases that can exchange and aggregate data through a data integration process or software application. In an exemplary embodiment, at least a portion of the data repository may be hosted within a cloud data center. In some cases, the data repository may be hosted on a single computer, server, storage device, cloud server, etc. In some other cases, the data repository may be hosted on a series of networked computers, servers, or devices. In some cases, the data repository may be hosted on tiers of data storage, including local, regional, and central.

[0027]

[0034] In some cases, various components of system 100 may execute software or firmware stored in non-transitory computer-readable media to implement various processing steps. Various components and processors of system 100 may be implemented by one or more computing devices, including, but not limited to, circuits, computers, cloud-based processing units, processors, processing units, microprocessors, mobile computing devices, and / or tablet computers. In some cases, various components of system 100 (e.g., document processor 120, machine learning processor 130, interface engine 140, presentation engine 150) may be implemented on a shared computing device. Alternatively, components of system 100 may be implemented on multiple computing devices. In some implementations, various modules and components of system 100 may be implemented as software, hardware, firmware, or a combination thereof. In some cases, various components of XBRL tag proposal and validation system 100 may be implemented by software or firmware executed by a computing device.

[0028]

[0035] The various components of the system 100 may communicate or be coupled by a communication interface, such as a wired or wireless interface. The communication interface includes, but is not limited to, any wired or wireless short-range and long-range communication interface. Short-range communication interfaces include, for example, a local area network (LAN), Bluetooth® standard, IEEE 802.11 standard, and the like. The long-range communication interface may be, for example, an interface conforming to a known communication standard such as IEEE 802.11, ZigBee, or a similar standard such as based on the IEEE 802.15.4 standard or other public or proprietary wireless protocols. AN), a cellular network interface, a satellite communication interface, etc. The communication interface can be in a private computer network such as an intranet, or can be on a public computer network such as the Internet.

[0029]

[0036] FIG. 2A illustrates an example flow diagram of an XBRL tag proposal system according to certain embodiments of the present disclosure. Aspects of an embodiment of method 200A may be performed, for example, by components of an XBRL tag proposal system (e.g., components of XBRL tag proposal / validation system 100 of FIG. 1). One or more steps of method 200A are optional and / or may be modified by one or more steps of other embodiments described herein. Additionally, one or more steps of other embodiments described herein may be added to method 200A. In some embodiments, an XBRL tag proposal system receives multiple XBRL datasets (210A). In some cases, the system may select a first subset of the XBRL datasets (215A). In some cases, the first subset of the XBRL datasets is selected based on the completeness of the XBRL datasets. In some cases, the first subset of the XBRL datasets is selected based on the taxonomy of the XBRL datasets. In some cases, a subset of the XBRL datasets is used for testing (e.g., one-third of the XBRL datasets). In some cases, at least some of the plurality of XBRL data sets each include a row header. In some designs, the row header is represented by a vector. In some designs, the row header is represented by a dense vector. In one example, the dense vector condenses information in the text used in the row header into a numerical form that may be usable by a machine learning model. In some cases, at least some of the plurality of XBRL data sets each include a table type. In some cases, at least some of the plurality of XBRL data sets each include an industry category.

[0030]

[0037] In some embodiments, the XBRL tag suggestion system trains a machine learning model using a first subset of the XBRL dataset (220A) or the XBRL dataset. The training process may use, for example, the process shown in FIG. 2C . In some cases, the XBRL dataset or subset used for training data is compiled to include raw header data along with corresponding XBRL data and metadata. The metadata may include SIC codes, table types, etc. The machine learning model may include any suitable machine learning model, deep learning model, etc. In some embodiments, the machine learning model includes at least one of a decision tree, a random forest, a support vector machine, a convolutional neural network, a recurrent neural network, etc. Optionally, the system may test the trained machine learning model using a second subset of the XBRL dataset. In one embodiment, the first subset of the XBRL dataset and the second subset of the XBRL dataset do not have any overlapping datasets. In another embodiment, the first subset of the XBRL dataset and the second subset of the XBRL dataset have at least one overlapping dataset.

[0031]

[0038] In some embodiments, the machine learning model includes a neural network. In some cases, the neural network includes multiple layers. In one embodiment, the neural network includes at least an input layer, a concatenation layer, and a dropout layer, an example of which is shown in FIG. 4. In some embodiments, the XBRL tag suggestion system receives 230A documents, such as business documents, reports, conformance documents, etc. In some cases, the system may receive the documents by interfacing with other systems. In some cases, the system may receive the documents by retrieving them through a data repository (e.g., XBRL data repository 150 of FIG. 1). In some embodiments, the system is coupled to an electronic filing system by a software interface to receive the documents. In some embodiments, the system is coupled to an electronic filing system by a software interface to receive the documents. It is configured to use a set of predetermined protocols. In some embodiments, the software interface includes at least one of an application programming interface and a web services interface.

[0032]

[0039] In some cases, the system normalizes (235A) the incoming document and / or normalizes relevant portions of the document. In some cases, the system performs one or more steps to preprocess the document. In some embodiments, the system identifies tagged sections within the document. In one embodiment, the tagged sections are identified based on known document formats. In one embodiment, the system can extract and aggregate potentially relevant XBRL metadata, tags, and relationships from the tagged sections, including, for example, tagged concepts, concept values, numeric measures, units of measure, unit numerators, unit denominators, row headers, column headers, document types, filter categories, sibling and parent tags, accession numbers, filing dates, XBRL taxonomies, company CIKs ("Securities and Exchange Commission Issued Codes"), company SICs, etc. In some cases, processing a single filing involves processing a set of six (6) or more files. In some cases of iXBRL ("inline XBRL"), the system can extract fact information from xhtml files along with html row headers. In some cases of XBRL, the system may predict the table type. In some cases of traditional XBRL, the table type may be provided to the system as part of an XBRL outline, which is a combination of a schema, labels, and a presentation linkbase. In some cases of traditional XBRL, the system may determine the html row header by matching an XBRL outline section with an html table. In some cases, the SIC code is taken from the document source (e.g., SEC) rather than being part of the filing. In some cases, the SIC codes are grouped into a small set of industry groups / categories.

[0033]

[0040] The XBRL tag suggestion system then uses the trained machine learning model to predict 240A one or more XBRL tags associated with the document. In some embodiments, the system predicts one or more XBRL tags for each tagged section. Figure 3 is an illustrative example of a document 300 having predicted / suggested XBRL tags 301. In some embodiments, the system determines 245A one or more confidence values, each of which is associated with a respective confidence value.

[0034]

[0041] In some embodiments, the XBRL tag suggestion system applies a multivariate function to the document to determine the confidence value. In one example, the multivariate function is a softmax function. In one embodiment, the softmax function may be applied within a layer of a neural network. In some examples, the output of the neural network corresponds to a vector of probabilities that sum to one by applying the softmax function.

[0035]

[0042] In some embodiments, the XBRL tag suggestion system generates 250A a representation indicating one or more XBRL tags and one or more trust values. An example is shown in FIG. 5, which is an illustrative example of a graphical user interface 500. As shown, the representation includes suggested XBRL tags 510, each represented by an associated individual label, and trust values ​​520. In some embodiments, each suggested XBRL tag is associated with a trust value. For example, XBRL tag 510_1 is associated with trust value 520_1, XBRL tag 510_2 is associated with trust value 520_2, XBRL tag 510_3 is associated with trust value 520_3, and XBRL tag 510_4 is associated with trust value 520_4. In some cases, the representation includes an indicator representing one or more trust value categories, e.g., high, medium, and low categories.

[0036]

[0043] In some embodiments, the system may receive input from a user (255A). In one embodiment, the system may receive input selecting a suggested XBRL tag from one or more suggested XBRL tags associated with a document and / or document section. In some embodiments, the system may determine a set of selected XBRL tags (260A) and generate an output of selected XBRL tags (265A). In one embodiment, the system may determine the set of selected XBRL tags based at least in part on trust values. In one example, each of the set of selected XBRL tags has the highest trust value of the trust values ​​corresponding to a group of suggested XBRL tags associated with the same document section. In some embodiments, the system may determine the set of selected XBRL tags based at least in part on user input. The selected XBRL tags may be used to tag a document.

[0037]

[0044] FIG. 2B illustrates an example flow diagram 200B of an XBRL tag validation system according to certain embodiments of the present disclosure. Aspects of an embodiment of method 200B may be performed, for example, by components of an XBRL tag proposal system (e.g., components of XBRL tag proposal / validation system 100 of FIG. 1 ). One or more steps of method 200B are optional and / or may be modified by one or more steps of other embodiments described herein. Additionally, one or more steps of other embodiments described herein may be added to method 200B. In some embodiments, an XBRL tag validation system receives multiple XBRL datasets (210B). In some cases, the system may select a first subset of the XBRL datasets (215B). In some cases, the first subset of the XBRL datasets is selected based on the completeness of the XBRL datasets. In some cases, the first subset of the XBRL datasets is selected based on the taxonomy of the XBRL datasets. In some cases, a subset of the XBRL dataset is used for testing (e.g., one-third of the XBRL dataset). In some cases, at least some of the multiple XBRL datasets each include row headers and at least one of the following types of tables. In some designs, the row headers are represented by vectors. In some designs, the row headers are represented by dense vectors.

[0038]

[0045] In some embodiments, the XBRL tag suggestion system trains a machine learning model using a first subset of the XBRL dataset (220B) or the XBRL dataset. The machine learning model may include any suitable machine learning model, deep learning model, etc. In some embodiments, the machine learning model includes at least one of a decision tree, a random forest, a support vector machine, a convolutional neural network, a recurrent neural network, etc. In some cases, the system can train the machine learning model using the XBRL dataset. The training process may use, for example, the process shown in FIG. 2C. Optionally, the system can test the trained machine learning model using a second subset of the XBRL dataset. In some cases, the XBRL dataset or subset used for training data is compiled to include raw header data along with corresponding XBRL data and metadata. The metadata may include SIC codes, table types, etc. In one embodiment, the first subset of the XBRL dataset and the second subset of the XBRL dataset do not have any overlapping datasets. In another embodiment, the first subset of XBRL datasets and the second subset of XBRL datasets have at least one overlapping dataset.

[0039]

[0046] In some embodiments, the machine learning model includes a neural network. In some cases, the neural network includes multiple layers. In one embodiment, the neural network includes at least an input layer, a concatenation layer, and a dropout layer, an example of which is shown in FIG. 4. In some embodiments, the system receives 230B an XBRL document associated with one or more assigned tags. In some cases, the system may receive the XBRL document by interfacing with another system. In some cases, the system may receive the XBRL document by a data repository (e.g., XBRL data repository 150 of FIG. 1). In some embodiments, the system may receive the XBRL document by retrieving the document via a software interface. In some embodiments, the system is coupled to an electronic filing system by a software interface to receive the XBRL document. In some embodiments, the system is configured by the software interface to use a set of predetermined protocols. In some embodiments, the software interface includes at least one of an application programming interface and a web services interface.

[0040]

[0047] In some embodiments, the XBRL tag validation system is configured to analyze the XBRL document using the trained machine learning model to generate one or more suggested XBRL tags (235B). In some embodiments, the XBRL document includes multiple sections, and each assigned tag is associated with a respective section of the XBRL document. In some cases, the system analyzes each section of the multiple sections and determines one or more suggested XBRL tags for each section. In some embodiments, the system determines a confidence value for each suggested tag of the one or more suggested tags (240B). In some cases, the system applies a multivariate function to generate the confidence value.

[0041]

[0048] In some embodiments, the XBRL tag validation system can compare (245B) one or more assigned tags to one or more suggested tags to generate a comparison result. The comparison result can include one or more assigned tags that match one or more suggested tags. The comparison result can include one or more assigned tags that do not match any suggested tags. In one embodiment, the comparison result can include, for each assigned tag and for each tagging section associated with the one or more suggested tags, whether the individual assigned tag matches any one of the one or more suggested tags.

[0042]

[0049] In some embodiments, the XBRL tag validation system may determine a tag confidence value for each of the assigned tags based on the comparison results (250B). In some embodiments, the tag confidence value of the assigned tag is set equal to the matching suggested tag (e.g., 95%). In some embodiments, the tag confidence value of the assigned tag is set equal to the matching suggested tag (e.g., 95%) of the one or more suggested tags for the corresponding tagged section. In one embodiment, if the assigned tag does not match any one of the one or more suggested section tags for the corresponding tagged section, the tag confidence value of the assigned tag is set to a logical low. In one embodiment, a logical low is represented by 0%.

[0043]

[0050] Optionally, the XBRL tag validation system can determine categories for each tag trust value (255B). In one embodiment, the tag trust categories can include a high trust category, a medium trust category, and a low trust category. In one embodiment, each tag trust category is associated with a predetermined range. For example, the high trust category corresponds to trust values ​​in the range of 40% to 100%, the medium trust category corresponds to trust values ​​in the range of 10% to 40%, and the low trust category corresponds to trust values ​​in the range of 0% to 10%. In another embodiment, the tag trust categories can include more than three categories. In yet another embodiment, the tag trust categories can include two categories.

[0044]

[0051] In some embodiments, the XBRL tag validation system may generate a representation of the XBRL document, assigned XBRL tags, tag trust values, and / or tag trust categories (260B). In one embodiment, the representation of the tag trust values ​​includes an indication of the tag trust value's corresponding category. In one example, the tag trust categories are represented by colors. For example, high trust categories are represented by green, medium trust categories are represented by yellow, and low trust categories are represented by red. Figure 6 shows one illustrative example of a representation of an XBRL document with assigned XBRL tags 610 and tag trust values ​​620. In one embodiment, an indication of the tag trust category for each tag trust value is included in the representation.

[0045]

[0052] 2C illustrates an example flow diagram 200C of a training process for a neural network used within an XBRL tag proposal / validation system, according to certain embodiments of the present disclosure. Aspects of an embodiment of method 200C may be performed, for example, by components of an XBRL tag proposal system (e.g., components of XBRL tag proposal / validation system 100 of FIG. 1 ). One or more steps of method 200C are optional and / or may be modified by one or more steps of other embodiments described herein. Additionally, one or more steps of other embodiments described herein may be added to method 200C. In some embodiments, the system receives / collects multiple XBRL datasets (210C). In some cases, the received / collected datasets are divided into training and validation subsets (215C).

[0046]

[0053] In some embodiments, the training process repeats the following process over several epochs. In the context of artificial neural networks, an epoch refers to one cycle through the complete training data set. Within each epoch, the following steps occur: First, a training subset is fed through the neural network to generate predictions for XBRL tags (220C). A loss function is used to calculate the effectiveness of the current neural network weights (225C). In some embodiments, the loss function determines a categorical cross-entropy loss. Optionally, a validation subset is fed through the neural network to generate predictions for XBRL tags (230C). The system can use the loss function to calculate the effectiveness of the current neural network weights (235C). The system can check whether predictions using the training and / or validation subsets meet certain training criteria (240C). In some embodiments, the training criteria relate to the failure of an iteration to reduce the loss on the validation subset. For example, the training criteria can be when training does not improve the effectiveness of the neural network by a predetermined number of epochs. In some embodiments, the training criteria relate to the effectiveness of the neural network above a predetermined threshold (e.g., 95%). If the training criteria are not met, the neural network weights are updated (245C), e.g., using predictions and loss values ​​generated on the training set, and the training process is repeated (220C). If the training criteria are met, the training process is stopped (250C).

[0047]

[0054] 4 shows an illustrative example of a neural network architecture 400 according to certain embodiments of the present disclosure. In this example, the neural network architecture 400 includes an input layer 410, an embedding layer 420, a flattening layer 430, a concatenation layer 440, a dropout layer 450, and a dense layer 460. One or more layers of the neural network architecture 400 are optional and / or may be modified by one or more layers / functions, including other embodiments described herein. Additionally, one or more layers may be added to the neural network architecture 400. In some cases, an input layer 410 captures input to the neural network model. In the illustrated example, the neural network includes multiple input layers 410 (e.g., a Row_Header input layer, a SIC_Bin input layer, and a Table_Type input layer).

[0048]

[0055] In some implementations, the embedding layer 420 may be used to compress the input into a smaller set and is configured to embed a constant vector associated with the input. In some cases, the embedding vectors of the embedding layer 420 are updated during training. In some cases, the flattening layer 430 converts multidimensional inputs into a single vector. In some embodiments, the concatenation layer 440 is configured to concatenate the flattened vectors for each input so that each input datum can be represented as one long flattened vector. In some cases, the dense layer 460 represents a densely connected neural network layer. In some cases, the dropout layer 450 represents a neural network layer that has the ability to probabilistically remove and re-add nodes within the layer during training. In some cases, this The dropout process helps to avoid memorizing the training data too closely, thus increasing the model's ability to generalize from the data.

[0049]

[0056] Various modifications and alternatives to the disclosed embodiments will be apparent to those skilled in the art. The embodiments described herein are illustrative examples. Unless otherwise specified, features of one disclosed example may also be applied to all other disclosed examples. It should also be understood that U.S. patents, patent application publications, and other patent and non-patent literature referenced herein are incorporated by reference to the extent they do not contradict the above disclosure.

Claims

1. 1. A method implemented on a computer system having one or more processors and a memory, comprising: receiving an XBRL document associated with one or more assigned XBRL tags; analyzing the XBRL document using a trained machine learning model to generate one or more suggested XBRL tags and determine one or more corresponding trust values, each suggested XBRL tag of the one or more suggested XBRL tags being associated with a trust value; comparing the one or more assigned XBRL tags with the one or more proposed XBRL tags to generate a comparison result; and determining a tag trust value associated with each assigned XBRL tag of the one or more assigned XBRL tags based on the comparison; A method comprising:

2. generating a representation of said tag confidence value; The method of claim 1 further comprising:

3. generating a representation of the XBRL document, the representation of the XBRL document including a representation of the one or more assigned XBRL tags and corresponding tag confidence values; The method of claim 1 further comprising:

4. The method of claim 2 , wherein the representation of the tag confidence value includes an indication of a category of the tag confidence value.

5. 2. The method of claim 1, wherein the XBRL document includes a plurality of sections, each assigned XBRL tag of the one or more assigned XBRL tags being associated with a respective section of the XBRL document, and each section of the plurality of sections being associated with one or more suggested XBRL tags.

6. 6. The method of claim 5, wherein determining a tag confidence value includes setting the tag confidence value of an assigned XBRL tag to logic low if the assigned XBRL tag does not match any one of one or more proposed XBRL tags for the section corresponding to the assigned XBRL tag.

7. The method of claim 1 , wherein the trained machine learning model comprises a trained neural network having multiple layers.

8. The method of claim 1 , wherein the trained machine learning model is trained using multiple XBRL datasets.

9. The method of claim 8 , wherein at least some of the plurality of XBRL data sets each include at least one of a row header and a table type.

10. 10. The method of claim 9, wherein the row header is represented by a dense vector.

11. 1. A method implemented on a computer system having one or more processors and a memory, comprising: receiving a plurality of XBRL data sets; training a machine learning model using the plurality of XBRL datasets; Receiving documents; and using the trained machine learning model to predict one or more XBRL tags associated with the document; Including, At least some of the plurality of XBRL data sets each include at least one of a row header and a table type. method.

12. The method of claim 11 , wherein the row header is represented by a vector.

13. The method of claim 11 , wherein the machine learning model comprises a neural network having multiple layers.

14. identifying one or more tagged sections within the document, each tagged section associated with one or more XBRL tags; The method of claim 11 further comprising:

15. determining one or more trust values, each XBRL tag of the one or more XBRL tags being associated with a respective trust value; The method of claim 11 further comprising:

16. 16. The method of claim 15, wherein determining one or more trust values ​​comprises applying a multivariate function to the XBRL document to determine the one or more trust values.

17. generating a representation indicating a label associated with the one or more XBRL tags and the one or more confidence values; 16. The method of claim 15, further comprising:

18. The method of claim 17 , wherein the representation includes an indicator representing a category of the one or more confidence values.

19. determining a set of selected XBRL tags from said one or more XBRL tags; The method of claim 11 further comprising:

20. generating a representation indicating the set of selected XBRL tags; 20. The method of claim 19 further comprising:

Citation Information

Patent Citations

  • XBRL report checking method and device

    CN111695326A

  • Identifying and suggesting classifications for financial data according to a taxonomy

    US20130117268A1

  • XBRL comparative reporting

    US20170069020A1

  • Systems and methods for automated taxonomy migration in an XBRL document

    US8825614B1