Discovering the values of an entity's metrics from a non-normalized dataset
A machine-learned engine processes SEC filings to predict equity metrics for non-public companies using regular expressions and neural networks, addressing the challenge of inconsistent reporting and enhancing market transparency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2026-03-04
AI Technical Summary
Existing systems struggle to accurately and timely predict the value of equity metrics for non-public companies due to the lack of standardized identification processes and inconsistent reporting practices, making it difficult for investors to determine fair market value.
A machine-learned engine processes reported data from various sources, including SEC filings, to harmonize non-standard data and predict the value of equity metrics for non-public companies by using regular expressions and fuzzy logic to identify and extract relevant data points, and then utilizes neural networks to forecast prices.
The system provides up-to-date and accurate predictions of non-public company equity values, enhancing transparency in private markets and reducing the message-to-execution ratio for transactions.
Smart Images

Figure 2026507726000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 483,586, filed February 7, 2023, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure relates to an engine for predicting the value of a metric. [Background technology]
[0003] A mutual fund is a collection of investment securities acquired according to a specific strategy. Mutual funds are managed by a manager, who sells holdings or purchases new securities to keep the mutual fund aligned with its investment strategy. Mutual funds are regulated by the U.S. Securities and Exchange Commission (SEC). For example, the SEC requires mutual funds to report their holdings (lists of securities) on a quarterly basis. One purpose of reporting is to provide transparency to funds. Specifically, these reports allow fund investors to gather information about whether the fund is adhering to its investment strategy. The form for submitting such reports is currently known as the N-PORT. Accordingly, registered management investment companies use Form N-PORT to file periodic (e.g., monthly, quarterly) reports of fund information and quarterly information on portfolio holdings. At least some of the reports are made public as a time snapshot of investment performance. [Brief explanation of the drawings]
[0004] The detailed description and implementation of the present invention will be explained with the aid of the accompanying drawings.
[0005] [Figure 1]FIG. 1 is a system diagram illustrating components for predicting metrics based on periodically reported non-standard data from different sources.
[0006] [Figure 2] 1 is a flowchart illustrating a fund discovery process performed by the system of the disclosed technology.
[0007] [Figure 3] 1 is a flowchart illustrating a process for predicting the value of an equity metric for a target entity.
[0008] [Figure 4] 1 is a flowchart illustrating the process of matching data items from a non-standard dataset of reported returns to predict the value of an equity metric for a target entity.
[0009] [Figure 5A] 1 shows a diagram of a user interface managed by the disclosed technology for a subscriber of the platform. [Figure 5B] 1 shows a diagram of a user interface managed by the disclosed technology for a subscriber of the platform. [Figure 5C] 1 shows a diagram of a user interface managed by the disclosed technology for a subscriber of the platform.
[0010] [Figure 6] FIG. 1 is a block diagram illustrating an example of a computer system in which at least some operations described herein may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0011] The technology described herein will become more apparent to those skilled in the art upon review of the detailed description in conjunction with the drawings. The embodiments or implementations illustrating aspects of the invention are shown by way of example, and like references may indicate similar elements. While the drawings show various implementations for illustrative purposes, those skilled in the art will recognize that alternative implementations may be employed without departing from the principles of the technology. Thus, while specific implementations are shown in the drawings, the technology is susceptible to various modifications.
[0012] The disclosed technology includes a system configured to predict an unknown value of an asset based on reported data posted to a central repository. The reported data is posted to the repository by various sources, which are asset aggregators. In one example, the repository may include the entire U.S. Securities and Exchange Commission (SEC) database of reported data. The reported data includes standard data for public entities and non-standard data for non-public (e.g., private) entities. The standard data includes a fixed value that is determined regardless of the asset aggregator, while the non-standard data varies depending on the aggregator. In other words, the aggregator of the asset for the non-public entity is the source of variability in the non-standard data.
[0013] An asset's fixed value is assigned to a public entity and held by different aggregators but has the same fixed value. In contrast, an asset's variable value is determined by the aggregator and known to the asset's issuer, but is not uniform among aggregators holding the same asset. The variable value data is used to train a machine-learned engine, which is then used to predict values that can be used to acquire assets from non-public entities at comparable values given the data found in the reports.
[0014] The disclosed technology improves upon previous systems by processing reported data with extensive coverage of private issuers in a timely manner to predict data that harmonizes non-standard data. In one example, a machine-learned engine can predict the "mark" of a non-public company based on recently reported data on mutual fund holdings. Due to the nature of filings and SEC regulations that limit what mutual funds (e.g., aggregators) can hold in their portfolios, these marks are unknown and difficult to aggregate regularly and in a timely manner.
[0015] In one example, the system can automatically check reports containing target data daily, which are then extracted and transformed on the same day and used to predict the marks of private companies. In another example, the system can capture new funds and check them as soon as they begin filing with the SEC. Thus, the data points have up-to-date information about issuers of interest. The system can perform processes for matching and filtering issuer names, which addresses the lack of any recognized or standardized identification process for private entities. In another example, a computer-implemented process can predict the marks of non-public entities based on quarterly reported Form NPORT-P filings of mutual fund holdings. A repository receives and stores data in reports communicated over one or more computer networks from various fund manager computer systems. The system can screen reports for target data that is extracted and processed to predict the marks of private entities.
[0016] The report includes data about the fund, such as equity metric values for public companies. In particular, the report includes metric values for the quantity and price of securities held by the fund manager for the company. The report may also include other data for non-public companies representing equity holdings. However, the data for non-public companies may not specify a publicly known price per equity unit. That is, while the price per equity value for public companies is publicly available, the same metric data is not publicly immutable for each non-public company. As a result, the metric values for non-public companies are unknown because they are not defined at any time except at the time of a particular transaction. Therefore, any buyer of a private company's stock lacks a way to determine fair market value (FMV).
[0017] Equity holdings of non-public companies held by multiple fund managers are periodically reported to a repository along with holdings of public companies. Reports of different fund entities reported at different times are communicated to a common repository via a communications network. For example, a monthly report of a first fund entity is issued and communicated to the repository via a computer network, and another report of a second fund entity is issued monthly and communicated to the repository via a computer network.
[0018] The repository publishes a portion of reported data that aggregates multiple marks for various entities. As a result, the metric values for public companies shown in the reported data include a quantity per equity share, such as the price per quantity paid for the shares. The metric values for public companies are known independently of the report. In contrast, different fund managers' reports may show values for holdings in non-public companies, whose values are unknown independently of the report. As a result, the value of a private company's stock is unknown to buyers because its value is undefined at the time. Thus, the report includes marks for public shares, which equate to a per-share value and can represent the aggregate value of private equity holdings. The disclosed technology therefore processes the data in the report to predict a mark for private holdings.
[0019] In one example, Form N-PORT is used by registered management investment firms (also referred to herein as "fund manager services" or "services") to submit monthly portfolio holdings reports. The SEC may use the information provided in the reports in its regulatory, enforcement, examination, disclosure review, inspection, and policymaking roles. Fund managers must report information about their portfolios and each portfolio holding as of the last business day or last calendar day of each month on a quarterly basis. More specifically, the SEC requires Form N-PORT reports for each month in a fiscal quarter to be submitted to the SEC within 60 days of the end of that fiscal quarter (as opposed to filing each monthly report within 30 days of the end of each month). Respondents to the collection of information contained in this form are not required to respond after the end of that fiscal quarter unless a currently valid OMB control number (SEC2940(8 / 22)) appears on the form. The report must disclose portfolio information calculated by the fund for net asset value at the end of the reporting period. This technology can also extract and convert data as soon as it is available in posted reports. That is, as soon as a mutual fund filing reports a new valuation for a particular issuer, the technology can capture that valuation data point and make it available on the platform to facilitate trading of assets from the same issuer, which may be the most recent data point available for that issuer.
[0020] The disclosed technology can identify, extract, and transform data points from one or more reports, aggregate the data points, and process or train a machine learning engine on the aggregated data to predict metrics for equity shares of non-public companies not included as marks in the reports. In one example, an autonomous program (e.g., a bot) on the Internet or another network can interact with a repository's network portal to target specific funds containing data about specific non-public companies. Through this processing, the technology can cover all funds that file Form NPORT-P reports and hold equity in the companies of interest. Thus, the machine learning engine is configured to find marks for non-public issuers that are relatively less than marks for public issuers. Furthermore, the amount (i.e., total value) of the total funds of non-public issuers is relatively less than the total value of public issuers. For example, the marks and / or total value of non-public funds may be 0.1% to 0.5% of the marks and / or total value of public funds.
[0021] In one example, a computer-implemented process has two parts. Specifically, when a new report for a target fund is identified, fund data is collected in a first process and provided to a second process that predicts the mark of non-public equity stocks based in part on the data contained in the new report. The technology checks to find the most recent filed report for the particular fund using an automated process that compares the filing date of the identified report with the date of the last processed report. That most recent filing is then retrieved for further processing. If no newer report is found, the process can still automatically identify, extract, and aggregate useful data from the filing.
[0022] The technique retrieves the most recent reports for aggregators with key identifiers matching the listed key identifiers and generates a data table containing equity metric data extracted from the reports. In particular, data extracted from different funds' reports can be aggregated in a table format and stored in a database, which can be updated periodically (e.g., quarterly) and / or as new filings are submitted. This allows the technique to extract data points and instantly train a machine-learned engine to accurately predict marks immediately after reports are submitted. The extraction process is performed daily, and new data can be constantly added from new fund reports. Metrics for non-public companies are predicted based on data derived from the database. For example, the technique can construct a list containing key identifiers for fund profiles (e.g., funds) and / or aggregators (e.g., fund managers) that issue reports containing equity metrics. The technique selects a key identifier to search a repository of reports containing data on equity held for private entities, but each containing distinct values for volume metrics for public entities rather than private entities.
[0023] Non-public issuers are not required to have key identifiers used to identify and extract public issuer data. Thus, reports do not include key identifiers for non-public issuers that are used to identify and extract public issuer data. Instead, non-public issuer data included in reports includes arbitrary data in fields that normally contain key identifiers. To address this deficiency, the system uses regular expressions (regex) and / or fuzzy algorithms to filter non-public issuer data in reports that match targets of interest on the list.
[0024] A regex is a sequence of characters that specifies a matching pattern in text. For example, patterns are used by string search algorithms for "find" or "search and replace" operations on strings, or for input validation. Thus, a regex algorithm obtains a pattern (or filter) that describes the set of strings that match the pattern. In other words, the regex algorithm accepts a certain set of strings and rejects the rest. Using regex algorithms, systems find patterns in reports that match data from non-public issuers.
[0025] The technology can also find data for a target private entity within the report based on fuzzy logic and derive values for each volume metric for the target private entity. The technology can report predicted values for non-public issuers to a user. The predicted data can be used to train and inform a user before buying or selling a particular asset (e.g., a stock) available for exchange on a platform marketplace managed by the system. The platform can be accessed on an electronic device that presents a control element, which can trigger a transaction to begin based on the predicted value for each volume metric. That is, a predicted price per share for a private company can be presented to prompt a user to begin a large purchase of equity shares of the private company at or near the predicted price.
[0026] FIG. 1 is a system diagram illustrating components of a system 100 configured to forecast metrics based on periodically reported data from disparate aggregator sources. As shown, the system includes a repository 102 that receives reports from disparate sources 104-1 and 104-2 (collectively referred to as "sources 104" and individually referred to as "source 104"). The repository may include a single storage location for all data from the sources 104. This model is utilized to create a single source of truth, providing significant advantages for visibility, collaboration, and consistency within data management. For example, the repository may include one or more servers in a data center connected to the disparate sources 104 via one or more communication networks. In another example, the repository may be a distributed network of storage systems that store reports provided by the sources 104.
[0027] An example of a report includes Form N-PORT, an SEC filing requiring registered investment companies to submit details of their portfolio holdings quarterly, along with monthly breakdowns. An example of a report includes one or more data having a standardized structure for processing by the repository 102. The repository 102 can make the report available to the public or other interested parties via an online interface. The interface can include a network portal managed by the repository 102 for access by subscribers or the general public. For example, the repository 102 can manage a web portal accessible to public users to access the information contained in the reports provided by the sources 104.
[0028] Sources 104 may include one or more servers managed by a fund manager. In one example, sources 104 aggregate information about equity holdings for public and non-public holdings in reports sent to repository 102. In that example, sources 104 are managed by the fund manager. Sources 104-1 and 104-2 are independently controlled to upload reports periodically. For example, reports can be uploaded from sources 104 to repository 102 on a weekly, monthly, or quarterly basis. Once received, the reports can be broken down into data that is searchable and publicly available. In one example, only a portion of the reports are publicly available. For example, repository 102 can receive reports from each source 104 on a monthly basis, but only make quarterly reports publicly available.
[0029] The publicly available data may include reports or portions of reports. Each report and / or asset issuer is associated with a key identifier that is used to map the issuers of the assets included in the report. The key identifier is unique to each source of the respective report. For example, a fund report from a fund manager includes a key identifier that uniquely identifies the fund and / or fund manager. Thus, all reports from the same fund manager include the same key identifier. Fund reports may also be time-stamped to indicate when the report was generated or sent to the repository 102. In this manner, the most recent report for a particular fund may be identified. For example, the machine learning engine may compare the timestamp of a report for the same fund manager to reports previously retrieved from the repository 102.
[0030] The system 100 includes one or more scripts 106 configured to discover, collect, and transform data obtained from the repository 102 into data used to predict unknown metric values. The data included in the reports goes through a discovery process 108, a collection process 110, and a transformation process 112 to generate metric data. The discovery process 108, the collection process 110, and the transformation process 112 are described in further detail below in FIGS. 2 and 3. In some embodiments, a machine-learned engine 114 processes target data from the reports to predict metric values for non-public entities. The target data may include training data stored on a storage device 116 for use in improving subsequent predictions of the metric data.
[0031] A "machine learned" engine can include one or more models, where a model refers to a construct trained using training data to make predictions or provide probabilities for new data items (e.g., private stock prices), regardless of whether the new data items are included in the training data. For example, training data for supervised learning may include items with various parameters and assigned classifications. New data items may have parameters that a model can use to assign a classification to the new data items. As another example, a model may be a probability distribution resulting from an analysis of training data, such as the likelihood of n-grams occurring in a given language based on an analysis of a large corpus from that language. Examples of models include neural networks, support vector machines, Parzen windows, Bayesian, clustering, reinforcement learning, probability distributions, decision trees, decision tree forests, etc. Models are configurable for a variety of situations, data types, sources, and output formats.
[0032] In some implementations, the machine-learned engine 114 may include a neural network with multiple input nodes that receive data from reports and / or outputs of scripts executed to perform the discovery process 108, collection process 110, and transformation process 112. Thus, the input nodes may correspond to functions that receive inputs and generate results. These results may be provided to one or more levels of intermediate nodes, each of which generates further results based on a combination of the results of lower-level nodes. A weighting factor may be applied to the output of each node before the result is passed to the next layer of nodes. At the final layer (the "output layer"), one or more nodes may generate a value that classifies the input, which, once the model is trained, may be used to predict an unknown metric value. In some implementations, such neural networks, known as deep neural networks, may have multiple layers of intermediate nodes with different configurations and may include a combination of models that are convolutional, receiving different portions of the input and / or inputs from other parts of the deep neural network, or partially using outputs from previous iterations of applying the model as further inputs to generate a result for the current input.
[0033] The machine learning engine 114 can be trained using supervised learning, where the training data includes processed or raw data from reports as input and desired output, such as a metric value for successful transactions of private stocks. A representation of the metric can be provided to the model for a predicted metric value. The output from the model can be compared to a desired output for that metric value, and based on that comparison, the model can be modified, such as by changing weights between nodes of the neural network or parameters of a function used at each node of the neural network (e.g., applying a loss function). After applying each data in the training data and modifying the model in this manner, the model can be trained to predict new metric values.
[0034] A user device, such as a desktop computer 118, a laptop computer, a handheld mobile device, or other device with a display, can present a user interface 120, which includes control elements that enable actionable processes based on actionable data or predictive metric data. For example, the actionable data can include predictive metric data (e.g., prices of equity shares of a private company) presented on the user interface 120. A user can submit a predicted price to a private issuer 122 of equity shares. Examples of control elements can include a button, a slider, or another graphic element. For example, a user can adjust one slider to a desired quantity of private shares and another slider to a proposed price, where the slider has a range (e.g., + / - 10%) for the predicted price. Thus, a user can transact directly with a private equity issuer (e.g., source 104-1).
[0035] In one example, the reported marks are presented in a graph on the interface to compare the mark price with other price indicators, such as Indication of Interest (IOI) prices, previous trading prices, and funding round prices. Users can then make an informed decision to trade the issuer's assets on the marketplace platform. The forecast data can be provided in multiple formats. For example, the data may be presented on the marketplace for users to see a sample of the latest price data from a fund manager. In another example, the data platform presents a complete and detailed view of historical data and graphs of price indicators. In yet another example, an application programming interface (API) can provide users with a complete dataset (e.g., mark prices for over 20,000 private issuers from over 300 mutual funds). In this manner, the API enables users to perform their own analysis of the reported dataset.
[0036] The disclosed technology can thus bring transparency to private markets and provide users with new price guidance to private issuers from a direct source to investors. The technology can also reduce the message-to-execution ratio, which corresponds to the number of messages required to execute instructions on a private issuer's assets. That is, fewer electronic messages are required to identify metric values and complete the execution of a transaction to purchase private shares because the predicted price of the private shares is more likely to be accepted by the seller to complete the transaction. In other words, communication between buyers and sellers is reduced, thereby reducing network resource utilization and congestion in the communication network.
[0037] FIG. 2 is a flowchart illustrating a fund discovery process 200 performed by the system of the disclosed technology. Process 200 can be performed periodically and frequently to ensure that up-to-date data is extracted to consistently cover any new funds that have recently been created and hold the issuer of interest or have just added the issuer of interest to their holdings. Process 200 can update a list of regex for fund portfolios that hold equity in non-public companies of interest. The portfolios are managed by one or more fund management services that aggregate the assets of the portfolios. That is, a management service is an aggregator that manages portfolios that aggregate equity shares of different entities into funds. An example of a fund management service includes issuers of mutual funds or exchange-traded funds (ETFs). Each fund management service has a unique identifier that distinguishes it from other services that manage other mutual funds.
[0038] Each fund management service can issue a variety of different funds, each with a unique key identifier. The key identifier can include a string of characters or another combination of elements that uniquely identifies a particular service or portfolio from others. For example, a particular fund can be identified based on a combination of the service's key identifier and the key identifier of the portfolio managed by the service. Examples of different entities include public entities (e.g., public companies) and non-public entities (e.g., private companies). A mutual fund portfolio can include equity metrics for public companies, such as the quantity and value per unit of equity shares held by the fund's issuer. A mutual fund can also hold equity in private companies, allowing the management service to report the mutual fund's holdings even though the public price per share of the non-public entity is undefined.
[0039] At 202, key identifiers are selected for non-public entities. For example, key identifiers for private companies of interest are identified to predict their metric values (e.g., price per equity share) based on reports published to the repository from management services that hold equity in the non-public entities. For example, key identifiers for two services that hold equity in a particular private company of interest are identified. Another key identifier for a different service that holds equity in another private company of interest is also identified. Thus, key identifiers for different private issuers are selected.
[0040] At 204, key identifiers for portfolios including equity of the non-public entity of interest are collected. In one example, the script uses the key identifiers of the non-public companies to search the websites of various management services to identify key identifiers for portfolios including equity of the non-public company of interest. In one example, the script is executed by a software agent that collects key identifiers for management services and key identifiers for their portfolios including equity of the non-public company. For example, key identifiers identifying mutual funds may be collected from the management service's website or a third-party service that maintains key identifiers. In one example, the key identifiers are unique to particular funds managed by different services. For example, a fund management service may manage 10 funds, where only three include equity of the non-public company of interest. The collected key identifiers may be for the three funds that include equity of the non-public company of interest. Key identifiers for the remaining funds that do not include equity of the non-public company of interest are excluded.
[0041] At 206, the script compares the collected key identifiers to an existing list of key identifiers used to monitor the repository for reporting metric data. For example, the key identifiers are compared to the existing list to determine whether a new key identifier is missing from the list and should be added, or whether the new key identifier is erroneously recorded in the list. In one example, the list is stored in a database, mapping key identifiers to fund names.
[0042] At 208, the updated list of key identifiers is communicated to a software agent configured to monitor the repository for reports of portfolios identified based on the key identifiers. Thus, the collected key identifiers of the fund and / or fund management service are compared to key identifiers currently known and in use by the software agent to search the repository for reports.
[0043] 3 is a flowchart illustrating a process 300 for predicting the value of an equity metric of a target entity (e.g., a private company). Process 300 can be performed by one or more servers of a computer system coupled to a repository via one or more networks (e.g., the Internet). In one example, a non-transitory computer-readable storage medium stores instructions that, when executed by at least one data processor of the system, cause the system to perform the functions described in process 300. The first stage of process 300 begins at 302 with fund identification.
[0044] At 302, the system compiles a list containing key identifiers for the fund management service's fund portfolio, as described with respect to FIG. 2. The portfolio includes equity metrics or related data for two types of entities (e.g., public entities and non-public entities). The list can be processed by a script executed by a software agent that searches reports stored in a repository. In one example, a bot executes a script that searches the repository for reports issued by the fund's management service that include equity information of interest for non-public entities. For example, the bot can enter the key identifier of a fund and / or its fund manager into a search field in the repository's web portal to search for matching keys associated with metric data included in recently posted reports.
[0045] At 304, a specific key identifier for a particular portfolio is selected from the list. The selected key identifier is included in a query submitted to a field of the repository to retrieve relevant reports. The repository stores multiple separate reports for different funds communicated over one or more computer networks (e.g., the Internet) from servers of multiple fund management services. The reports include distinct values per volume metric for each public entity and equity data for non-public entities of interest, but exclude unique values per volume metric for non-public entities. The reports stored in the repository are communicated to the repository periodically (e.g., monthly) from servers of multiple management services over one or more computer networks; in some cases, only a portion of the reports (e.g., only quarterly reports) are made publicly available. The second stage of process 300 begins at 306 with retrieving fund declarations.
[0046] At 306, the repository is monitored based on the key identifiers on the list to search for specific reports generated by specific management services for portfolios that match specific keys on the list. For example, a bot can generate a query string that is entered into a search field on the repository's website. The query is used to search for reports from management services that manage fund portfolios with matching keys. At step 308, the bot can recursively select the next key identifier on the list of keys to monitor the next portfolio report, and so on. Thus, one or more key identifiers are included in one or more queries to search the repository for reports issued by one or more management services. The third stage of process 300 begins at 310 with performing an extraction process.
[0047] At 310, the system collects fund reports and associated metadata that allows it to identify the most recent report for a particular portfolio. The entire portfolio, or a portion thereof, is retrieved from the repository. The report can be identified by looking up a key identifier and comparing the report's timestamp to identify the most recent report among a group of reports submitted by the same management service, or by comparing the report's timestamp to the current date or the last date of the report previously retrieved from the repository.
[0048] At 312, a data table is generated and / or populated with equity metric data for non-public entities extracted from the reports retrieved from the repository. The data table aggregates equity metric data for non-public entities extracted from the reports. The data table may also aggregate data for public entities extracted from the reports in addition to data from non-public entities. In one example, an XML file is generated and populated with the data extracted from the reports. The system may also execute a script to process the table file (e.g., the XML file) to determine where to select relevant data from all the data in the report at 314. In one example, a machine learning engine may be trained to identify relevant portions of the report. The fourth stage of process 300 begins at 316 and involves performing an issuer identification and cleaning process to transform the extracted data for predicting equity metrics.
[0049] At 316, a script is executed to find target data of non-public entities of interest within the data table, for example, based on a fuzzy logic matching process. For example, names used to identify issuers of private stocks are ambiguous between filings of different aggregators. For example, the names of private issuer stocks can include characters or be omitted, making string matching impossible. This can result because the same company may have a public name that is different from its legal name, and different aggregators may use one name or the other. In fact, completely random names that are not human-recognizable may be used to identify related issuers.
[0050] A fuzzy logic matching process can find similar but not identical entries that represent non-public entities of interest. In one example, key identifiers for target non-public entities are vectorized and compared to other vectorized keys in the data to identify the target data. For example, text in a data table is matched based on a particular vector key assigned to a particular non-public entity. The matching process can identify issuers using data other than the issuer's name. For example, the matching process can identify target issuers using data indicating the issuer's country, the exchange rate associated with the issuer, or any number of multiple dimensions.
[0051] At 318, security features are optionally cleared from the target data of the target non-public entity in the data table. In one example, clearing security features includes performing text and pattern recognition to determine security types and remove unnecessary information from the target data of the non-public entity of interest.
[0052] At 320, a value per volume metric is predicted for the non-public entity of interest based on the target data extracted from the report. In one example, the value per volume metric for the non-public entity of interest is predicted by processing the non-public entity's equity metric data with a machine learning engine including a model generated and trained based on data extracted from the report, as described above. The output of the machine learning engine includes a predicted value per volume metric for the non-public entity. In another example, the value of the non-public entity is determined from one or more reports of multiple funds issued by one or more management services. The total unit equity value of the non-public entity held by each aggregator is analyzed to predict or estimate a value based on reports from different aggregators. For example, values may be averaged for the same mutual fund or multiple mutual funds. Thus, the predicted equity metric value is estimated by dividing the total value by the total unit value in one or more reports issued by one or more aggregators. In another example, data points for unit equity value are weighted differently for different aggregators. The output may include a range of unit equity values or a specific value. The fifth stage of process 300 begins at 322 with performing the upload process.
[0053] At 322, the system causes one or more electronic devices to present actionable information or actionable control elements based on the predicted metric data of the non-public entity, as described above. In one example, execution of the actionable control element causes communication of a message configured to initiate a trade of one or more equity units of the non-public entity at the predicted value per volume metric. Additional analytics providing insight into the target non-public entity can also be derived and presented to the user on the electronic device.
[0054] 4 is a flowchart illustrating a process 400 performed by the platform's computing resources for matching data items from non-standard datasets extracted from reported filings. The extracted data items are used to predict or estimate the value of equity metrics for target entities (e.g., non-public issuers) available to platform subscribers. Process 400 may operate iteratively to update or refine values for target issuers available to subscribers. For example, process 400 may operate iteratively to aggregate data items for a target issuer obtained over time and / or obtained from additional or different aggregators. Thus, process 400 may aggregate data for the same issuer from the same source and / or new sources.
[0055] Process 400 can improve platform performance and computational efficiency by pulling only data items of target issuers for pre-processing (e.g., sorting, filtering, and extraction). Rather than discarding data items of non-target entities in reported filings, the platform stores the raw data in a repository. Thus, the platform can pull raw data items from the repository when an issuer is added as a new target of process 400. That is, the raw data can be processed later to extract data items for the new target. Furthermore, the platform can process the raw data of newly recognized identifiers of target issuers to update or fine-tune the values of the target issuer. For example, the platform can discover strings identifying target issuers that were not considered in previous iterations of process 400. Thus, process 400 reduces processing by curating data items for target issuers while maintaining available raw data to expand target issuers and / or to expand identifiers for existing target issuers.
[0056] The ingest pipeline 402 is a source of datasets that are processed to group data items of matching target entities and predict or estimate values of metrics for those entities. In one example, the dataset includes equity information for public and private companies, as well as identification information for issuers of the equity. A central index key (CIK) table 404 can store key identifiers for aggregators that hold assets and related information for the issuers. The contents of the CIK table 404 can be obtained from a repository, such as the SEC's computer system, to identify entities (e.g., companies, individuals) that have filed disclosures with the SEC. Information from the CIK table 404, including the key identifiers, is provided to the ingest pipeline 402. Additionally, information about issuers (e.g., private companies) is stored in an issuer table 422 and provided to the ingest pipeline 402.
[0057] The SEC API 406 is operable to search the SEC EDGAR archive repository for recently disclosed SEC filings and access related corporate documents. In particular, the SEC API 406 can discover and analyze audited and unaudited financial statements from 10-Q and 10-K filings, extract text content from EDGAR documents, convert filings into different formatting file types such as PDF, Word, or Excel, and stream SEC filing data in real time. Thus, the CIK table 404 can store key identifiers and associated data items extracted from the streamed SEC filing data, where the key identifiers are those of entities (e.g., mutual fund holders) that filed disclosures with the SEC.
[0058] The ingest pipeline 402 and SEC API 406 provide data sets to the extract component 408, which functions to extract target data items from the reports as soon as they are available from the SEC. The extract component 408 stores the raw data sets obtained from the ingest pipeline 402 and SEC API 406 in a repository in a raw bucket 410 and can provide the extracted target data to the transform component 412. In one example, the ingest pipeline 402 can check daily to see if an aggregator (e.g., a mutual fund) has filed a new N-PORT return. When a new return is discovered, code executed by the extract component 408 creates a direct URL that can access the XML format of the return. The extract component 408 then downloads and stores each return as an XML file, along with the metadata for that return. A raw XML file is created and the extracted data is stored in a file named “[name] / ... <yyyy-mm-dd> / <cik> / data / <accession-number>.xml.” In one example, each declaration is stored in a database and rescanned for new issuers added to the platform.
[0059] The transformation component 412 can convert the extracted data from the declaration into a readable and queryable table format. The transformation component 412 can also perform cleaning operations on the extracted data items. This processing can include correcting or removing incorrect, corrupted, improperly formatted, duplicate, or incomplete data items. When combining multiple data items from different sources or sources collected at different times, there are many opportunities for data to be duplicated or mislabeled, which the transformation component 412 can repair. In one example, the transformation component 412 can convert the extracted data items into a parquet data format that includes fields of interest. Parquet data is a columnar data file format designed for efficient data storage and retrieval.
[0060] The table creation component 414 creates a refined table 418 that can store all data points from the filing. The transformation component 412 can add metadata (e.g., fund name, filing date, filing number) to the raw bucket 410 and / or refined bucket 416 for future use. The extracted fields from the data items can be inserted as data records into the table created by the table creation component 414. Thus, the table creation component 414 stores a parquet dataset containing values for the fields of interest. The transformation component 412 can read and / or extract data directly from the extraction component 408 and / or raw bucket 410 using a Python library. In one example, the Python library includes Panas, which is used to manipulate the datasets. It can have functions for analyzing, cleaning, exploring, and manipulating the extracted data.
[0061] The refinement bucket 416 stores the transformed data in parquet format. The transformed data is easier to query compared to pre-transformed data and can therefore be used to quickly discover unknown values of metrics. The platform can read the data from the file and create a table that is updated with every new declaration that is received. The matching process also uses the refinement bucket 416 to read data items and extract matching data items. The refinement table 418 provides a table view for querying the data, analyzing the data, or downloading the data (e.g., using AWS Athena to read the table). In one example, the table contains all records from the declaration and all fields that can be reused for various purposes as needed.
[0062] Matching engine 420 obtains data from table creation component 414 and issuer table 422. Issuer table 422 includes one or more tables that store information about non-public issuers identified in declarations. For example, issuer table 422 may aggregate identifiers and metric values of non-private issuers collected over time and used later to discover and update the values of their equity metrics. Issuer table 422 is synchronized to track match attempts with known issuers and backfills for new issuers. Thus, for example, matching engine 420 can match data items extracted from table creation component 414, refinement table 418 through table creation component 414, and / or refinement bucket 416 based on data for known issuers stored in issuer table 422. Adding a new issuer to the issuer table 422 triggers a search of the refinement table 418 using the matching engine 420, which then adds the identified data to the matched table 426 via the transformation component 424. The matching engine 420 generates multiple Glue jobs to process the issuers in parallel. A Glue job encapsulates scripts that connect to source data, process it, and then write it to a data target. Typically, a job executes extract, transform, and load (ETL) scripts. A job can also execute generic Python scripts (Python shell jobs). In one example, the matching engine 420 runs regex or fuzzy logic algorithms to match data for the same issuer retrieved from the table creation component 414, the refinement table 418, and the refinement bucket 416 with the issuer table 422. The matching engine 420 can use parallel multiprocessing to reduce job execution time. In one example, the matching engine 420 performs string matching on millions of data points.To reduce overall processing time, matching engine 420 can run multiple instances of the same job simultaneously.
[0063] In one example, the matching engine includes a regex processor that converts regular expressions into an internal representation that can be executed and matched against strings representing the text being searched. One possible approach is to build a non-deterministic finite automaton (NFA), which is then made deterministic. The resulting deterministic finite automaton (DFA) is then executed against the target text string to recognize substrings that match the regular expression. Thus, the regex processor can match regular expressions for target non-private issuers from the issuer table 422 with data items obtained by the matching engine 420 from the creation table component 414 or other sources. Regex algorithms can be used to preprocess strings before a matching job is performed. Regular expressions can be used to extract the exact equity type for a particular issuer and / or for a particular equity type. In another example, the platform uses the RapidFuzz algorithm, a fast string matching library for Python and C++. The fast string matching library uses the Levenshtein distance to find the closest similarity between two strings and identify data items or the same issuer.
[0064] The following example illustrates string transformation and matching: The platform matches data for one issuer (using multiple aliases and attributes) against, for example, 12 million records of different names and aliases used by different funds for the same issuer. The fuzzy nature of the algorithm addresses the issue that private issuer marks are not identified with a specific identifier number but are scattered within an aggregator's portfolio. [Table 1]
[0065] As described above, data items of the same issuer are obtained from various reports. After the matching process is complete, issuer table 422 updates or creates a table containing data of only the issuer of interest (e.g., the target issuer). That table can then be used by matching engine 420 to search for and aggregate data items of the same issuer.
[0066] Process 400 optionally includes another transformation component 424 coupled to matching engine 420 to perform transformation jobs that refine the matched data before it is made available for subscriber consumption. For example, transformation component 424 can assign internal issuer names, derive price marks, and clean up security names before releasing the data to subscribers.
[0067] The matched table component 426 stores the data of identified matches. The matched refinement table component 428 is the final component that serves to load the platform's output for subscriber consumption. The output may include a prediction or estimate of the value of a metric for the equity of a non-public issuer. The discovered value may be estimated based on a numerical calculation, such as, for example, the average metric value for the equity of a particular issuer.
[0068] Process 400 can generate estimates of price marks for private securities despite the lack of a recognized or standardized method of referencing or identifying private issuers in mutual fund filings (e.g., the lack of a standard identifier). Process 400 can also disambiguate variable references to common private issuers, thereby resolving the problem of mutual funds using different methods to refer to or identify private securities. Process 400 can estimate price marks for issuers of interest as soon as the fund files an N-PORT with the SEC.
[0069] In one embodiment, process 400 analyzes a dataset of over 18 million unique data points for private issuers. For example, process 400 can analyze over 31,000 filings from over 2,500 mutual funds over a four- or five-year period. The process identifies over 600,000 individual securities of the target private issuers from the over 12 million individual securities held by the mutual funds. The 600,000 individual securities are used to perform price analysis and other historical data analysis. Currently, mutual funds are required to limit the aggregation of their illiquid assets to less than 15%, and depending on the type of fund, a small percentage of these are typically allocated to private issuer securities, making identifying the necessary securities a challenging process. The platform aggregates price marks but also aggregates different fields (attributes) related to each private issuer. Estimates of the total number of mutual funds filed with the SEC currently stand at approximately 10,594 mutual funds. So, with four filings per year per mutual fund and approximately 1,160 securities per filing, that's 49,183,523 different data points (securities) per year. If the scope of private issuers covered includes more than 2,600 names, process 400 can identify any security associated with one of these issuers within the 50 million data points within the SEC.
[0070] 5A-5C illustrate screen displays of a user interface managed by the disclosed technology for platform subscribers. In particular, FIG. 5A illustrates a screen display 500A of a user interface showing a price comparison of mutual fund marks. A subscriber can use this screen display to visualize a comparison of fund prices to another price index for a selected timeline. For example, screen display 500A includes a graph 502A plotting a comparison between the average price reported for all funds over time and funding round prices. A price index can be selected in section 504A.
[0071] FIG. 5B shows a screen display 500B of a user interface presenting historical market data. The screen display includes a graph 502B that plots a comparison between the average price reported for all funds, funding round prices over time, and the average price reported by a selected fund for a specific issuer. In this example, a subscriber compares how a BlackRock fund assigns prices to a specific issuer with the prices of all other major funds (e.g., Fidelity, Franklin Templeton, etc.). Screen display 500B includes a section 504B that allows a subscriber to select a specific fund family and / or fund for analysis and presents the analyzed data, including average values for all funds over a period of time, in graph 502B. A subscriber can further drill down to review specific fund data for each listed fund.
[0072] 5C shows a screen display 500C of a user interface with a table showing various mutual funds being analyzed. The screen display includes a section 504C that presents price metrics for different funds for a selected time period (e.g., "2023-Q2"). Subscribers can view more detailed information about specific securities via links to source filings in the table and select and view section 506C, which contains information about preferred and / or common securities.
[0073] Computer Systems
[0074] FIG. 6 is a block diagram illustrating an example of a computer system 600 capable of implementing at least some operations described herein. FIG. 6 is a block diagram illustrating an example of a computer system 600 capable of implementing at least some operations described herein. As illustrated, computer system 600 may include one or more processors 602, a main memory 606, a non-volatile memory 610, a network interface device 612, a video display device 618, input / output devices 620, a control device 622 (e.g., a keyboard and pointing device), a drive unit 624 including a storage medium 626, and a signal generating device 630 communicatively coupled to a bus 616. Bus 616 represents one or more separate physical buses and / or point-to-point connections connected by appropriate bridges, adapters, or controllers. Various common components (e.g., cache memory) have been omitted from FIG. 6 for clarity. Instead, computer system 600 is intended to represent a hardware device in which the components shown or described with respect to the illustrative example, as well as any other components described herein, may be implemented.
[0075] Computer system 600 can take any suitable physical form. For example, computing system 600 can share an architecture similar to that of a server computer, a personal computer (PC), a tablet computer, a mobile phone, a game console, a music player, a wearable electronic device, a network-connected (“smart”) device (e.g., a television or home assistant device), an AR / VR system (e.g., a head-mounted display), or any electronic device capable of executing a set of instructions that specify the action(s) to be taken by computing system 600. In some implementations, computer system 600 can be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC), or a distributed system such as a mesh of computer systems or includes one or more cloud components in a network. Where appropriate, one or more computer systems 600 can perform operations in real time, near real time, or batch mode.
[0076] Network interface devices 612 enable computing system 600 to broker data within network 614 with entities external to computing system 600 via communication protocols supported by computing system 600 and the external entities. Network interface devices 612 include network adapter cards, wireless network interface cards, routers, access points, wireless routers, switches, multi-layer switches, protocol converters, gateways, bridges, bridge routers, hubs, digital media receivers, and / or repeaters, as well as all wireless elements mentioned herein.
[0077] Memory (e.g., main memory 606, non-volatile memory 610, machine-readable medium 626) may be local, remote, or distributed. Although shown as a single medium, machine-readable medium 626 may include multiple media (e.g., centralized / distributed databases and / or associated caches and servers) that store one or more sets of instructions 628. Machine-readable (storage) medium 626 may include any medium capable of storing, encoding, or carrying a set of instructions for execution by computing system 600. Machine-readable medium 626 may be non-transitory or include a non-transitory device. In this context, non-transitory storage media may include a tangible device, meaning that the device has a concrete physical form, although the device can change its physical state. Thus, for example, non-transitory refers to a device that remains tangible despite changes in this state.
[0078] While the embodiments have been described in the context of a fully functional computing device, various examples may be distributed as a program product in various forms. Examples of machine-readable storage media, machine-readable media, or computer-readable media include recordable-type media such as volatile and non-volatile memory devices 610, removable flash memory, hard disk drives, optical disks, and transmission-type media such as digital and analog communications links.
[0079] Generally, the routines executed to implement the examples herein may be implemented as part of an operating system or as a specific application, component, program, object, module, or sequence of instructions (collectively referred to as a "computer program"). A computer program typically includes one or more instructions (e.g., instructions 604, 608, 628) that are configured at various times in various memory and storage devices within a computing device(s). When read and executed by processor 602, the instruction(s) cause computing system 600 to perform operations that implement elements comprising various aspects of the present disclosure.
[0080] remarks
[0081] The terms "example," "embodiment," and "implementation" are used interchangeably. For example, references to "one example" or "an example" in this disclosure may, but do not necessarily, refer to the same embodiment; such references refer to at least one of the embodiments. Appearances of the phrase "in one example" do not necessarily all refer to the same example, nor do they refer to separate or alternative examples that are mutually exclusive of other examples. Features, structures, or characteristics described in connection with an example may be included in other examples of the disclosure. Furthermore, various features are described that may be exhibited by some examples and not by other examples. Similarly, various requirements are described that may be requirements of some examples but not other examples.
[0082] The terms used in this specification should be interpreted in their broadest reasonable manner, even when used in conjunction with certain specific embodiments of the present invention. Terms used in this disclosure generally have their ordinary meanings in the art, in the context of this disclosure, and in the specific context in which each term is used. The description of alternative words or synonyms does not exclude the use of other synonyms. Whether a term is detailed or discussed herein should not be given special importance. The use of highlighting does not affect the scope and meaning of a term. Furthermore, it should be understood that the same thing can be said in more than one way.
[0083] Unless the context clearly requires otherwise, throughout the specification and claims, terms like "comprise," "comprising," and the like should be construed in an inclusive sense, i.e., "including, but not limited to," rather than an exclusive or exhaustive sense. As used herein, the terms "connected," "coupled," or any variation thereof, mean any direct or indirect connection or coupling between two or more elements. As used herein, the terms "connected," "coupled," or any variation thereof, mean any direct or indirect connection or coupling between two or more elements, and the coupling or coupling between elements may be physical, logical, or a combination thereof. Furthermore, the terms "herein," "above," "below," and words of similar import, when used in this application, refer to this application as a whole and not to any particular portions of this application. Where the context allows, words in the above detailed description using singular or plural numbers can also include plural or singular numbers, respectively. The word "or," when referring to a list of two or more items, encompasses all of the following interpretations of the word: any item in the list, all items in the list, and any combination of items in the list. The term "module" broadly refers to a software component, a firmware component, and / or a hardware component.
[0084] While specific examples of techniques are described above for illustrative purposes, those skilled in the art will recognize that various equivalent modifications are possible within the scope of the invention. For example, while processes or blocks are presented in a given order, alternative implementations may perform routines having steps or use systems having blocks in a different order, and some processes or blocks may be deleted, moved, added, subdivided, combined, and / or modified to provide alternative or subcombinations. Each of these processes or blocks may be implemented in a variety of different ways. Also, while processes or blocks are shown as being performed serially in time, these processes or blocks may instead be performed or implemented in parallel, or may be performed at different times. Furthermore, any specific numbers described herein are merely examples, and alternative implementations may use different values or ranges.
[0085] The details of the disclosed embodiments may vary considerably in particular embodiments while still being encompassed by the disclosed teachings. As noted above, certain terms used in describing features or aspects of the present invention, as redefined herein, should not be construed as limiting the disclosed technology to the specific characteristics, features, or aspects of the technology with which they are associated. In general, the terms used in the following claims should not be construed as limiting the disclosed technology to the specific examples disclosed herein, unless such terms are explicitly defined in the detailed description above. Thus, the actual scope of the technology encompasses not only the disclosed examples, but also all equivalent ways of practicing or implementing the technology under the scope of the claims. Some alternative embodiments may include additional or fewer elements than the above-described embodiments.
[0086] All of the above patents and applications and other references, as well as any that may be listed in accompanying application documents, are incorporated herein by reference in their entirety, except for any subject matter disclaimer or disclaimer, and except to the extent the incorporated material contradicts the language of the disclosure herein, in which case the language of the disclosure will control. Aspects of the present technology can be modified to employ the systems, functions, and concepts of the various references above to provide further implementations of the present technology.
[0087] To reduce the number of claims, certain aspects of the invention are presented below in the form of several claims, but the applicant contemplates other forms of various aspects of the invention. For example, claimed aspects may be described in means-plus-function form or in other forms, such as embodied on a computer-readable medium. Claims intended to be interpreted as means-plus-function claims use the words "means for." However, the use of the term "for" in other contexts is not intended to invoke a similar interpretation. The applicant reserves the right to pursue additional claims after the filing of this application, either in this application or a continuation application. < / cik> < / yyyy-mm-dd>
Claims
1. 1. A method for discovering a value for an equity metric of a non-public entity, comprising: Ingesting a non-standard dataset of reports disclosed by an aggregator of a combination of public and non-public entity equities, the non-standard dataset is streamed from a public repository using an application programming interface (API); the ingesting, wherein the report includes variable values of equity metrics issued by public and non-public issuers and maintained by the aggregator; extracting data items from the non-standard data based on one or more key identifiers of one or more aggregators and non-public issuers of equity; said extracting, wherein said data items include variable values for a particular metric of a particular equity of a particular non-public issuer; executing a script to convert the extracted data items into a standard format; the executing, wherein the script includes a library configured to read the extracted data items; populating a data table with the converted data items of the particular non-public issuer in the standard format; spawning a plurality of computing jobs to process in parallel the transformed data items in the data table for the particular non-public issuer; spawning, wherein each job encapsulates a script and executes a regex or fuzzy logic algorithm configured to perform a string matching process to match data items of a common issuer in the data table; Discovering values of the metrics for each equity of the particular non-public issuer based on the output of the plurality of jobs; The method comprising:
2. causing an electronic device to present an actionable control element associated with the value of the metric for each equity of the non-public entity; 10. The method of claim 1, wherein execution of the actionable control element results in communication of a message configured to initiate a trade of one or more equity units of the non-public entity at a value per unit of quantity.
3. Before spawning the plurality of computing jobs, selecting a particular key identifier for a particular aggregator that holds equity units for the particular public issuer or the particular non-public issuer; adding the particular key identifier to a query configured to search a Central Index Key (CIK) table for a particular data item that matches the particular key identifier; the adding, wherein the specific data items are extracted using the query that searches the CIK table; The method of claim 1 further comprising:
4. Detecting new reports published to the public repository that contain key identifiers that match the one or more key identifiers of the one or more aggregators and the non-public issuer; updating the values of the metrics for each equity of the particular non-public issuer based on new data items extracted from the new report; and The method of claim 1 further comprising:
5. 1. A method for predicting an equity metric of a target entity, comprising: and configuring the list to include specific keys for fund portfolios issued by specific management services of equity metrics for two types of entities; said two types of entities include public entities and non-public entities; selecting the particular keys for the particular fund portfolios to include in a query; the selecting, wherein the query is configured for input to a repository to search for reports matching the particular key; monitoring the repository based on the query for a particular report generated by the particular management service having a key matching the particular key; the repository is configured to store a plurality of separate reports for different funds communicated over one or more computer networks from a plurality of management service servers; Each report contains a separate value for each quantity metric for each public entity; each report includes equity metric data for the non-public entity but excludes publicly available values for each volume metric for the non-public entity; retrieving a particular report for the particular fund portfolio that has a timestamp subsequent to other reports stored in the repository for the fund portfolio; populating a data table with equity metric data for the non-public entities extracted from the particular reports; inputting the data, wherein the data table includes additional data about the non-public entity aggregated from additional reports about the particular fund portfolio and additional fund portfolios; Finding target data for the target non-public entity in the data table based on fuzzy logic matching similar but non-identical entries representing the target non-public entity; clearing security features from the target data of the target non-public entity in the data table; predicting a value for each quantity metric of the target non-public entity based on the target data of the target non-public entity with the security features cleared; causing an electronic device to present an actionable control element associated with the predicted value for each quantity metric of the target non-public entity; execution of the actionable control element results in communication of a message configured to initiate a trade of one or more equity units of the target non-public entity at the predicted value per volume metric; The method comprising:
6. configuring the list to include the particular key; selecting the specific key for the target non-public entity of interest; collecting keys for one or more fund portfolios containing equity data of the target non-public entities of interest; said collecting, wherein a script is executed to navigate between websites of management services that host said one or more fund portfolios; comparing the collected keys to keys stored in a current list configured to monitor a fund portfolio; adding keys collected to monitor reports on funds of interest to said list of keys; The method of claim 5 , comprising:
7. 6. The method of claim 5, further comprising recursively selecting a next key on the list of keys to monitor a next report for a next fund portfolio.
8. the equity metric relates to equity shares of a public or non-public company; said equity metric of any public company is public; The method of claim 5 , wherein the equity metric of any non-public company is private.
9. The method of claim 5 , wherein the equity metric relates to an equity share of a public or non-public entity.
10. selecting the particular key to monitor the repository; The method of claim 5 , further comprising configuring a query to include one or more keys to search the repository of one or more reports generated by one or more management services.
11. reports are periodically communicated from the plurality of management services to the repository via the one or more computer networks; The method of claim 5 , wherein only a portion of the reports communicated over the one or more computer networks are publicly available.
12. a plurality of reports each including a total equity value of the target non-public entity maintained by a management service, and forecasting a value per said quantity metric of the target non-public entity; processing the equity metric data of the target private entity with a machine learning engine that is generated and trained based on reports that include data of the private entity; outputting a predicted value for each of the quantitative metrics of the target non-public entity as an output of the machine learning engine; The method of claim 5 , comprising:
13. predicting a value for each of the quantitative metrics for the target non-public entities; determining a total value held in the target non-public entity's fund portfolio; forecasting the total unit value of the target non-public entities held in the fund portfolio; Including, The method of claim 5 , wherein the predicted value for each quantitative metric is estimated by dividing the sum value by the sum unit value.
14. retrieving the particular report, comparing timestamps of the reports for the particular fund portfolio to a current time; identifying a most recent report as the particular report for the particular fund portfolio based on the comparison; The method of claim 5 , comprising:
15. populating said data table with said non-public entity data extracted from said particular report; generating an xml file; populating the xml file with entries for the data of the non-public entities; The method of claim 5 , comprising:
16. The method of claim 5 , wherein the data table includes data about the public entities aggregated from reports in addition to the data from the non-public entities.
17. Discovering the target data of the target non-public entity based on fuzzy logic; matching text in the data table based on the particular key; The method of claim 5 , wherein the particular key is for the particular non-public entity.
18. clearing the security feature, The method of claim 5 , comprising performing text recognition and pattern recognition to determine security types and filter out unnecessary information from the targeted data of the targeted non-public entities.
19. The method of claim 5 , further comprising deriving additional analytics that provide insight into the target non-public entity.
20. Discovering the target data of the target non-public entity in the data table based on fuzzy logic; and comparing the key for the target non-public entity with the keys in the data table.
21. 1. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one data processor of a system, cause the system to: monitoring a centralized repository for reports with specific keys matching specific fund portfolios; the centralized repository stores reports generated by EquityMetric's multiple management services for two types of entities; said monitoring, wherein each report includes a value for each quantity metric for entities of a first type but not for entities of a second type; retrieving a particular report for the particular fund portfolio; the retrieving, wherein the particular report is the most recent report among other reports stored in the centralized repository for the particular fund portfolio; generating a data table containing equity metric data for entities of a first type extracted from the particular report; generating, wherein the data table is aggregated to include additional data for the first type of entity extracted from additional reports collected periodically from the centralized repository; identifying target data for a target entity in the data table based on a machine learning engine; the identifying the target entity is an entity of a first type; predicting a value for each quantitative metric for the target entity based on the target data; causing an electronic device to present the predicted values for each quantity metric for each of the target entities; The non-transitory computer-readable storage medium.
22. Monitoring the centralized repository includes: constructing the list to include keys for the fund portfolios of each of the equity metrics of the two types of entities; recursively selecting a key from said list for monitoring a fund portfolio; 22. The non-transitory computer-readable storage medium of claim 21, comprising causing:
23. The two types of entities include public companies and private companies, and the system further comprises:
22. The non-transitory computer-readable storage medium of claim 21, wherein the non-transitory computer-readable storage medium is configured to cause a display device to present a graphical element configured to trigger execution of a transaction with a private company based on the predicted value for each volume metric of the private company.
24. at least one processor; at least one non-transitory memory storing instructions, the instructions, when executed by the at least one hardware processor, causing the system to: monitoring a centralized repository of data files containing keys for particular fund portfolios; the centralized repository stores a plurality of data files generated by a plurality of EquityMetric's management services; said monitoring, wherein each report includes values for each quantity metric for entities of a first type but not for entities of a second type; retrieving a particular report for the particular fund portfolio; the particular report being the most recent report among other reports stored in the centralized repository for the particular fund portfolio; aggregating the equity metric data for the first type of entities extracted from the particular reports into a set of data tables containing additional data for the first type of entities; identifying target data for a target entity in the data table; the identifying, wherein the target entity is an entity of a first type; deriving a value for each quantitative metric of the target entity based on the target data.