Domain specific language interpreter and interactive visual interface for high speed screening
The domain-specific language interpreter and interactive visual interface facilitate efficient data exploration and screening by allowing users to enter and edit symbols and operators, providing immediate feedback and AI-assisted pattern identification, addressing the inefficiencies in existing data screening methods.
Patent Information
- Application Number
- JP2025167141
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-24
- Filing Date
- 2025-10-03
- Publication Date
- 2026-01-27
AI Technical Summary
Sifting through large amounts of data to find valuable information is challenging, as exemplified by tasks such as recruiters finding job candidates, home buyers identifying homes, or investors screening securities, where existing screening strategies are often difficult and inefficient.
A domain-specific language interpreter and interactive visual interface that allows users to freely enter and edit symbols and operators, providing immediate feedback and enabling rapid refinement of screening strategies, with AI assistance for identifying patterns and trends.
Enables rapid discovery and exploration of data, allowing users to visualize screening criteria effects and perform complex operations with minimal technical expertise, integrating ideation, screening, research, testing, execution, and monitoring in a unified workflow.
Smart Images

Figure 2026012710000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 029,556, filed May 24, 2020, entitled "Domain-Specific Language Interpreter and Interactive Visual Interface for Rapid Screening," listing Sara Itani as the inventor. The entire contents of all priority documents referenced in the above-referenced applications and in the Application Data Sheet filed herewith are incorporated herein by reference in their entirety for all purposes.
[0002] FIELD OF THE INVENTION The present disclosure is directed to improved systems and methods that enable users of computing systems to explore and filter data in domain-specific attribute-rich data sets to test filtering strategies and discover targets of particular interest. [Background technology]
[0003] Sifting through large amounts of data to find something of value is often a difficult task. Metaphors like "looking for a needle in a haystack" and "searching high and low" describe the challenges, for example, of a recruiter finding candidates to interview for a job, a home buyer identifying a small number of homes to consider viewing before buying, or an investor screening securities to select a group of stocks worth considering as an investment.
[0004] For example, stock markets present investors with a vast number of securities, with a vast amount of information available about each security from a vast number of sources. To make the process of selecting securities of interest more manageable, investors may screen the securities against several criteria to narrow down the list. The result is a smaller list, and whether the remaining securities are more promising depends on the investor's selection of criteria and their sophistication in selecting and validating them. Even effectively screening securities for useful, worthless, misleading, etc., based on the vast number of criteria available for narrowing down the list can be difficult.
[0005] Screening strategies may attempt to reflect the investment philosophy of the investor, or they may be largely treated as a formality or used sporadically to support intuition. [Brief explanation of the drawings]
[0006] [Figure 1] 1 illustrates an exemplary user interface of a rapid screening system showing a multi-line editor and grid view display configured for stock screening, according to one embodiment. [Figure 2] 1 illustrates a calculation routine for a rapid screening system according to one embodiment. [Figure 3A] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening showing revisions within a multi-line editor according to one embodiment. [Figure 3B] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening showing revisions within a multi-line editor according to one embodiment. [Figure 4] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing a dialog for creating a custom population of strains according to one embodiment. [Figure 5A]1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing domain-specific flexible text matching and completion suggestions according to one embodiment. [Figure 5B] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, illustrating data tag searching according to one embodiment. [Figure 6A] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, illustrating filtering on criteria according to one embodiment. [Figure 6B] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, illustrating filtering on criteria according to one embodiment. [Figure 7] 1 illustrates an example user interface of a rapid screening system configured for stock screening showing expressions assigned to custom variable names according to one embodiment. [Figure 8A] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing the simultaneous renaming of multiple references to a custom variable name according to one embodiment. [Figure 8B] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing the simultaneous renaming of multiple references to a custom variable name according to one embodiment. [Figure 8C] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing the simultaneous renaming of multiple references to a custom variable name according to one embodiment. [Figure 9] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, illustrating domain-specific syntax error handling according to one embodiment. [Figure 10]1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing a conversion function according to one embodiment. [Figure 11] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing automated graphical display of sequence data according to one embodiment. [Figure 12] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing the automatic display of links to 10-K filings according to one embodiment. [Figure 13] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing a selective display of companies that hold patents, according to one embodiment. [Figure 14] 1 illustrates an example user interface of a rapid screening system configured for stock screening showing filtering on text found in 10-K filings according to one embodiment. [Figure 15] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing grouping of results according to one embodiment. [Figure 16A] 1A-1C illustrate exemplary formulations for a prior art system and corresponding exemplary expressions for a rapid screening system configured for stock screening, illustrating improved usability according to one embodiment. [Figure 16B] 1A-1C illustrate exemplary formulations for a prior art system and corresponding exemplary expressions for a rapid screening system configured for stock screening, illustrating improved usability according to one embodiment. [Figure 17A] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, illustrating backtesting according to one embodiment. [Figure 17B]1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, illustrating backtesting according to one embodiment. [Figure 18] 1 illustrates an example modified alert graph of a rapid screening system configured for stock screening, according to one embodiment. [Figure 19A] 1 illustrates exemplary AI capabilities for introspection in a high-speed screening system configured for stock screening, according to one embodiment. [Figure 19B] 1 illustrates an exemplary AI function for forecasting in a high-speed screening system configured for stock screening, according to one embodiment. [Figure 20] 1 illustrates an exemplary AI function for regime change detection in a high-speed screening system configured for stock screening, according to one embodiment. [Figure 21] 1 illustrates an exemplary AI function for optimizing blending in a high-speed screening system configured for stock screening, according to one embodiment. [Figure 22] 1 illustrates an exemplary AI function for feature suggestions in a high-speed screening system configured for stock screening, according to one embodiment. [Figure 23] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, illustrating the creative use of operators according to one embodiment. [Figure 24] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing automatic formatting according to one embodiment. [Figure 25] 1 illustrates an exemplary user interface of a rapid screening system configured for stock screening showing point-in-time status reporting of forecasts according to one embodiment. [Figure 26]1 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing a situation report historical forecast graph according to one embodiment. [Figure 27] FIG. 1 is a block diagram illustrating some of the components typically incorporated into computing systems and other devices in which the technology may be implemented. [Figure 28] FIG. 1 is a schematic diagram and data flow diagram illustrating various components of an exemplary concurrent server interaction for backtesting according to one embodiment. [Figure 29] FIG. 1 is a schematic diagram illustrating various components of an exemplary server system for implementing a rapid screening system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0007] This application discloses improved systems and methods that enable users of computing systems specially configured for a domain to explore and filter data in attribute-rich data sets to test screening strategies and discover targets of particular interest.
[0008] The disclosed technology utilizes a novel, intuitive approach that includes a domain-specific language interpreter and an interactive visual interface that both immediately responds to symbols and operators entered in the domain-specific language. The technology includes an input interface, such as a multi-line editor, that allows a user to freely enter, modify, add, insert, subtract, change, and otherwise edit the entered symbols and operators at any time. The domain-specific language interpreter continuously processes the entered symbols and operators as they are updated. The interactive visual interface also includes a grid view that displays live results that update according to the current content of the input interface.
[0009] The technology further provides fast, simultaneous evaluation of selected targets against historical data and benchmarks, allowing strategies to be backtested in seconds. The technology also includes artificial intelligence (AI) machine learning capabilities to assist users, for example, in identifying strategies or the factors driving them, finding similar targets, and considering different evaluation criteria.
[0010] For example, as applied to securities information, the systems and methods disclosed herein provide an improved approach to screening securities.
[0011] Together, various aspects of the disclosed technology offer high ease of use with a short learning curve and provide immediate feedback that enables rapid refinement of iterative search strategies for exploring, visualizing, and filtering structured and / or unstructured data. The domain-specific language and associated interpreter and interactive visual interface bring together data exploration, querying, visualization, and the goal at hand in one central workflow. The structure of the domain-specific language and the immediate, user-friendly feedback provided by the rapid screening system combine to promote iterative discovery and exploration, allowing users to visualize the effects of screening criteria and making it easy enough for users with only business spreadsheet experience to rapidly perform complex screening operations. Thus, those without technical expertise who want to find answers can obtain them through the presently disclosed technology without relying on a separate person or team with such expertise. As a further result, users can explore options and refine their thinking in real time based on the functionality provided by the technology.
[0012] Additionally, the disclosed technology enables a fundamental shift in the process for ideation, screening, research, testing, execution, and monitoring of strategies (e.g., investment strategies). In the past, each of these processes was handled separately. The disclosed technology uniquely integrates them all. It replaces disparate processes performed by different people with limited and separate feedback loops and brings everything together in one language, tools, and interface with a unified and immediate feedback loop.
[0013] The techniques described in this disclosure that provide the basis for the disclosed rapid screening system include advances and insights in compilers, human-computer interface or interaction (HCI), programming language design, database engineering, distributed computing, machine learning, quantitative analysis, and finance. As a result, the improvements in this disclosure as a whole combine improvements in several different technologies and would not generally be apparent to one skilled in any one technology.
[0014] Reference will now be made in detail to the description of the embodiments as illustrated in the Figures. While the embodiments are described in connection with the drawings and associated description, there is no intent to limit the scope to the embodiments disclosed herein. On the contrary, the intent is to cover all alternatives, modifications, and equivalents. In alternative embodiments, different or additional input interfaces (e.g., drop-down menu selectors or natural language processing) may be added to or combined with those shown without limiting the scope of the embodiments disclosed herein. For example, the embodiments described below are primarily described in the context of stock screening. However, the embodiments described herein are illustrative examples and are not intended to limit the disclosed technology to any particular application, domain, body of knowledge, type of search, or computing platform.
[0015] The phrases "in one embodiment," "in various embodiments," and "in some embodiments" are used repeatedly. Such phrases do not necessarily refer to the same embodiment. The terms "comprising," "having," and "including" are synonymous. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. It should also be noted that the term "or" is used in its sense to include "and / or" unless the content clearly dictates otherwise.
[0016] The disclosed rapid screening systems and methods can take a variety of form factors. Figures 1-29 show several different configurations and designs. The rapid screening systems shown are not an exhaustive list; in other embodiments, syntax can be rearranged (e.g., "only" or "limit" instead of "filter" or "~"), or the input editor or result display can be configured in a different arrangement. However, the details of such optional implementations need not be exhaustively shown to describe exemplary embodiments.
[0017] 1 illustrates an exemplary user interface 100 of a rapid screening system showing a multi-line editor 101 and a grid view display 102 configured for stock screening, according to one embodiment. The user interface 100 may be provided by one or more rapid screening system computing devices, such as a local or remote server, described in more detail below with reference to FIGS. 27, 28, and 29. In the illustrated user interface 100, the rapid screening system provides the user with a multi-line editor 101 having user-entered symbols 110, 120, 130, 140, 150, 160 in a series of lines numbered 1 through 6.
[0018] Each of rows 1-6 includes symbols 110-160 that are part of a domain-specific language (in this case, a language specific to the domain of securities). Symbols 110-160 represent one or more of: a population of identifiers to start from, where each identifier represents a member of the population; each data value associated with the identifier; a data tag for the data value; an operator to operate on the data value; and / or filter criteria to restrict or narrow the population of identifiers that are considered.
[0019] The specification defines a domain-specific language. A domain-specific language is domain-focused and therefore may include constraints on general-purpose programming tools or languages. General-purpose languages like Python or Structured Query Language (SQL), in contrast, are cross-domain programming languages, not domain-specific languages. Most of the syntax in a domain-specific language is domain-specific terminology, and those skilled in the domain will recognize most of the syntax in a domain-specific language. This means that domain-specific languages contain little programming-specific syntax and are not suitable for general-purpose programming. It also means that users can obtain results without any "programming," and high-speed screening systems display results that are semantically meaningful and therefore human-readable, even if the user enters just a single data tag.
[0020] In various embodiments, the domain-specific language includes symbols that represent either data associated with the identifiers (including the results of operations on the data) or filtering operations to select among the identifiers.
[0021] In some embodiments, the domain-specific language does not include non-domain-specific keywords, in which case the operators are all symbols (such as "+" or brackets "["..."]"). As a result, all alphabetic text that the user enters into the multi-line editor 101 can be understood to be meaningful, for example, a data tag. In some embodiments, the domain-specific language does not include non-domain-specific keywords other than transformation functions.
[0022] Each data tag represents a data value associated with an identifier (e.g., a security that may be identified by a unique security ID and / or a recognizable stock ticker symbol), including a value calculated from other values, such as the result of an expression. In many embodiments, the data value is also associated with a date, a date range, or a set of dates. Data value types may include, for example, numeric values (e.g., a security's latest price or return on equity (“ROE”)), string values (e.g., a country or sector), arrays of multiple values (e.g., the last four quarterly ROEs or recent news headlines), data structures / collections / serializations of JSON (JavaScript Object Notation) or XML (eXtensible Markup Language) or HTML (HyperText Markup Language) or CSV (comma-separated values) data (e.g., information about a 10-K filing that includes dates and hyperlinks), etc.
[0023] An expression is a finite, well-formed combination of data tags and operations in a domain-specific language, and a line of user input can be interpreted as one expression or multiple expressions. Each expression entered by the user is evaluated to produce one of the above types of data values for each identifier (e.g., concatenating or otherwise combining strings to produce a new string, or calculating a normalized and weighted blend of ROE and return on assets ("ROA") to produce a new numeric value). An expression can be, for example, a "+" or "* ". Thus, expressions result in data values that are calculated from other data values, including from other expressions.
[0024] A transform is an expression that generates metadata that characterizes the data values of an identifier against a standard, such as over time, against a standard distribution, or compared to other securities (e.g., mean, rank, quintile, normalization, trend stability, etc.).
[0025] A variable name represents any collection of data values (e.g., the value of any expression the user wants to compute) that the user assigns to any user-chosen variable name, and this variable name becomes a custom data tag that represents that data value for easy later reference. In this way, users can extend the domain-specific language.
[0026] The universe of identifiers (e.g., securities) to consider and screen may include predefined sets (such as all stocks, all members of the S&P 500®, all foreign securities, all large U.S. capitalization stocks, corporate or government bonds, etc.) and / or user-defined custom sets or templates.
[0027] Matching operators provide, for example, numeric comparison operators (e.g., <, <=, ==, >=, >, != / <>) (e.g., Revenue > 10M) (in various embodiments, abbreviated units can be treated and / or displayed, e.g., M = million or k = thousands), or text comparison operators such as "contains" or "is" or "?" (TickerSymbol = "AAPL"? or Latest10K = "china tariffs"). In some embodiments, operators such as "contains" are implemented without alphabetic text. For example, using bracket notation to indicate an "include," "has," or "contains" relationship, the syntax could be "Filings10k["china tariffs"]". This provides an intuitive representation of one thing (the phrase "china tariffs") within another (a collection of company 10-K filings) and avoids non-domain-specific keywords in an otherwise domain-specific language.
[0028] A filtering operator selects identifiers that match certain data values or expression criteria according to one or more matching operators (e.g., ~RoE>0 or Price<100). In various embodiments, filtering results in displaying a reduced-size population or a smaller portion of the initially selected population, displaying data values that correspond to the criteria used for filtering.
[0029] In various embodiments, the domain-specific language is processed through an environment or tool, such as the rapid screening system shown in Figure 1, that includes an interactive multi-line editor configured to accept user input (e.g., from a keyboard, speech recognition or natural language processing, on-screen buttons, drop-down menus, and / or context menu item selection from among various input options, some of which may be displayed in place of the editor), a parser and interpreter configured to interpret the input according to specifications defining the domain-specific language, a server engine configured to obtain structured and / or unstructured data relevant to the domain according to the input, and a visualization tool configured to provide a live-updating results display or grid view as the user enters input into the interactive editor. In various embodiments, each expression and / or data tag entered into the editor corresponds to a column of information displayed in the grid view.
[0030] Unlike existing systems, the resulting data is presented in a manner that can be meaningfully interpreted and easily visualized by the user. This reduces the time a user must spend ensuring the correct data is pulled down and potentially cleansed. This also makes it easier to discern patterns and relationships. Thus, a rapid screening system helps inform investment strategies for users, and unlike traditional screeners, the disclosed system and method allows for rapid discovery and exploration of data, testing different ideas to validate or invalidate those ideas, flexible generation of custom indicators and metrics, a repeatable process, identification of trends, commonalities, and characteristics, and actionable insights that can be directly applied to strategies.
[0031] Returning to FIG. 1 , in the first row 110, the user has entered "$UnitedStatesAll." In this example (within the securities domain), "$" is a signifier indicating a population of identifiers, which is all U.S. stocks. As used herein, an identifier represents a member of a population, a particular security, in the example shown. Thus, the user has chosen to begin by considering all U.S. stocks. Within the securities domain, the user may alternatively choose to consider other types of securities, such as, for example, corporate or government (e.g., municipal) bonds, options, mutual funds, etc., stocks from other countries, stocks traded on specific exchanges, industry-specific securities, fixed income financial products, distressed debt, derivative securities, etc. This disclosure generally uses the terms "stock," "share," or "company" as convenient shorthand, but the intent of this disclosure is to encompass all securities bearing those terms.
[0032] Below the multi-line editor 101, a grid view display 102 is shown. In other embodiments, the arrangement of the multi-line editor 101 and the grid view display 102 may be different, such as placing the editors below, side by side, or in completely different windows or screens. The grid view display 102 includes a series of columns 115, 123, 125, 133, 135, 145, and 155. Each of the columns corresponds directly to a symbol or expression in the multi-line editor 101. For example, the illustrated "Ticker" column 115 corresponds to the "$UnitedStatesAll" population of all U.S. stocks. The "Ticker" column 115 displays a series of rows, one row for each identifier associated with the population of all U.S. stocks. Thus, each U.S. stock is identified in the "Ticker" column 115 by its ticker symbol. The ticker is merely a display name; in various implementations, the technology uses a unique identifier to unambiguously identify each security. For example, each member of the selected population may be identified by its full name and exchange, or by some other unique identifier, such as, for example, a security ID. This is particularly useful for supporting international stocks. In other embodiments of the present technology, the grid view display 102 may be arranged differently, such as as a nested table where each column represents a security (or other screening target) and each row represents an attribute of the security, or other equivalent arrangements within the scope of this disclosure.
[0033] In various implementations of the displayed column and row format, the data values in each column are sortable, such as by user interaction with the column header (e.g., mouse clicks to sort, reverse, or restore sort order or context menus). In various implementations, the columns themselves are rearrangeable. For example, the grid view display 102 may provide controls for a column to be shifted left or right relative to another column, dragged to a different position between columns (e.g., by a mouse click-and-drag operation), hidden or closed (or undisplayed), or minimized (e.g., as an icon indicating a collection or grouping of data that can be re-expanded). In some embodiments, when a column of information is moved or removed within the grid view, the interface displays an indication in the editor with the corresponding expression and / or data tag, such as a control to enable the column to be shown again or a notation indicating that the column is displayed out of its original order. In some embodiments, the system also updates the code in the editor when a column of information is moved or removed within the grid view.
[0034] Continuing with the multi-line editor 101, on the second row 120, the user enters "ReturnOnEquityPct|Standardize=>zScoreROE." While the specific syntax of the illustrated example is described in more detail below, the present technology encompasses various equivalent alternatives and is not limited to the exact domain-specific language shown. In this example, the user first selects or enters the data tag "ReturnOnEquityPct." As used herein, a data tag represents an attribute of each of the members of a given population; the data tag labels or references a specific data value associated with each identifier. For example, in a population of all U.S. stocks, ReturnOnEquityPct refers to the percentage return on equity for each company with shares in that population. Similarly, as shown in this example on the third row 130 of the multi-line editor 101, ReturnOnAssets refers to the return on total assets for each such company.
[0035] Thus, in grid view display 102, column 123 shows a "Return On Equity Pct" data tag in a header row at the top of the column and displays data values listing the percent return on equity for each ticker symbol shown in the corresponding row. In the example shown, the screen runs for the current date. In various embodiments, the system allows the screen to run for prior "dates."
[0036] Rows 120 and 130 begin with the ReturnOnEquityPct and ReturnOnAssets data tags and proceed to expressions that transform each of them by normalizing them to a normal distribution, generating a z-score. In some embodiments, the normalization function first culls values to reduce the influence of outliers, replaces null values with the median, and so on. The illustrated expressions then assign those z-scores to custom-named variables or data tags "zScoreROE" and "zScoreROA," respectively. Thus, column 125, under the heading "z-score ROE," displays the normalized z-score for each stock's return on equity percentage (represented by its ticker symbol in column 115). Column 135, under the heading "z-score ROA," similarly displays the normalized z-score for each stock's return on assets.
[0037] Line 140 in multi-line editor 101 is an expression that adds the zScoreROE and zScoreROA data tags and assigns them to a new variable or data tag "zScoresAdded." Thus, column 145 in grid view 102 has the heading "z Scores Added" and displays the new data value for each stock identifier that has a ticker symbol in column 115. The data value in column 145 is the sum of the data values in columns 125 and 135.
[0038] In this example, obvious rounding can be observed, and the displayed value is automatically limited to two decimal places to enhance readability during the screening process, while the system can operate on the exact numbers underlying the displayed data value. In other embodiments, the use of significant figures or other approaches can be employed to provide ease of use. In some embodiments, user actions such as hovering over a data value, right-clicking, copy-pasting, or long-pressing can reveal additional details about the value.
[0039] Continuing with multi-line editor 101, on fifth row 150, the user has entered the expression "zScoresAdded|SplitQuintiles=>bucket." As shown in the function suggestion pane and explanatory text 151, the SplitQuintiles function in this example sorts the data values of zScoresAdded across the current set of identifiers (e.g., security IDs, even if a recognizable ticker symbol is displayed) into quintiles, with the lowest quintile assigned a value of "1" and the highest quintile assigned a value of "5" (or vice versa in some systems). Thus, among all U.S. stocks, companies in the bottom 20% of Return on Equity Percentage (normalized) + Return on Assets (normalized) have a value of "1" in column 155, companies in the top 20% have a value of "5" in column 155, and other companies have a value of "2," "3," or "4" in column 155. In grid view 102, the heading of column 155 is "bucket" because in multi-line editor 101, the expression in row 150 assigns the quintile data tag to the custom variable "bucket."
[0040] In line 6 160, the user enters the expression "~bucket==5." The tilde in this expression represents a filtering operation in the syntax of this illustrative example. This operation allows the user to filter, limit, reduce, or narrow the set of identifiers previously displayed according to matching criteria. In this case, when the rapid screening system processes the filtering expression in line 160 in the multi-line editor 101, the grid view 102 no longer displays all identifiers (or stock ticker symbols, for convenience) in the full range of all U.S. stocks. Instead, only stock tickers for which the custom variable "bucket" has a data value 165 of "5," i.e., the highest quintile of assets for the combined normalized return on equity percentage and return on assets, are included in the updated grid view 102.
[0041] Thus, the disclosed rapid screening system allows users to perform advanced screening of large amounts of multi-attribute data, much more easily than previous systems have allowed, with minimal syntax and unprecedented, continuous, and effective immediate feedback and results.
[0042] Figure 2 shows a computational routine 200 of a rapid screening system according to one embodiment. In various embodiments, the computational routine 200 is executed by one or more rapid screening system computing devices, such as a local or remote server, described in more detail below with reference to Figures 27, 28, and 29. The computational routine 200 begins at start block 201.
[0043] FIG. 2 and the subsequent flow diagrams and schematic diagrams are representative and do not show every function, step, or data exchange; instead, they provide an understanding of how a system may be implemented. Those skilled in the art will recognize that some functions may be repeated, modified, omitted, or supplemented, and that other (less significant) aspects not shown may be readily implemented. Those skilled in the art will understand that the blocks shown in FIG. 2 and in each of the schematic diagrams described below may be modified in various ways. For example, while processes or blocks are presented in a given order, alternative implementations may perform the routines in a different order, and some processes or blocks may be rearranged, deleted, moved, added, sub-divided, combined, and / or modified to provide alternative or sub-combinations. Each of these processes or blocks may be implemented in a variety of different ways. Also, while processes or blocks are at times shown as being performed sequentially, these processes or blocks may instead be performed or implemented in parallel, or may be performed at different times. 2 and other schematic diagrams are of a type well known in the art and may themselves comprise sequences of operations that need not be described herein. Those skilled in the art will be able to create source code, microcode, program logic arrangements, etc., or implement the disclosed techniques based on the figures and detailed description provided herein.
[0044] At block 215, the computation routine 200 generates an interface for user input and / or display of rapidly screened results. For example, in one embodiment used as an example to illustrate various computations described below, the computation routine 200 provides an interactive editor (such as the multi-line editor 101 of FIG. 1 ) and a results grid view (such as the grid view 102 of FIG. 1 ). In various embodiments, the rapid screening system causes the interface for user input and / or display of results to be presented on a client device remote from the rapid screening system server.
[0045] At block 225, the computational routine 200 obtains input, for example, via the interactive editor provided at block 215, and parses the input with respect to a domain-specific language specification or grammar. The domain-specific language specification or grammar may, for example, contain terms that represent data tags from a particular domain, such that a user with domain knowledge can use that knowledge to discover terms in the domain-specific language. For example, in the field of securities, particularly stocks, "ROE" is a common shorthand for "return on equity," which is a measure of profitability or financial performance calculated, for example, by dividing annual take-home pay by shareholder's equity (assets minus liabilities). Thus, if a rapid screening system is configured for such a domain and receives an input of "roe" from a user, the computational routine 200 can parse the input as a request for a data tag named "ROE" or "ReturnOnEquity" using text pattern matching.
[0046] In various implementations, parsing the input in block 225 includes performing a lexical analysis of the input to identify symbols or tokens from a domain-specific language specification or grammar in the input, and performing a syntactic analysis to turn the input symbols into an abstract syntax tree (AST).
[0047] At block 235, the computation routine 200 handles parsing errors. For example, a user may type "roe" when the domain-specific language does not contain a data tag named "roe" or "ROE." In some embodiments, the computation routine 200 logs any errors that prevent further progress through the computation routine 200 for the row with the error (including, for example, errors in other stages such as data retrieval) and loops back to block 225 to process additional user input, such as corrections. In various embodiments, the computation routine 200 proceeds as much as possible despite the error, preserving results that existed before the error was encountered.
[0048] In some embodiments, the computation routine 200 identifies unrecognized input to the user, such as by highlighting it. In some embodiments, the computation routine 200 may attempt to infer the closest match in the domain-specific language, for example, replacing the unrecognized input with the closest valid match or suggesting a set of potential replacements or completions for the unrecognized input. In various embodiments, the computation routine 200 continues to parse the user's input and process parser-recognized symbols.
[0049] At block 245, the computation routine 200 obtains any data needed to perform the operation on the data tags parsed at block 225. For example, if the parsed expression requires a set of data values that are not already loaded into memory, the computation routine 200 identifies one or more data sources associated with the required data values and loads the required information from at least one of the one or more data sources. In some embodiments, the computation routine 200 obtains the data at a separate stage or on the fly.
[0050] In some embodiments, the computing routine 200 loads or preloads and caches snapshots of slowly changing (in terms of volatility or stability of amplitude and / or frequency of change), recently displayed, and / or frequently requested data values for a set or subset of identifiers and / or data tags to speed up processing when user input requires such data values. In some embodiments, the computing routine 200 identifies relevant data tags (e.g., based on user history, popularity across users, availability, etc.), prefetches them (e.g., in parallel), caches the data locally, and provides results quickly to the user.
[0051] In block 255, the computation routine 200 interprets the parsed input. For example, if the user inputs an expression, the computation routine 200 identifies the data values referenced by the parsed symbols, determines which operations to perform, and performs the necessary operations. For example, the computation routine 200 may interpret the AST and evaluate each parsed expression. In some embodiments, the computation routine 200 operates row by row. In some embodiments, the computation routine 200 evaluates each input row in order.
[0052] In some embodiments, the computing routine 200 interprets and executes inputs continuously (e.g., without waiting for the user to stop typing) as soon as the input is analyzed, so that the interface is updated as soon as the input is entered and results are obtained, and the user does not need to finish entering the input before determining the execution of the completed input. Results may be considered “live” if they are updated immediately or after a short delay whenever the user input changes (and optionally whenever the underlying data changes). Providing live results increases the learnability of the interface in addition to facilitating the learnability and exploration of the data set itself. In some embodiments, the computing routine 200 interprets and / or executes the inputs after a short delay (e.g., a “debounce” period of about one-tenth of a second to about one to five seconds, or as soon as the user pauses while typing) that allows the user to confirm, correct, or revise the input before the rapid screening system has finished processing the input. In some embodiments, the computing routine 200 applies such a “debounce” period to the display of results, as described further below. Such a delay can help the computing routine 200 present results when the user is ready, improving the user's perception of receiving immediate responsive results.
[0053] In various embodiments, the computation routine 200 provides incremental interpretation, interpreting and executing only statements, expressions, or lines affected by user input. For example, when the computation routine 200 receives, parses, and interprets the fifth line of an expression, the computation routine 200 may leave the contents of the first four lines unchanged to minimize processing time. By interpreting incrementally and combining new or changed elements with previous results when possible, the high-speed screening system maximizes the responsiveness of the interface.
[0054] In various embodiments, a domain-specific language is fast for the computation routine 200 to interpret (255) due to a combination of several reasons: the interpreter can exploit its design for concurrency and data sharding (e.g., across time, data items, and user-specific patterns); the language is domain-specific and therefore highly specialized and optimized for the task rather than being suitable for general-purpose computing; only results that need to be visible to the user are sent to the client; and it is designed for incremental interpretation, allowing the language itself to understand the meaning of user actions and optimize the query planner. Furthermore, back-end systems are configured to provide fast results based in part on constraints on the domain-specific language, as further described below with reference to Figures 27-29.
[0055] In block 265, after the input is parsed and interpreted, the computation routine 200 updates, removes, and / or adds categories to the interface for display of results corresponding to the interpreted input from the interface for user input. For example, in an example grid view embodiment, when the rapid screening system interprets an expression that newly invokes a data tag, the computation routine 200 adds a column with a header associated with that data tag in the grid view. In some embodiments, the computation routine 200 generates new column header names for expressions not explicitly named by the user.
[0056] In some embodiments, as the computation routine 200 is acquiring data as described in block 245, the computation routine 200 may display a column of data tags and fill in data values as they are received. In some embodiments, the computation routine 200 displays a column of each calculation or newly introduced data tag in each statement or expression (e.g., each line in a multi-line editor), providing transparency and understanding (and aiding in debugging) because all of the intermediate steps are visible. In some embodiments, the computation routine 200 displays one column for each statement or expression and hides intermediate calculations by default. When the user edits or deletes an expression, the computation routine 200 correspondingly modifies or removes data tags and data values that are no longer included in the input. In various embodiments, the computation routine 200 incrementally updates the result display by combining new or changed elements with previous results to minimize processing and rendering time and therefore maximize the responsiveness of the interface. In some embodiments, the computation routine 200 acquires data as described in block 245 and displays results only when enough data has been received to display to the user. For example, rather than displaying headings over nearly empty columns in a grid view display while the results are being obtained, the computation routine 200 can assemble the results (e.g., most or all of the result values, or the result values that the user can see first) before updating the display. Adding complete or nearly conflicting results to the display all at once when the values are ready, rather than piecemeal, can help the computation routine 200 present results in a manner that improves the user's sense of receiving immediate results. Generally, results that are presented within a few seconds of the user completing an expression or line of input are perceived as "immediate." Users may discount startup delays, for example, due to the initial loading of complex data sets, when determining whether an interface will provide immediate responsiveness in presenting results thereafter.
[0057] In block 275, after the input is parsed and interpreted, the computation routine 200 filters the displayed identifiers corresponding to the selected population of identifiers and the active filter criteria. For example, in an example grid view embodiment, when the rapid screening system interprets the expression filtering data values associated with the data tags, the computation routine 200 determines which identifiers match the filtering criteria and displays only the matching identifiers and the data values associated with the matching identifiers. In some embodiments, the computation routine 200 continuously builds a result set with each input row or expression entry. In some embodiments, as result data is obtained, the computation routine 200 inserts the result set into a database for display (e.g., a client may initially request a subset of the data visible to the user, allowing the client to actually display that subset of the entire result set). In some embodiments, the computation routine 200 can use results from an existing session with the user.
[0058] In block 285, the computation routine 200 determines whether additional input has been received, for example, via an interface for user input. In some embodiments, the computation routine 200 may process the additional input (including modifying or deleting a previous input anywhere in the previous input) asynchronously and without waiting for other blocks (e.g., obtain data (245)) to complete, or may cancel processing of a previous input in response to new or changed input to ensure fast display of recently requested data. In other words, the computation routine may determine whether additional input has been received at any time. If additional input has been received, the computation routine 200 loops back to the input analysis block 225 to process the next input symbol, if any.
[0059] The calculation routine 200 ends at end block 299 .
[0060] 3A-3B illustrate an exemplary user interface 300 of a rapid screening system configured for stock screening, showing revisions within a multi-line editor according to one embodiment.
[0061] 3A, the first line 310 in the multi-line editor 301 specifies the "$UnitedStatesSmallAndMidCap" identifier, i.e., the population of U.S. stocks classified as small and mid-capitalization, and the second line 320 is a data tag for "CompanyName." Thus, in the grid view 302, the first column 315 displays the stock ticker symbols (or unique security IDs, for example, as appropriate identifiers) for the stocks in the specified population, and the second column 325 displays the associated company name for each identifier. While the illustrated example focuses on U.S. stocks, international securities could alternatively be selected.
[0062] In FIG. 3B, the first row 312 is different, now specifying the "$UnitedStatesLargeCap" identifier, i.e., the population of U.S. stocks classified as large capitalization. The second row 320 remains unchanged and is a data tag for "CompanyName." In the grid view 302, the revised first column 317 displays a new set of stock ticker symbols (instead of identifiers) for stocks in the specified population, and the second column 327 displays the associated company name for each identifier. In this example, the rapid screening system accepts modifications to the inputs on the first rows 310-312, even after the inputs have been entered and processed on the second row 320. The rapid screening system seamlessly processes such changes, completely replacing the field of identifiers under consideration without requiring the user to delete and revert the input or otherwise start over. In some embodiments, the rapid screening system allows users to undo changes (e.g., through an undo stack or with keyboard strokes such as Control-z) and caches results, allowing users to rapidly move back and forth between dynamic results. By providing flexible, freely editable multi-line input and by continually refreshing updates in response to changing user input, the disclosed rapid screening system allows users to powerfully and rapidly explore alternatives to identify desired strategies in ways not previously possible.
[0063] 4 illustrates an exemplary user interface of a rapid screening system configured for stock screening, showing a dialog 400 for creating a custom population of stocks according to one embodiment. A user may wish to create custom criteria for the starting population of identifiers under consideration. For example, a real estate agent may specialize in housing type or location (e.g., urban apartment complexes or suburban single-family homes), or an investor may focus on an industry category or company size. In various embodiments, the rapid screening system provides a convenient interface for defining new or custom populations using a variety of criteria that can be optimized based on frequent use.
[0064] In the example shown, dialog 400 prompts the user to provide a name 410 for a custom population of stocks with selected aspects 420, such as a specified country, market capitalization range, liquidity minimum, or Global Industry Classification Standard (GICS®) industry category. In some embodiments, available aspects 420 may include a combination of factors, such as a set of countries. In some embodiments, a custom population can be composed of selected companies. Such a population can be created by looking up names or loading a list of identifiers, such as names in a portfolio. In this example, dialog 400 also allows the user to select an index for a benchmark, such as the S&P 500. After receiving a selection defining the population of identifiers, dialog 400 allows the user to select “Create Universe” 430 and then refer to the custom population by the selected name.
[0065] In some embodiments, the rapid screening system allows created populations to be shared with others for collaboration between users, for example, in a shared office environment. Similarly, in some embodiments, the rapid screening system allows custom data tags, expressions, and entire screening sessions to be shared collaboratively. In some embodiments, custom populations of identifiers can be edited or deleted after creation. In some embodiments, to improve the performance of the rapid screening system in creating new populations, data regarding population creation criteria (e.g., for stocks, country, market capitalization, industry, etc.) is stored in a separate database snapshot that is updated periodically (e.g., in-memory data) so that the information is rapidly available to clients without having to be pulled from a database of securities information. In some embodiments, the domain-specific language itself can be used to define the populations.
[0066] FIG. 5A illustrates an exemplary user interface 500 of a rapid screening system configured for equity screening, illustrating domain-specific flexible text matching and completion suggestions according to one embodiment. As described above with reference to FIG. 2, the rapid screening system can flexibly match inputs against domain-specific terms to provide accurate, on-the-fly term matching, even when the input contains errors or is otherwise not an exact match. For example, in the multi-line editor 501, the second line input 520 is "returonequit." The input separates the "n" from the end of "return" (or, by another interpretation, includes an extra "o" in "return" and omits the word "on" entirely) and the "y" from the end of "equity." Nevertheless, the disclosed rapid screening system displays a data tag suggestion pane and explanatory text window 521 at the cursor and highlights matching characters of the likely intended input data tag. The multi-line editor 501 allows the user to select one of the provided options to immediately replace the incomplete and misspelled input. In contrast, conventional systems such as spreadsheet formulas require character-complete text and unintuitive cell number cross-referencing that is plagued by typos, and general-purpose tools such as Excel® spreadsheets cannot provide semantic error handling and domain-specific term completion. By providing flexible text matching and domain-specific completion suggestions, rapid screening systems enable users to screen more easily and rapidly than previously possible.
[0067] FIG. 5B illustrates an exemplary user interface 550 of a rapid screening system configured for stock screening, showing data tag searching according to one embodiment. Data explorer dialog 551 allows a user to find input data tags intended for immediate use. In various domains, the number of available data tags can be overwhelming to a user, reaching thousands or tens of thousands across tens or hundreds of thousands of securities across unlimited data sources. The exemplary data explorer dialog 551 provides fields for company name 555, search text 560, and data package or data provider 565. In the example shown, the user searches for data tags related to the term "EBIT" (earnings before interest and taxes) for Microsoft Corporation, available from a data package titled S&P Global—Fundamental Data. In some embodiments, the user-entered text is assisted with an auto-fill feature to aid in rapid discovery. Matching data tags are listed in results box 570, showing the name, description, and value (for the current quarter and the following 12 months) for each data tag. Other embodiments include other properties or complete example snippets. In some embodiments, the search text can match any of the information about the data tags, including values (e.g., a specific annual growth rate). A convenient user interface element allows the user to select desired data tags by double-clicking 575 a cell in the results box 570 to pin it and using button 580 to copy the selected data tag or tags. For example, double-clicking the 8.163900 cell 571 pins the data tag "Ebit10YrCagrPct," and double-clicking the 8.566600 cell 572 pins the data tag "Ebit10YrCagrPctTtm" in the illustrated example, allowing one or both data tags to be easily copied to the clipboard and pasted into an expression.By providing such domain-specific data discovery tools, the rapid screening system enables users to discover relevant data tags, explore available data tags, and screen data more easily, rapidly, and effectively than previously possible.
[0068] 6A-6B illustrate exemplary user interfaces 600A-600B of a rapid screening system configured for stock screening, illustrating filtering on criteria according to one embodiment. As described above with reference to both FIGS. 1 and 2, the rapid screening system allows a user to filter, limit, reduce, or narrow a set of identifiers according to expressive matching criteria. In the example shown, user interface 600A displays U.S. large capitalization stocks in column 615 of grid view 602. The input 620 in the second multi-line editor 601 is "ReturnOnEquityPctTtm," which is a data tag for each company's return on equity percentage for the trailing 12-month period. The value for "Return On Equity Pct Ttm" is displayed in column 625 in grid view 602.
[0069] Referring to user interface 600B, the multi-line editor 601 input 622 on the second line, here "~ReturnOnEquityPctTtm>40," adds a tilde ("~") filtering operator to the beginning of the line and adds the comparison condition ">40" after the data tag. Thus, grid view 602 no longer displays all identifiers (or their stock ticker symbols) in the specified universe of all U.S. large capitalization stocks. Instead, only the stock tickers in column 617 for companies whose company's return on equity percentage for the trailing 12-month period is greater than 40 are included in the updated grid view 602 in column 627. Because the ReturnOnEquityPctTtm data for all identifiers displayed in column 615 has already been loaded by the system (and displayed in column 625), such a limiting operation can be accomplished very quickly.
[0070] In some embodiments, expression operators such as "~" (or "filter," "restrict," "only," etc.) are required (e.g., to improve user readability). In some embodiments, expression operators may be allowed, but filter operations may be inferred from the presence of comparison operators. When filtering is inferred, or for all filtering operations, the rapid screening system can distinguish lines or sentences (e.g., by decorating sentences or applying text formatting, background color, etc., and / or by inserting explicit operators that are omitted by the user and inferred by the system).
[0071] In some embodiments, the rapid screening system provides additional filtering related operators, such as an "OR" operator, a "NOT" operator, and / or enumeration (all X's in Y), etc.
[0072] FIG. 7 illustrates an exemplary user interface 700 of a rapid screening system configured for stock screening, showing expressions assigned to custom variable names, according to one embodiment. As described above with reference to FIG. 1, the rapid screening system allows users to create new custom variable names or data tags. For example, the rapid screening system allows any data tag to be given a new alias, allowing any expression (e.g., modifying a data value associated with a data tag or combining multiple data values) to be named so that it can be usefully labeled and conveniently referenced again. The name can be applied, for example, to a code snippet or an actual variable (thus not needing to be recalculated). Thus, in conjunction with the ability to define custom populations as described above with reference to FIG. 4, users can extend the domain-specific language to meet their own needs, for example, by defining populations, data tags, data sources, and / or transformations.
[0073] Statement 740 shown in multi-line editor 701 assigns the sum of two data tags, "ReturnOnEquityPct" and "ReturnOnAssets," to a custom named variable or data tag, "MyCustomIndex." Thus, column 745 under the heading "My Custom Index" displays the sum of two company data values, Return on Equity Percentage and Return on Assets, for each stock ticker symbol displayed in grid view 702. In some embodiments, the rapid screening system uses PascalCase (all words in compound names written in uppercase) as a convention for naming data tags for user readability, and the results display interface (e.g., headings in grid view 702) automatically adds spaces in appropriate places in data tags (e.g., between lowercase letters and following uppercase letters) to enhance readability. Other equivalent conventions for compound names in data tags without spaces include camelCase (all words after the first are capitalized), kebab-case (dashes separate words), and snake_case (underscores separate words). In some embodiments, the rapid screening system uses dot notation, e.g., "ReturnOnEquity.TTM," or employs parentheses to name variables. In some embodiments, spaces are allowed within data tags, or the system parses input with spaces, to identify unambiguous matching data tags (or suggest likely options when the input is ambiguous in the context of the domain-specific language and operators). By linking variable names with display output, the technique encourages non-programmers to naturally write maintainable "code," since they can recognize whether headings in the output look correct.
[0074] In some embodiments, the system includes a natural language user interface, such as a character recognition interface or a speech recognition interface. Natural language processing (NLP) often fails because human language is not sufficiently constrained to produce consistently reliable results. The disclosed technology provides a domain-specific language as an intermediate layer, which can significantly improve results. For example, a natural language request such as "Find all companies that mention China tariffs in their 10K filings" can be processed by a language model (e.g., GPT-3) to generate an expression in the domain-specific language. Because the domain-specific language is semantically meaningful, a human can understand the output, verify that the expression is doing what the natural language statement requested, and edit it if necessary. Because expressions in the domain-specific language are highly expressive within a short space, they are easily verifiable by users, providing a more meaningful and easily reliable mode of interaction.
[0075] In contrast to most programming languages, where you declare a variable and then assign a value to it, in the illustrated embodiment, the domain-specific language includes a post-result assignment operator, psychologically encouraging the user to explore various expressions (which are immediately interpreted as the user types or modifies or replaces them) and then assign the results (which are immediately displayed) to the user-named variable. This also ensures that expressions are "completed" faster; for example, the formula or statement "A=awesome(param)" does not complete until the end of the statement, but the expression "param|awesome=>A" completes throughout its structure.
[0076] 8A-8C illustrate example user interfaces 800A, 800B, and 800C of a rapid screening system configured for stock screening, illustrating the simultaneous renaming of multiple references to a custom variable name, according to one embodiment. In the illustrated example, in entry 830 on line 3 of multi-line editor 801, a user defines a custom variable or data tag named "StandrRoa." Additionally, in entry 840 on line 4 of multi-line editor 801, a user enters an expression that references the custom "StandrRoa" variable or data tag. In FIG. 8A, the rapid screening system displays a data tag suggestion pane and explanatory text window 841 at the cursor on line 4, highlighting "StandrRoa" as a recognized user-defined variable. In FIG. 8B, a highlighted selection in a context menu 842 in multi-line editor 801 provides the user with the ability to "change all occurrences" of the custom variable name. In FIG. 8C, multi-line editor 801 allows the user to simultaneously change "StandrRoa" to "StandardRoa" in multiple places. Thus, the fast screening system allows for consistent global changes to variable names without breaking any of the variable references. In contrast, traditional in-domain search systems do not provide the ability to reorganize text input.
[0077] FIG. 9 shows an example user interface 900 of a rapid screening system configured for stock screening, illustrating domain-specific syntax error handling according to one embodiment. As described above with reference to both FIGS. 2 and 5, the rapid screening system is error-tolerant when parsing input. For example, if the input is incomplete, the rapid screening system may attempt to infer the intended match in the domain-specific language and, for example, replace the unrecognized input with the closest valid match or suggest a set of potential replacements or completions for the unrecognized input. Additionally, malformed input (e.g., a syntax error such as a symbol that does not match any operator or data tag in the domain-specific language) does not crash the rapid screening system or halt processing. In various embodiments, the rapid screening system can continue parsing the remainder of the input (both before and after the error), process parser-recognized symbols, and display previously successful results.
[0078] In the illustrated example 900, the user enters a typo or incomplete input 920, "ReturnOn," on line 2 of the multi-line editor 901 after previously entering, for example, the valid data tag "ReturnOnEquityPct." The rapid screening system displays a data tag suggestion pane and an explanatory text window 921 at the cursor and highlights matching characters of the likely intended input data tag. Simultaneously, the multi-line editor 901 displays a text warning 922 indicating "Error: invalid syntax on line 2" (or, for example, "Error: no data tag present on line 2"), margin text decoration 923, and a red zigzag underline that highlights the location of the unrecognized and unprocessed input 920. In some embodiments, error handling includes applying domain-specific information so that the error message is domain-specific.
[0079] Meanwhile, the remainder of the grid view 902 is unaffected. The population of US Large Capitalization Stocks 915 remains displayed, along with the corresponding data values for the data tags and displayed identifiers in the remainder of the multi-line editor 901. In some embodiments, previously displayed data (e.g., "ROE Pct" column 925) remains shown until input that the rapid screening system can process is entered in its place. In some embodiments, data is not displayed for symbols that the rapid screening system cannot process.
[0080] 10 illustrates an exemplary user interface 1000 of a rapid screening system configured for stock screening, showing transformation functions according to one embodiment. In the illustrated example, a pipe symbol ("|") in the second line of input 1020 of the multi-line editor 1001 indicates a transformation function or operation. In this example, five transformation functions are listed in the operator suggestion pane and explanatory text window 1021: average (when applied to an array of numbers, produces, for example, the arithmetic mean), Rank (compares the data values of each identifier and ranks them either high to low or low to high); Quintiles (compares the data values for each identifier and categorizes them into five buckets of 20% each to generate a quintile number), Normalization (comparing the data value of each identifier to a normalized normal distribution and generating a z-score), and Trend stability (which when applied to an array of numbers produces a number indicating whether the trend is positive, negative, or neutral) is listed.
[0081] In various embodiments, the rapid screening system may provide additional or different transformation functions (e.g., median function, decile function, etc.). For example, a set of natural language processing transformation functions may allow a user to input an expression such as "NewsRecent["lawsuit"]|Sentiment=>LawsuitNewsSentiment" to evaluate the tone of news coverage about a company involved in litigation.
[0082] FIG. 11 illustrates an exemplary user interface 1100 of a rapid screening system configured for stock screening, showing an automated graphical display of array data according to one embodiment. In the illustrated example, the input 1120 on the second line of the multi-line editor 1101 is “ReturnOnEquityPct[-11q:0q],” which is a data tag representing each company’s return on equity percentage for each of the past 12 quarters (i.e., from 11 quarters ago to the present). In some embodiments, an expression involving a data tag associated with an array of values does not require explicit array notation, except to display a selected subset of the array data. For example, two data tags representing array values can be added together or otherwise manipulated in an expression. In some embodiments, an alternative or shorthand notation such as “ReturnOnEquityPct[12q]” can concisely indicate the number of periods (e.g., years, quarters, months, weeks, or days) to date. In some embodiments, the system automatically aligns the array data. For example, if a data tag referencing time series A is separated by a data tag referencing time series B, the system can automatically align the current value by forward filling.
[0083] In this example, the disclosed rapid screening system displays the sequence data as user-friendly compact graphs 1125, one per identifier (a stock ticker symbol is actually displayed). The compact graphs generated by the rapid screening system include shading above or below the horizontal axis for each data value in the result sequence. Thus, the rapid screening system visually reveals positive and negative values, making it easy for users to identify trends over time. Additionally, user actions on one of the compact graphs 1125, such as hovering, right-clicking, or long-pressing over the graph or a data point on the graph, may reveal one or more of the underlying data values 1127 in the result sequence for the identifier. In some embodiments, the compact graphs 1125 include or display key information, such as boundaries or x-axis labels (e.g., dates).
[0084] In some embodiments, the rapid screening system provides similar functionality or extensibility for user-defined functions and data tags. For example, data tags or functions are extensible, allowing programmers to write custom code underlying the data tag or function (e.g., for custom visualization of complex data) and make it available to non-programmers who use the data tag or function. For example, in some embodiments, naming a custom variable or data tag ending in "Score" can automatically format and color-code associated data values according to quartiles, deciles, ranks, etc. Similarly, in some embodiments, naming a custom variable or data tag ending in "Trend" can automatically prompt the system to display any associated sequence data in a graph, and / or naming a custom variable or data tag ending in "Pie" can automatically prompt the system to display any associated sequence data in a pie chart. In some embodiments, the system can be configured to graph any sequence data. In various embodiments, formatting and / or color-coding of results can be achieved via any interface, even interactively. For example, in some embodiments, columns may be grouped in the display to show group headings above a collection of subheadings.
[0085] FIG. 12 illustrates an exemplary user interface 1200 of a rapid screening system configured for stock screening, showing the automatic display of links to 10-K filings according to one embodiment.
[0086] Similar to FIG. 11 , the rapid screening system interface grid view 1202 includes a column 1225 displaying a compact graph representing an array of data values corresponding to the data tag in the input 1220 in the second row of the multi-row editor 1201. In this case, the input 1220 is "ReturnOnEquityPct[-3q:0q]," a data tag representing each company's return on equity percentage for each of the past four quarters (i.e., from three quarters ago to the present). The same input 1220 generates a "trend stability" transformation of the array data values, providing a numerical indicator 1245 of the trend in each case. As shown, the fourth row of input 1240 limits the displayed identifiers to those with a strongly positive (greater than 0.8) return on equity trend over the past four quarters, so the trends are all upward. The third row is blank, demonstrating the ability of the multi-row editor 1201 in the illustrated embodiment to seamlessly parse discrete inputs.
[0087] Another way to apply the disclosed techniques to identify positive stability trends is to compare stability trends over two time periods. For example, the following expression establishes a custom data tag named "RoeStabilityPrevious" that represents the trend stability for return on equity percentage from one year ago to four months ago, and a custom data tag named "RoeStabilityRecent" that represents the trend stability for return on equity percentage from four months ago to today: (ReturnOnEquityPct[-11q:-4q])|TrendStability=>RoeStabilityPrevious (ReturnOnEquityPct[-3q:0q])|TrendStability=>RoeStabilityRecent
[0088] By subtracting one trend from the other, the difference can be determined and filtered. (RoeStabilityRecent-RoeStabilityPrevious)=>RoeTrendDifference ~RoeTrendDifference>1
[0089] Or alternatively, the absolute values of previous and current trends can be filtered to identify companies whose ROE has recently spiked, for example. ~RoeStabilityRecent>0.6 ~RoeStabilityPrevious<0.6
[0090] The fifth line of input 1250 in multi-line editor 1201 adds the data tag "Filings10k" to grid view 1202, displaying information about each company's most recent Form 10-K filing with the Securities and Exchange Commission (SEC). In the illustrated embodiment, the 10-K filing information includes the date of the most recent available filing and a hyperlink to a copy of the filing online at the SEC so that a screener can read or download the document directly. Thus, a rapid screening system can be configured to handle complex data types, e.g., JSON-formatted data and metadata, and thereby display results in a user-accessible format to present a greater amount of useful information to the screener than previously possible.
[0091] The sixth line of input 1260 in multi-line editor 1201 adds the data tag "GicsSector" to grid view 1202, displaying the GICS sector for each of the displayed identifiers (stock ticker symbols).
[0092] FIG. 13 illustrates an exemplary user interface 1300 of a rapid screening system configured for stock screening, showing a selective display of companies that hold patents, according to one embodiment. In the multi-line editor 1301, an input 1330 on line 3 and an input 1340 on line 4 use the data tag “PatentsIssued” to identify companies that have been awarded one or more patents. The rapid screening system can, for example, search a database of issued patents to identify patents whose owner or assignee matches the name of the company (in some implementations, this corresponds to a loose match and / or related entities). In the illustrated example, in input 1330, the expression “PatentsIssued[“autonomous vehicle”, “autonomous driving”]” is interpreted to identify companies that have been awarded one or more patents (assigned the data tag AutonomousDrivingPatents) that contain either the phrase “autonomous vehicle” or the phrase “autonomous driving.” Similarly, the expression "PatentsIssued["electric vehicle"]" in input 1340 is interpreted by the rapid screening system to identify companies that have been issued one or more patents containing the phrase "electric vehicle" (which are assigned the data tag ElectricVehiclePatents).
[0093] In grid view 1302, column 1335 displays autonomous driving patents, and column 1345 displays electric vehicle patents, showing patent titles and issue dates, and providing links to the patent documents. In some embodiments, the screening system provides a total count of patents (or currently active patents) found by the search. For example, the electric vehicle pure play patent score can be expressed as "PatentsIssuedCount["Electric Vehicles"] / PatentsIssuedCount" or a similar expression. In the illustrated example, input 1370 on line 7 of multi-line editor 1301 indicates filtering using the expression "~AutonomousDrivingPatents>0," and input 1380 on line 8 filters using the expression "~ElectricVehiclePatents>0" so that the only companies included in grid view 1302 (listed in company ticker column 1315) are companies with at least one patent in each category. In other words, the listed company holds both the autonomous driving patent 1375 and the electric vehicle patent 1385 (although not necessarily the autonomous electric vehicle patent).
[0094] In various embodiments, the rapid screening system can be configured to obtain, link, and filter content from a wide range of data sources. For example, a “NewsRecent” data tag can cause the grid view 1202 to display recent news headlines about each company from various news sources. Similar to the Form 10-K links in FIG. 12 and the patent titles in FIG. 13, each news headline for a given identifier links to the entire article, while effectively providing summary information directly in the search results. In some embodiments, the system allows users to filter recent news articles containing terms of interest. For example, an expression such as “NewsRecent[“autonomous driving”]” can allow users to identify companies with few published patents in a field compared to media engagement in that field, or vice versa. In this way, the rapid screening system allows screeners to directly observe a company's technological development activity or read news about targets of interest and get a subjective impression of coverage related to such targets to complement numerical data analysis or to identify breaking news that may move the market related to certain securities.
[0095] FIG. 14 illustrates an exemplary user interface 1400 of a rapid screening system configured for stock screening, illustrating filtering on text found in 10-K filings, according to one embodiment. Similar to FIG. 12 , the rapid screening system of FIG. 14 includes a compact trend graph and the data tag “Filings10k.” However, rather than providing links to Form 10-K filings, FIG. 14 uses that data tag to filter identifiers based on the content of those 10-K filings. In particular, the fifth line of input 1450 in the multi-line editor 1401 in the illustrated example indicates filtering by the expression “~Filings10k contains “china tariffs”.” Thus, the list of companies 1415 in the results grid view 1402 is limited based on multiple criteria: a large U.S. company 1410 with significant growth in ROE over the past four quarters 1420, 1425, 1440, 1445, and whose most recent Form 10-K filing mentioned China 1450 and tariffs 1455.
[0096] In grid view 1402, the "10k Filings" column 1455 displays relevant matching text from the 10-K, with matching search terms highlighted, rather than simply providing a link to the entire Form 10-K document. In some embodiments, the highlighted text provides a link to the source document, or more specifically, to the cited portion of the source document.
[0097] As noted above, in some embodiments, operators such as "contains" are implemented without alphabetic text. For example, using bracket notation, the syntax could be "Filings10k["china tariffs"]" as an equivalent example, similar to the "PatentsIssued["electric vehicle"]" example in FIG. 13. In some embodiments, the "contains" operator functions as a transformation, e.g., "Filings10k|contains "china tariffs"," making the step explicit and enabling filtering on results, although the filtering itself can be implicit. By providing the ability to filter identifiers based on the content of documents such as Form 10-K filings or patents, the disclosed rapid screening system enables a powerful new approach to screening and provides the ability to synthesize across heterogeneous data sets.
[0098] FIG. 15 illustrates an exemplary user interface 1500 of a rapid screening system configured for stock screening, illustrating result grouping according to one embodiment. Similar to FIG. 12, the rapid screening system of FIG. 15 includes the data tag "GicsSector" in the sixth line of input 1560 of multi-line editor 1501. Thus, the results grid view 1502 displays the GICS sector for each of the displayed identifiers (stock ticker symbols). However, in FIG. 15, the "Gics Sector" column 1565 heading is used to group the identifiers according to their categorical data values.
[0099] In the illustrated embodiment, a "GICS Sector" heading is also displayed 1506 as a set row group at the top of the grid view 1502. Thus, the left side of the grid view 1502 lists a series of GICS sector categories 1566. Each category can be expanded to show identifiers of companies in that sector, or contracted to show the number of companies in that sector that meet the active population filtering criteria. In this example, the criterion of recent strong ROE growth among large U.S. companies creates a list of primarily information technology stocks. These interface features of the disclosed rapid screening system enable a screener to identify trends, explore strategies for identifying targets of interest, and easily see the impact of trying alternative strategies.
[0100] As demonstrated in each of the above descriptions, the flexible customizability, ease of use, fast feedback, and power of the disclosed rapid screening system enable, for example, complex and expressive screening ("all companies with increasing and stable return on equity (RoE) over the past three years"), unique perspectives ("all companies that mention China tariffs in their most recent 10-K"), identifying discontinuities or dissonances in data ("all companies with increasing stock sentiment but stable stock prices"), discovering aggregation trends ("sectors with increasing sentiment in earnings are called Q&A"), and benchmarking ("how does MSFT perform against all the screens we've built?"). Notably, the screens are configurable, allowing users to reference one screen from another or "unit test" a company across multiple screens. This facilitates building a mosaic of perspectives. The disclosed system and method provides flexible, structured perspectives not previously available.
[0101] 16A-16B show an example formulation 1600A for a prior art system and a corresponding example expression 1600B for a high-speed screening system configured for stock screening, illustrating improved usability according to one embodiment.
[0102] 17A-17B show exemplary user interfaces 1700A-1700B of a rapid screening system configured for stock screening, illustrating backtesting according to one embodiment. In FIG. 17A, a highlighted selection in a context menu 1772 in a multi-line editor 1701 provides a user with the ability to "backtest" the performance of a set of criteria against past market results.
[0103] Backtesting refers to testing a model, such as a set of screening criteria, against historical data to identify stocks worth buying at any given time. By applying the same criteria to historical data, the screener can determine whether the strategy would have performed at another time when different stock ticker identifiers might have met the data tag criteria being tested. Backtesting allows a user to determine how well a set of criteria would have performed if used consistently over a historical period. The backtester creates virtual positions according to the user-selected criteria, runs the user's strategy over time, and records the results. Prior to the present disclosure, backtesting was typically expensive, took hours or weeks, and was limited to narrow data sets. The disclosed rapid screening system technology takes approximately a fraction of a second to 5-10 seconds to generate backtest results 1700B and can backtest strategies expressed using the full vocabulary of a domain-specific language without requiring significant configuration or use of separate tools. An exemplary concurrent server interaction for backtesting is described in more detail below with reference to FIG. 28. In some embodiments, backtesting is performed using the same subset of code that the system uses for screening, which is made possible by domain-specific context.
[0104] FIG. 17B shows a collection of backtest results 1700B for the criteria shown in the multi-line editor 1701 of FIG. 17A. The backtests shown are equally weighted. In some embodiments, the rapid screening system is configured to optionally run factor-weighted backtests (rather than simply equal weighting), such as by selecting (e.g., right-clicking) the factors to weight and then selecting Backtest. In various embodiments, the backtest results 1700B include report generation, including various approaches to charting and listing the results of the backtest. For example, the backtests can provide statistics on how selected investments according to selected screening criteria would have performed compared to one or more benchmarks, over a selected period, and / or considering returns over time. The disclosed technology provides unprecedented accessibility to concurrent programming execution, which was simply unavailable to users with conventional technology.
[0105] FIG. 18 illustrates an exemplary change alert graph 1800 of a rapid screening system configured for stock screening, according to one embodiment. In some embodiments, the rapid screening system may be configured to perform automated daily screening and send results to users, for example, via email. For example, if there are any significant changes (e.g., companies added or removed from a list, new sources of risk, related news articles, etc.), the rapid screening system can automatically highlight them to save users time and reduce availability bias. In the illustrated graph 1800, a collection of stock screening results is shown with identifiers listed vertically and a numeric horizontal axis. The change alert graph 1800 provides a visualization of changes between yesterday's (blue circle) and today's (orange circle) values, allowing viewers to instantly see what has changed and by how much.
[0106] In some embodiments, the technology provides users with the ability to generate similarity graphs comparing any two sets of screening results.
[0107] 19A illustrates an exemplary AI function 1900 for introspection in a high-speed screening system configured for stock screening, according to one embodiment. In the illustrated example, the disclosed technology uses artificial intelligence (AI) machine learning (ML) to introspect and improve existing screening practices.
[0108] First, factors currently used in existing screening models (e.g., to rank potential investments) are identified. For example, a hypothetical screening model expressed in the domain-specific language of the disclosed rapid screening system might be (0.4 * ReturnOnEquity+0.4 * ReturnOnAssets+0.2 * DebtToEquity) => CustomFactor, and your existing screener sorts potential investments by CustomFactor from high to low, prioritizing those queries by what is most important to them).
[0109] To train a black-box AI model to replicate existing practice, we ask the question, "Given the inputs, ROE, ROA, and DTE, and the corresponding input data and results, can we build a model that replicates those screening results?" We then train many ML models, using the same sample inputs as the existing model and the existing and predicted ranking outputs to verify that the model reflects current practice. After training a model, we introspect the model, i.e., look inside the black box, to determine the relative feature importance for the model. Introspection may determine, for example, that a factor is redundant or unexpectedly heavily weighted. For example, a user may believe that ROE and ROA were weighted equally in the screening process, but if it were possible to build a model that predicted the user's intended model output by primarily using ROA, AI function 1900 can help the user understand how their screening process actually works and improve it or offer alternative ways to achieve the same results.
[0110] One of the most challenging parts of training an ML model is the feature selection process. Typically, the expert users who select the features are different from the data scientists or quantitative researchers who are implementing the model. The disclosed technology enables the inference of a screener's feature selection parameters from a user's use of a domain-specific language when constructing an existing screen. Analyzing the use of such a domain-specific language is not possible with existing screeners because other similarly accessible screeners are not expressive enough to say what the user is looking for, and it is not possible with existing languages that are not sufficiently structured or constrained to allow semantic meaning to be inferred from the user's use of the tool.
[0111] Another challenge is overfitting, especially when dealing with financial data. In introspection use cases, it is not important that the generated ML model is overfitted because we are not actually trying to predict future events, but only analyze existing events. By leveraging the properties of overfitted models, the disclosed technology goes against conventional practice and teaching. Overfitting is typically considered undesirable and highly counterintuitive to those skilled in the art.
[0112] FIG. 19B illustrates an exemplary AI function 1950 for prediction in a high-speed screening system configured for stock screening, according to one embodiment. This is a different approach to improving AI screening than the introspection ML model described above with respect to FIG. 19A. The approach in this case is to use machine learning to train a set of classifiers that take as input the universe of companies at a given time and a selected subset of example companies. The example companies could be a list of individually selected individual identifiers or those that meet some filtering criteria (e.g., ROE above a threshold, or company names containing "hotel"). The example companies in the subset could be "preferred" targets, companies with characteristics the screener wants to avoid, or some other category. Approaches to developing ML models could include bootstrapping, random forests, AutoML, or other ML models, preferably models that allow the trainer to infer relative feature weights. The trainer can exercise some control over the resulting screening model, for example, by training and averaging more models. Classifiers are trained to generate example companies, all of which achieve the same solution but in different ways. They can therefore predict the filters in which an example company would have been found. This can reveal user preferences by showing which data tags and weights are correlated or uncorrelated. In some implementations, the result is the generation of probability distributions and feature weights.
[0113] Unlike typical recommendation engines that are trained on one subset of a given set of similarities and operate on another presumed similar subset, this training model creates a layer of abstraction, describing a pocket that a user may believe is well-known (e.g., a selected company in a selected industry). In addition to helping to quantify it, the model allows the user to apply their understanding of that pocket to other pockets or populations, such as other industries, countries, and / or time periods.
[0114] These classifiers can be used individually or as part of an ensemble. Ensembles allow us to reduce overfitting to the prediction context, and some models allow us to further introspect on specific weights. Other implementations include combining classifiers generated from different "runs" of this same algorithm. The resulting prediction module (combination of classifiers) can be applied in several different ways, or the user can be allowed to machine-tune or iterate on feature weights, for example.
[0115] 20 illustrates an exemplary AI function 2000 for regime change detection in a high-speed screening system configured for stock screening, according to one embodiment. In finance, regime changes are associated with abrupt changes in financial market behavior, such as may be associated with cycles of economic activity that change between expansion and contraction. Understanding regime changes is important because it allows investment managers to react to systematic market changes.
[0116] Regime change models are typically defined as time series models whose parameters can take on different values in each of a fixed number of "regimes." Regime change models typically include some set of predefined factors that they look for and some fixed definitions of regimes. The key idea is usually to define "what a regime is" and "what are the functions that affect the regime." However, two problems with current approaches are:
[0117] 1. Models can tell you about possible regime changes given our understanding of historical and predefined regimes, but the key question for investors to answer is not "Is the world changing?" but "Is the world changing in ways that matter to this investment strategy?"
[0118] 2. They require input variables to be defined in advance and do not adapt to changes in the forces affecting the market.
[0119] To address custom regime change detection, we can continue to build on the AI / ML approach described above with reference to FIG. 19A. Based on the AI in FIG. 19A, we infer and build custom regime change models without explicit instructions from the investment manager. We perform the same procedure, but this time, instead of looking at one snapshot time window, we can observe many windows in sequence. Thus, we build a set of ML models for each period (e.g., quarter).
[0120] Again, it is not important that the ML model we build is overfitted. The output is the impact of the defined factors / features over time. If there is a change in the impact of a feature that is important to your investment process, it is an early warning signal that the factors you rely on may have shifted, and therefore you may want to consider adapting your approach.
[0121] The ability to infer regime change models from this little information currently does not exist. The combination of language expressiveness and semantic structure / constraints, as well as the application of ML, helps make that ability more powerful. The disclosed technology allows domain expert decision makers to combine it with their own judgment and iterate with the AI to improve the results.
[0122] Referring back to FIG. 19B, a predictive AI model can be applied to use intuition about known regimes of the past to infer similar choices to make in the current regime. For example, if a user believes that today's regime is similar to the regime a company experienced in 1990, the model can be used to highlight companies similar to a selected subset of companies from 1990 (such as companies that grew in that regime). That is, a model trained to predict targets of interest in 1990 can now be applied to today. Thus, with an input of a set of names and dates, the output would be a set of similar names with their associated scores in the current and screening sets.
[0123] As another example, the current state of the world is generally evaluated with respect to whether factors such as "value," "growth," "quality," and "momentum" favor stocks currently performing better. However, a user may be unsure, for example, whether a quality category actually applies to that user's "pocket" (e.g., a limited subset of securities) of the world. For example, finding companies that performed well in 2007 and / or 2003 requires too much data to hold in one's head. Thus, traditional investors must rely on intuition. An AI model can essentially codify that intuition, and where the idea of "this feels like a past time" is not concrete, the model makes "now feels like the time" concrete and quantifiable.
[0124] Additionally, this AI model approach can be applied across industries or countries, as well as across dates or asset classes (e.g., the US market in 2003 vs. China today), enabling transfer learning and generating insights into cross-connections that may otherwise remain undiscovered.
[0125] The disclosed AI model training does not just use the past to predict the future via an unadjustable black box; it adds user input of experience-based intuition about the past, quantifies regimes (defined on the fly, not pre-classified), defines a forecasting model based on that, and provides a rationale for identifying characteristics that performed well then and, if the user's intuition is correct, can perform equally well now. Rather than attempting to tell which pre-defined regime we are in today without providing actionable insights (e.g., which companies we might consider buying given a similar climate), the model provides recommendations of companies today that are similar to companies that would have been of interest in the past in past climates (e.g., based on their track record / revenue or other qualities).
[0126] In some embodiments, relevant "features" for training the model can be inferred from the code specified in the editor (e.g., by analyzing data tags, including user-defined data tags and expressions). Typically, a data science or machine learning expert performs this functionality engineering. The disclosed technology enables non-technical users unfamiliar with machine learning to effectively collaborate with machine learning AI screening models.
[0127] This facilitates a continuous feedback loop between the user and the AI screening model, allowing the user to effectively perform feature engineering to improve the AI screening model by modifying the text in the multi-line editor. In this way, the user can effectively iterate on the specified features fed into the model to better find targets of interest.
[0128] FIG. 21 illustrates an exemplary AI function 2100 for optimizing blending in a high-speed screening system configured for stock screening, according to one embodiment.
[0129] Based on the approach disclosed above in connection with Figures 19A and 20, we use these introspection and / or regime change insights to optimize future screening models. If ROE outperforms, i.e., if the screener believes ROE and ROA are equally important, but the ML model infers that ROE is more likely to identify more companies of interest, we can highlight companies in the generated list that we may want to focus more on.
[0130] Similarly, based on the approach disclosed above in connection with Figures 19B and 20, a screen that models the high-performing securities in a portfolio can be applied to identify additional screening criteria and / or weightings by finding a set of screens to generate names for those securities.
[0131] In a variation of the above approach, predictive AI models can also be applied to obtain herding and / or minimum consensus risk assessments for names in a screen. Given screen and date inputs, the output is two lists of identifiers: names that hit the fewest screens of other screens, and names that hit the most screens of other screens. The first represents the minimum consensus risk. The second represents the herding risk.
[0132] So, for example, if you generate 100 screens that help you find that set of companies (i.e., if there are 100 ways to arrive at the same solution), and 75 of those screens agree that MSFT is a good pick, the takeaway depends on your confidence that MSFT is the kind of company you want to buy. Is it representative of what you value? Or do you wonder if it reflects how independent your perspective is? That is, if the screens reveal that a large portion of what you do is what traditional growth investors do, and a small portion is your "secret sauce," then you have the ability to introspect and act on it to adjust your biases, your perspective, and / or your actions.
[0133] In another variation, predictive AI models can be used to identify near misses, e.g., ticker names that did not hit any names in the user's original screen or portfolio, but did hit many of the AI-learned screens. Such companies may warrant action to investigate.
[0134] FIG. 22 illustrates an exemplary AI function 2200 for feature suggestions in a high-speed screening system configured for stock screening, according to one embodiment.
[0135] Based on the approach disclosed above in connection with Figures 19A, 20, and 21, we can suggest features (e.g., data tags) to add to and features to remove from the set of screening criteria. For example, we recommend adding another feature that is simply random noise, setting it as a baseline for comparing the impact of other factors on the existing screening criteria, and removing features below that threshold. For example, if DTE has less impact than random noise in the shadow model, you might think that a signal is embedded, but in fact it is not. Because we have a strong semantic understanding of the goals users are trying to achieve, we can statically and / or dynamically analyze them to make suggestions to improve their strategy performance.
[0136] The disclosed AI model and interface provides a simple yet powerful level of abstraction because user input is domain-specific and has more context about the type of problem the user is trying to solve. User input allows for inference of user preferences within the domain. The interface is non-technical and conversational, encouraging users to iterate based on the results and facilitating "feature" engineering.
[0137] Additionally, these AI models contrast with regular quantitative modeling, where goals are typically very rigidly defined. The disclosed models allow us to target qualitative and / or hard-to-state goals, which makes them feasible and useful for more experienced discretionary investors.
[0138] FIG. 23 shows an example user interface of a rapid screening system configured for stock screening, illustrating the creative use of operators according to one embodiment. In a multi-line editor 2301, an input 2330 on line 3 applies the transformation "RankHighToLow" to US Large Capitalization Companies 2310 based on their current return on assets and assigns the ranking metadata to the variable "RoaRank" 2335. An expression 2360 on line 6 incorporates ternary operators, emojis, and string concatenation to generate an easy-to-read Roa Indicator column 2365. The ternary operator functions like a short if-then statement of the form "(X?Y:Z)," where if expression X is true, the output of the ternary operator is Y; otherwise, the output is Z. Thus, for any company with an RoA rank of 75 or greater, the result is a green checkbox emoji character; for any company with an RoA rank of 76 or less, the result is a red X mark emoji character. The output pictogram characters are concatenated with a space character and a return on assets value to display the Roa Indicator column 2365, which highlights stocks with the highest ROA among US large capitalization stocks.
[0139] 24 shows an exemplary user interface 2400 of a rapid screening system configured for stock screening, illustrating automatic formatting according to one embodiment. In the example shown, applied to a population of U.S. large capitalization stocks 2410, three calculations determine a "Value" factor 2450 / 2455, a "Growth" factor 2460 / 2465, and a "Quality" factor 2470 / 2475. In lines 10, 11, and 12, each of the three factors has a "SplitQuintiles" transformation applied (2451, 2461, 2471) to generate metadata that assigns a quintile scale of 1 to 5 to the factor score for each identifier.
[0140] The quintile scores for each factor are assigned to variables or custom data tags whose names end in "Score." As a result, the rapid screening system specifically processes the values associated with these variables by displaying them in a grid view 2402 using "heat map" color coding. As shown, "1" quintile scores are shown in cells colored dark red, "2" quintile scores are shown in cells colored dark orange, "3" quintile scores are shown in cells colored yellow, "4" quintile scores are shown in cells colored light green, and "5" quintile scores are shown in cells colored dark green.
[0141] This and other automated formatting of displayed values offers a contrast to conventional screening systems, providing greater flexibility and expressiveness while making it easier for untrained users to understand the displayed results.
[0142] 25 illustrates an exemplary user interface 2500 of a rapid screening system configured for stock screening showing point-in-time status reporting of forecasts according to one embodiment. The time series status reporting model provides not only research and screening capabilities, but also insight and powerful analysis that leverages analysts' forecasts over time and compares them to a company's actual performance.
[0143] In the illustrated example, in the first row 2510 of the multi-row editor interface 2501, the notation "# / Model / SituationRepor" invokes the situation report mode or model (many other equivalent approaches, such as a button, drop-down menu, or voice command, could also be used). Some controls are not shown, including an input field for entering a list of companies or population identifier (in this case, the five companies shown in the "Company Name" column 2515) and a date field with a calendar drop-down. The multi-row editor interface 2501 includes a collection of data tags representing analyst expectations in several areas, such as "EbitConsensusMean" 2530, which reflects the consensus average of analysts' forecasts of a company's earnings before profits and taxes.
[0144] The contents of each cell in the status report model are not simple values; each is a summary of a complex set of data. To generate a summary status report, the rapid screening system spawns a computing process for each cell in the grid display 2502 to download and process the underlying information. The data is interpolated, smoothed, and then displayed. The displayed value represents the current (for the selected calendar date) forecast change rate for the selected metric for the selected company. For example, in the example shown, for Bed Bath & Beyond Inc. 2531, the Ebit consensus average 2535 value is 53 (2545). This indicates a slight positive slope for the consensus average, implying that forecasts will remain flat or increase very slightly for EBIT.
[0145] In the example shown, cells are colored along a red-yellow-green spectrum (based on standard deviation to match the human eye / expectations) based on the curve slope of the smoothed forecast. Thus, values between 0 and 10 correspond to the fastest decline, and values from 90 to 100 correspond to the fastest increase relative to the maximum absolute slope. The EBIT consensus average of 2535 for Microsoft Corporation 2532 is very high at 95 (2546), and is therefore shaded a bright green, indicating that the consensus forecast is for Microsoft's EBIT to increase nearly as rapidly. At the date of this disclosure, the computational complexity of the analyses shown makes them impossible to perform on a personal computer.
[0146] In addition to displaying current trends, the situation report provides controls that allow navigation, synthesis, and contextual awareness of analyst forecasts that were previously unavailable. For example, by adjusting the calendar date, users can easily run forecasts at earlier times and compare how the state of the consensus forecast has changed or is changing. Additionally, green arrow 2580 or red arrow 2585 indicate that a company's actual forecast is exceeding or falling short of the expected consensus forecast, which may signal a change. Additionally, each cell is a link to a larger graph that displays more complete information over time, as further described below with reference to FIG. 26.
[0147] FIG. 26 illustrates an exemplary user interface 2600 of a rapid screening system configured for stock screening, showing a status report historical forecast graph 2601, according to one embodiment. This graph 2601 is displayed when a user selects cell 2545 from FIG. 25, and details are provided after the numbers displayed in the initial status report table. Graph 2601 includes a stepped line 2610 representing the actual consensus forecast for Bed, Bath, & Beyond's EBIT over the displayed time frame. In addition, a smoothed line 2620 is shown, which is also the source of the slope calculation described above with reference to FIG. 25. Around or near the actual and smoothed forecasts is a shaded region 2640 representing the forecast interval of the analyst's forecast. When the actual forecast 2610 falls outside the forecast interval 2640, the status report table (shown in FIG. 25) marks the excursion, which may indicate a change from a previously determined forecast. Legend 2650 displays the exact value for a given date.
[0148] The situation report graph applies to any time series data, not just estimates. When applied to estimates in the example shown, it highlights large changes in analyst expectations and consistent movements beyond the forecast. For example, a dy (smoothed) graph runs along the bottom half of the chart, showing the rate of change in forecast. Thus, in the example shown, analyst EBIT expectations rose 2630 around October 2018 and fell 2635 around November 2019. This makes it easier to identify inflection points in analyst expectations without requiring coding or other special technical expertise from the user. This visualization of consensus forecasts reduces the learning curve and time to insight by adding companies to the list, allowing users to obtain immediate results. Thus, the disclosed technology can easily alert users, potentially even before the market reacts. Other embodiments allow users to screen this information in the same manner as shown and described above.
[0149] 27 is a block diagram illustrating some of the components typically incorporated in computing systems and other devices in which the present technology can be implemented. In the illustrated embodiment, computer system 2700 includes a processing component 2730 that controls the operation of computer system 2700 according to computer-readable instructions stored in memory 2740. Processing component 2730 may be any logic processing unit, such as, for example, one or more central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc. Processing component 2730 may be a single processing unit or multiple processing units in an electronic device, or may be distributed across multiple devices. Aspects of the present technology may be embodied within a special-purpose computing device or data processor specifically programmed, configured, or constructed to execute one or more of the computer-executable instructions described in detail herein.
[0150] Aspects of the present technology may also be practiced in distributed computing environments where functions or modules are performed by remote processing devices that are linked through a communications network, such as a local area network (LAN), a wide area network (WAN), or the Internet. In a distributed computing environment, modules may be located in both local and remote memory storage devices. In various embodiments, computer system 2700 may comprise one or more physical and / or logical devices that collectively provide the functionality described herein. In some embodiments, computer system 2700 may comprise one or more replicated and / or distributed physical or logical devices. In some embodiments, computer system 2700 may comprise one or more computing resources provided by a “cloud computing” provider, such as Amazon® Elastic Compute Cloud (“Amazon EC2®”), Amazon Web Services® (“AWS®”), and / or Amazon Simple Storage Service™ (“Amazon S3™”) offered by Amazon.com, Inc. of Seattle, Washington; Google Cloud Platform™ and / or Google Cloud Storage™ offered by Google Inc. of Mountain View, California; Windows Azure™ offered by Microsoft Corporation of Redmond, Washington; etc.
[0151] Processing component 2730 is connected to memory 2740, which can include a combination of temporary and / or persistent storage, and both read-only memory (ROM) and writeable memory (e.g., random access memory or RAM, CPU registers, and on-chip cache memory), writeable non-volatile memory such as flash memory or other solid-state memory, hard drives, removable media, magnetically or optically readable disks and / or tape, nanotechnology memory, synthetic biological memory, etc. Memory is not a propagating signal separate from the underlying hardware, and thus memory and computer-readable storage media do not refer to a transitory propagating signal per se. Memory 2740 includes data storage containing programs, software, and information, such as operating system 2742, application programs 2744, and data 2746. Operating system 2742 of computer system 2700 can include, for example, Windows®, Linux®, Android, iOS®, and / or an embedded real-time operating system. Application programs 2744 and data 2746 may include software and databases, including data structures, database records, other data tables, etc., configured to control computer system 2700 components, process information (e.g., to optimize program code data), and communicate and exchange data and information with remote computers and other devices.
[0152] The computer system 2700 can include an input component 2710 that receives input from user interaction and provides input to the processor 2730, typically mediated by a hardware controller that interprets raw signals received from input devices and communicates the information to the processor 2730 using known communication protocols. Examples of the input component 2710 include a keyboard 2712 (with physical or virtual keys), a pointing device (such as a mouse 2714, joystick, dial, or eye-tracking device), a touchscreen 2715 that detects contact events when touched by a user, a microphone 2716 that receives audio input, and a camera 2718 for still photo and / or video capture. The computer system 2700 can also include various other input components 2710, such as a GPS or other position-determining sensor, a motion sensor, a wearable input device (e.g., a wearable glove-type input device) with an accelerometer, a biometric sensor (e.g., a fingerprint sensor), an optical sensor (e.g., an infrared sensor), a card reader (e.g., a magnetic stripe reader or a memory card reader), etc.
[0153] The processor 2730 may also be connected, for example, directly or through a hardware controller, to one or more various output components 2720. The output device may include a display 2722 on which text and graphics are displayed. The display 2722 may be, for example, an LCD, LED, or OLED display screen (such as a desktop computer screen, a handheld device screen, or a television screen), an electronic ink display, a projection display (such as a heads-up display device), and / or a display integrated with a touch screen 2715 that serves as both an input device and an output device that provides graphical and textual visual feedback to a user. The output device may also include a speaker 2724 for playing audio signals, a haptic feedback device for tactile output such as vibration, etc. In some implementations, the speaker 2724 and microphone 2716 are implemented by a combined audio input / output device.
[0154] In the illustrated embodiment, computer system 2700 further includes one or more communications components 2750. The communications components may include, for example, a wired network connection 2752 (e.g., one or more of an Ethernet port, a cable modem, a Thunderbolt cable, a FireWire cable, a Lightning connector, a Universal Serial Bus (USB) port, etc.) and / or a wireless transceiver 2754 (e.g., one or more of a Wi-Fi transceiver, a Bluetooth transceiver, a near field communication (NFC) device, a wireless modem or cellular radio utilizing GSM, CDMA, 3G, 4G, and / or 5G technology, etc.). Communications component 2750 is suitable for communications between computer system 2700 and other local and / or remote computing devices directly via wired or wireless peer-to-peer connections and / or indirectly via communications links and networking hardware such as switches, routers, repeaters, electrical and optical cables, light emitters and receivers, wireless transmitters and receivers, etc. (which may include the Internet, public or private intranets, local or extended Wi-Fi networks, cell towers, plain old telephone systems (POTS), etc.) Computer system 2700 further includes power 2760, which may include battery power and / or utility power for the operation of various electrical components associated with computer system 2700.
[0155] Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with the present technology include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile phones, wearable electronics, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, etc. While computer systems configured as described above are typically used to support the operations of the present technology, those skilled in the art will understand that the present technology can be implemented using various types and configurations of devices and with various components. It is not necessary to illustrate such infrastructures and implementations to describe exemplary embodiments.
[0156] FIG. 28 is a schematic and data flow diagram 2800 illustrating various components of exemplary concurrent server interactions for backtesting, according to one embodiment. A user interacts with the rapid screening system, such as via a web application or other client device 2810 (e.g., running local software or accessing software-as-a-service (SaaS)). The user requests a backtest 2805, for example, as described in detail above with reference to FIG. 17. The rapid screening system may compute less complex results as well, such as modifying screening criteria and other portfolio analyses. The client web application or device 2810 sends a request 2815 to a session server 2820, for example, a host device that manages a user database on a remote server. The request 2815 may include, for example, subscriber authentication or session information to ensure the subscriber is authenticated (e.g., charged for the backtesting service) and / or a date and code that specifies the parameters of the backtest, such as the criteria to test and the benchmark to test. In various embodiments, the parameters are already associated with the user's session and available on the server, so identification of the current session and backtest request can be achieved with minimal data transmission. In some embodiments, the client interface device or web application 2810 sends one or more triggers to the session server 2820 for when a new view, backtest, or other portfolio analysis request occurs and subscribes to updates received when the analysis is completed.
[0157] When a user is validated by session server 2820, session server 2820 sends request 2825 to main function 2830, which runs on a server, such as one or more cloud computing instances. Main function 2830 manages the processing and assembly of backtest requests overall, including executing the request using concurrent computing resources (e.g., triggering lambda function 2840) to efficiently parallelize the backtest calculations. Lambdas are simple yet can be spawned in large numbers, allowing the system to spawn many in parallel and simultaneously perform large tasks in short bursts. For example, in a backtest, a high-speed screening system can decompose the calculation into discrete time periods, such as each year in the backtest and each individual month within each year. This allows for significant improvements in backtest speed. In various embodiments, the domain-specific language is designed for concurrent processing and data sharding (e.g., across time, data items, and user-specific patterns).
[0158] In the illustrated embodiment, lambda function 2840 obtains data from one or more financial data databases 2850 in materialized views 2845 to contain snapshots of data and time. In some embodiments, the snapshots in materialized views 2845 also include a recent set of historical values, causing some repetition but allowing for very rapid and comparable access to custom trends over the past few years, for example. Materialized view example data 2848 includes a date, an identifier, a pair of values (the value for the following 12 months and the value for the current quarter), and a pair of historical value sets (the value for the following 12 months and the value for the current quarter). In various embodiments, materialized views 2848 are saved to speed access when the same or different backtests are run, allowing users to switch back and forth between results of different backtests without requiring recalculation. In various embodiments, the screening and / or backtesting functionality can be plugged into many types of data stores and is not limited to databases 2850; for example, application programming interfaces (APIs) can also be effectively used.
[0159] Lambda function 2840 returns to main function 2830, which delivers the complete results 2855 of the backtest or other analysis to session server 2820. In various embodiments, after the query is executed, the complete results 2855 are stored server-side. This means that the rapid screening system does not need to constantly re-run queries if the data itself does not change (e.g., sorting, grouping, etc.), opening up the possibility of incremental queries.
[0160] In some embodiments, lambda function 2840 updates client interface device or web application 2810 with the subscribed results. In some embodiments, session server 2820 provides result identifier 2857 to client interface device or web application 2810. Client interface device or web application 2810 may then request partial results 2865 (e.g., only those necessary to visibly render the results on the user's display). Session server 2820 then delivers the requested portion of results 2875 to client interface device or web application 2810.
[0161] FIG. 29 is a schematic diagram 2900 illustrating various components of an exemplary server system for implementing a rapid screening system according to one embodiment. Users interact with the rapid screening system, such as via a web application or other client device 2910. The web application 2910 communicates with a session server 2920, for example, via server endpoints 2905, 2908, for requests such as screening, backtesting, or performing time-series predictive analysis. Requests to the server endpoint 2905 are handled by the server on a one-to-one basis, with the session server 2920 inserting, for example, one row per request into a session database. For example, when processing a screening session for a user who enters an expression, the web application 2910 provides the entered expression to the session server 2920, which can initiate processing for the expression. For example, whenever a row is updated, added, or deleted in the editor, the web application 2910 notifies the session server 2920, which initiates a request to a backend database 2940 or a compiler / interpreter 2930. In some embodiments, backtest requests from web application 2910 to session server 2920 are also handled on a 1:1 basis.
[0162] Requests to the server endpoint 2908 are handled by the server on a one-to-many basis, with the session server 2920 inserting multiple rows per request into the session database, for example. For example, in response to a time series analysis request, the session server 2920 may initiate a separate process for each company and forecast combination (i.e., each cell in a results table) rather than one process for the entire request. In some embodiments, the web application 2910 client initiates a series of separate requests, and in some embodiments, the session server 2920 receives the analysis request and spawns the necessary subtasks. In either case, the separate tasks may generate a collection of results for each cell rather than, for example, a single JSON blob, allowing the results to be assembled asynchronously.
[0163] Other components of session server 2920 include authentication service 2921, which ensures that requesters are authorized and that data is secure; authentication service 2921 may be linked to user data 2922. User database 2922 may additionally store a user's code (e.g., ensuring that screening expressions are automatically saved and available across sessions), preferences, and / or authentication credentials, and may provide a caching layer. Session data service 2923 manages user session data, including caching screening results 2934. For example, for a user running a screen, session data service 2923 may insert lines of code in the session database as the user writes code in a domain-specific language. In some embodiments, the session data service polls for updates. When requested data is loaded (e.g., from compiler / interpreter 2930), session data service 2923 caches the data and updates web application 2910 with the results.
[0164] Additionally, session server 2920 can store context data 2924. Context data 2924 can include, for example, custom populations (each including, for example, the name of the population, parameters defining the population (e.g., minimum market cap, country, specific company name, etc.)) and / or custom indicators (each including, for example, the name of the indicator and a time series of returns (which can be uploaded by users)). Context data 2924 can be shared between users to enable easier collaboration and minimize duplication or synchronization issues.
[0165] In the illustrated embodiment, session server 2920 is separate from compiler / interpreter 2930. Storing session data separately can provide advantages. For example, it can make the data more easily auditable for user analytics and / or telemetry. Additionally, it can allow for reporting on granular data such as field mappings of not only packages but also which data is most used by which people, which can help reduce costs. Thus, the structure of the disclosed technology allows for introspection and analysis that is generally not possible when performing traditional back-end database queries.
[0166] Additionally, because the compiler / interpreter 2930 is a separate server, the domain-specific language provides context for what the user is trying to achieve, and the expected results are verifiable and replicable, the disclosed system allows system administrators to make language feature updates without the traditional constraints of a programming language. This allows new features to be delivered to users faster, because the structure of the domain-specific language and its contextual understanding can allow administrators to ensure that they never break (as is typical with new versions of traditional languages).
[0167] Compiler / interpreter 2930 is the workhorse responsible for executing the actual requests for screening searches, backtesting, and / or analytical modeling. The schematic is a simplified logical representation, and one skilled in the art will understand that compiler / interpreter 2930 compiler / interpreter can, by way of example, comprise distributed computing resources to perform the necessary data processing, load balancing / queue management, etc. The illustrated separation of session server 2920 from compiler / interpreter 2930 is also advantageous for compiler / interpreter 2930, which does not need to understand users, but does need to understand data sources.
[0168] In the illustrated embodiment, the screening model 2935 includes or references a grammar 2931 and uses it to parse the user's code (e.g., as a whole) into an abstract syntax tree (AST) 2932. The screening model 2935 evaluates each line 2933 (e.g., in order), including the expressions in each line. Line by line, it interprets the AST. The screening model 2935 connects to one or more databases (including a cache, if available) to obtain data and continually builds a result set 2934 row by row. The screening model 2935 inserts the result set 2934 into the session data service 2923 results database.
[0169] In the illustrated embodiment, the backtest model 2936 runs through backtest-specific logic 2937 and shares components of the screening model 2935. The backtest model 2936 leverages the screening model 2935 as well as various database optimizations (including snapshots and offloading some calculations closer to the data store). When a backtest is run, it is often rebalanced over time according to a set of buy / sell criteria. These criteria can be expressed as screening criteria. Thus, when a user configures a screen, they can immediately run a backtest based on that screening code with just one click, without any additional configuration.
[0170] The compiler / interpreter 2930 then leverages its semantic understanding of the user's intent (to run a backtest) to optimize the code before executing it as part of the rebalancing step in the backtest model 2936. For example, if a data tag or expression is determined to be unused within the context of the backtest, perhaps with remaining columns from an exploratory screening session, or is calculated solely for display purposes to the end user, the backtest model 2936 can strip or otherwise exclude that code before running the backtest because it does not significantly affect the filtering criteria or factor weights of the backtest. In contrast, backtests performed using a generalized programming language lack this semantic understanding (that some of the data pulled down or some of the analysis performed does not have a meaningful impact on the current action the user is attempting to take), and therefore this type of optimization cannot be performed automatically using conventional means.
[0171] In various embodiments (e.g., to run backtests over multiple years), the backtest model 2936 decomposes the backtest into subcomponents and parallelizes them (e.g., by month, or to calculate revenue, etc.) 2938.
[0172] Additional models 2939 (e.g., time series analysis processing models) are similarly executed by compiler / interpreter 2930 to generate a set of results that are sent to session server 2920. With the minimum context required, the same domain-specific language is applicable to all models provided by compiler / interpreter 2930.
[0173] In some embodiments, to facilitate caching, requests can be run twice through the compiler / interpreter 2930: once to collect and prefetch data tags, and a second time to actually evaluate the request. This can allow the data store to effectively cache the results on the first run-through.
[0174] Backend data can include both unoptimized data stores 2945 and optimized databases 2940. For example, commonly accessed data can be optimized for quick retrieval, especially for the most common queries. The optimized database 2940 component can include, for example, lightweight read-access tables 2941 for different data collections and point-in-time snapshots 2942 (e.g., for "today" and "end of month" going back a certain number of years) to reduce query times. Unstructured data (e.g., 10-K filings) can be processed to be more optimized. However, optimization is not binary; for example, patent information contains a large amount of unstructured text, but filings can also be processed to link them to companies (rather than just linking to static data), such as through name similarities. Thus, such databases do not need to strictly match security names or identifiers and can nevertheless be "joined" in other ways.
[0175] Optimization is not simply a tool for providing improved performance, but a benefit provided by the structure of the disclosed system. Traditionally, someone running a query against a data set (e.g., running a backtest request) has no control over the underlying database, which is a large data store that must satisfy queries from users with different objectives and deal with generalized programming languages to access the database. The disclosed technology, in contrast, imposes significant constraints on use cases and access patterns while maintaining high expressiveness, which enables optimizations that produce high-performance results and a more user-friendly interface.
[0176] In some embodiments, the data store and other elements of the system are fully extensible and open to integration. For example, remote access is possible to any database with API adherence, and any provider can implement the interface. The data itself is domain-specific, where the language is specific, but the data server "works" if it conforms to the specified interface. Similarly, extensibility can be applied to collections of data tags (in a normalized format, whether defined in a domain-specific language, custom, derived, or self-referential), additional transformations, custom reporting templates (e.g., backtesting or screening results), results (e.g., returned in a normalized JSON format), views (e.g., to display an array of strings as a bulleted HTML list), and views (e.g., providing a separate session database endpoint for each model). For example, views can be returned to the web application 2910 as HTML or in data structures, allowing the data to be rendered specifically based on its content or context. Thus, for example, URLs can be displayed as links, arrays of values can be treated as graphs, and values of variables whose names end with "score" or "ranking" can be displayed as a color gradient in a heatmap. Essentially, any data tag, any function, any model, and any report can be integrated into the disclosed fully extensible system. The basic operators of the domain-specific language can remain the same, but new data types, operations, and transformations can be customized for different domains or clients.
[0177] While specific embodiments have been illustrated and described herein, those skilled in the art will understand that substitutions may be made for the specific embodiments shown and described without departing from the scope of the present disclosure.
[0178] For example, while various embodiments are described above with respect to rapid screening systems and / or services provided to remote clients by one or more servers, in other embodiments, screening methods similar to those described herein may be employed locally on a client computer to find and display results within a local or remote corpus.
[0179] Similarly, while various embodiments are described above with respect to screening systems or services that enable screening of stocks or other securities (e.g., debt instruments, cryptocurrencies, etc.), other embodiments may use similar techniques to enable screening or sifting of data in other domains of expertise, such as real estate, advertising, healthcare diagnostics, drug research, employment, fantasy sports, movies, scientific data, photography, etc. As a few examples, backtesting can run simulations of different drugs under a given set of conditions, a screening system applied to the field of geological data can improve the search for oil drilling sites, a system applied to real estate can find undervalued home purchase opportunities, and a system applied to cancer screening can quickly perform complex analyses on samples, providing an improved ability to identify trends and promising avenues for investigation. In other embodiments, a variety of other uses can be made of the disclosed technology. This application is intended to cover any adaptations or variations of the embodiments discussed herein.
Claims
1. 1. A system for processing natural language input into domain-specific screening criteria for screening a set of identifiers, each identifier representing a member of a domain, each member having a plurality of attributes, the system comprising: a processor; and a memory; the processor and memory providing a natural language input interface; receiving a natural language input via the natural language input interface; processing the natural language input through a model; The model is trained to process natural language sentences into domain-specific symbols, where the domain-specific symbols are a plurality of domain-specific data tags, each data tag associated with one or more time-linked attribute values for each of a set of a plurality of said identifiers; a plurality of operations that can be applied to time-linked attribute values associated with the data tags; processing the natural language input with the model; a first data tag of the plurality of domain-specific data tags; a first operation of the plurality of operations; and providing a reference interface; inputting the first domain-specific screening criteria generated by said processing into the criteria interface; executing the first domain-specific screening criterion, wherein said executing includes: determining a first subset of identifiers from the plurality of identifiers on which the first domain-specific screening criterion is performed; loading a time-linked attribute value associated with the first data tag from a data source; applying the first operation to the loaded time-linked attribute values to generate result values for a second subset of identifiers included in the first subset, each result value being associated with an identifier in the second subset of identifiers; providing an output display interface; presenting the result values for the second subset of identifiers in the output display interface, said presenting including, for each identifier in the second subset of identifiers, presenting an associated value of the result values in the output display interface; whereby the output display interface corresponds to content of the criteria interface including at least the first domain-specific screening criterion; and A system that is configured to:
2. The system of claim 1 , wherein the model is a language model or a language processing model.
3. The system of claim 2 , wherein the model is a generative pre-trained Transformer or a large language model (LLM) with at least 1 billion parameters.
4. 2. The system of claim 1, wherein the processing the natural language input with the model includes recognizing at least a portion of the natural language input that corresponds more closely to the first data tag in the plurality of domain-specific data tags than to another data tag in the plurality of domain-specific data tags.
5. 10. The system of claim 1, wherein in response to an input error, the model generates a domain-specific completion suggestion or error correction, and the natural language input interface or the reference interface displays the domain-specific completion suggestion or error correction, or in response to an input error, the model generates a domain-specific completion suggestion or error correction based on a determined textual or semantic similarity of the input error to domain-specific data tags.
6. 2. The system of claim 1, wherein data tags are associated with a series of time-linked attribute values over time, and for each of a plurality of the identifiers in the set, the presented result values comprise an array, a time series, a trend, or a graph of the series of time-linked attribute values over time.
7. The system of claim 1 , wherein the processing the natural language input with the model generates multiple domain-specific screening criteria from a single natural language sentence.
8. (a) the model processes the received natural language input continuously without waiting for the user to stop, or (b) the model processes the received natural language input after the user pauses for at least a determined debounce period and enters an indication of completion or makes a selection from among multiple options, or 2. The system of claim 1, wherein executing the first domain-specific screening criteria and presenting the resulting values in the output display interface includes executing the contents of the criteria interface including the first domain-specific screening criteria (c) continuously as they are entered into the criteria interface, or (d) after a user pauses for at least a determined debounce period and makes a selection from among a plurality of options or inputs an indication of completion.
9. 10. The system of claim 1, wherein inputting the first domain-specific screening criterion into the criteria interface comprises presenting one or more of an icon, a button, a selection or drop-down menu, a statement, an expression, or a mathematical formula representing the first domain-specific screening criterion.
10. The system of claim 1 , wherein the criteria interface includes a link or control for changing the order of the first domain-specific screening criteria relative to a second domain-specific screening criteria.
11. 2. The system of claim 1, wherein data tags are associated with data tag descriptive text, and wherein inputting the first domain-specific screening criteria generated by the processing into the criteria interface includes the data tag descriptive text for the first data tag, or a link or control for displaying the data tag descriptive text for the first data tag.
12. 2. The system of claim 1, wherein an operation is associated with an operation description text, and wherein inputting the first domain-specific screening criteria generated by the processing into the criteria interface includes the operation description text for the first operation or a link or control for displaying the operation description text for the first operation.
13. 10. The system of claim 1, wherein the criteria interface includes a link or control for editing the domain-specific screening criteria or a link or control for displaying an editing interface that allows a user to edit the first domain-specific screening criteria.
14. The system of claim 13 , wherein the link or control or the editing interface comprises a single-line editor, a multi-line editor, a drop-down menu, a context menu item, a button, a drag-and-drop interface, or the natural language input interface.
15. The system of claim 13 , wherein the natural language input interface, the criteria interface, or the editing interface displays the natural language input as text.
16. 16. The system of claim 15, further comprising: enabling the first domain-specific screening criteria entered into the criteria interface to be hidden as an intermediate computation, step, or layer between the natural language input and the output display interface.
17. the system is configured to (a) receive user input selecting the set of identifiers, and wherein the determining selects the selected set of identifiers as the first subset of identifiers on which the first domain-specific screening criterion is performed, or (b) receive a zeroth domain-specific screening criterion that precedes the first domain-specific screening criterion, and wherein the determining selects a subset of identifiers resulting from the zeroth domain-specific screening criterion as the first subset of identifiers on which the first domain-specific screening criterion is performed, or 2. The system of claim 1, wherein the first domain-specific screening criteria includes a data tag that references an output from the zeroth domain-specific screening criteria, such that the result value for the second subset of identifiers is based on an output from the zeroth domain-specific screening criteria.
18. The system of claim 1 , further configured to receive user input selecting one or more predefined or custom collections of identifiers from one or more data sources as the set of identifiers.
19. automatically identifying a data source associated with said first data tag without requiring a user to explicitly specify the data source; or Identifying multiple data sources from which time-linked attribute values will be loaded, or The system of claim 1 , further comprising presenting (i) data provenance or attribute information or (ii) a hyperlink to a data source or source record.
20. The first domain-specific screening criteria includes a statement, an expression, or a mathematical formula, and the performing includes: further comprising parsing the first domain-specific screening criteria sentence, expression, or formula against the plurality of domain-specific data tags and the plurality of operations; The system of claim 1 , wherein the parsing comprises recognizing the first data tag and the first operation in a sentence, expression, or formula of the first domain-specific screening criteria.
21. 2. The system of claim 1, further comprising: assigning the resulting value for the second subset of identifiers to a new automatically named or user named data tag, wherein the system recognizes the new automatically named or user named data tag as one of the plurality of domain specific data tags in a domain specific symbol, and wherein second domain specific screening criteria can reference the new automatically named or user named data tag.
22. The system of claim 1 , wherein the first operation generates the result value for each identifier in the first subset such that the second subset includes the same identifiers as the first subset.
23. 23. The system of claim 22, wherein the first operation is a transformation that produces a result value comprising metadata characterizing time-linked attribute values associated with data tags, the transformation being one of averaging by mean or median, ranking, grouping, bucketing into quintiles or deciles, normalization, or an indication of trend or stability of trend, or sentiment.
24. 2. The system of claim 1, wherein the first operation filters the first subset of identifiers on which the first domain-specific screening criteria is performed such that the second subset of identifiers is a proper subset of the first subset that contains fewer identifiers than the first subset, the output display interface before the presenting includes identifiers that are included in the first subset and not included in the second subset, and the output display interface after the presenting includes only identifiers that are included in the second subset.
25. 25. The system of claim 24, wherein the first operation comprises one or more of matching, comparing, containing or including, mentioning, top, highest, low, lowest, increasing, decreasing, or enumerating.
26. The system of claim 1 , further comprising: modifying the presentation of the result values when a user modifies the content of the criteria interface.
27. the output display interface is presented on a page together with the natural language input interface or the reference interface; or The system of claim 1 , wherein the output display interface is presented on a separate page from the natural language input interface or the reference interface such that the natural language input interface or the reference interface is hidden when the output display interface is presented.
28. the first domain-specific screening criteria input into the criteria interface, or data tags contained in the first domain-specific screening criteria, correspond to cells, columns, rows, tables, clusters, windows, or other visual collections of information presented in the output display interface; and The system of claim 1 , wherein the presenting in the output display interface comprises inserting the cell, column, row, table, cluster, window, or other visual collection of the result values in the output display interface.
29. and further configured to backtest the criteria input to the criteria interface, including the first domain-specific screening criteria, by applying the criteria to a period of historical data of the set of identifiers, the backtesting comprising: determining a backtest subset of identifiers from the set of identifiers according to the criteria entered into the criteria interface, including the first domain-specific screening criteria; loading, from the data source, historical values for the period of time associated with data tags of the criteria entered into the criteria interface, including the first data tag; applying the criteria operations, including the first operation entered into the criteria interface, to the loaded historical values over the period of historical data to generate backtest result values over the period of historical data for the backtest subset of identifiers; The system of claim 1 , wherein the output display interface presents a display of how the backtest subset of identifiers determined according to the criteria input to the criteria interface, including the first domain-specific screening criteria, performed over the period of historical data.
30. 1. A computer software information processing method comprising: a system having a processor and memory training a model to generate, from a natural language sentence, domain-specific symbols constituting domain-specific screening criteria for screening a set of identifiers; The domain specific symbol is a plurality of domain-specific data tags, each data tag associated with one or more time-linked attribute values for each of a set of a plurality of said identifiers; a plurality of operations that can be applied to time-linked attribute values associated with the data tags; The method, wherein the domain-specific screening criteria include a data tag of the plurality of domain-specific data tags and an operation of the plurality of operations.