Domain-specific language interpreter and interactive visual interface for rapid screening

Through a domain-specific language interpreter and interactive interface, combined with a multi-line editor and grid view, the problem of data screening complexity in the existing technology is solved, and fast and easy-to-use data screening and strategy optimization in the securities market is achieved.

CN116235144BActive Publication Date: 2025-08-19QUEXOTIC LAB INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180061410.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-24
Filing Date
2021-05-25
Publication Date
2025-08-19
Estimated Expiration
2041-05-25

AI Technical Summary

Technical Problem

When screening large amounts of data to find valuable goals, it is difficult for existing technologies to explore and screen data efficiently and intuitively, especially in the securities market, where the process of screening securities is complex and relies on complex screening standards and verification processes.

Method used

It adopts a domain-specific language interpreter and interactive visual interface, combined with a multi-line editor and grid view, allowing users to enter and edit symbols and operators, provide real-time results display, and uses artificial intelligence features for high-speed parallel evaluation and backtrack testing.

Benefits of technology

A fast and easy-to-use data filtering process is realized, and users can instantly feedback and refine screening strategies, reducing their dependence on technical expertise and improving the efficiency and accuracy of data exploration and screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116235144B_ABST
    Figure CN116235144B_ABST
Patent Text Reader

Abstract

The present application discloses an improved system and method for allowing users of a computing system configured specifically for a domain to explore and filter data in a rich attribute data set to test screening strategies and discover targets of particular interest. The technology includes an input interface such as a multi-line editor that allows users to enter, modify, add, insert, subtract, change, and otherwise freely edit entered symbols and operators at any time. A domain-specific language interpreter continuously processes the entered symbols and operators as they are updated. The interactive visual interface also includes a grid view that displays real-time results that are updated based on the current content of the input interface.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 029,556, filed on May 24, 2020, entitled “Domain-Specific Language Interpreter and Interactive Visual Interface for Rapid Screening,” which lists Sara Itani as an inventor. The entire contents of the aforementioned application and the entire contents of all priority documents cited in the accompanying Application Data Sheet are incorporated herein by reference in their entirety for all purposes. Technical Field

[0003] The present disclosure relates to improved systems and methods that allow users of computing systems to explore and filter data in domain-specific, attribute-rich datasets to test screening strategies and discover targets of particular interest. Background Art

[0004] Sifting through vast amounts of data to find something of value is often a daunting task. Metaphors like "finding a needle in a haystack" and "scouring the waters" convey the challenge of, for example, a hiring manager finding a group of candidates to interview for a job, a homebuyer identifying a small number of homes to consider viewing before buying, or an investor screening securities to select a group of stocks worth considering as investments.

[0005] For example, the stock market offers investors a vast array of securities, each with extensive information available from numerous sources. To make the process of selecting securities of interest more manageable, investors can screen securities based on a number of criteria to narrow the list. The result is a smaller list; whether the remaining securities are more promising depends on the investor's choice of criteria and the level of sophistication in selecting and verifying those securities. Given the vast array of available criteria for narrowing the list (some useful, some worthless, and some ultimately misleading), even effectively screening securities can be challenging.

[0006] Screening strategies may attempt to reflect an investor's investment philosophy or may be viewed generally as a form of or occasionally used to support a hunch. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 An exemplary user interface of a rapid screening system is illustrated, showing a multi-line editor and grid view display configured for equity screening, according to one embodiment.

[0008] Figure 2 An operating routine of a rapid screening system according to one embodiment is illustrated.

[0009] Figures 3A to 3B An exemplary user interface of a rapid screening system configured for equity screening is illustrated, showing modifications within a multi-line editor, according to one embodiment.

[0010] Figure 4 An exemplary user interface of a rapid screening system configured for equity screening is illustrated, showing a dialog box for creating a custom equity range, according to one embodiment.

[0011] Figure 5A An exemplary user interface of a rapid screening system configured for equity screening showing domain-specific flexible text matching and completion suggestions is illustrated according to one embodiment.

[0012] Figure 5B An exemplary user interface of a rapid screening system configured for equity screening showing data tag exploration is illustrated according to one embodiment.

[0013] Figures 6A to 6B An exemplary user interface of a rapid screening system configured for equity screening showing filtering according to criteria is illustrated, according to one embodiment.

[0014] Figure 7 An exemplary user interface of a rapid screening system configured for equity screening is illustrated, showing expressions assigned to custom variable names, according to one embodiment.

[0015] Figures 8A to 8C An exemplary user interface of a rapid screening system configured for equity screening is illustrated, showing simultaneous renaming of multiple references to custom variable names, according to one embodiment.

[0016] Figure 9 An exemplary user interface of a rapid screening system configured for equity screening showing domain-specific syntax error handling is illustrated, according to one embodiment.

[0017] Figure 10 An exemplary user interface of a rapid screening system configured for equity screening showing a transformation function is illustrated according to one embodiment.

[0018] Figure 11 An exemplary user interface of a rapid screening system configured for equity screening showing automatic graphical display of array data is illustrated according to one embodiment.

[0019] Figure 12 An exemplary user interface of a rapid screening system configured for equity screening is illustrated, showing the automatic display of a link to a 10-K report, according to one embodiment.

[0020] Figure 13 An exemplary user interface of a rapid screening system configured for equity screening showing the selective display of patent-holding companies is illustrated according to one embodiment.

[0021] Figure 14 An exemplary user interface of a rapid screening system configured for equity screening is illustrated, showing filtering of text found in a 10-K report, according to one embodiment.

[0022] Figure 15 An exemplary user interface of a rapid screening system configured for equity screening showing grouping of results is illustrated according to one embodiment.

[0023] Figure 16A to Figure 16B Exemplary formulas for a prior art system and corresponding exemplary expressions for a rapid screening system configured for equity screening, showing improved ease of use, according to one embodiment, are illustrated.

[0024] 17A to 17B An exemplary user interface of a rapid screening system configured for equity screening showing backtesting is illustrated according to one embodiment.

[0025] Figure 18 An exemplary change alert diagram of a rapid screening system configured for equity screening is illustrated, according to one embodiment.

[0026] Figure 19A Illustrated are exemplary AI features for introspection in a rapid screening system configured for equity screening, according to one embodiment.

[0027] Figure 19B Exemplary AI features for making predictions in a rapid screening system configured for equity screening are illustrated, according to one embodiment.

[0028] Figure 20 Exemplary AI features for regime change detection in a rapid screening system configured for stock screening are illustrated, according to one embodiment.

[0029] Figure 21 Exemplary AI features for optimizing mixing in a rapid screening system configured for equity screening are illustrated, according to one embodiment.

[0030] Figure 22 Illustrated are exemplary AI features for feature suggestions in a rapid screening system configured for equity screening, according to one embodiment.

[0031] Figure 23An exemplary user interface of a rapid screening system configured for equity screening, showing the inventive use of operators, according to one embodiment is illustrated.

[0032] Figure 24 An exemplary user interface of a rapid screening system configured for equity screening showing automatic formatting is illustrated according to one embodiment.

[0033] Figure 25 An exemplary user interface of a rapid screening system configured for equity screening showing a predicted point-in-time situation report is illustrated according to one embodiment.

[0034] Figure 26 An exemplary user interface of a rapid screening system configured for equity screening showing a scenario report historical forecast graph is illustrated according to one embodiment.

[0035] Figure 27 is a block diagram illustrating some of the components that are typically incorporated into computing systems and other devices in which the present technology may be implemented.

[0036] Figure 28 is a schematic and data flow diagram illustrating several components of an exemplary concurrent server interaction for backtesting according to one embodiment.

[0037] Figure 29 is a schematic diagram illustrating several components of an exemplary server system for implementing a rapid screening system according to one embodiment. DETAILED DESCRIPTION

[0038] The present application discloses improved systems and methods that allow users of domain-specifically configured computing systems to explore and filter data in attribute-rich datasets to test screening strategies and discover targets of particular interest.

[0039] The disclosed technology utilizes a novel, intuitive approach that includes a domain-specific language interpreter and an interactive visual interface that together respond instantly to symbols and operators entered in the domain-specific language. The technology includes an input interface, such as a multi-line editor, that allows a user to enter, modify, add, insert, subtract, change, and otherwise freely edit the entered symbols and operators at any time. The domain-specific language interpreter continuously processes the entered symbols and operators as they are updated. The interactive visual interface also includes a grid view that displays real-time results that are updated based on the current contents of the input interface.

[0040] The technology also provides high-speed, parallel evaluation of selected targets against historical data and benchmarks, enabling strategies to be backtested in seconds. The technology also includes artificial intelligence (AI) machine learning features to help users identify strategies or the factors driving them, find similar targets, and consider different evaluation criteria.

[0041] For example, when applied to securities information, the systems and methods disclosed herein provide improved methods for screening securities.

[0042] In summary, various aspects of the disclosed technology provide high ease of use with a short learning curve and provide immediate feedback to enable rapid refinement of iterative search strategies for exploring, visualizing, and filtering structured and / or unstructured data. The domain-specific language and associated interpreter and interactive visual interface integrate data exploration, querying, visualization, and upcoming goals into a central workflow. The structure of the domain-specific language and the immediate, user-friendly feedback provided by the rapid screening system combine to encourage iterative discovery and exploration, thereby enabling users to visualize the effects of their screening criteria and making it easy enough for users with only commercial spreadsheet experience to quickly perform complex screening operations. Therefore, people without technical expertise who want to find answers can get answers through the currently disclosed technology without having to rely on different people or teams with such expertise. A further result is that users can explore options and refine their ideas in real time based on the features provided by the technology.

[0043] Furthermore, the disclosed technology allows for a fundamental shift in the process of idea generation, screening, research, testing, execution, and monitoring of strategies (e.g., investment strategies). In the past, each of these processes would be handled separately. The disclosed technology uniquely synthesizes them all. The technology replaces disjoint processes performed by different people with limited, separate feedback loops, and integrates them into a single language, tool, and interface, forming an integrated, immediate feedback loop.

[0044] The technologies described in this disclosure that provide the basis for the disclosed rapid screening system include advances and insights in compilers, human-computer interfaces or interactions (HCI), programming language design, database engineering, distributed computing, machine learning, quantitative analysis, and finance. Thus, the improvements disclosed herein as a whole combine improvements in several different fields and are generally not obvious to one of ordinary skill in any one field.

[0045] Reference is now made in detail to the description of the embodiments as shown in the accompanying drawings. Although embodiments are described in conjunction with the accompanying drawings and related descriptions, there is no intention to limit the scope to the embodiments disclosed herein. On the contrary, it is intended to cover all alternatives, modifications and equivalents. In alternative embodiments, different or additional inputs (e.g., drop-down menu selectors or natural language processing) can be added to or combined with those input interfaces illustrated without limiting the scope to the embodiments disclosed herein. For example, the embodiments set forth below are primarily described in the context of equity screening. However, the embodiments described herein are illustrative examples and in no way limit the disclosed technology to any particular application, field, subject of knowledge, search type, or computing platform.

[0046] The phrases "in one embodiment," "in various embodiments," "in some embodiments," etc. are used repeatedly. Such phrases do not necessarily refer to the same embodiment. The terms "comprising," "having," and "including" are synonymous unless the context dictates otherwise. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should also be noted that the term "or" is generally used in its sense including "and / or" unless the context clearly dictates otherwise.

[0047] The disclosed rapid screening systems and methods may take on a variety of form factors. Figures 1 to 29 Several different arrangements and designs are shown. The illustrated rapid screening system is not an exhaustive list; in other embodiments, the syntax can be rearranged (e.g., "only" or "limit" instead of "filter" or "~"), or the input editor or result display can be formed in a different arrangement. However, it is not necessary to show these optional specific implementation details in detail to describe the exemplary embodiments.

[0048] Figure 1 An exemplary user interface 100 of a rapid screening system according to one embodiment is illustrated, showing a multi-line editor 101 and a grid view display 102 configured for equity screening. The user interface 100 may be provided by one or more rapid screening system computing devices such as a local or remote server, hereinafter referred to as Figure 27 、 Figure 28 and Figure 29 Described in more detail, in the illustrated user interface 100, the quick screening system provides the user with a multi-line editor 101 in which the user has entered symbols 110, 120, 130, 140, 150, 160 on a series of lines numbered 1-6.

[0049] Each of rows 1-6 contains symbols 110-160 that are part of a domain-specific language (in this case, a language specific to the securities domain). Symbols 110-160 represent one or more of: a starting universe of identifiers, with each identifier representing a member of the universe; data labels for data values, with each data value associated with an identifier; operators for operating on data values; and / or filtering criteria for limiting or narrowing the universe of identifiers to be considered.

[0050] The specification defines a domain-specific language. A domain-specific language is focused on a domain and can therefore include constraints relative to general-purpose programming tools or languages. In contrast, general-purpose languages such as Python or Structured Query Language (SQL) are cross-domain programming languages rather than domain-specific languages. Most of the syntax in a domain-specific language is domain-specific terminology, so that a person of ordinary skill in the art will recognize most of the syntax in a domain-specific language. This means that domain-specific languages contain very little programming-specific syntax and are not suitable for general-purpose programming. This also means that users can get results without doing any "programming": even if the user only enters a data label, the quick screening system will display the result because it is semantically meaningful and therefore human-readable.

[0051] In various embodiments, the domain-specific language includes symbols for representing data associated with identifiers (including results of operations on data) or filtering operations for selecting among identifiers.

[0052] In some embodiments, the domain-specific language does not contain non-domain-specific keywords. In this case, operators are symbols (such as "+" or brackets "[" ... "]"). Therefore, all alphanumeric text entered by the user in the multi-line editor 101 can be understood as meaningful, such as data labels. In some embodiments, the domain-specific language does not contain non-domain-specific keywords other than transformation functions.

[0053] The data tags each represent a data value associated with an identifier (e.g., a security that can be identified by a unique security ID and / or a recognizable stock symbol), which includes a value calculated based on other values such as the result of an expression. In many embodiments, the data value is also associated with a date, a time span, or a series of dates. The type of value can include, for example, a numeric value (e.g., the most recent price of a security or return on equity ("ROE")), a string value (e.g., a country or industry), an array of multiple values (e.g., the ROE for each of the past four quarters or recent news headlines), a data structure / collection / serialization such as JSON (JavaScript Object Notation) or XML (Extensible Markup Language) or HTML (Hypertext Markup Language) or CSV (Comma Separated Values) data (e.g., information about a 10-K report, including its date and hyperlinks), etc.

[0054] An expression is a finite, well-formed combination of data tags and operations in a domain-specific language; a line of user input can be interpreted as one or more expressions. Each expression entered by the user is evaluated, generating one of the aforementioned types of data values for each identifier (for example, concatenating or otherwise combining strings to generate a new string, or calculating a normalized weighted combination of ROE and return on assets ("ROA") to generate a new numeric value). Expressions can use standard mathematical operators, such as "+" or "*." Thus, expressions produce data values that are calculated based on other data values, including other expressions.

[0055] A transformation is an expression that generates metadata that characterizes an identifier relative to a referenced data value, such as over time, relative to a standard distribution, or relative to other securities (e.g., average, tier, quintile, normalization, trend stability, etc.).

[0056] Variable names represent any set of data values that the user assigns to any user-selected variable name (e.g., the value of any expression the user wants to evaluate), which then becomes a custom data label representing that data value for easy reference later. In this way, users can extend the domain-specific language.

[0057] The universe of identifiers (e.g., securities) to be considered and screened may include predefined sets (all stocks, S&P all members of , all foreign securities, all U.S. large-cap stocks, corporate or government bonds, etc.) and / or user-defined custom collections or templates.

[0058] Matching operators provide, for example, numeric comparators (e.g., <, <=, ==, >=, >, != / <>) (e.g., revenue > 10M) [in various embodiments, shorthand units may be handled and / or displayed, e.g., M for millions or k for thousands] or text comparators, such as "contains" or "is" or "?" (e.g., TickerSymbol? "AAPL" or Latest10K contains "China tariffs"). In some embodiments, operators like "contains" are implemented without alphabetic text. For example, using bracket notation to indicate an "includes," "has," or "contains" relationship, the syntax might be "Filings 10k ["china tariffs"]". This provides an intuitive representation of one thing (the phrase "China tariffs") within another thing (a set of company 10-K reports) and avoids non-domain-specific keywords in an otherwise domain-specific language.

[0059] The filter operator selects identifiers that match certain data value or expression criteria according to one or more matching operators (e.g., ~RoE>0 or Price<100). In various embodiments, filtering causes the display of a reduced-size universe or a smaller portion of the initially selected universe to be displayed, and data values corresponding to the criteria used for filtering are displayed.

[0060] In various embodiments, by means such as Figure 1 The invention provides an environment or tool such as the rapid screening system shown to process domain-specific languages, the environment or tool including an interactive multi-line editor: the interactive multi-line editor is configured to accept user input (e.g., from a keyboard, speech recognition or natural language processing, on-screen buttons, drop-down menus, and / or context menu item selections from various input options, some of which can be displayed in place of the editor); a parser and interpreter configured to interpret the input according to a specification defining the domain-specific language; a server engine configured to obtain domain-related structured and / or unstructured data based on the input; and a visualization tool configured to provide a real-time updating result display or grid view as the user enters the input into the interactive editor. In various embodiments, each expression and / or data label entered in the editor corresponds to a column of information displayed in the grid view.

[0061] Unlike existing systems, the resulting data is presented in a way that can be meaningfully interpreted and easily visualized by users. This reduces the time users must spend ensuring they are pulling down the right data and potentially cleaning it. It also makes it easier to discover patterns and relationships. Thus, the rapid screening system helps users instantiate investment strategies: unlike traditional screeners, the disclosed system and method allows for rapid discovery and exploration of data; testing of different ideas to validate or invalidate them; flexible generation of custom indicators and indices; repeatable processes; identification of trends, commonalities, and characteristics; and actionable insights that can be directly applied to strategies.

[0062] return Figure 1 , on the first line 110, the user has entered "$UnitedStatesAll". In this example (within the securities field), the "$" is an identifier indicating a universe, and the universe is all U.S. stocks. As used herein, an identifier represents a member of a universe - in the illustrated example, a specific security. Therefore, the user has chosen to start by considering all U.S. stocks. Within the securities field, the user can alternatively choose to consider, for example, other types of securities such as corporate or government (e.g., municipal) bonds, options, mutual funds, etc.; equities in other countries; stocks traded on specific exchanges; industry-specific securities; fixed-income financial products; non-performing loans; derivative securities; and so on. Although this disclosure generally uses the terms "equity," "stock," or "company" as a convenient shorthand, it is the intent of this disclosure to cover all securities using these terms.

[0063] Below the multi-line editor 101, a grid view display 102 is illustrated. In other embodiments, the multi-line editor 101 and grid view display 102 may be arranged differently; for example, the editors may be below each other, side by side, or in different windows or screens entirely. Grid view display 102 includes a series of columns 115, 123, 125, 133, 135, 145, 155. Each column corresponds directly to a symbol or expression in the multi-line editor 101. For example, the depicted "Stock Symbol" column 115 corresponds to the "$UnitedStatesAll" universe for all US stocks. "Stock Symbol" column 115 displays a series of rows, one row for each identifier associated with the universe for all US stocks. Thus, each US stock is identified by its stock symbol in "Stock Symbol" column 115. The stock symbol is merely a display name; in various implementations, the technology uses a unique identifier to unambiguously identify each security. For example, each member of the selected universe can be identified by its full name and exchange, or, for example, by some other unique identifier such as a security ID. This is particularly useful for supporting international stocks. In other embodiments of the present technology, the grid view display 102 can be arranged differently, such as with each column representing a security (or other screening target) and each row representing a security's attributes, nested tables, or other equivalent arrangements, all of which are within the scope of the present disclosure.

[0064] In various implementations of the illustrated column and row formats, the data values in each column are sortable, such as by user interaction with the column header (e.g., by clicking a mouse to sort, reverse, or reverse the sort order, or through a context menu). In various implementations, the columns themselves are reorderable. For example, the grid view display 102 may provide controls for a column to be moved left or right relative to another column, to be dragged (e.g., by a mouse click and drag operation) to a different position between columns, to be hidden or closed (or re-shown), or to be minimized (e.g., as an icon indicating a re-expandable data set or grouping). In some embodiments, when a column of information is moved or removed in the grid view, the interface displays an indication in the editor with a corresponding expression and / or data label, such as a control allowing the column to be re-shown or a symbol indicating that the column is not displayed in the original order. In some embodiments, when a column of information is moved or removed in the grid view, the system also updates the code in the editor.

[0065] Continuing in the multi-line editor 101, on the second line 120, the user has entered "ReturnOnEquityPctStandardize=>zScoreROE". Although the specific syntax of the illustrated example will be described in more detail below, the technology encompasses various equivalent alternatives and is not limited to the exact domain-specific language illustrated. In this example, the user first selected or entered the data label "ReturnOnEquityPct". As used herein, a data label represents an attribute of each member in a given universe; a data label marks or refers to a specific data value associated with each identifier. For example, in the universe of all U.S. stocks, ReturnOnEquityPct means, for each company with stock in the universe, the return on equity percentage. Similarly, as shown on the third line 130 in the multi-line editor 101 in this example, ReturnOnAssets means, for each such company, the return on assets.

[0066] Thus, in the grid view display 102, column 123 shows a "Return on Equity Pct" data label in a header row at the top of the column, and displays data values listing the return on equity percentage for each stock symbol shown in the corresponding row. In the illustrated example, the screen is run starting from the current date. In various embodiments, the system also allows the screen to be run from a previous "current day."

[0067] Rows 120 and 130 begin with the ReturnOnEquityPct and ReturnOnAssets data labels and then proceed with expressions that transform each of them by normalizing them to a normal distribution, thereby producing z-scores. In some embodiments, the normalization function first winsorizes the values to reduce the influence of outliers, replaces null values with the median, etc. The illustrated expressions then assign these z-scores to custom named variables or data labels "zScoreROE" and "zScoreROA," respectively. Thus, column 125, under the heading "z-score ROE," displays the normalized z-score of the return on equity percentage for each stock (represented by the stock symbol in column 115); column 135, under the heading "z-score ROA," similarly displays the normalized z-score of the return on assets for each stock.

[0068] Row 140 in multi-line editor 101 is an expression that adds the zScoreROE and zScoreROA data labels and assigns them to a new variable or data label, "zScoresAdded." Column 145 in grid view 102 thus has the title "z-scores added" and displays the new data value for each stock identifier with the stock symbol in column 115. The data value in column 145 is the sum of the data values in columns 125 and 135.

[0069] In this example, significant rounding can be observed: the displayed value is automatically limited to two decimal places to enhance readability during the screening process, while the system can operate on the exact digits below the displayed data value. In other embodiments, significant digits or other methods can be used to improve usability. In some embodiments, user actions such as hovering, right-clicking, copying and pasting, or long-pressing a data value can reveal additional details about the value.

[0070] Continuing in the multi-line editor 101, on the fifth line 150, the user has entered the expression "zScoresAddedSplitQuintiles=>bucket". As shown in the function suggestion pane and explanatory text 151, the SplitQuintiles function in this example sorts the zScoresAdded data values across the current set of identifiers (e.g., security IDs, even where identifiable stock symbols are displayed) into quintiles, where, for example, the lowest quintile is assigned a value of "1" and the highest quintile is assigned a value of "5" (or vice versa in some systems). Thus, among all U.S. stocks, companies that rank in the bottom 20% for the combined return on equity percentage (normalized) plus return on assets (normalized) have a value of "1" in column 155, while companies in the top 20% have a value of "5" in column 155, and other companies have values of "2", "3", or "4" in column 155. In grid view 102 , column 155 is titled “bucket” because, in multi-line editor 101 , the expression in row 150 assigned the quintile data labels to the custom variable “bucket”.

[0071] On the sixth row 160, the user has entered the expression "~bucket == 5". In the syntax of this illustrative example, the tilde in this expression represents a filtering operation. This operation enables the user to filter, limit, reduce, or narrow down the previously displayed set of identifiers based on matching criteria. In this case, once the quick screening system processes the filtering expression in row 160 in the multi-line editor 101, the grid view 102 no longer displays all identifiers (or, for convenience, stock symbols) in the universe of all U.S. stocks. Instead, only the stock symbols for which the custom variable "bucket" has a data value of "5" 165 (i.e., the top quintile of the normalized combined return on equity and return on assets) are included in the updated grid view 102.

[0072] Thus, the disclosed rapid screening system allows users to perform complex screening of large amounts of multi-attribute data, and to do so far more easily than previous systems have allowed, using minimal syntax, and with unprecedented continuous, efficient, immediate feedback and results.

[0073] Figure 2 An operational routine 200 of a rapid screening system according to one embodiment is illustrated. In various embodiments, the operational routine 200 is provided by one or more rapid screening system computing devices, such as local or remote servers, hereinafter referred to as Figure 27 、 Figure 28 and Figure 29 The operating routine 200 begins at a start block 201 .

[0074] Figure 2 The flowcharts and diagrams that follow are representative and may not show all functions, steps, or data exchanges; rather, they provide an understanding of how the system can be implemented. Those skilled in the relevant art will recognize that some functions may be repeated, changed, omitted, or supplemented, and that other (less important) aspects not shown may be easily implemented. Those skilled in the art will understand that Figure 2 The blocks shown in each of the schematic diagrams discussed below can be changed in a variety of ways. For example, when processes or blocks are presented in a given order, alternative implementations can perform the routines in a different order, and some processes or blocks can be rearranged, deleted, moved, added, subdivided, combined, and / or modified to provide alternatives or sub-combinations. Each of these processes or blocks can be implemented in a variety of different ways. In addition, although processes or blocks are sometimes shown as being performed continuously, these processes or blocks can alternatively be performed or implemented in parallel, or can be performed at different times. Figure 2 Some of the blocks and other schematic diagrams depicted in the drawings are of a type well known in the art, and they themselves may include sequences of operations that do not need to be described herein. One of ordinary skill in the art can create source code, microcode, program logic arrays, etc., or implement the disclosed technology based on the figures and detailed description provided herein.

[0075] In block 215, the operating routine 200 generates an interface for user input and / or display of results to be quickly filtered. For example, in one embodiment, the operating routine 200 provides an interactive editor (such as Figure 1 Multi-line editor 101) and result grid view (such as Figure 1 Grid view 102 of ), wherein this embodiment will be used as an example to illustrate various operations described below. In various embodiments, the rapid screening system enables an interface for user input and / or result display to be presented on a client device remote from the rapid screening system server.

[0076] In box 225, the operating routine 200 obtains input, for example, via the interactive editor provided in box 215, and parses the input according to a domain-specific language specification or grammar. The domain-specific language specification or grammar can contain, for example, terms that represent data tags from a particular domain, so that a user with knowledge of the domain can use that knowledge to discover terms in the domain-specific language. For example, in the securities field, and particularly in the stock field, "ROE" is a common shorthand for "return on equity," which is a measure of profitability or financial performance calculated by, for example, dividing annual net income (assets minus liabilities) by shareholders' equity. Thus, if the rapid screening system is configured for such a domain and receives an input of "roe" from a user, the operating routine 200 can parse the input into a request for a data tag named "ROE" or "ReturnOnEquity" using text pattern matching.

[0077] In various implementations, parsing the input in block 225 includes performing lexical analysis on the input to identify symbols or tokens from a domain-specific language specification or grammar in the input, and performing syntactic analysis to transform the input symbols into an abstract syntax tree (AST).

[0078] In block 235, the operation routine 200 processes the parsing error. For example, a user may type "roe" when the domain-specific language does not contain a data tag named "roe" or "ROE." In some embodiments, the operation routine 200 records any errors (including, for example, errors at other stages such as data retrieval) that prevent the operation routine 200 from further progressing with respect to the row with the error and loops back to block 225 to process additional user input such as corrections. In various embodiments, the operation routine 200 processes as much as possible despite the error and retains the results that existed before the error was encountered.

[0079] In some embodiments, the operation routine 200 identifies the unrecognized input for the user, such as by highlighting it. In some embodiments, the operation routine 200 may attempt to infer the closest match in the domain-specific language and, for example, replace the unrecognized input with the closest valid match, or propose a set of possible replacements or supplements for the unrecognized input. In various embodiments, the operation routine 200 continues to parse the user's input and process the symbols recognized by the parser.

[0080] In block 245, the operation routine 200 obtains any data necessary to perform operations on the data tags resolved in block 225. For example, if the resolved expression requires a set of data values that have not yet been loaded into memory, the operation routine 200 identifies one or more data sources associated with the required data values and loads the required information from at least one of the one or more data sources. In some embodiments, the operation routine 200 obtains data at another stage or on the fly.

[0081] In some embodiments, the operating routine 200 loads or preloads and caches snapshots of slowly changing (in terms of volatility or stability in magnitude and / or frequency of change), recently displayed, and / or frequently requested data values for a set or subset of identifiers and / or data tags to speed up processing when a user's input requires such data values. In some embodiments, the operating routine 200 identifies relevant data tags (e.g., based on user history, popularity across users, availability, etc.), pre-fetches them (e.g., in parallel), and locally caches the data to quickly provide results to the user.

[0082] In block 255, the operation routine 200 interprets the parsed input. For example, if the user entered an expression, the operation routine 200 identifies the data value to which the parsed symbol refers, determines what operation should be performed, and performs the required operation. For example, the operation routine 200 can interpret the AST and evaluate each parsed expression. In some embodiments, the operation routine 200 operates on a line-by-line basis. In some embodiments, the operation routine 200 evaluates each input line in order.

[0083] In some embodiments, once the input is parsed, the operation routine 200 interprets and executes the input on a continuous basis (e.g., without waiting for the user to stop typing) so that the interface will update as quickly as possible when the input is entered and the result is obtained, and the user does not have to complete the entry input before deciding to execute the completed input. Whenever the user input changes (and optionally whenever the underlying data changes), if the result is updated immediately or after a short delay, the result can be considered "real-time". Providing real-time results not only facilitates learning and exploring the data set itself, but also increases the learnability of the interface. In some embodiments, the operation routine 200 interprets and / or executes the input after a short delay (e.g., from about one tenth of a second to about 1-5 seconds of "de-bounce" period, or as long as the user pauses when typing), to allow the user to finally complete or correct or modify the input before the rapid screening system completes processing the input. In some embodiments, the operation routine 200 applies such a "de-bounce" period to the result display, as described further below. This delay can help the operation routine 200 present the results when the user is ready, thereby improving the user's perception of receiving immediate response results.

[0084] In various embodiments, the operation routine 200 provides incremental interpretation, interpreting and executing only the statements, expressions, or lines affected by the user input. For example, if the operation routine 200 receives, parses, and interprets the expression on the fifth line, the operation routine 200 may keep the contents of the first four lines unchanged to minimize processing time. By incremental interpretation and by combining new or changed elements with previous results when possible, the fast screening system maximizes the responsiveness of the interface.

[0085] In various embodiments, the domain-specific language is fast for the operation routine 200 interpretation 255 due to a combination of the following reasons: because the interpreter can exploit its concurrent and data-sharding design (e.g., across time, data items, and user-specific schemas); the language is domain-specific and therefore highly specialized and optimized for tasks rather than for general-purpose computation; only results that need to be visible to the user are sent to the client; it is designed for incremental interpretation; and the language itself allows for semantic understanding of user behavior, which can optimize the query planner. In addition, the back-end system is arranged to provide fast results based in part on constraints on the domain-specific language, as described below with reference to Figures 27 to 29 Further discussion.

[0086] In block 265, after the input has been parsed and interpreted, the operation routine 200 updates, removes, and / or adds categories to the interface for displaying results corresponding to the interpreted input from the interface for user input. For example, in the example grid view embodiment, when the quick screening system interprets an expression that newly calls a data label, the operation routine 200 adds a column with a title associated with the data label to the grid view. In some embodiments, the operation routine 200 generates new column header names for expressions that are not explicitly named by the user.

[0087] In some embodiments, while the operation routine 200 is acquiring data (as described in block 245), the operation routine 200 may display a column of data labels and populate the data values as they are received. In some embodiments, the operation routine 200 displays a column for each calculation or newly introduced data label in each statement or expression (e.g., each line in a multi-line editor), thereby providing transparency and understanding (and aiding debugging) because all intermediate steps are visible. In some embodiments, the operation routine 200 displays a column for each statement or expression and hides intermediate calculations by default. When the user edits or deletes an expression, the operation routine 200 accordingly changes or removes data labels and data values that are no longer included in the input. In various embodiments, the operation routine 200 incrementally updates the result display by combining new or changed elements with previous results to minimize processing and rendering time, thereby maximizing the responsiveness of the interface. In some embodiments, the operation routine 200 acquires data as described in block 245 and displays the results only when enough data has been received to display to the user. For example, rather than displaying headers above mostly empty columns in a grid view display as results are obtained, the operating routine 200 can compile the results (e.g., most or all of the result values, or most or all of those result values that will be initially visible to the user) before updating the display. Adding complete or mostly complete results to the display immediately when the values are ready, rather than adding them piecemeal as they load, can help the operating routine 200 present the results in a way that improves the user's perception of receiving immediate results. Generally speaking, results that are presented within several seconds of a user completing an expression or input line are perceived as "immediate." When determining whether an immediate response is provided in presenting the results later, the user can ignore startup delays, such as for the initial loading of a complex data set.

[0088] In box 275, after the input has been parsed and interpreted, the operation routine 200 filters the displayed identifiers corresponding to the selected identifier universe and the active filtering criteria. For example, in the grid view embodiment example, when the quick screening system interprets an expression that filters data values associated with a data tag, the operation routine 200 determines which identifiers match the filtering criteria and displays only the matching identifiers and the data values associated with the matching identifiers. In some embodiments, the operation routine 200 continuously builds a result set with an entry for each input row or expression. In some embodiments, when the result data is obtained, the operation routine 200 inserts the result set into the database for displaying the results (for example, the client requests a subset of the data that will be initially visible to the user so that the client can actually display the subset of the entire result set). In some embodiments, the operation routine 200 can use results from an existing session with the user.

[0089] In frame 285, operation routine 200 determines whether for example receives additional input via the interface that is used for user input.In some embodiments, operation routine 200 can asynchronously process additional input (being included in any position in previous input and amend or delete previous input) and do not wait for other frames (for example, obtain data 245) to finish, or can cancel the processing of previous input to ensure the quick display of the data of recent request with the input that responds to new or change.In other words, operation routine can determine whether to receive additional input at any time.If receive additional input, then operation routine 200 loops back to input parsing frame 225 to process next input symbol (if any).

[0090] The operating routine 200 ends in end block 299 .

[0091] Figure 3A 、 Figure 3B An exemplary user interface 300 of a rapid screening system configured for equity screening is illustrated, showing modifications within a multi-line editor, according to one embodiment.

[0092] exist Figure 3A In the example, row 1 310 in multi-line editor 301 specifies the "$UnitedStatesSmallAndMidCap" identifier universe, i.e., U.S. stocks classified as small and medium-sized, and row 2 320 is a data label for "CompanyName." Thus, in grid view 302, first column 315 displays the stock symbols (or, for example, unique security IDs as appropriate identifiers) for the stocks in the specified universe, and second column 325 displays the company name associated with each identifier. While the illustrated example focuses on U.S. stocks, international securities can also be selected.

[0093] exist Figure 3B, row 1 312 is different: it now specifies the universe of "$UnitedStatesLargeCap" identifiers, i.e., U.S. stocks classified as large-cap. Row 2 320 remains unchanged, containing the data label "CompanyName." In grid view 302, the modified first column 317 displays a new set of stock symbols (instead of identifiers) for the stocks within the specified universe, and the second column 327 displays the company name associated with each identifier. In this example, the quick screening system accepts modifications to the input on rows 1 310-312 even after the input has been entered and processed on row 2 320. The quick screening system seamlessly handles such changes (completely replacing the universe of identifiers under consideration) without requiring the user to delete back from the end of the input to the point of the change or otherwise start over. In some embodiments, the quick screening system allows the user to undo changes (e.g., by undoing the stack, or using a keyboard key such as Control-z) and allows the user to quickly switch back and forth between dynamic results by caching the results. By providing flexible, freely editable multi-line input and by continuously refreshing updates in response to changing user input, the disclosed rapid screening system allows users to powerfully and quickly explore alternatives to identify a desired strategy in a manner not previously possible.

[0094] Figure 4 An exemplary user interface 400 of a rapid screening system configured for equity screening according to one embodiment is illustrated, showing a dialog box for creating a custom equity range. A user may wish to create custom criteria for the starting universe of identifiers to be considered. For example, a real estate agent may specialize in the type or location of housing (e.g., downtown apartment units or exurban single-family homes), or an investor may focus on industry categories or company sizes. In various embodiments, the rapid screening system provides a convenient interface for defining new or custom universes using various criteria, which can be optimized based on frequent use.

[0095] In the illustrated example, dialog box 400 prompts the user to give a name 410 to a custom universe of, for example, stocks having a selected aspect 420, such as a specified country, market capitalization range, minimum liquidity, or Global Industry Classification Standard (GICS). ) Industry category. In some embodiments, available facets 420 may include a combination of factors, such as a group of countries. In some embodiments, a custom universe may be composed of selected companies. Such a universe may be created by looking up a name or loading a list of identifiers (e.g., names in a portfolio). In this example, dialog box 400 also allows the user to select an index as a benchmark, such as the S&P 500. After receiving the selection to define a universe of identifiers, dialog box 400 allows the user to select "Create Universe" 430 and then reference the custom universe by selecting a name.

[0096] In some embodiments, the rapid screening system allows created universes to be shared with others, for example, in a shared office environment, for collaboration between users. Similarly, in some embodiments, the rapid screening system allows custom data labels, expressions, and entire screening sessions to be shared collaboratively. In some embodiments, custom identifier universes can be edited or deleted after creation. In some embodiments, to enhance the performance of the rapid screening system when creating new universes, data related to the universe creation criteria (e.g., the country, market capitalization, industry, etc. of the stock) is stored in a separate database snapshot (e.g., stored in memory data) that is updated regularly, so that the information is quickly available to customers without having to be retrieved from a securities information database. In some embodiments, the domain-specific language itself can be used to define the universe.

[0097] Figure 5A An exemplary user interface 500 of a rapid screening system configured for equity screening, showing domain-specific flexible text matching and completion suggestions, is illustrated, according to one embodiment. Figure 2 As discussed, the rapid screening system can flexibly match inputs against domain-specific terms to provide accurate, real-time term matches even when the input contains errors or is otherwise not an exact match. For example, in the multi-line editor 501, the line 2 input 520 is "returonequit." This input omits the "n" at the end of "return" (or, according to another interpretation, includes an additional "o" in "return" and omits the word "on" entirely), and also omits the "y" at the end of "equity." However, the disclosed rapid screening system displays a data label suggestion pane and an explanatory text window 521 at the cursor, highlighting matching letters for possible intended input data labels. The multi-line editor 501 allows the user to select one of the provided options to immediately replace the unfinished and misspelled input. In contrast, previous systems such as spreadsheet formulas require character-complete text and unintuitive cell number cross-references that can be broken by typos, while general-purpose tools such as Spreadsheets are unable to provide semantic error handling and domain-specific term completion. By providing flexible text matching and domain-specific completion suggestions, the rapid screening system enables users to screen more easily and quickly than before.

[0098] Figure 5B An exemplary user interface 550 for a rapid screening system configured for equity screening, according to one embodiment, is illustrated, showing data tag exploration. A data exploration dialog box 551 allows a user to discover prospective input data tags for immediate use. In various fields, the number of available data tags can be enormous for a user, numbering in the thousands or tens of thousands, across unlimited data sources, and spanning tens or hundreds of thousands of securities. The exemplary data exploration dialog box 551 provides fields for a company name 555, search text 560, and a data package or data provider 565. In the illustrated example, a user searches for data tags related to the term "EBIT" (earnings before interest and taxes) for Microsoft Corporation, which are available from a data package named "S&P Global - Fundamental Data." In some embodiments, an auto-entry feature is used to assist with user-entered text to facilitate rapid discovery. Matching data tags are listed in a results box 570, which shows the name, description, and value (for the current quarter and the past twelve months) for each data tag. Other embodiments include additional attributes or even complete exemplary snippets. In some embodiments, the search text can match any information about the data tag, including the value (e.g., a specific annual growth rate). A convenient user interface element allows the user to select a desired data label by double-clicking 575 a cell in the result box 570 to lock it and copying the selected data label or labels using button 580. For example, in the illustrated example, double-clicking the 8.163900 cell 571 would lock the data label "Ebit10YrCagrPct", and if double-clicking the 8.566600 cell 572 would lock the data label "Ebit10YrCagrPctTtm", either or both of the data labels could be easily copied to the clipboard for pasting into an expression. By providing such domain-specific data exploration tools, the rapid screening system allows the user to discover relevant data labels, explore available data labels, and screen data more easily, quickly, and efficiently than previously possible.

[0099] Figure 6A 、 Figure 6B Illustrated are exemplary user interfaces 600A-600B of a rapid screening system configured for equity screening, showing filtering based on criteria, according to one embodiment. Figure 1 and Figure 2As discussed, the quick screening system allows users to filter, limit, reduce, or narrow down a set of identifiers based on expressive matching criteria. In the illustrated example, user interface 600A displays U.S. large-cap stocks in column 615 of grid view 602. Input 620 on row 2 of multi-line editor 601 is "ReturnOnEquityPctTtm," a data label representing the return on equity percentage for each company over the past 12 months. The value for "ReturnOnEquityPctTtm" is displayed in column 625 of grid view 602.

[0100] Turning to user interface 600B, input 622 on row 2 of multi-line editor 601 is now "~ReturnOnEquityPctTtm>40," with the tilde ("~") filter operator added at the beginning of the line and the comparison condition ">40" added after the data label. As a result, grid view 602 no longer displays all identifiers (or their stock symbols) in the specified universe of all U.S. large-cap stocks. Instead, only the stock symbols in column 617 for companies whose return on equity for the past twelve months was greater than 40 are included in the updated grid view 602 in column 627. Because the ReturnOnEquityPctTtm data for all identifiers displayed in column 615 has already been loaded by the system (and displayed in column 625), this restriction can be accomplished quickly.

[0101] In some embodiments, expression operators such as "~" (or "filter", "limit", "only", etc.) are required (e.g., to enhance user readability). In some embodiments, while expression operators may be allowed, a filtering operation may be inferred based on the presence of such an operator. Where filtering is inferred, or for all filtering operations, the quick screening system may distinguish between rows or statements (e.g., by modifying the statement or applying text formatting, background color, etc.; and / or by inserting expression operators where they were omitted by the user and inferred by the system).

[0102] In some embodiments, the rapid screening system provides additional filter-related operators, such as an "or" operator, a "not" operator, and / or enumeration (all x in Y), etc.

[0103] Figure 7 An exemplary user interface 700 of a rapid screening system configured for equity screening is illustrated, showing expressions assigned to custom variable names, according to one embodiment. Figure 1As discussed, the Quick Screening system allows users to create new custom variable names or data labels. For example, the Quick Screening system allows any data label to be given a new alias, and any expression (e.g., modifying the data value associated with a data label, or combining multiple data values) to be named so that it can be usefully labeled and easily referenced again. The name can be applied to, for example, a code snippet or an actual variable (so it does not need to be recalculated). Therefore, in conjunction with the above reference Figure 4 With the ability to define custom universes, users can extend the domain-specific language to meet their own needs, such as defining universes, data labels, data sources, and / or transformations.

[0104] Statement 740 illustrated in multi-line editor 701 assigns the sum of two data labels "ReturnOnEquityPct" and "RemrnOnAssets" to a custom named variable or data label "MyCustomIndex". Thus, column 745 displays the sum of the two data values (Return On Equity Percentage and Return On Assets) for each stock symbol displayed in grid view 702 under the heading "My Custom Index". In some embodiments, the quick screening system uses PascalCase (where all words of a compound name are capitalized) as a convention for naming data labels to facilitate user reading, and the result display interface (e.g., the header row in grid view 702) automatically adds spaces in appropriate locations in the data labels (e.g., between lowercase letters and subsequent uppercase letters) to improve readability. Other equivalent conventions for compound names without spaces in data labels include camelCase (where all words after the first word are capitalized), kebab-case (where dashes separate words), and snake case (where underscores separate words). In some embodiments, the quick screening system names variables using dot notation, such as "ReturnOnEquity.TTM" or using brackets. In some embodiments, spaces are allowed in data labels, or the system parses the input with spaces to identify unambiguous matching data labels (or suggests possible options if the input is ambiguous within the context of the domain-specific language and operators). By linking variable names to displayed output, this technique encourages non-programmers to naturally write maintainable "code" because they can tell whether the titles in the output look correct.

[0105] In some embodiments, the system includes a natural language user interface, such as a written text or speech recognition interface. Natural language processing (NLP) often fails because human language is not limited enough to obtain consistent and reliable results. Because the disclosed technology provides a domain-specific language as an intermediate layer, the results can be greatly improved. For example, a natural language request to "find all companies that mentioned China tariffs in the 10K report" can be processed by a language model (e.g., GPT-3) to generate an expression in a domain-specific language. Because the domain-specific language is semantically meaningful, humans can understand the output, verify that the expression is in accordance with the requirements of the natural language statement, and edit it if necessary. Because expressions in domain-specific languages are very expressive in a short space, they are easily verified by users, providing a more meaningful and more trustworthy interaction mode.

[0106] In contrast to most programming languages that declare variables and then assign values to them, in the illustrated embodiment, the domain-specific language includes assignment operators that operate after the result, psychologically encouraging the user to explore various expressions (which are immediately interpreted as the user types, amends, or replaces them) and then assign the results to the user-named variables (which are immediately displayed). This also ensures that the expression is "completed" more quickly: for example, while the formula or statement "A=awesome(param)" is not completed until the end of the statement, the expression "param|awesome=>A" is complete throughout its construction.

[0107] Figures 8A to 8C Exemplary user interfaces 800A, 800B, and 800C of a rapid screening system configured for equity screening according to one embodiment are illustrated, showing simultaneous renaming of multiple references to custom variable names. In the illustrated example, in input 830 in row 3 of the multi-line editor 801, the user has defined a custom variable or data label named "StandrRoa." Additionally, in input 840 in row 4 of the multi-line editor 801, the user has entered an expression that references the custom variable or data label "StandrRoa." Figure 8A , the quick screening system displays a data label suggestion pane and descriptive text window 841 at the cursor in row 4, highlighting "StandrRoa" as the identified user-defined variable. Figure 8B In the context menu 842 of the multi-line editor 801, the highlighted option provides the user with the ability to "change all occurrences" of a custom variable name. Figure 8CIn the example above, the multi-line editor 801 allows the user to change "StandrRoa" to "StandardRoa" in multiple locations simultaneously. Thus, the rapid screening system can make consistent global changes to variable names without breaking any variable references. In contrast, conventional in-domain search systems do not provide any ability to restructure text entries.

[0108] Figure 9 An exemplary user interface 900 of a rapid screening system configured for equity screening, showing domain-specific syntax error handling, is illustrated, according to one embodiment. Figure 2 As discussed with FIG5 , the rapid screening system is fault-tolerant when parsing input. For example, if the input is imperfect, the rapid screening system can attempt to infer the expected match in the domain-specific language and, for example, replace the unrecognized input with the closest valid match, or propose a set of possible replacements or supplements for the unrecognized input. In addition, malformed input (e.g., a syntax error such as a symbol that does not match any operator or data label in the domain-specific language) will not crash the rapid screening system or stop processing. In various embodiments, the rapid screening system continues to parse the remainder of the input (before and after the error) and process the symbols recognized by the parser, and can continue to display previously successful results.

[0109] In the illustrated example 900, a user enters a typographical error or incomplete input 920 "Returone" in line 2 of the multi-line editor 901, for example, after having previously entered a valid data label "ReturnOnEquityPct". The quick screening system displays a data label suggestion pane and an explanatory text window 921 at the cursor, highlighting matching letters for the input data label that may be expected. At the same time, the multi-line editor 901 displays a text warning 922 labeled "Error: Invalid syntax for line 2" (or, for example, "Error: No data label exists for line 2"), as well as margin text decorations 923 and a red diagonal underline to highlight the location of the unrecognized and unprocessed input 920. In some embodiments, error handling includes applying domain-specific information so that error messages are domain-specific.

[0110] Meanwhile, the rest of grid view 902 is unaffected; the universe of U.S. Large Cap stocks 915 remains displayed, along with the data labels and corresponding data values for the displayed identifiers in the rest of multi-line editor 901. In some embodiments, previously displayed data (e.g., "Return on Equity Pct" column 925) remains displayed until input that the quick screening system can process is entered in its place; in some embodiments, data for symbols that the quick screening system cannot process is not displayed.

[0111] Figure 10 An exemplary user interface 1000 of a rapid screening system configured for equity screening, according to one embodiment, is illustrated, showing transformation functions. In the illustrated example, a vertical bar symbol ("") in input 1020 in row 2 of multi-line editor 1001 represents a transformation function or operation. In this example, five transformation functions are listed in the operator suggestion pane and descriptive text window 1021:

[0112] average (when applied to a numeric array, produces, for example, the arithmetic mean);

[0113] rank (compare the data values of each identifier and rank them, high to low or low to high);

[0114] quintile (compares the data values for each identifier and sorts them into five buckets of 20% each, and produces quintiles);

[0115] standardize (compares the data value for each identifier to a standardized normal distribution and produces a z-score); and

[0116] trend stability (when applied to a numeric array, produces a number indicating whether the trend is positive, negative, or neither).

[0117] In various embodiments, the rapid screening system can provide additional or different transformation functions (e.g., median function, decile function, etc.). For example, a set of natural language processing transformation functions allows a user to enter an expression such as "NewsRecent["lawsuit"]|Sentiment=>LawsuitNewsSentiment" to determine the gist of news reports about companies involved in litigation.

[0118] Figure 11An exemplary user interface 1100 for a rapid screening system configured for equity screening, according to one embodiment, is illustrated, showing automatic graphical display of array data. In the illustrated example, input 1120 on line 2 of multi-line editor 1101 is "ReturnOnEquityPct[-11q:0q]," which is a data label representing the return on equity percentage for each company for each of the past 12 quarters (i.e., from 11 quarters ago to the present). In some embodiments, expressions involving data labels associated with arrays of values do not require explicit array notation, except to display a selected subset of the array data. For example, two data labels representing array values can be added together or otherwise manipulated in an expression. In some embodiments, alternative or shorthand notation such as "ReturnOnEquityPct[12q]" can compactly indicate the number of periods up to now (e.g., years, quarters, months, weeks, or days). In some embodiments, the system automatically aligns array data. For example, if a data label referencing time series A is divided by a data label referencing time series B, the system can automatically align the current value by forward filling.

[0119] In this example, the disclosed quick screening system displays the array data as user-friendly compact graphs 1125, with one graph corresponding to each identifier (with actual displayed stock symbols). The compact graphs generated by the quick screening system include shading above or below the horizontal axis for each data value in the result array. Thus, the quick screening system makes positive and negative values visually obvious and makes it easy for users to discern trends over time. In addition, user actions in one of the graphs in the compact graphs 1125 (such as hovering over the graph, right-clicking, or long-pressing the graph or a data point on the graph) can reveal one or more values in the underlying data values 1127 in the result array of the identifier. In some embodiments, the compact graph 1125 includes or displays key information, such as boundaries or x-axis labels (e.g., dates).

[0120] In some embodiments, the rapid screening system also provides similar functionality or scalability for user-defined functions and data labels. For example, data labels or functions can be scalable so that programmers can write custom code (e.g., custom visualization for complex data) that is the basis of a data label or function and make it available to non-programmers who use the data label or function. For example, in some embodiments, naming a custom variable or data label ending with "Score" can automatically set the format of the associated data value and color-code it according to quartiles, deciles, grading, etc. Similarly, in some embodiments, naming a custom variable or data label ending with "Trend" can automatically prompt the system to display any associated array data in the graph, and / or naming a custom variable or data label ending with "Pie" can automatically prompt the system to display any associated array data in a pie chart. In some embodiments, the system can be configured to draw any array data. In various embodiments, formatting the results and / or color-coding them can be implemented via any interface, even interactively. For example, in some embodiments, columns can be grouped in the display to display a group title above a group of subtitles.

[0121] Figure 12 An exemplary user interface 1200 of a rapid screening system configured for equity screening is illustrated, showing the automatic display of a link to a 10-K report, according to one embodiment.

[0122] Similar to Figure 11 , the quick screening system interface grid view 1202 includes columns 1225 displaying compact graphs representing arrays of data values corresponding to the data labels in input 1220 in row 2 of the multi-line editor 1201. In this case, input 1220 is "ReturnOnEquityPct[-3q:0q]," which is a data label representing the return on equity percentage for each company in each of the past four quarters (i.e., from three quarters ago to the present). The same input 1220 generates a "trend stability" transformation of the array data values, giving a numerical indicator 1245 of the trend in each case. As shown, the trend is upward because input 1240 in row 4 limits the displayed identifiers to those whose return on equity trend over the past four quarters was very positive (greater than 0.8). Row 3 is blank, illustrating the ability of the multi-line editor 1201 in the illustrated embodiment to seamlessly parse discontinuous input.

[0123] Another way to apply the disclosed technology to identify positive stability trends is to compare stability trends over two time periods. For example, the following expression creates a custom data label named "RoeStabilityPrevious" that represents the trend stability of the return on equity percentage from one year ago to four months ago, and a custom information label named "RoeStabilityRecent" that represents the trend stability of the return on equity percentage from four months ago to today:

[0124] (ReturnOnEquityPct[-11q:-4q])|TrendStability=>RoeStabilityPrevious

[0125] (ReturnOnEquityPct[-3q:0q])|TrendStability=>RoeStabilityRecent

[0126] By subtracting one trend from another, differences can be determined and filtered, for example:

[0127] (RoeStabilityRecent-RoeStabilityPrevious)=>RoeTrendDifference

[0128] ~RoeTrendDifference>1

[0129] Or, in an alternative, one can filter on the absolute values of the previous and current trends, for example, to identify companies whose ROE has recently surged:

[0130] ~RoeStabilityRecent>0.6

[0131] ~RoeStabilityPrevious<0.6

[0132] Input 1250 in row 5 of multi-line editor 1201 adds the data label "Filings10k" to grid view 1202, thereby displaying information about each company's most recent 10-K report filed with the U.S. Securities and Exchange Commission (SEC). In the illustrated embodiment, the 10-K report information includes the date of the most recently available report, as well as a hyperlink to the SEC's online copy of the report so that screeners can directly read or download the document. Thus, the rapid screening system can be configured to handle complex data types, such as JSON-formatted data and metadata, so that results can be displayed in a user-accessible format, thereby presenting screeners with a greater amount of useful information than previously possible.

[0133] Input 1260 in row 6 of the multi-line editor 1201 adds the data label "GicsSector" to the grid view 1202, thereby displaying the GICS sector for each displayed identifier (stock symbol).

[0134] Figure 13 An exemplary user interface 1300 for a rapid screening system configured for equity screening according to one embodiment is illustrated, showing a selective display of companies holding patents. In a multi-line editor 1301, input 1330 in row 3 and input 1340 in row 4 use the data label "PatentsIssued" to identify a company that has been granted one or more patents. For example, the rapid screening system can search a database of issued patents and identify patents whose owners or assignees match the company name (accommodating non-exact matches and / or related entities in certain specific implementations). In the illustrated example, in input 1330, the expression "PatentsIssued["autonomous vehicle", "autonomous driving"]" is interpreted as identifying a company that has been granted one or more patents containing the phrase "autonomous driving vehicle" or the phrase "autonomous driving" (which is assigned the data label AutonomousDrivingPatents). Similarly, the expression "PatentsIssued["electric vehicle"]" in input 1340 is interpreted by the rapid screening system as identifying a company (which is assigned the data tag ElectricVehiclePatents) that has been issued one or more patents containing the phrase "electric vehicle."

[0135] In the grid view 1302, column 1335 displays autonomous driving patents, and column 1345 displays electric vehicle patents, showing the patent titles and publication dates and providing links to these patent documents. In some embodiments, the quick screening system provides a total number of patents (or currently valid patents) found by the search. For example, the pure-play patent score for electric vehicles can be expressed as "PatentsIssuedCount["ElectricVehicles"] / PatentsIssuedCount" or a similar expression. In the illustrated example, input 1370 in row 7 of the multi-line editor 1301 shows filtering using the expression "~AutonomousDrivingPatents>0", and input 1380 in row 8 filters using the expression "~ElectricVehiclePatents>0" so that the companies included in the grid view 1302 (as listed in the company stock code column 1315) are only companies with at least one patent in each category. In other words, the company shown has both 1,375 autonomous driving patents and 1,385 electric vehicle patents (though not necessarily any autonomous electric vehicle patents).

[0136] In various embodiments, the quick screening system can be configured to obtain, link to, and filter content from a wide range of data sources. For example, a "NewsRecent" data tag can cause the grid view 1202 to display the most recent news headlines about each company from various news sources. Figure 12 The Form 10-K link and Figure 13 Patent titles for each news headline about a given identifier are linked to the full article, while usefully providing summary information directly in the search results. In some embodiments, the system allows the user to filter for recent news articles containing terms of interest. For example, an expression such as "NewsRecent["autonomous driving"]" can allow the user to identify companies with a small number of published patents in a field, compare their media engagement in that field, and vice versa. In this way, the rapid screening system enables the screener to directly observe active areas of corporate technological development or read news about targets of interest, and obtain a subjective impression of coverage about such targets to supplement numerical data analysis, or identify breaking news that may affect the market related to certain securities.

[0137] Figure 14 An exemplary user interface 1400 of a rapid screening system configured for equity screening is illustrated, showing filtering of text found in a 10-K report, according to one embodiment. Figure 12 , Figure 14The quick screening system in includes a compact trend chart and data label "Filings10k". However, Figure 14 Instead of providing a link to a Form 10-K report, the data tag is used to filter identifiers based on the contents of their 10-K reports. Specifically, in the illustrated example, input 1450 in row 5 of the multi-line editor 1401 shows filtering using the expression "~Filings10k contains "chinatariffs"". Thus, the list of companies 1415 in the results grid view 1402 is limited based on multiple criteria: US large-cap companies 1410 whose ROE has grown dramatically 1440, 1445 over the previous four quarters 1420, 1425, and whose last Form 10-K report mentioned China and price lists 1450, 1455.

[0138] In grid view 1402, "Report 10k" column 1455 displays relevant matching text from the 10-K, with matching search terms highlighted, rather than simply providing a link to the entire Form 10-K document. In some embodiments, the highlighted text provides a link to the source document, or more specifically, a link to the cited portion of the source document.

[0139] As mentioned above, in some embodiments, an operator similar to "contains" is implemented without alphabetic text. For example, using bracket notation, the syntax could be "Filings10k["china tariffs"]" as an equivalent example, similar to Figure 13 ” example in . In some embodiments, the “contains” operator is used as a transform, e.g., “Filings1ok|contains“china tariffs””, making this step explicit and allowing filtering of the results; the filtering itself can be implicit. By providing the ability to filter identifiers based on the content of documents such as Form 10-K reports or patents, the disclosed rapid screening system enables powerful new screening methods and provides the ability to synthesize across different data sets.

[0140] Figure 15 An exemplary user interface 1500 of a rapid screening system configured for equity screening is illustrated, showing grouping of results, according to one embodiment. Figure 12 , Figure 15 The quick screen system in includes a data label "GicsSector" in input 1560 of row 6 of multi-line editor 1501. Thus, the result grid view 1502 displays the GICS sector for each displayed identifier (stock symbol). However, in Figure 15 , the "Gics Department" column 1565 heading is used to group the identifiers according to their category data values.

[0141] In the illustrated embodiment, a "GICS Sector" heading is also displayed 1506 as a row group at the top of the grid view 1502. Thus, a series of GICS sector categories 1566 are listed on the left side of the grid view 1502. Each category can be expanded to display the identifiers of the companies in that sector, or collapsed to display the number of companies in that sector that meet the active global filter criteria. In this example, the criteria of strong recent ROE growth among U.S. large-cap companies yields a list of primarily information technology stocks. These interface features of the disclosed rapid screening system allow screeners to identify trends, explore strategies for identifying targets of interest, and easily see the impact of trying alternative strategies.

[0142] As demonstrated in each of the examples above, the flexible customizability, ease of use, rapid feedback, and power of the disclosed rapid screening system allow, for example, complex, expressive screens ("all companies with increasing and stable return on equity (RoE) over the past three years"); unique perspectives ("all companies that mentioned China tariffs in the past 10-K"); identifying discontinuities or inconsistencies in the data ("all companies whose brand popularity increased but whose stock prices remained flat"); discovering overall trends ("sectors whose popularity increased in earnings conference call Q&A"); and benchmarking ("how did MSFT perform across all the screens we built?"). Notably, screens can be combinable, so a user can reference one screen from another, or "unit test" a company across multiple screens. This makes it easy to build a mosaic of perspectives. The systems and methods of the present disclosure provide flexible, structured perspectives that were previously unavailable.

[0143] Figure 16A to Figure 16B Illustrated are an exemplary formula 1600A for a prior art system and a corresponding exemplary expression 1600B for a rapid screening system configured for equity screening, showing improved ease of use, according to one embodiment.

[0144] 17A to 17B Illustrated are exemplary user interfaces 1700A-1700B of a rapid screening system configured for equity screening showing backtesting, according to one embodiment. Figure 17A , the highlighted selection in the context menu 1772 in the multi-line editor 1701 provides the user with the ability to "backtest" the performance of a set of standards relative to past market results.

[0145] Backtesting refers to testing a model (such as a set of screening criteria) to determine which stocks are worth buying at any given time based on historical data. By applying the same criteria to past data, the screener can determine whether the strategy would have been effective at another time when different stock symbol identifiers may have met the data label criteria being tested. Backtesting allows the user to determine how a set of criteria would have performed if applied consistently over a period of time in the past. The backtester creates hypothetical positions based on the user's selected criteria, executes the user's strategy over time, and records the results. Prior to the present disclosure, backtesting was typically expensive and slow (hours or weeks) and was limited to limited data sets. The disclosed rapid screening system technology takes from a fraction of a second to five to ten seconds to generate a backtest result 1700B and can backtest strategies expressed using the full vocabulary of a domain-specific language without requiring extensive configuration or the use of separate tools. Reference below Figure 28 Exemplary concurrent server interactions for backtesting are described in more detail.In some embodiments, backtesting is implemented using a subset of the same code that the system uses for screening, which is made possible by the domain-specific context.

[0146] Figure 17B Shown Figure 17A A set of backtest results 1700B for the criteria shown in the multi-line editor 1701 of FIG. The illustrated backtest is equally weighted; in some embodiments, the quick screening system is configured to optionally run a factor-weighted backtest (rather than just equally weighted), such as by selecting (e.g., right-clicking) the factors to be weighted and then selecting a backtest. In various embodiments, backtest results 1700B include report generation that includes various methods for plotting and listing the results of the backtest. For example, a backtest can provide statistics on how investments selected according to selected screening criteria performed over a selected time period compared to one or more benchmarks and / or calculate returns over time. The disclosed technology provides unprecedented access to concurrent programming execution that is simply not available to users using conventional techniques.

[0147] Figure 18An exemplary change alert diagram 1800 of a rapid screening system configured for equity screening according to one embodiment is illustrated. In some embodiments, the rapid screening system can be configured to perform automatic daily screening and, for example, send the results to the user via email. For example, if there are any important changes (e.g., companies added or deleted from a list, new sources of risk, relevant news reports, etc.), the rapid screening system can automatically highlight them to save user time and reduce availability bias. In the illustrated chart 1800, a set of results for stock screening is shown with vertically listed identifiers and a digital horizontal axis. Change alert diagram 1800 provides a visualization of the change between yesterday's (blue circle) and today's (orange circle) values, allowing viewers to immediately see what has changed and how much.

[0148] In some embodiments, the technology provides the user with the ability to generate similar graphs comparing any two sets of screening results.

[0149] Figure 19A An exemplary AI feature 1900 for introspection in a rapid screening system configured for equity screening is illustrated, according to one embodiment. In the illustrated example, the disclosed technology uses artificial intelligence (AI) machine learning (ML) to introspect and improve existing screening practices.

[0150] First, identify factors currently used in existing screening models (e.g., to rank possible investments). For example, a hypothetical screening expressed in the domain-specific language of the disclosed rapid screening system might be (0.4*RetumOnEquity+0.4*RetumOnAssets+0.2*DebtToEquity)=>CustomFactor, where existing screeners rank possible investments from high to low based on CustomFactor to prioritize their queries based on what is most important to them.

[0151] To train a black-box AI model to replicate existing practice, we asked the question: “Knowing that the inputs are ROE, ROA, and DTE, and the corresponding input data and results, can we build a model that replicates their screening results?” We then trained a number of ML models with sample inputs identical to those used by existing models, existing ratings, and outputs of predicted ratings to verify that the model reflects current practice. After training a model, we can introspect it (i.e., look inside the black box) to determine the relative feature importance to the model. Introspection can determine, for example, that factors are redundant or accidentally overweighted. For example, a user may believe that ROE and ROA are equally weighted in their screening process; but if a model can be built that predicts the output of the user’s expected model by using primarily ROA, AI Features 1900 can help the user understand how their screening process actually works and improve it, or provide alternative ways to obtain the same results.

[0152] One of the most difficult parts of training an ML model is the feature selection process. Often, the expert user who selects features is different from the data scientist or quantitative researcher who implements the model. The disclosed technology allows inferring a screener's feature selection parameters from the domain-specific language used by the user when constructing an existing screen. Analyzing the use of this domain-specific language is impossible in any existing screener because there are no other similarly accessible screeners that fully express what the user wants, nor are there sufficiently structured or constrained languages in the existing language to allow us to infer semantic meaning from the user's use of the tool.

[0153] Another challenge is overfitting, particularly when working with financial data. For introspection use cases, it doesn't matter if the resulting ML model is overfitted, as we're not actually predicting future events, merely analyzing existing ones. By exploiting the properties of overfitted models, the disclosed techniques depart from conventional practice and teaching. Overfitting is generally considered undesirable and would be extremely counterintuitive to those of ordinary skill in the art.

[0154] Figure 19B Illustrated is an exemplary AI feature 1950 for making predictions in a rapid screening system configured for equity screening, according to one embodiment. This is different from the above with respect to Figure 19AMethods for improving AI screening using introspective ML models. In this case, the method uses machine learning to train a set of classifiers that take as input the universe of companies at a point in time and a selected subset of exemplary companies. The exemplary companies can be a list of individually selected single identifiers or identifiers that meet certain filtering criteria (e.g., ROE is above a threshold, or the company name contains "hotel"). The exemplary companies in the subset can be "preferred" targets, companies with characteristics that the screener wishes to avoid, or other categories. Methods for developing ML models can include bootstrapping, random forests, AutoML, or other ML models, preferably models that allow the trainer to infer relative feature weights. The trainer can exert some control over the resulting screening model, for example, by training more models and averaging them. The classifiers are trained to produce exemplary companies; they all achieve the same solution in different ways. Therefore, they can predict filters that will find those exemplary companies. This can reveal user preferences by showing which data labels and weights are correlated or anti-correlated. In some implementations, the result is the generation of probability distributions and feature weights.

[0155] Unlike typical recommendation engines that are trained on a subset of a predefined set of similar things and then operate on another assumed similar subset, this trained model creates a layer of abstraction, describing a scope that users may believe they are familiar with (e.g., a selected company in a selected industry). In addition to helping quantify, the model also allows users to apply their understanding of that scope to other scopes or universes, such as other industries, countries, and / or time.

[0156] These classifiers can be used individually or as part of an ensemble. Ensembles allow us to reduce overfitting to the prediction context. And some models allow us to further introspect the specific weights. Other implementations include combining classifiers generated from different "runs" of the same algorithm. For example, the resulting prediction module (combination of classifiers) can be applied in a variety of different ways, or it can enable the user to use the machine to adjust or iterate on the feature weights.

[0157] Figure 20 An exemplary AI feature 2000 for detecting regime changes in a rapid screening system configured for stock screening, according to one embodiment, is illustrated. In finance, regime changes are associated with sudden changes in financial market behavior, such as might be associated with cycles of economic activity between expansion and recession. Understanding regime changes is important because it enables investment managers to react to systemic market changes.

[0158] Institutional change models are typically defined as time series models in which parameters are allowed to take different values in each of a fixed number of "institutions." Institutional change models typically include some predefined set of factors they seek, along with some fixed definition of institutions. The key ideas are often defining what an institution is and what characteristics influence it. However, two problems with current approaches are:

[0159] 1. While models can tell you about likely regime changes given our historical and predefined understanding of institutions, the important question for investors to answer is not “Is the world changing?” but “Is the change in the world significant for this investment strategy?”

[0160] 2. They require predefined input variables and are not adaptable to changes in the forces affecting the market.

[0161] We can continue to refer to the above Figure 19A Building upon the AI / ML approach described above to solve the problem of custom system change detection. Figure 19A Based on AI, we infer and construct custom models of regime change without explicit instructions from investment managers. We perform the same process, but this time we can observe multiple windows sequentially rather than a single snapshot. Therefore, we build a set of ML models for each time period (e.g., quarterly).

[0162] Again, it doesn't matter if the ML model we built is overfitting. The output is the impact of the defined factors / features over time. If the impact of a feature that is important to your investment process changes, it's an early warning sign that the factors you rely on may have changed, and you may need to consider adjusting your approach.

[0163] The ability to infer institutional change models from this small amount of information currently doesn't exist. The combination of language expressibility and semantic structure / constraints, along with the application of ML, helps enhance this capability. The disclosed technique allows domain expert decision makers to combine this with their own judgment and leverage AI to iterate together to improve outcomes.

[0164] Return Reference Figure 19B, predictive AI models can be applied to use intuition about past known regimes to infer similar choices to make in the current regime. For example, if a user believes that today's regime is similar to the one companies experienced in 1990, the model can be used to highlight companies that are similar to a subset of the companies selected in 1990 (e.g., companies that thrived under that regime). That is, a model trained to predict targets of interest in 1990 can now be applied today. Thus, by inputting a set of names and dates, the output will be a set of similar names with scores associated with them in the current and set screenings.

[0165] As another example, the current state of the world is often assessed based on whether stocks are currently favored to outperform the market based on factors such as “value,” “growth,” “quality,” and “momentum.” But a user may be unsure whether a quality category, for example, actually applies to the “universe” of that user’s world (e.g., a limited subset of securities). For example, finding companies that performed well in 2007 and / or 2003 would require too much data to remember; therefore, traditional investors must fall back on intuition. AI models can essentially codify this intuition: where the idea that “this feels like a time in the past” is not concrete, the model makes “now feels like then” concrete and quantifiable.

[0166] Furthermore, this AI model approach can be applied across industries or countries, as well as dates or asset classes (e.g., the U.S. market in 2003 versus the Chinese market today), enabling transfer learning and generating insights about cross-connections that might otherwise remain undiscovered.

[0167] The disclosed AI model training does not simply use the past to predict the future via an unadjustable black box; it adds user input of experience-based intuition about the past, quantifies the regime (defined on the fly, not pre-categorized), defines a predictive model based on that regime, and provides a basis for identifying characteristics that worked well then and, if the user's intuition was correct, may also work well now. The model provides recommendations of current companies that are similar to companies of past interest in past environments (e.g., based on their performance / returns or other qualities), rather than trying to tell you what predefined regime you are in today, while not providing actionable insights (e.g., which companies you might consider buying given similar environments).

[0168] In some embodiments, relevant "features" for training a model can be inferred from the code specified in the editor (e.g., by analyzing data labels (including user-defined data labels) and expressions). Typically, a data science or machine learning expert would perform this feature engineering. The disclosed technology enables non-technical users who are not familiar with machine learning to effectively collaborate with machine learning AI screening models.

[0169] This facilitates a continuous feedback loop between the user and the AI screening model, where the user can effectively perform feature engineering to improve the AI screening model by changing the text in the multi-line editor. In this way, the user can effectively iterate on the specified features fed into the model and better find targets of interest.

[0170] Figure 21 Illustrated is an exemplary AI feature 2100 for optimizing mixing in a rapid screening system configured for equity screening, according to one embodiment.

[0171] Based on the above Figure 19A and Figure 20 We use these introspective and / or institutional change insights to optimize future screening models based on the disclosed methods. If ROE outperforms—that is, if screeners consider ROE and ROA to be equally important, but the ML model infers that ROE might identify more interesting companies—we can highlight companies in your generated list that you might want to focus more on.

[0172] Similarly, based on the above Figure 19B and Figure 20 The disclosed method of modeling screening of high performing securities in your portfolio can be applied to identify additional screening criteria and / or weightings by finding a set of screens to generate the names of these securities.

[0173] In a variation of the above method, the predictive AI model can also be applied to obtain herd and / or minimum consensus risk estimates for names in a screen. Given the input of a screen and a date, the output is two lists of identifiers: the names that appear least frequently in other screens, and the names that appear most frequently in other screens. The first represents minimum consensus risk; the second represents herd risk.

[0174] So, if we generate, say, 100 screens to help you find this group of companies (i.e., if there are 100 ways to get to the same solution), then if 75 of those screens say MSFT is a good choice, then the payoff depends on your confidence that MSFT is the kind of company you want to buy. Does it represent what you value? Or do you reflect and question how independent your view is? In other words, if the screens reveal that most of what you do is what traditional growth investors do, but a small percentage is your "secret sauce," then you have the ability to introspect and act accordingly to adjust your biases, your views, and / or your behavior.

[0175] In another variation, predictive AI models can be used to identify near misses: for example, a stock ticker that didn’t appear in a user’s initial screen or portfolio, but did appear in many AI-learned screens. Such a company might need to take action to investigate.

[0176] Figure 22 Illustrated is an exemplary AI feature 2200 for feature suggestions in a rapid screening system configured for equity screening, according to one embodiment.

[0177] Based on the above Figure 19A 、 Figure 20 and Figure 21 With the disclosed method, we can recommend features to add to a set of screening criteria (e.g., data labels) as well as features to remove. For example, we add one or more features that are just random noise, set it as a baseline for comparing the impact of other factors in the existing screening criteria, and recommend removing features that fall below that threshold. For example, if the impact of DTE in our shadow model is less than random noise, one might think that the signal is being incorporated, but in fact it is not. Because we have a strong semantic understanding of what the user is trying to accomplish, we can analyze statically and / or dynamically to make recommendations to improve the performance of their strategy.

[0178] The disclosed AI model and interface provide a simple yet powerful level of abstraction. Because user input is domain-specific, we gain more context about the type of problem the user wishes to solve. User input allows for inference of user preferences within the domain. The interface is non-technical and interactive, encouraging users to iterate based on their results and easily perform feature engineering.

[0179] Furthermore, these AI models contrast with common quantitative modeling, where the objectives are often very rigidly defined. The disclosed models allow us to target qualitative and / or difficult-to-describe objectives, making them feasible and useful for discretionary investors who rely more on experience.

[0180] Figure 23An exemplary user interface for a rapid screening system configured for equity screening, according to one embodiment, illustrates the creative use of operators. In a multi-line editor 2301, line 3 input 2330 applies a "RankHighToLow" transformation based on the current return on assets of US large-cap companies 2310 and assigns the ranking metadata to a variable "RoaRank" 2335. Line 6 expression 2360 combines a ternary operator, an emoji, and string concatenation to produce an easy-to-read Roa indicator column 2365. The ternary operator functions like a short if-then statement of the form "(X?Y:z)," where if the expression X is true, the output of the ternary operator is Y, otherwise it is z. Thus, for any company with an Roa rating of 75 or lower, the result is a green checkbox emoji, and for any company with an Roa rating of 76 or higher, the result is a red X-mark emoji. The output emoji is concatenated with a space character and the return on assets value to display the Roa indicator column 2365, highlighting the stocks with the highest Roa among US large-cap stocks.

[0181] Figure 24 An exemplary user interface 2400 of a rapid screening system configured for equity screening, according to one embodiment, is illustrated, showing automatic formatting. In the illustrated example, applied to the universe of US large-cap stocks 2410, three calculations determine a "value" factor 2450 / 2455, a "growth" factor 2460 / 2465, and a "quality" factor 2470 / 2475. In rows 10, 11, and 12, the three factors are each subjected to a "SplitQuintiles" transformation 2451, 2461, and 2471, generating metadata that assigns a quintile of 1 to 5 to each identifier's factor score.

[0182] The quintile scores for each factor are assigned to variable or custom data labels ending in "Score." Therefore, the quick screening system specifically treats these values by displaying the values associated with those variables in grid view 2402 using "heat map" color coding. As shown, the "1" quintile score is displayed in a dark red cell; the "2" quintile score is displayed in a dark orange cell; the "3" quintile score is displayed in a yellow cell; the "4" quintile score is displayed in a light green cell; and the "5" quintile score is displayed in a dark green cell.

[0183] This and other automatic formatting of displayed values provides a contrast to conventional screening systems and makes displayed results easier to understand for untrained users while providing greater flexibility and expressiveness.

[0184] Figure 25An exemplary user interface 2500 of a rapid screening system configured for equity screening, showing a forecasted point-in-time situation report, according to one embodiment, is illustrated. The time series situation report model not only provides exploration and screening capabilities, but also provides insight and powerful analysis that leverages analysts' forecasts over time and compares them to a company's actual performance.

[0185] In the illustrated example, in multi-line editor interface 2501, in row 12510, the symbol "# / Model / SituationReport" invokes the situation report mode or model (many other equivalent methods may be used, such as buttons, drop-down menus, voice commands, etc.). Some controls are not shown, including an input field for entering a list of companies or global identifiers (in this case, the five companies displayed in the "Company Name" column 2515) and a date field with a calendar drop-down list. Multi-line editor interface 2501 includes a set of data labels (such as "EbitConsensusMean" 2530) representing analysts' expectations in several areas. This data label reflects the consensus mean of analysts' earnings before interest and taxes forecasts for the companies.

[0186] The contents of each cell in the situation report model are not simple values, but rather summaries of complex data sets. To generate the summary situation report, the quick screening system generates calculations for each cell in the grid display 2502 to download and process the underlying information. The data is interpolated and smoothed before being displayed. The displayed values represent the current (for the selected calendar date) forecast rate of change for the selected indicator for the selected company. For example, in the illustrated example, for Bed Bath & Beyond Company 2531, the Ebit consensus average 2535 has a value of 53 2545. This indicates that the consensus average has a slightly positive slope, meaning that the forecast EBIT will remain the same or increase slightly.

[0187] In the illustrated example, cells are colored along a red-yellow-green spectrum (based on standard deviation, to match the human eye / expectation) based on the slope of the smoothed predicted curve. Thus, values between 0 and 10 correspond to the fastest declines, while values between 90 and 100 correspond to the fastest increases relative to the maximum absolute slope. The consensus average 2535 for Microsoft Corporation's 2532 EBIT is very high at 95 2546 and is therefore colored a bright green, indicating that the consensus forecast is that Microsoft's EBIT will continue to grow sharply. As of the date of this disclosure, the computational complexity of the analysis shown would make it impossible to run on a personal computer.

[0188] In addition to displaying current trends, the Outlook report provides controls that allow for previously unavailable navigation, synthesis, and contextual awareness of the analyst's forecasts. For example, by adjusting the calendar date, the user can easily make forecasts for previous times and compare how the status of the consensus forecast has changed or is changing. In addition, a green arrow 2580 or a red arrow 2585 indicates whether the company's actual forecast has exceeded or fallen short of the expected consensus forecast, which may indicate a change. In addition, each cell is a link to a larger graph showing more complete information over time, as referenced below. Figure 26 Further described.

[0189] Figure 26 An exemplary user interface 2600 of a rapid screening system configured for equity screening according to one embodiment is illustrated, showing a situation report historical forecast graph 2601. When a user selects Figure 25 When cell 2545 is selected in the chart 2601 is shown and shows the details behind the numbers shown in the initial situation report table. The chart 2601 includes a stepped line 2610 which represents the actual consensus forecast for Bed, Bath & Beyond's EBIT over the displayed timeframe. In addition, a smoothed line 2620 is shown which is also referenced above. Figure 25 The source of the slope calculation described above. Around or near the actual and smoothed forecasts is the shaded area 2640, which represents the forecast interval of the analyst's forecast. When the actual forecast 2610 is outside the forecast interval 2640, the situation report table ( Figure 25 ) marks a deviation because it may indicate a change from a previously determined expectation. Legend 2650 shows the exact value for a given date.

[0190] The trend report chart is applicable to any time series data, not just estimates. As applied to the estimates in the illustrated example, it highlights large changes in analyst forecasts that exceeded the forecast, as well as consistent changes. For example, the dy (smoothed) chart extends along the lower half of the chart, showing the rate of change of the forecast. Thus, in the illustrated example, the analyst's EBIT is expected to rise by 2630 around October 2018 and fall by 2635 around November 2019. This makes it easier to identify inflection points in the analyst's expectations without requiring coding or other special technical expertise from the user. This visualization of the consensus forecast reduces the learning curve and shortens the time to insight by adding companies to the list, allowing users to get immediate results. Therefore, the disclosed technology can make it easy for users to see warnings, perhaps even before the market reacts. Other embodiments enable users to screen this information in the same manner as described above.

[0191] Figure 2727 is a block diagram illustrating some components that are typically incorporated into computing systems and other devices that can implement the present technology. In the illustrated embodiment, computer system 2700 includes a processing component 2730 that controls the operation of computer system 2700 according to computer-readable instructions stored in memory 2740. Processing component 2730 can be any logical processing unit, such as one or more central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc. Processing component 2730 can be a single processing unit or multiple processing units in an electronic device, or distributed across multiple devices. Aspects of the present technology can be embodied in a dedicated computing device or data processor that is specially programmed, configured, or constructed to perform one or more of the computer-executable instructions described in detail herein.

[0192] Aspects of the present technology may also be practiced in a distributed computing environment, where functions or modules are performed by remote processing devices that are linked through a communications network, such as a local area network (LAN), a wide area network (WAN), or the Internet. In a distributed computing environment, modules may be located in both local and remote memory storage devices. In various embodiments, the computer system 2700 may include one or more physical and / or logical devices that collectively provide the functionality described herein. In some embodiments, the computer system 2700 may include one or more replicated and / or distributed physical or logical devices. In some embodiments, the computer system 2700 may include one or more computing resources provided from a "cloud computing" provider, such as those provided by Amazon.com, Inc. of Seattle, Washington. Elastic Compute Cloud(“Amazon ), Amazon Web Services (“ ”) and / or Amazon Simple Storage Service TM (“Amazon S3 TM ”); Google Cloud Platform provided by Google Inc., Mountain View, CA TM and / or Google Cloud Storage TM ; Windows provided by Microsoft Corporation in Redmond, Washington etc.

[0193] Processing component 2730 is connected to memory 2740, which may include temporary and / or permanent storage devices and a combination of both read-only memory (ROM) and writable memory (e.g., random access memory or RAM, CPU registers, and on-chip cache memory), writable non-volatile memory such as flash memory or other solid-state memory, hard drives, removable media, magnetically or optically readable disks and / or tapes, nanotechnology memory, synthetic biology memory, and the like. Memory is not a propagating signal that is independent of the underlying hardware; therefore, memory and computer-readable storage media do not themselves refer to transient propagating signals. Memory 2740 includes data storage devices containing programs, software, and information, such as an operating system 2742, applications 2744, and data 2746. Computer system 2700 operating system 2742 may include, for example Android TM 、 Applications 2744 and data 2746 may include software and databases (including data structures, database records, other data tables, etc.) configured to control computer system 2700 components, process information (e.g., to optimize program code data), communicate and exchange data and information with remote computers and other devices, etc.

[0194] The computer system 2700 may include an input component 2710 that receives input from user interaction and provides input to the processor 2730, typically mediated by a hardware controller that interprets the raw signals received from the input device and transmits information to the processor 2730 using a known communication protocol. Examples of input components 2710 include a keyboard 2712 (with physical or virtual keys), a pointing device (such as a mouse 2714, a joystick, a dial, or an eye-tracking device), a touch screen 2715 that detects contact events when touched by a user, a microphone 2716 that receives audio input, and a camera 2718 for still photo and / or video capture. The computer system 2700 may also include various other input components 2710, such as a GPS or other location-determining sensor, a motion sensor, a wearable input device with an accelerometer (e.g., a wearable glove-type input device), a biometric sensor (e.g., a fingerprint sensor), an optical sensor (e.g., an infrared sensor), a card reader (e.g., a magnetic stripe reader or a memory card reader), and the like.

[0195] The processor 2730 may also be connected to one or more various output components 2720, for example, directly or via a hardware controller. The output device may include a display 2722 on which text and graphics are displayed. The display 2722 may be, for example, an LCD, LED, or OLED display screen (such as a desktop computer screen, a handheld device screen, or a television screen), an electronic ink display, a projection display (such as a head-up display device), and / or a display integrated with a touch screen 2715 that acts as an input device and an output device that provides graphical and textual visual feedback to the user. The output device may also include a speaker 2724 for playing audio signals, a tactile feedback device for tactile output such as vibration, and the like. In some implementations, the speaker 2724 and the microphone 2716 are implemented by a combined audio input-output device.

[0196] In the illustrated embodiment, the computer system 2700 also includes one or more communication components 2750. The communication components may include, for example, a wired network connection 2752 (e.g., one or more of an Ethernet port, a cable modem, a Thunderbolt cable, a FireWire cable, a Lightning connector, a Universal Serial Bus (USB) port, etc.) and / or a wireless transceiver 2754 (e.g., one or more of a Wi-Fi transceiver, a Bluetooth transceiver, a near field communication (NFC) device, a wireless modem or cellular radio utilizing GSM, CDMA, 3G, 4G, and / or 5G technology). The communication components 2750 are adapted to communicate between the computer system 2700 and other local and / or remote computing devices directly via a wired or wireless peer-to-peer connection and / or indirectly via communication links and networking hardware such as switches, routers, repeaters, cables and optical fibers, optical transmitters and receivers, radio transmitters and receivers, etc. (which may include the Internet, a public or private intranet, a local or extended Wi-Fi network, a cellular tower, a plain old telephone system (POTS), etc.). Computer system 2700 also includes a power supply 2760 , which may include battery power and / or utility power for operating the various electrical components associated with computer system 2700 .

[0197] Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with the technology include, but are not limited to, personal computers, server computers, handheld or laptop devices, cellular phones, wearable electronic devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. Although a computer system configured as described above is typically used to support the operation of the technology, those skilled in the art will appreciate that the technology can be implemented using devices of various types and configurations and having various components. It is not necessary to show such infrastructure and specific implementation details to describe exemplary implementations.

[0198] Figure 28 FIG28 is a schematic and data flow diagram illustrating several components of an exemplary concurrent server interaction for backtesting, according to one embodiment. A user interacts with the rapid screening system, such as via a web application or other client device 2810 (e.g., running local software or accessing Software as a Service (SaaS)). The user requests a backtest 2805, for example, as described in detail above with reference to FIG17. The rapid screening system can operate similarly for less complex results, such as modifying screening criteria and other portfolio analysis. The client web application or device 2810 sends a request 2815 to a session server 2820 (e.g., a host device that manages a user database on a remote server). Request 2815 may include, for example, subscriber authentication or session information to ensure the subscriber is authorized (and, for example, to charge for backtesting services), and / or a date and code specifying backtest parameters (such as the standard to be tested and the benchmark to be tested). In various embodiments, the parameters are already associated with the user's session and available on the server, so identification of the current session and backtest request can be accomplished with minimal data transfer. In some embodiments, the client interface device or web application 2810 sends one or more triggers to the session server 2820 for when a new screening, backtesting, or other portfolio analysis request occurs, and subscribes to receive updates when the analysis is complete.

[0199] When the user has been authenticated by session server 2820, session server 2820 sends a request 2825 to a main function 2830, running, for example, on a server such as one or more cloud computing instances. Main function 2830 manages the processing and assembly of backtesting requests as a whole, including using concurrent computing resources (e.g., triggering lambda functions 2840) to execute requests to efficiently parallelize backtesting calculations. Lambda functions are simple but can accomplish a lot because the system can spawn many tasks in parallel, executing a large number of tasks simultaneously in a short period of time. For example, in backtesting, a fast screening system can break down calculations into discrete time periods, such as each year in the backtest and the individual months within each year. This allows for significant improvements in backtesting speed. In various embodiments, a domain-specific language is designed for concurrency and data sharding (e.g., across time, data items, and user-specific schemas).

[0200] In the illustrated embodiment, lambda function 2840 obtains data from one or more financial data databases 2850, where materialized view 2845 contains snapshots across data and time. In some embodiments, materialized view 2845 snapshots also include a recent set of historical values, which may result in some duplication but enables, for example, extremely fast and comparable access to customized trends over the past few years. Materialized view example data 2848 includes a date, an identifier, a pair of values (for the past twelve months and the current quarter), and a pair of sets of historical values (for the past twelve months and the current quarter). In various embodiments, materialized view 2848 is saved to speed up access when running the same or another backtest; this allows users to switch back and forth between the results of different backtests without recalculating. In various embodiments, screening and / or backtesting functionality can be plugged into a variety of data stores and is not limited to database 2850; for example, an application programming interface (API) can also be effectively used.

[0201] The lambda function 2840 returns to the main function 2830, and the main function 2830 delivers the full results 2855 of the backtest or other analysis to the session server 2820. In various embodiments, after running the query, the full results 2855 are stored on the server side. This means that if the data itself does not change (e.g., classification, grouping, etc.), the fast screening system does not need to constantly rerun the query, and also opens up the possibility of incremental querying.

[0202] In some embodiments, lambda function 2840 is updated with the results to which client interface device or web application 2810 has subscribed. In some embodiments, session server 2820 provides result identifier 2857 to client interface device or web application 2810. Client interface device or web application 2810 can then request partial results 2865 (e.g., only the results necessary to render a visual result on the user's display). Session server 2820 then delivers the requested portion of results 2875 to client interface device or web application 2810.

[0203] Figure 29 29 is a schematic diagram illustrating several components of an exemplary server system for implementing a rapid screening system according to one embodiment. A user interacts with the rapid screening system, such as via a web application or other client device 2910. The web application 2910 communicates with a session server 2920 via, for example, server endpoints 2905, 2908 to request operations such as screening, backtesting, or performing time series forecasting analysis. The server processes requests to the server endpoint 2905 on a 1:1 basis, with the session server 2920 inserting one row per request into a session database, for example. For example, when processing a screening session for a user who enters an expression, the web application 2910 provides the entered expression to the session server 2920, which can initiate processing of the expression. For example, whenever a row is updated, added, or removed in the editor, the web application 2910 notifies the session server 2920, which initiates a request to the backend database 2940 or the compiler / interpreter 2930. In some embodiments, backtesting requests from the web application 2910 to the session server 2920 are also processed on a 1:1 basis.

[0204] The server processes requests to server endpoint 2908 on a 1:many basis, with session server 2920, for example, inserting multiple rows into the session database for each request. For example, in response to a time series analysis request, session server 2920 can initiate a separate process for each combination of company and forecast (i.e., each cell in the results table), rather than launching a single process for the entire request. In some embodiments, the web application 2910 client initiates a series of separate requests; in some embodiments, session server 2920 receives the analysis request and generates the requested subtasks. In either case, the separate tasks can produce a set of results on a per-cell basis, rather than, for example, a single JSON blob, allowing the results to be assembled asynchronously.

[0205] Other components of the session server 2920 include an authentication service 2921 to ensure that the requester is authorized and that the data is secure; the authentication service 2921 can be linked to user data 2922. The user database 2922 can additionally store the user's code (e.g., to ensure that screening expressions are automatically saved and available across sessions), preferences, and / or authentication credentials, and can provide a caching layer. The session data service 2923 manages the user's session data, including caching screening results 2934. For example, for a user who is running a screening, when the user writes code in a domain-specific language, the session data service 2923 can insert a row for the code in the session database. In some embodiments, the session data service polls for updates. When the requested data is loaded (e.g., from the compiler / interpreter 2930), the session data service 2923 caches the data and updates the web application 2910 with the results.

[0206] Additionally, session server 2920 may store context data 2924. Context 2924 may include, for example, custom universes (each including, for example, the name of the universe, parameters defining the universe (e.g., minimum market capitalization, country, specific company name, etc.)) and / or custom indexes (each including, for example, the name of the index and a time series of returns (which a user may upload)). Context data 2924 may be shared between users to allow for easier collaboration and minimize duplication or synchronization issues.

[0207] In the illustrated embodiment, session server 2920 is separate from compiler / interpreter 2930. Storing session data separately can offer advantages. For example, it can allow data to be more easily audited for user analytics and / or telemetry. Furthermore, it can report which data is being used most and by whom, not just by which package, and can also report on granular data such as field mappings, which helps reduce costs. Thus, the structure of the disclosed technology enables introspection and analysis that is typically not possible when running conventional backend database queries.

[0208] Furthermore, because the compiler / interpreter 2930 is a separate server, and because the domain-specific language provides context about what the user is trying to accomplish, and the expected results are verifiable and reproducible, the disclosed system enables system administrators to make language feature updates without the traditional constraints of programming languages. This enables faster delivery of new features to users, because the structure of the domain-specific language and its contextual understanding can allow administrators to ensure that they will never break (which is typical for new versions of traditional languages).

[0209] Compiler / interpreter 2930 is the workhorse responsible for executing actual requests for screening searches, backtesting, and / or analytical modeling. This diagram is a simplified logical representation; those skilled in the art will appreciate that, by way of example, compiler / interpreter 2930 may include distributed computing resources to perform the required processing, load balancing / queue management, and the like. The illustrated separation of session server 2920 from compiler / interpreter 2930 also benefits compiler / interpreter 2930, which does not need to understand the user, but does need to understand the data source.

[0210] In the illustrated embodiment, the screening model 2935 contains or references a grammar 2931, which it uses to parse the user's code (e.g., as a whole) into an abstract syntax tree (AST) 2932. The screening model 2935 evaluates each line 2933, including the expressions in each line (e.g., in order). For each line, it interprets the AST. The screening model 2935 connects to one or more databases (including caches, if available) to obtain data and successively builds a result set 2934 for each line. The screening model 2935 inserts the result set 2934 into the session data service 2923 result database.

[0211] In the illustrated embodiment, backtesting model 2936 is run through backtest-specific logic 2937 and shared screening model 2935 components. Backtesting model 2936 utilizes screening model 2935 as well as various database optimizations (including snapshots and offloading some calculations closer to the database). When a backtest is run, it is typically rebalanced over time according to a set of buy / sell criteria. These criteria can be expressed as screening criteria. Therefore, if a user writes a screening, they can immediately run a backtest based on that screening code with a single click, without any additional configuration.

[0212] The compiler / interpreter 2930 then leverages its semantic understanding of the user's intent (they want to run a backtest) to optimize the code before it is executed as part of the rebalancing step in the backtesting model 2936. For example, if a data label or expression is determined to be unused within the context of the backtest (perhaps a leftover column from an exploratory screening session, or calculated solely for display purposes to the end user), the backtesting model 2936 can remove or otherwise exclude that code before executing the backtest, as it will not meaningfully affect the filtering criteria or factor weights for the backtest. In contrast, backtests run using general-purpose programming languages lack this semantic understanding (i.e., certain data being pulled down or certain analysis being performed does not have any meaningful impact on the actions the user is currently attempting to perform), and therefore such optimizations cannot be automatically performed using conventional means.

[0213] In various embodiments (e.g., for running backtests spanning multiple years), the backtesting model 2936 breaks the backtest into subcomponents and parallelizes them 2938 (e.g., by month, or calculating returns, etc.).

[0214] Additional models 2939 (e.g., time series analysis processing models) are similarly executed by compiler / interpreter 2930 and produce result sets that are sent to session server 2920. With minimal context required, the same domain-specific language applies to all models provided by compiler / interpreter 2930.

[0215] In some embodiments, to facilitate caching, a request may be run through the compiler / interpreter 2930 twice: the first time to collect and prefetch data tags, and the second time to actually evaluate the request. This may enable the data store to effectively cache the results of the first run.

[0216] The back-end data may include both an unoptimized data repository 2945 and an optimized database 2940. For example, commonly accessed data may be optimized for fast retrieval, particularly for the most common queries. The optimized database 2940 components may include, for example, lightweight read-access tables 2941 for different data sets, and point-in-time snapshots 2942 (e.g., for "today" and "end of month," dating back to a specific year) to reduce query times. Unstructured data (e.g., 10-K reports) may be processed to become more optimized. However, optimization is not binary; for example, patent information includes large amounts of unstructured text, but reports can also be processed to link them to companies (rather than just to static data), such as via name similarity. Therefore, such databases do not have to have exact matches to security names or identifiers, and they can still be "joined" in other ways.

[0217] Optimization not only provides tools for improving performance, but also provides advantages provided by the structure of the disclosed system. Typically, the person running queries against a dataset (e.g., running a backtesting request) has no control over the underlying database, which is a large data repository that must satisfy queries from users with different goals and handle general-purpose programming languages that access the database. In contrast, the disclosed technology imposes significant constraints on usage and access patterns while maintaining a high degree of expressiveness, an optimization that allows for high-performance results and a more user-friendly interface.

[0218] In some embodiments, the data repository and other elements of the system are fully extensible and open to integration. For example, remote access to any database with API compliance is possible, and any provider can implement the interface. The data itself is language-specific to the domain, but if the data server adheres to the specified interface, it "just works." Similarly, extensibility can be applied to: sets of data tags (in a standardized format, whether defined in a domain-specific language, customized, exported, or even self-referencing); additional transformations; custom report templates (e.g., results of backtesting or screening); results (e.g., returned in a standardized JSON format); views (e.g., displaying an array of strings as a bulleted HTML list); and even models (e.g., providing a separate session database endpoint for each model). For example, views can be passed back to the web application 2910 in the form of HTML or data structures, which allows data to be presented specifically based on its content or context. Thus, for example, URLs can be displayed as links, arrays of values can be viewed as charts, and the values of variables whose names end with "Score" or "Ranking" can be displayed with a heat map color gradient. In essence, any data tag, any function, any model, and any report can be integrated into the disclosed fully extensible system. While the basic operators for the domain-specific language can remain the same, new data types, processing, and transformations can be customized for different domains or customers.

[0219] While specific implementations have been illustrated and described herein, those skilled in the art will appreciate that alternative and / or equivalent implementations may be substituted for the specific embodiments shown and described without departing from the scope of the present disclosure.

[0220] For example, although various embodiments are described above in terms of rapid screening systems and / or services provided by one or more servers to remote clients, in other embodiments, screening methods similar to those described herein may be employed locally on a client computer to find and display results within a local or remote corpus.

[0221] Similarly, although various embodiments are described above in terms of screening systems or services capable of screening stocks or other securities (e.g., debt securities, virtual currencies, etc.), other embodiments may use similar techniques to allow screening or filtering of data in another professional field, such as real estate, advertising, medical diagnosis, pharmaceutical research, employment, fantasy sports, movies, scientific data, photos, etc. For example, backtesting can simulate different drugs under a given set of conditions; a screening system applied to the field of geological data can improve the search for oil drilling sites; a system applied to real estate can find undervalued home purchase opportunities; a system applied to cancer screening can provide an improved ability to immediately perform complex analysis in a sample and identify trends and promising avenues of research. In other embodiments, various other applications of the disclosed technology may be made. This application is intended to cover any modifications or variations of the embodiments discussed herein.

Claims

1. A computing device for interpreting a domain-specific language for interactive exploration, filtering, and analysis of dynamic data sets, the device comprising a processor and a memory storing instructions that, when executed by the processor, configure the device to: Provide a multi-line editor input interface that allows users to enter and edit input on any line at any time; Provides a grid view display interface; On a continuous basis, when the user enters input including a first expression in the multi-line editor input interface: parsing the first expression with respect to the domain-specific language, wherein the domain-specific language comprises a plurality of data tags and a plurality of operations, each data tag being associated with at least one value for each of a plurality of identifiers of the data set, and each operation being applicable to the value associated with the data tag; wherein the parsing comprises identifying data tags in the domain-specific language and operations provided in the domain-specific language in the first expression; Executing the parsed expression, wherein the executing comprises: determining a first subset of identifiers from the plurality of identifiers to which the parsed expression is applied; identifying one or more data sources associated with the identified data tags; loading a value associated with the recognized data tag from at least one of the identified data sources; and applying the operation to the loaded value to produce result values of a second subset of identifiers, wherein each result value is associated with an identifier of the second subset of identifiers; and updating the grid view display interface with the real-time display of the result values for the second subset of identifiers, wherein the updating comprises, for each displayed identifier, adding an associated value of the result value to the grid view display interface, The grid view display interface is immediately updated according to the current content of the multi-line editor input interface.

2. The computing device of claim 1 , further comprising: User input specifying a universe of identifiers is received via the multi-line editor input interface, such that the first subset of identifiers to which the parsed expression is applied is determined to be the specified universe of identifiers.

3. The computing device of claim 1 , further comprising: An expression including a previous expression is received on multiple lines in the multi-line editor input interface, such that the first subset of identifiers to which the parsed expression is applied is determined to be a subset of identifiers derived from the previous expression.

4. The computing device of claim 1, wherein the first expression comprises data tags that reference output from an existing screen, such that the result value of the second subset of identifiers is based on the output from the existing screen. 5 . The computing device of claim 1 , wherein the identifying one or more data sources associated with the recognized data tags is performed automatically without requiring the user to explicitly specify a data source. 6 . The computing device of claim 1 , further comprising assigning the resulting values of the second subset of identifiers to new automatically named data tags.

7. The computing device of claim 1, wherein the operating comprises assigning the resulting values of the second subset of identifiers to new user-named data tags.

8. The computing device of claim 7, further comprising: Another expression containing the new user-named data tag is parsed such that the new user-named data tag is used as one of the plurality of data tags in the domain-specific language.

9. The computing device of claim 1, wherein the operation produces a result value for each identifier in the first subset of identifiers such that the second subset of identifiers contains the same identifiers as the first subset of identifiers.

10. A computing device according to claim 9, wherein the operation is a transformation that produces a result value, the result value includes metadata characterizing the value associated with the data tag, and wherein the transformation is one of: averaging by mean or median, grading, sorting by quintiles or deciles, standardizing, or indicating a trend or trend stability.

11. The computing device of claim 1 , wherein the operation filters the first subset of identifiers to which the expression is applied so that the second subset of identifiers is a proper subset of the first subset of identifiers, the proper subset containing fewer identifiers than the first subset of identifiers; and wherein, Updating the grid view display interface further includes displaying the second subset of identifiers in place of the first subset of identifiers, or removing identifiers not included in the second subset of identifiers from the grid view display interface. 12 . The computing device of claim 1 , wherein each expression or data label entered in the multi-line editor input interface corresponds to a column of information displayed in the grid view display interface.

13. The computing device of claim 12 , wherein updating the grid view display interface comprises inserting a column of the result value in the grid view display interface; and further comprises removing the corresponding column from the grid view display interface when the user deletes a row or a data label from the multi-row editor input interface.

14. The computing device of claim 1 , wherein updating the grid view display interface comprises: Displaying a heat map of values in the grid view display interface based on transformations or data label names.

15. The computing device of claim 1 , wherein the data tags are associated with structured data or a value array for each identifier, and wherein updating the grid view display interface comprises determining a type of the structured data or value array and automatically displaying a graphic of consecutive values or a link to a source document in the grid view display interface.

16. The computing device of claim 1 , wherein parsing, executing, and updating on a continuous basis comprises: When the user edits the first expression in the multi-line editor input interface, the displayed result value is changed to reflect the user's editing.

17. A computer-readable storage medium having instructions stored thereon, wherein when executed by a processor, the instructions configure the processor to: Providing a multi-line editor user input interface that allows the user to enter and edit input on any line at any time; Provides a grid view display interface; On a continuing basis, while the user enters input comprising a first expression in the multi-line editor user input interface: parsing the first expression relative to a domain-specific language, wherein the domain-specific language comprises a plurality of data tags and a plurality of operations, each data tag being associated with at least one value for each of a plurality of identifiers of a data set, and each operation being applicable to a value associated with the data tag; wherein the parsing comprises identifying data tags in the domain-specific language and operations provided in the domain-specific language in the first expression; Executing the parsed expression, wherein the executing comprises: determining a first subset of identifiers from the plurality of identifiers to which the parsed expression is applied; identifying one or more data sources associated with the identified data tags; loading a value associated with the recognized data tag from at least one of the identified data sources; and applying the operation to the loaded value to produce result values of a second subset of identifiers, wherein each result value is associated with an identifier of the second subset of identifiers; and updating the grid view display interface with the real-time display of the result values for the second subset of identifiers, wherein the updating comprises, for each displayed identifier, adding an associated value of the result value to the grid view display interface, The grid view display interface is immediately updated according to the current content of the multi-line editor user input interface.

18. The computer-readable storage medium of claim 17, further comprising a natural language user interface configured to generate the first expression in the domain-specific language based on a natural language input and insert the first expression into the multi-line editor user input interface.

19. The computer-readable storage medium of claim 17, further comprising: Provides user-selectable options to perform backtesting; And in response to the user selecting the option: performing the backtest, the backtest selecting securities based on parsed expressions of historical data; wherein the multi-line editor user input interface continues to allow the user to enter and edit input on any line at any time; and Wherein performing said backtesting does not affect parsing, execution and updating on a continuous basis.

20. The computer-readable storage medium of claim 17, wherein in response to a typographical error or incomplete input in the first expression, the multi-line editor user input interface displays semantically correct error handling suggestions.

Citation Information

Patent Citations

  • Information searching and access authorization method

    CN101221566A

  • Expanded search and find user interface

    CN101263495A