Regular Expression Generation Based on Positive and Negative Pattern Matching Examples
By generating regular expressions to automatically process and format large-scale data sets, the problem of inefficient data preprocessing in the prior art is solved, and efficient and accurate data processing is achieved.
Patent Information
- Application Number
- CN201980035772.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-11
- Filing Date
- 2019-06-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-06-12
AI Technical Summary
The prior art is inefficient in processing preprocessing and formatting of large-scale data sets, and it is difficult to effectively process data from different data sources and formats, resulting in poor signal-to-noise comparison of data analysis and affecting the accuracy of the results.
Extract, reformat, or modify data by generating regular expressions to automatically identify and match patterns in the data. Specific methods include generating regular expressions using the longest general subsequence (LCS) algorithm, processing positive and negative examples, and receiving input data through the user interface to generate efficient regular expressions.
It realizes rapid and automatic generation of regular expressions, improves the efficiency and accuracy of data processing, can effectively process data in different formats and structures, and reduces the need for manual intervention.
Smart Images

Figure CN112262390B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority under 35 U.S.C.§119(e) to U.S. Provisional Patent Application No. 62 / 684,498, filed on June 13, 2018, entitled “AUTOMATED GENERATION OF REGULAR EXPRESSIONS”, and this application also claims the benefit of priority under 35 U.S.C.§119(e) to U.S. Provisional Patent Application No. 62 / 749,001, filed on October 22, 2018, entitled “AUTOMATED GENERATION OF REGULAR EXPRESSIONS”. The entire contents of U.S. Provisional Patent Application Nos. 62 / 684,498 and 62 / 749,001 are incorporated herein by reference for all purposes. Background of the Invention
[0003] Big data analytics systems can be used for predictive analytics, user behavior analytics, and other advanced data analytics. However, before any data analytics can be effectively performed to provide useful results, the initial data set may need to be formatted into a clean and curated data set. This data loading typically poses challenges to cloud - based data repositories and other big data systems, where data from various different data sources and / or data streams can be compiled into a single data repository. Such data can include structured data in a variety of different formats, semi - structured data according to different data models, and even unstructured data. Repositories of such data typically include data representations in a variety of different formats and structures, and may also include duplicate data and error data. When analyzing these data repositories for reporting, predictive modeling, and other analytics tasks, a poor signal - to - noise ratio of the initial data set can lead to inaccurate or useless results.
[0004] Many current solutions to the data formatting and pre - processing problem include manual and ad - hoc processing to clean and organize the data so that it is processed into a common format before performing data analytics. While these manual processes may be effective for some smaller data sets, such processes can be inefficient and impractical when attempting to pre - process and format large - scale data sets. Summary of the Invention
[0005] Aspects described herein provide various techniques for generating regular expressions. As used herein, a "regular expression" may refer to a character sequence that defines a pattern, which can be used to search for matches within a longer input text string. In some embodiments, a symbolic wildcard matching language can be used to compose regular expressions, and the patterns defined by the regular expressions can be used to match strings and / or extract information from strings provided as input. In various embodiments described herein, a regular expression generator implemented as a data processing system can be used to receive and display input text data, receive a selection of a particular subset of characters of the input text via a client user interface, and then generate one or more regular expressions based on the selected subset of characters. After generating one or more regular expressions, a regular expression engine can be used to match the patterns of the regular expressions against one or more data sets. In various embodiments, data that matches the regular expressions can be extracted, reformatted, modified, etc. In some cases, additional columns, tables, or other data sets can be created based on the data that matches the regular expressions.
[0006] According to certain aspects described herein, a regular expression generator implemented via a data processing system can generate regular expressions based on a determined longest common subsequence (LCS) shared by different sets of one or more regular expression codes. Regular expression codes (which can also be referred to as category codes) can include, for example, L for letters in the English alphabet, N for digits, Z for spaces, P for punctuation, and S for other symbols. Each set of one or more regular expression codes can be transformed from a different sequence of one or more characters received as input data via a user interface. Regular expression codes that are excluded from the LCS can be represented as optional and / or alternative. In some embodiments, a regular expression code can be associated with a minimum number of occurrences of that regular expression code. Additionally or alternatively, a regular expression code can be associated with a maximum number of occurrences of the regular expression code. For example, a set of category codes can include L<0,1> to indicate that a particular portion of the LCS includes letters at most once (if any). As discussed in more detail below, generalizing the input data into an intermediate regular expression code (IREC) can provide various technical advantages, including using very little input data, enabling regular expressions to be generated almost instantaneously that are not affected by false positive or false negative matches in data not yet seen.
[0007] According to additional aspects described herein, a regular expression can be generated based on input data that includes three or more character sequences. When three or more character sequences are identified as input data, a regular expression generator that identifies the longest common subsequence (LCS) of the character sequences can result in an exponential increase in run time. To identify the LCS of all character sequences in a high-performance manner, the regular expression generator can perform the LCS algorithm on each different combination of two character sequences. A fully connected graph can be generated based on the results of the LCS algorithm, where each graph node represents a different character sequence and the length of each graph edge corresponds to the LCS of the nodes that define the graph edge. Then, the order of selection of the character sequences can be determined by performing a depth-first traversal on a minimum spanning tree of the fully connected graph.
[0008] Other aspects described herein relate to generating a regular expression based on input that includes both positive character sequence examples and negative character sequence examples. Positive examples can refer to character sequences that match the regular expression to be generated, while negative examples can refer to character sequences that do not match the regular expression to be generated. In some embodiments, when both positive and negative examples are received, the regular expression generator can identify discriminators or shortest subsequences of one or more characters that distinguish the positive examples from the negative examples. The discriminator selected can be the shortest sequence (e.g., represented by a class code) and can be positive or negative such that positive examples will match and negative examples will not match. Then, the discriminator can be hard-coded into the regular expression generated by the regular expression generator. In some cases, the shortest subsequence can be included in the prefix or suffix portion of the negative example.
[0009] Additional aspects described herein relate to one or more user interfaces through which input data can be provided to generate regular expressions. In some embodiments, the user interface can be displayed at a client device communicatively coupled to a regular expression generator server. The user interface can be programmatically generated by the server, by the client device, or by a combination of software components executing at the server and the client. Input data received through the user interface can correspond to a user's selection of one or more character sequences, which can represent positive or negative examples. In some cases, the user interface can support input data that includes the selection of a first character sequence within a second character sequence. For example, the user can highlight one or more characters within a larger previously highlighted character sequence, and a second user selection can provide context for the larger first user selection. This enables the input data to be provided to the regular expression generator in a more targeted manner and provides "context" to the regular expression generator such that it can generate regular expressions that avoid false positives. In response to a user selecting a character sequence via the user interface, the regular expression generator can generate and display a regular expression. For example, when the user highlights a first character sequence, the regular expression generator can generate and display a regular expression that matches the first character sequence and other similar character sequences (e.g., in accordance with the user's intent for the matching sequence). When the user highlights a second character sequence, the regular expression generator can generate an updated regular expression that includes the first and second character sequences. Then, when the user highlights a third character sequence (e.g., within the first or second sequence), the regular expression generator can update the regular expression again, and so on. Figure 1 When the user highlights the second character sequence, the regular expression generator can generate an updated regular expression that includes the first and second character sequences. Then, when the user highlights a third character sequence (e.g., within the first or second sequence), the regular expression generator can update the regular expression again, and so on.
[0010] According to additional aspects described herein, regular expressions can be generated based on the longest common subsequence from one or more input sequence examples and can also handle characters that occur only in some of the examples. To handle characters that occur only in some of the input examples, spans can be defined, where the minimum and maximum occurrences of the regular expression code are tracked. In cases where the span may not exist in all given input examples, the minimum occurrence can be set to zero. These minimum and maximum numbers can then be mapped to the regular expression heavy syntax. The longest common subsequence (LCS) algorithm can operate on spans of characters obtained from the input examples, including "optional" spans (e.g., of minimum length zero) that do not occur in every input example. As discussed below, consecutive spans can be merged during the execution of the LCS algorithm. In such cases, the LCS algorithm can also operate recursively on these optional spans when the additional optional spans carried are no longer consecutively present.
[0011] Other aspects described herein relate to combinatorial search, where the LCS algorithm performed by a regular expression generator can be run multiple times to generate a "correct" regular expression (e.g., a regular expression that correctly matches all given positive examples and appropriately excludes all given negative examples), and / or to generate multiple correct regular expressions from which the most desirable or best regular expression can be selected. In some embodiments, the LCS algorithm can generally be performed on the input examples from right to left to generate a regular expression. However, for comparison purposes and to find alternative regular expressions, the LCS algorithm can be performed backward (e.g., in the left-to-right direction) on the input examples separately. For example, an example character sequence received as user input can be reversed before they are run through the LCS algorithm, and then the result from the LCS algorithm can be reversed back (including the original text fragments). Additionally, in some embodiments, the LCS algorithm can be run multiple times by the regular expression generator, in the normal character sequence order and the reverse order, anchored at the beginning of a line, anchored at the end of a line, and not anchored at the beginning or end of a line. Thus, in some cases, the LCS algorithm can be performed at least six times like this, and the shortest successful regular expression can be selected from these executions. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a block diagram showing components of an exemplary distributed system for generating regular expressions, in which various embodiments can be implemented.
[0013] Figure 2 is a flowchart showing a process for generating a regular expression based on input received via a user interface, according to one or more embodiments described herein.
[0014] Figure 3 is a flowchart showing a process for generating a regular expression using the longest common subsequence (LCS) algorithm on a set of regular expression codes, according to one or more embodiments described herein.
[0015] Figure 4 is an example diagram showing a process for generating a regular expression using the longest common subsequence (LCS) algorithm on a set of regular expression codes based on two example character sequences, according to one or more embodiments described herein.
[0016] Figure 5 is a flowchart showing a process for generating a regular expression using the longest common subsequence (LCS) algorithm on a larger set of regular expression codes, according to one or more embodiments described herein.
[0017] Figure 6An example diagram that generates a regular expression based on five character sequence examples using the Longest Common Subsequence (LCS) algorithm on a set of regular expression codes, according to one or more embodiments described herein.
[0018] Figure 7 A flowchart that shows a process for determining an execution order for the Longest Common Subsequence (LCS) algorithm on a larger set of regular expression codes, according to one or more embodiments described herein.
[0019] Figure 8A and 8B A fully connected graph and a minimum spanning tree representation of the fully connected graph, which are shown for determining an execution order for the Longest Common Subsequence (LCS) algorithm on a larger set of regular expression codes, according to one or more embodiments described herein.
[0020] Figure 9 A flowchart that shows a process for generating a regular expression based on positive and negative character sequence examples, according to one or more embodiments described herein.
[0021] Figure 10A and 10B An example user interface screen that shows a regular expression generated based on positive and negative character sequence examples, according to one or more embodiments described herein.
[0022] Figure 11 A flowchart that shows a process for generating a regular expression based on a user data selection received within a user interface, according to one or more embodiments described herein.
[0023] Figure 12 A flowchart that shows a process for generating a regular expression and extracting data based on capture groups via a user data selection received within a user interface, according to one or more embodiments described herein.
[0024] Figure 13 An example user interface screen that shows a table data display, according to one or more embodiments described herein.
[0025] Figure 14 and 15 An example user interface screen that shows a regular expression and capture groups generated based on a data selection from a table display, according to one or more embodiments described herein.
[0026] Figure 16A and 16B An example user interface screen that shows a regular expression generated based on positive and negative example selections from a table display, according to one or more embodiments described herein.
[0027] Figure 17 According to one or more embodiments described herein, another example user interface screen is shown that generates a regular expression and capture groups based on a data selection from a table display.
[0028] Figure 18 According to one or more embodiments described herein, a flowchart is shown of a process for generating a regular expression including an optional span using a Longest Common Subsequence (LCS) algorithm.
[0029] Figure 19 According to one or more embodiments described herein, an example diagram is shown of generating a regular expression including an optional span using a Longest Common Subsequence (LCS) algorithm.
[0030] Figure 20 According to one or more embodiments described herein, a flowchart is shown of a process for generating a regular expression based on a combined execution of a Longest Common Subsequence (LCS) algorithm.
[0031] Figure 21 A block diagram is shown of components of an exemplary distributed system in which various embodiments of the present invention may be implemented.
[0032] Figure 22 A block diagram is shown of components of a system environment in which services provided by embodiments of the present invention may be provided as cloud services.
[0033] Figure 23 A block diagram is shown of an exemplary computer system in which embodiments of the present invention may be implemented. Detailed Description
[0034] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments of the present invention. However, it will be apparent to one of ordinary skill in the art that the embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form.
[0035] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the exemplary embodiments will provide those skilled in the art with an implementation description for implementing the exemplary embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the invention as set forth in the appended claims.
[0036] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, one of ordinary skill in the art will understand that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.
[0037] Note also that each embodiment can be described as a process which is depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process terminates when its operations are completed, but there may be other steps not included in a figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0038] The term "computer-readable medium" includes, but is not limited to, non-transitory media such as portable or fixed storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. A code segment or computer-executable instruction may represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0039] In addition, the embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks may be stored in a machine-readable medium. The processor may execute the necessary tasks.
[0040] This document describes various techniques (e.g., methods, systems, non-transitory computer-readable memories storing multiple instructions executable by one or more processors, etc.) for generating regular expressions corresponding to patterns identified in one or more input data examples. In certain embodiments, in response to receiving a selection of input data, one or more patterns in the input data are automatically identified, and a regular expression (or simply "regex") can be automatically and efficiently generated to represent the identified pattern. Such a pattern can be based on a character sequence (e.g., a sequence of letters, numbers, spaces, punctuation marks, symbols, etc.). Various embodiments are described herein, including methods, systems, non-transitory computer-readable storage media storing programs, code, or instructions executable by one or more processors, etc.
[0041] In some embodiments, a symbol wildcard matching language can be used to compose regular expressions in order to match strings and / or extract information from strings provided as input. For example, the first example regular expression [A-Za-z]{3}\d?\d,\d\d\d\d can match certain dates (e.g., April 3, 2018), and the second example regular expression [A-Za-z]{3}\d?\d,(\d\d\d\d) can be used to extract the year from a matched date. The input data received by a regular expression generator system can include, for example, one or more "positive" data examples, and / or one or more "negative" data examples. As used herein, a positive example can refer to a character sequence received as input that is to match a regular expression generated based on that input. In contrast, a negative example can refer to an input character sequence that is not to match a regular expression generated based on that input.
[0042] Many technical advantages can be achieved within the various embodiments and examples described herein. For example, certain techniques described in this disclosure can improve the speed and efficiency of the regular expression generation process (e.g., a regex solution can be generated in less than a second, and the user interface can be suitable for interactive real-time use). The various techniques described herein can also be deterministic, can not require training data, can produce a solution without any initial regular expression input, and can be fully automated (e.g., generate a regular expression without any human intervention). Additionally, the various techniques described herein do not limit the types of data inputs that can be effectively processed, and such techniques can improve the human readability of the resulting regular expressions.
[0043] Some embodiments described herein include one or more executions of the Longest Common Subsequence (LCS) algorithm. The LCS algorithm can be used as a difference engine in certain contexts (e.g., the engine behind the Unix "diff" utility), which is configured to determine and display the differences between two text files. In some embodiments, input data (e.g., strings and other character sequences) can be converted into abstract tokens, and then the abstract tokens can be provided as input to the LCS algorithm. Such abstract tokens can be, for example, tokens based on regular expression codes (e.g., Loogle codes or other character class codes) representing regular expression character classes. Various different examples of such codes are possible and may be referred to herein as "regular expression codes" or "intermediate regular expression codes" (IREC). For example, the input character sequence "May 3" can be converted into the IREC code "LLLZN", and then the tokenized string along with other tokenized strings can be provided to the LCS algorithm. In some embodiments, IREC (e.g., regular expression codes) that the input character sequence does not commonly possess can appear as an option (e.g., optional spans) in the finally generated regular expression. In certain embodiments, the regular expression code can be a category code based on the UNICODE category codes shown at https: / / www.regular-expressions.info / unicode.html#category. For example, the code L can represent letters, the code N can represent digits, the code Z can represent spaces, the code S can represent symbols, the code P can represent punctuation, and so on. For example, the code L can correspond to Unicode\p{L}, and the code N can correspond to Unicode\p{N}. This allows for a one-to-one mapping from the LCS output to the regular expression (e.g., \pN\pN\pZ\pL\pL can match "10am"), which may be beneficial for human readability. Additionally, these different categories may be disjoint or mutually exclusive. That is, in this example, the categories L, N, Z, P, and S can be disjoint, such that there may be no overlap between the members of the categories.
[0044] Additional technical advantages can be achieved in various embodiments, including more efficient generation of regular expressions based on the use of regular expression codes (e.g., category codes), spans, etc. By using such codes, no computational resources need to be wasted when the LCS algorithm successfully identifies all or substantially all characters in the input string as different. Other technical advantages provided by various embodiments herein include improved readability of the generated regular expressions, as well as support for positive and negative examples of the input data, and providing various advantageous user interface features (e.g., allowing the user to highlight text segments within larger character sequences or data units for extraction).
[0045] I. General Overview
[0046] Various embodiments disclosed herein relate to the generation of regular expressions. In some embodiments, a data processing system configured as a regular expression generator can generate a regular expression by identifying the longest common subsequence (LCS) shared by a set of different regular expression codes (e.g., category codes). Each set of regular expression codes can be transformed from a character sequence received as input data through a user interface. Among the technical advantages described herein, abstracting the input data into intermediate codes (e.g., regular expression codes, spans, etc.) enables the efficient generation of regular expressions using very little input data.
[0047] Figure 1FIG. 0 is a block diagram showing components of an exemplary distributed system for generating regular expressions, in which various embodiments may be implemented. As shown in this example, client device 120 may communicate with regular expression generator server 110 (or regular expression generator) and interact with a user interface to retrieve and display tabular data, and generate a regular expression based on a selection of input data (e.g., examples) via the user interface. In some embodiments, client device 120 may communicate with regular expression generator 110 via client web browser 121 and / or client-side regular expression application 122 (e.g., a client application that receives / consumes regular expressions generated by server 110). Within regular expression generator 110, requests from client device 120 may be received at a network interface via various communication networks and processed by an application programming interface (API) such as REST API 112. The user interface data model generator 114 component with regular expression generator 110 may provide server-side programming components and logic to generate and present various user interface features described herein. Such features may include the functionality of allowing a user to retrieve and display tabular data from data repository 130, select input data examples to initiate generation of a regular expression, and modify and / or extract data based on the generated regular expression. In this example, regular expression generator component 116 may be implemented to generate regular expressions, including converting an input character sequence into regular expression code and / or spans, performing algorithms (e.g., LCS algorithm) on the input data, and generating / simplifying regular expressions. The regular expressions generated by regular expression generator 116 may be transmitted by REST service 112 to client device 120, where the Javascript code on client browser 121 (or corresponding client application component 122) may then apply the regular expressions to each cell in the spreadsheet columns presented in the browser. In other cases, a separate regular expression engine component may be implemented on the server side to compare the generated regular expressions with the tabular data displayed on the user interface and / or other data stored in data repository 130 in order to identify matching / mismatching data on the server side. In various embodiments, matching / mismatching data may be automatically selected (e.g., highlighted) within the user interface and may be selected for extraction, modification, deletion, etc. Any data extracted or modified via the user interface based on regular expression generation may be stored in one or more data repositories 130. Additionally, in some embodiments, the generated regular expressions (and / or corresponding inputs to the LCS algorithm) may be stored in regular expression library 135 for future retrieval and use. In some embodiments, the generated regular expressions do not actually need to be stored in a "library" but may be incorporated into a "transformation script".For example, as described in more detail in U.S. Patent No. 10,210,246, which is incorporated herein by reference for all purposes, such a transformation script can include programs, code, or instructions executable by one or more processing units to transform received data. Other possible examples of transformation script actions can include "rename columns", "uppercase column data", or "infer gender from name and create a new column with gender", etc.
[0048] Figure 2 is a flowchart showing a process of generating a regular expression based on an input received via a user interface, according to one or more embodiments described herein. In step 201, the regular expression generator 110 can receive a request from the client device 120 to access the regular expression generator user interface and view specific data via the user interface. The request in step 201 can be received via the REST API 112 and / or a web server, an authentication server, etc., and the user's request can be parsed and authenticated. For example, a user within an enterprise or organization can access the regular expression generator 110 to analyze and / or modify transaction data, customer data, performance data, prediction data, and / or any other category of data that can be stored in the organization's data repository 130. In step 202, the regular expression generator 110 can retrieve and display the requested data via a user interface that supports generating a regular expression based on the selected input data. Various embodiments and examples of such a user interface are described in detail below.
[0049] In step 203, the user can select one or more input character sequences from the data displayed in the user interface provided by the regular expression generator 110. In some embodiments, the data can be displayed in tabular form within the user interface, including labeled columns with specific data types and / or data categories. In such a case, the selection of input data in step 203 can correspond to the user selecting a data cell, or selecting (e.g., highlighting) a single text segment within the data cell. However, in other embodiments, the regular expression generator 110 can support retrieving and displaying semi-structured and unstructured data via the user interface, and the user can select the input data for regular expression generation by selecting a character sequence from the semi-structured or unstructured data. As described in the examples below, the user selecting an input character sequence from the displayed tabular data is just one example use case. In other examples, the user (e.g., a software developer or advanced user who may be trying to write regular expressions for Linux command line tools such as grep, sed, or awk, etc.) can input examples from scratch rather than pick them from a spreadsheet.
[0050] In step 204, the regular expression generator 110 may generate one or more regular expressions based on the input data selected by the user in step 203. In step 205, the regular expression generator 110 may update the user interface, for example, to display the generated regular expressions and / or highlight the matching / mismatching data within the displayed data. In step 206 (which may be optional in some embodiments), the user interface may support the function of allowing the user to modify the underlying data based on the generated regular expressions. For example, the user interface may support the following features: allowing the user to filter, modify, delete, or extract specific data fields from tabular data based on whether these fields match the regular expression. Filtering or modifying the data may include modifying the underlying data stored in the repository 130, and in some cases, the extracted data may be stored in the repository 130 as a new column and / or a new table.
[0051] Although these steps illustrate a general and high-level overview of an exemplary user interaction with the user interface of the regular expression generator 110, various additional features and functions may be supported in other embodiments. For example, in some embodiments, a regular expression code (or category code) may be associated with a minimum occurrence count of the code. Additionally or alternatively, the regular expression code may be associated with a maximum occurrence count of the code. As an example, a set of regular expression codes may include the code L<0,1> to indicate that a particular part of the LCS includes the letter at least zero times and at most one time.
[0052] Furthermore, in some embodiments, the input data may include three or more character sequences. In such embodiments, techniques may be used to determine the order in which to perform the LCS algorithm on the three or more character sequences, so that the resulting regular expression can be generated in a high-performance manner to avoid an exponential increase in the running time caused by three or more input character sequences. Instead, the regular expression generator 110 may perform the LCS algorithm on two character sequences at a time and may determine the order for selecting the pair of character sequences based on a graph. For example, a fully connected graph may indicate that a first execution of the LCS algorithm (e.g., LCS1) should be performed on sequence 1 and sequence 3, and then a second execution of the LCS algorithm (e.g., LCS2) should be performed on LCS1 and sequence 2, and so on. The graph may be a fully connected graph having nodes representing the character sequences and edges connecting the nodes to represent the length of the LCS shared by the connected nodes. Each node in the graph may be connected to every other node in the graph, and the order for selecting the character sequences may be determined by performing a depth-first traversal of the minimum spanning tree of the graph.
[0053] In other embodiments, input data can be provided in a variety of different ways via a user interface. For example, the input data can indicate a first user selection of one or more characters within a second user selection of a character set. For example, a user can highlight characters within a previously highlighted set of characters. Thus, the second user selection can provide context for the first user selection, which can enable the input data to be provided to the regular expression generator 110 in a more targeted manner. In some embodiments, the regular expression generator 110 can generate and display a regular expression in near real-time in response to each user selection. For example, when a user highlights a first character range, the regular expression generator 110 can display a regular expression representing the first character range. Then, when the user highlights a second character range within the first character range, the regular expression generator 110 can update the displayed regular expression.
[0054] Additionally, in some embodiments, the regular expression generator 110 can generate a regular expression based on inputs that include positive and negative examples. As described above, positive examples can refer to character sequences to be included in the regular expression, while negative examples can refer to character sequences that will not be included in the regular expression. In such cases, the regular expression generator 110 can identify the shortest subsequence of one or more characters that separates the positive examples and the negative examples at a particular location. Then, the shortest subsequence can be hard-coded within the regular expression generated by the regular expression generator 110. In various examples, the shortest subsequence can include a prefix / suffix portion or an intermediate span within the negative examples.
[0055] Other examples for automatically generating regular expressions according to certain embodiments are described below. These examples can correspond to various specific possible implementations of the general techniques in Figure 2 and can be implemented with software (e.g., code, instructions, programs, etc.) executed by one or more processing units (e.g., processors, cores) of a corresponding system, hardware, or a combination thereof. The software can be stored on a non-volatile storage medium (e.g., on a storage device). The other examples described below are intended to be illustrative and non-limiting. Although these examples describe various processing steps that occur in a particular sequence or order, this is not intended to be restrictive. In certain alternative embodiments, these steps can be performed in a somewhat different order, or some steps can also be performed in parallel.
[0056] In some examples, user input received via a user interface (e.g., step 203) can include one or more “positive examples” for which the regular expression output is to match, and zero or more “negative examples” for which the regular expression output does not match. Optionally, one or more positive examples can be highlighted to select a particular range (or subsequence) of characters. In some cases, at step 204, the positive examples received via the user interface can be converted into spans of regular expression code (e.g., character class codes such as Unicode category codes). For each positive example, a series of spans can be generated. In some embodiments, a graph can be created where each vertex corresponds to one of the sequences of spans and the edge weights are equal to the length of the output of the LCS algorithm executed on the two sequences of spans corresponding to the endpoints of the edge. A minimum spanning tree can be determined for the graph. For example, in some embodiments, Prim's algorithm can be used to obtain the minimum spanning tree. A depth-first traversal can be performed on the minimum spanning tree to determine a traversal order, after which the LCS algorithm can be executed on the first two elements of the traversal. Then, by repeatedly executing the LCS algorithm on the output of the previous LCS iteration and the next current traversal element, each additional element of the traversal can be merged into the current LCS output one by one in order. Then, the final output of the LCS algorithm (which can be a sequence of spans) can be converted into a regular expression. In some embodiments, the conversion can be a one-to-one conversion, while certain optional adornments described herein may not correspond to a one-to-one conversion. Finally, at step 203, the resulting regular expression can be tested against all of the positive and negative examples received via the user interface. If any of the tests fail, the above process can be repeated using all of the failed positive examples and any negative examples.
[0057] II. Regular Expression Generation Using the Longest Common Subsequence Algorithm on Regular Expression Code
[0058] As described above, certain aspects described herein relate to generating a regular expression based on computing the longest common subsequence (LCS) shared by different sets of regular expression code corresponding to input data.
[0059] Figure 3 FIG. [FIGURE NUMBER] is a flow chart showing a process of generating a regular expression using the LCS algorithm on a set of regular expression code, according to one or more embodiments described herein. At step 301, a regular expression generator 110 can receive one or more character sequences as input data. As described above, in some examples, the input data can correspond to positive example data selected from tabular data displayed in a user interface, but it should be understood that in some embodiments, the user interface is optional and in various examples, the input data can correspond to any character sequence received from any other communication channel (e.g., non-user interface).
[0060] In step 302, each character sequence received in step 301 can be converted into a corresponding regular expression code. In various embodiments, the regular expression code can be a Loogle code, a Unicode category code, or any other character class code representing a regular expression character class. For example, the input character sequence "May 3" can be converted into the Loogle code "LLLZN". In some embodiments, the regular expression code can be a category code based on the UNICODE category code shown at https: / / www.regular-expressions.info / unicode.html#category. For example, the code L can represent a letter, the code N can represent a number, the code Z can represent a space, the code S can represent a symbol, the code P can represent a punctuation mark, and so on. For example, the code L can correspond to Unicode \p{L}, and the code N can correspond to Unicode \p{N}.
[0061] In step 303, the longest common subsequence can be determined based on the set of regular expression codes generated in step 302. In some embodiments, the LCS algorithm can be executed using two sets of regular expression codes as input. Various different features of the execution of the LCS algorithm (such as processing direction, anchoring, push space, merging low-cardinality spans, aligning common tokens, etc.) can be used in different embodiments. In step 304, a regular expression can be generated based on the output of the LCS algorithm. In some cases, step 304 can include capturing the output of the LCS algorithm in the regular expression code and converting the regular expression code into a regular expression. In step 305, the regular expression can be simplified and output, for example, by displaying the regular expression to the user via a user interface.
[0062] Figure 4 is an example diagram for generating a regular expression using the longest common subsequence (LCS) algorithm on a set of regular expression codes based on two character sequence examples. Therefore, Figure 4 shows an example of applying the processing discussed above in Figure 3 As shown in Figure 4 The regular expression in this example is generated based on two input strings: "iPhone 5" and "iPhone X". Each sequence in this example can be converted into a corresponding set of regular expression codes. Therefore, iPhone 5 can be converted into "LLLLLLZN", and iPhone X can be converted into "LLLLLLZL". As Figure 4As shown, these category codes are then provided as input to the LCS algorithm, which determines that both IRECs (or category codes) include six Ls and one Z. The category codes excluded from the LCS can be represented as optional and / or alternative. Thus, a regular expression containing two character sequences can be represented as: \pL{6}\pZ\pN?\pL?. In this example, the regular expression includes Unicode category codes (e.g., \pL represents letters, \pZ represents whitespace, and \pN represents digits). The curly braces containing the number 6 indicate six instances of the letter, and the question mark indicates that the trailing digit / letter is optional. Finally, the regular expression generator can perform a simplification process during which the regular expression is simplified by reinserting the common text fragment "iphone" back into the final regular expression, replacing the broader "\pL{6}" portion of the regular expression.
[0063] As shown in this example, an input string received by the regular expression generator 110 can be converted into "regular expression code" representing the generalized categories of the regular expression (which can also be referred to as "category codes"), and the LCS algorithm can operate on these regular expression codes. In some embodiments, Unicode category codes can be used for the regular expression code. For example, an input text string can be converted into a code representing the regex Unicode generalized categories (e.g., \pL represents letters, \pP represents punctuation, etc.). Figure 3 and Figure 4 This method as shown can be referred to as an indirect method. However, in other embodiments, a direct method can be used, where the LCS algorithm operates directly on the character sequences received as input.
[0064] In some embodiments, the indirect method can provide additional technical advantages as it does not require large amounts of training data and can generate valid regular expressions with relatively few input examples. This is because the indirect method employs heuristics to reduce the uncertainty in regular expression generation and eliminate potential false positives and false negatives. For example, when generating a regular expression based on the input strings "May 3" and "Apr 4", the direct method may require at least one example per month to generate a valid regular expression that matches the date pattern. Relying solely on these two examples, the direct method may generate a regex of "[AM][ap][yr]
[13] 1?". In contrast, based on Unicode general categories, the indirect method can generate a more effective regular expression of "\pL{3}\d{1,2}". Additionally, as described above, one of the technical advantages described herein includes efficiently generating regular expressions using very little input data (even potentially from a single example). For example, regarding generating a regular expression from a single example "am", the heuristics can determine whether to generate "am" or "\pL\pL" for regular expression generation. Both can arguably be correct, but the programmed heuristics can implement user preferences and / or standards to determine how to generate the best regular expression (e.g., whether it should also match "pm").
[0065] Additionally, the indirect method can also simplify the generated regular expression "\pL{3}\d{1,2}" to "[A-Za-z]{3}\d{1,2}" to make it more human-readable. This can be beneficial in some embodiments, such as when outputting to non-complex regular expression users who may not be familiar with Unicode expressions for regular expressions.
[0066] Furthermore, in some embodiments, when performing the LCS algorithm, instead of treating each character independently, sequential and equal regular expression codes can be converted into a span data structure (which can also be referred to as a span). In some cases, a span can include a representation of a single regular expression code (e.g., a Unicode general category code), as well as a repetition count range (e.g., a minimum number and / or a maximum number). The conversion from regular expression codes to spans can facilitate some different additional features described below, such as identifying alternatives (e.g., disjunctions), and can also help merge adjacent optional spans to further simplify the generated regular expression.
[0067] As described above, the LCS algorithm can be configured to store and retain underlying text segments within the input character sequence, which may be inserted back into the final regular expression, e.g. Figure 4the string "iphone" in []. By tracing the text segment that originally caused the class code assigned to the span, such an embodiment can allow literal text (e.g., am and pm) to be directly included in the generated regular expression, which can reduce false positives and make the regular expression output more human-readable.
[0068] III. Generating a regular expression using the longest common subsequence algorithm on regular expression code combinations
[0069] Additional aspects described herein relate to generating a regular expression based on input data including three or more strings (e.g., three or more individual character sequences). When three or more strings are identified as input data, the regular expression generator 110 can use performance optimization features, where the best order for sequence determination for the LCS algorithm is performed. As described below, the performance optimization features for more than two strings can include constructing a graph having vertices corresponding to each string and edge lengths / weights based on the size of the LCS output between each string and every other string. Then, these edge weights can be used to derive a minimum spanning tree, and a depth-first traversal can be performed to determine the order of the input strings. Finally, the determined input string order can be used to complete a series of LCS algorithms.
[0070] Figure 5 is a flowchart showing the process of generating a regular expression using the longest common subsequence (LCS) algorithm on a larger set of regular expression codes (e.g., three or more character sequences). Thus, steps 502 - 505 in this example can correspond to step 303 discussed above in Figure 3 However, since this example involves generating a regular expression based on three or more input character sequences, the LCS algorithm can be executed multiple times. For example, to avoid an exponential increase in the running time of three or more input strings, the LCS algorithm can be executed multiple times, where each execution is performed on only two input strings. For example, the regular expression generator 110 can perform an initial execution of the LCS algorithm on two strings (e.g., two input character sequences or two transformed regular expression codes), then can perform a second execution of the LCS algorithm on the output of the first LCS algorithm and a third string, then can perform a third execution of the LCS algorithm on the output of the second LCS algorithm and a fourth string, and so on.
[0071] To improve and / or optimize the performance of such an embodiment, it may be desirable to determine the best order for processing an LCS algorithm sequence for an input string (e.g., an input character sequence or a regular expression code). For example, obtaining a good order for the input string may affect the readability of the generated regular expression, e.g., by minimizing the number of alternative spans. To maintain the conciseness of the generated regex, the additional strings from the LCS to the current regex should preferably already be somewhat similar to the current regex (the intermediate result obtained by performing LCS on the strings already seen).
[0072] Accordingly, in step 501, a plurality (e.g., 3 or more) of input character sequences are converted into regular expression codes. In step 502, an order for processing the regular expression codes using the LCS algorithm is determined. The order determination in step 502 is discussed in more detail below with reference to Figure 7 In step 503, either the first two regular expression codes in the determined order (for the first iteration of step 503) or the next regular expression code in the determined order (for subsequent iterations of step 503) is selected. In step 504, the LCS algorithm is performed on two input strings corresponding to the regular expression code format. For the first iteration of step 504, the LCS algorithm is performed on the first two regular expression codes in the determined order, while for subsequent iterations of step 504, the LCS algorithm is performed on the next regular expression code in the determined order and the output of the previous LCS algorithm (which can also be a regular expression code of the same format). In step 505, the regular expression generator 110 determines whether there are additional regular expression codes in the determined order that have not yet been provided as input to the LCS algorithm. If so, the process returns to step 503 for another execution of the LCS algorithm. If not, then in step 506, a regular expression is generated based on the output of the last execution of the LCS algorithm.
[0073] Figure 6 is an example diagram of generating a regular expression based on five input character sequence examples. In this example, each input character sequence is converted into a regular expression code, and then the LCS algorithm is repeatedly executed based on the determined order of the regular expression codes. Thus, Figure 6 illustrates an example of applying the process discussed above Figure 5 In this example, the determined order of the five regular expression codes is code #1 to code #5, and each code is input into the LCS algorithm in the determined order to generate a regular expression output. The final regular expression output (Reg Ex#4) corresponds to the final regular expression generated based on all five input character sequences.
[0074] Figure 7is a flowchart showing a process for determining the execution order of a longest common subsequence (LCS) algorithm over a larger set of regular expression codes (e.g., three or more). Thus, as shown in this example, steps 701 - 704 may correspond to the order determination in step 502 discussed above. In step 701, the LCS algorithm may be run on each unique pair of regular expression codes corresponding to the input data, and the resulting output LCS may be stored for each execution. Thus, for k input data, this may represent all (k(k - 1)) / 2 possible string pairings on which the LCS algorithm is to be run, or k(k - 1) in some embodiments. For example, if k = 3 input character sequences are received, the LCS algorithm may be run three times in step 701; if k = 4 input character sequences are received, the LCS algorithm may be run six times in step 701; if k = 5 input character sequences are received, the LCS algorithm may be run 10 times in step 701, and so on. In step 702, a fully connected graph may be constructed from k nodes representing the strings, where the edge weights of the (k(k - 1)) / 2 edges are the lengths of the original LCS outputs between two nodes. In step 703, a minimum spanning tree may be derived from the fully connected graph in step 702. In step 704, a depth - first traversal may be performed on the minimum spanning tree. The output of this traversal may correspond to the order in which the regular expression codes are to be input into the LCS algorithm execution sequence.
[0075] Brief reference Figure 8A and 8B , Figure 5 shows an example of a fully connected graph generated based on k = 5 input character sequences received, and Figure 8B shows a minimum spanning tree representation for the fully connected graph.
[0076] In some embodiments, Figure 5 - 8B the method described in 2 may provide additional technical advantages regarding performance. For example, some traditional implementations of the LCS algorithm may exhibit a runtime performance of O(n k ), where n is the length of the string. Extending such an implementation to k strings instead of just 2 may result in an exponential runtime performance of O(n k ), as the LCS algorithm may need to search a k - dimensional space. Such traditional implementations of the LCS algorithm may not be performant enough or sufficient for a real - time online user experience.
[0077] As described above, the LCS algorithm can be executed (k(k-1)) / 2 times, where sometimes the duplicates are exactly the same as those seen before, because the LCS algorithm can be performed when the original input examples from the user have been converted into regex category codes. Therefore, memoization can be implemented in some cases, where a cache can be used to map previously seen LCS problems to previously worked LCS solutions.
[0078] IV. Regular Expression Generation Based on Positive and Negative Pattern Matching Examples
[0079] Additional aspects described herein relate to generating regular expressions based on input data corresponding to both positive examples and negative examples. As described above, a positive example can refer to an input data character sequence that is specified as an example string that should match the regular expression to be generated by the regular expression generator. In contrast, a negative example can refer to an input data character sequence that is specified as an example string that should not match the regular expression to be generated by the regular expression generator. As described below, in some embodiments, the regular expression generator 110 can be configured to identify the positions that distinguish the positive examples from the negative examples and the shortest character subsequence at those positions. Then, the shortest subsequence can be hard-coded into the generated regular expression such that the positive examples will match the regular expression while the negative examples will be excluded by the regular expression (e.g., will not match).
[0080] Figure 9 is a flowchart showing a process of generating a regular expression based on positive and negative character sequence examples. In step 901, the regular expression generator 110 can receive one or more input data character sequences corresponding to positive examples. In step 902, the regular expression generator 110 can generate a regular expression based on the received positive examples. Thus, steps 901-902 can include some or all of the steps performed in Figure 3 or Figure 5 above to generate a regular expression based on an input data character sequence.
[0081] In step 903, the regular expression generator 110 may receive an additional input data character sequence corresponding to a negative example. Thus, the negative example is specifically specified such that it does not match the regular expression generated in step 902. In some embodiments, the negative example received in step 903 may first be tested against the regular expression generated in step 902, and if it is determined that the negative example does not match the regular expression, no further action is taken. However, in this example, it may be assumed that at least one of the negative examples received in step 903 matches the regular expression generated in step 902. Thus, in step 904, a disambiguation position may be determined within the regular expression generated in step 902. In some embodiments, the disambiguation position may be selected as a prefix position (e.g., at the beginning of the regular expression) or a suffix position (e.g., at the end of the regular expression). For example, the regular expression generator 110 may determine a first number of characters needed at the prefix to distinguish the positive example from the negative example, and a second number of characters needed at the suffix to distinguish the positive example from the negative example. The regular expression generator 110 may then select the suffix or prefix based on the shortest number of replacement characters required. In some cases, for readability purposes, using the prefix as the disambiguation position may be preferred (e.g., emphasized). In other examples, the disambiguation position may be an intermediate span position that does not correspond to the prefix or suffix of the regular expression.
[0082] In step 905, the regular expression generator 110 may determine a replacement sequence for a custom character class that, when inserted into the regular expression at the determined position, can distinguish the positive example from the negative example. In some embodiments, in step 905, the regular expression generator 110 may retrieve text fragments from each of the positive and negative examples corresponding to the disambiguation position (or replacement position), and then use the text fragments to determine a discriminator that will be used as the replacement sequence for distinguishing the positive example from the negative example. Additionally, the discriminator replacement sequence determined in step 905 may include multiple different replacement sequences for the custom character class that may be replaced at the same or different positions within the regular expression.
[0083] As described above, in some cases, the determination of the replacement sequence in step 905 can be performed in conjunction with the determination of the disambiguation position (or replacement position) in step 904. For example, the regular expression generator 110 can determine one or more replacement sequences that can distinguish positive and negative examples at the first possible replacement position. The regular expression generator 110 can also determine one or more other replacement sequences that can distinguish positive and negative examples at a second different possible replacement position. In this example, when choosing between different possible replacement positions and the corresponding replacement sequences, the regular expression generator 110 can apply a heuristic formula to perform the selection based on one or more of the character size of the replacement position and the number and / or size of the corresponding replacement sequences. Finally, in step 906, the regular expression can be modified by inserting one or more determined replacement sequences into the determined position to replace a previous portion of the regular expression. In some cases, after modifying the regular expression in step 906, the positive examples and / or negative examples can be tested against the modified regular expression to confirm that the positive examples match and the negative examples do not match the regular expression.
[0084] Figure 10A and 10B are example user interface screens showing the generation of regular expressions based on positive and negative character sequence examples. Thus, Figure 10A and 10B The examples shown in can correspond to the user interface displayed during the execution of the Figure 9 process discussed above. In Figure 10A , the user provides three positive examples of the data input character sequence 1001, and the regular expression generator 110 generates a regular expression 1002 that matches each positive example. Then, in Figure 4 B, the user provides a negative example 1004, and the regular expression generator 110 generates a modified regular expression 1005 based on the two current sets of positive example 1003 and negative example 1004.
[0085] As described above, in some embodiments, when both positive and negative examples are received, the regular expression generator 110 can identify a discriminator or the shortest subsequence of one or more characters that distinguishes the positive examples from the negative examples. The discriminator selected can be the shortest sequence (e.g., represented as a class code) and can be either positive or negative such that positive examples will match and negative examples will not match. In some cases, the discriminator can correspond to a replacement subsequence that can then be hardcoded into the regular expression in step 905. As an example, in "[AL][a - z]+", [AL] is a positive discriminator, which, assuming it is applied to street suffixes, will match (or allow) "Alley", "Avenue", and "Lane", but will not match (or will not allow) any other words. As another example, in "[BC][o][a - z]+", [BC][o] is a positive discriminator, which consists of a sequence of two character classes that match "Boulevard" and "Court". As yet another example, in "[^A][a - z]+", [^A] can be a negative discriminator that does not allow "Alley" and "Avenue". In some cases, the algorithm generates a negative lookbehind to correctly discriminate. For example, (?<!Av)[A - Za - z]+ will exclude "Avenue" but allow "Alley".
[0086] As another example, if the user provides positive examples "202 - 456 - 7800" and "313 - 678 - 8900" and negative examples "404 - 765 - 9876" and "515 - 987 - 6570", then in some embodiments, the regular expression generator 110 can generate the regular expression "\d\d\d-\d\d\d-\d\d00". That is, based on determining that phone numbers ending with 00 distinguish the positive examples from the negative examples (e.g., assuming the goal is to match a regular expression for business phone numbers), a replacement character subsequence can be identified for the suffix of the regular expression. This is an example of a negative example of the suffix (or more specifically, an example of using a positive suffix to accommodate negative examples), but various other embodiments can support replacements at the prefix, suffix, or middle span positions. In an example of replacement at the middle span position, the character offsets within the span can be tracked and split at the middle span point.
[0087] To decide whether to use a prefix or a suffix, in some embodiments, a heuristic is employed, where the minimum score is selected among all combinations of k a and the prefix / suffix:
[0088]
[0089] where,
[0090] k a = The number of characters considered for disambiguating an affix (prefix or suffix)
[0091] |F p | = The number of unique text segments in the positive examples required for disambiguating the affix
[0092] |F n | = The number of unique text segments in the negative examples required for disambiguating the affix
[0093] |E p | = The number of (complete) positive examples provided by the user
[0094] |E n | = The number of (complete) negative examples provided by the user
[0095] In the above example, the heuristic is designed to favor shorter disambiguating text segments over longer ones (e.g., thus multiplying by k a ). The heuristic is also designed to favor prefixes over suffixes (e.g., thus a penalty of 0.1 for suffixes) to improve readability. Finally, the heuristic is designed to favor disambiguating longer prefixes or suffixes (e.g., replacement) over disambiguating by using a larger number of string segments (e.g., thus the square of the number of string segments to be replaced).
[0096] As described above, some embodiments may also support negative middle-span examples as well as negative backtracking examples and negative look-ahead examples.
[0097] Once the prefix / suffix and k (the number of characters to be disambiguated) are determined, the regular expression generator 110 can still determine how to represent that disambiguation in the generated regular expression. The generated regular expression can allow affixes (e.g., prefixes or suffixes) that look like positive examples and can exclude affixes that look like negative examples.
[0098]
[0099] If usePermissive is greater than zero, then content that looks like positive examples is allowed through by generating a regular expression that allows characters to be obtained one by one (for each character position) from the positive examples. In other cases, the regular expression generator 110 can take the approach of not allowing content that looks like negative examples by generating a regular expression that does not allow characters to be obtained one by one (for each character position) from the negative examples.
[0100] As another example, the regular expression generated for the positive example 8am and the negative example 9pm could be \d[^p]m. This uses the caret syntax. In some cases, the regular expression generator 110 can be configured to prefer shorter regular expressions, which are not only more readable to the user but also more likely to be correct. The reason is that frequently occurring characters are more likely to occur again in the future, so the focus should be on frequently occurring characters. If the unique character |F p | is less (the unique character is less because the characters that occur occur more frequently), then the reward in the heuristic is to put it in the denominator.
[0101] Referring again to the usePermissive example heuristic above, if there is only one positive example from the user, it is not a big deal to determine a unique positive affix. So, in this heuristic, low |E p | is penalized by putting it in the numerator (i.e., high |E p | is rewarded in this heuristic).
[0102] Additionally, in some embodiments, negative examples can be based on backtracking and / or lookahead. For example, if the user provides a positive example of "323-1234" and a negative example of "202-754-9876", then this involves using the regex backtracking syntax (?<! ) to exclude phone numbers with area codes.
[0103] In some cases, negative examples can also be based on optional spans. For example, the user can provide positive examples of "ab" and "a2b" and a negative example of "a3b". In this case, the example implementation may fail because it may try to distinguish only based on the required span, and the "2" digit is in the optional span. In this example, failure may refer to the situation where the generated regular expression matches all positive examples (correctly) and also matches one or more negative examples (incorrectly). In this case, the user can be warned of the failure, and the user can be provided with options via the user interface to manually fix the generated regular expression and / or remove some negative examples.
[0104] V. User Interface for Regular Expression Generation
[0105] Additional aspects described herein include several different features and functions related to regular expression generation within a graphical user interface. As described below, some of these features can include various options for user selection and highlighting of positive and negative examples, color coding for positive and negative examples, and multiple overlapping / nested highlights within data cells.
[0106] Figure 11It is a flowchart showing a process of generating a regular expression based on a user data selection received within a user interface. Figure 11 The example process in Figure 11 describes a process regarding a user interface that can be generated and displayed on the client device 120. In step 1101, in response to a request from the user via the user interface, the regular expression generator 110 can retrieve data (e.g., from the data repository 130) and present / display the data in tabular form within the graphical user interface. Although tabular data is used in this example, it should be understood that tabular data is not required to be used or displayed in other examples. For example, in some cases, the user can directly type in the raw data (instead of selecting data from the user interface). Additionally, when presenting data via the user interface, the data does not need to be in tabular form and can be unstructured data (e.g., a document) or semi-structured (e.g., a spreadsheet of unformatted / unstructured data items such as tweets or posts). In various examples, the tabular data can correspond to transaction data, customer data, performance data, prediction data, and / or any other category of data that can be stored in the data repository 130 of an enterprise or other organization. In step 1102, a user selection of the input data can be received via the user interface. For example, the selected input data can correspond to an entire data unit selected by the user, or a subsequence of characters within the data unit. In step 1103, the regular expression generator 110 can generate a regular expression based on the input data (e.g., the data unit or a portion thereof) received in step 1102. In step 1104, the user interface can be updated in response to the generation of the regular expression. In some cases, the user interface can simply be updated to show the generated regular expression to the user, while in other cases, the user interface can be updated in various other ways as discussed below. As shown in this example, the user can select multiple different input data character sequences via the user interface, and in response to each new input data received, the regular expression generator 110 can generate an updated regular expression that includes first and second (positive) examples of the character sequence. Then, when the user highlights a third character sequence (e.g., outside the two character sequences, or within the first or second character sequence), the regular expression generator 110 can update the regular expression again, and so on. In some embodiments, the regular expression generator 110 can execute the algorithm in real time (or near real time) such that a brand new regular expression is generated in response to each new keystroke or each new highlighted portion made by the user.
[0107] Therefore, as Figure 11As shown, in response to a user's selection of a character sequence via a user interface, the regular expression generator 110 can generate and display a regular expression. For example, when the user highlights a first character sequence, the regular expression generator can generate and display a regular expression representing the first character sequence. When the user highlights a second character sequence, the regular expression generator can generate an updated regular expression that includes the first and second character sequences. Then, when the user highlights a third character sequence (e.g., within the first or second sequence), the regular expression generator can update the regular expression again, and so on.
[0108] Figure 12 is another flowchart showing a process of generating a regular expression and extracting data based on capture groups via selection of user data received within a user interface. In step 1201, as discussed above in step 1101, the regular expression generator 110 can retrieve data (e.g., from the data repository 130) and present / display the data in tabular form within a graphical user interface. In step 1202, the regular expression generator 110 can receive a user's highlighted selection of a text snippet within a particular data cell. In step 1203, the regular expression generator 110 can generate a regular expression based on positive examples of the selected data cell, and in step 1204 can create a regular expression capture group based on the text snippet highlighted within the cell. In step 1205, the regular expression generator 110 can determine one or more additional cells in the displayed tabular data that match the generated regular expression, and in step 1206, can extract the corresponding text snippets from the additional cells that match the generated regular expression.
[0109] Thus, in addition to providing positive examples, the user can also select (e.g., via mouse text highlighting) a text snippet within any selected positive example. In response, the regular expression generator 110 can create a regular expression capture group to extract that text snippet from the example and corresponding snippets from all other matches in the text to which the regular expression is applied. Extracting text snippets from matching data cells can also include deletion and modification and can, in some cases, be used to create new data columns outside of existing columns in semi-structured or unstructured text.
[0110] Using an example where the user selects a positive data example, and if the user highlights the year, the regular expression generator 110 can generate the regular expression (? :[A - Z]{3}\s+\d\d,\s+|\d\d / \d\d)(\d\d\d\d). As shown in this example, the regular expression generator 110 has placed parentheses around the year and has also converted the old parentheses around the month and day (for alternatives) to "non - capturing" groups using the ?: regular expression syntax. In some embodiments, it may be necessary for the extraction / capture group to fall on a span boundary, and in such embodiments, the regular expression generator 110 can take the highlighted character range as input and expand it to include the nearest anchor span boundary. However, in other examples, the user interface can support mid - span extraction / capture.
[0111] In some embodiments, the user interface can support input data from the user that includes a selection of a first character sequence within a second character sequence. For example, the user can highlight one or more characters within a larger previously highlighted character sequence, and the second user selection can provide context for the larger first user selection. Such embodiments can enable the input data to be provided to the regular expression generator 110 in a more targeted manner.
[0112] In addition, in some examples, an operation can be initiated in response to a user selection (e.g., highlighting text) in the user interface and a dialog box can be opened. In some cases, the dialog box can be a non - modal dialog box, such as a floating toolbox window that does not prevent the user from interacting with the main screen. Depending on the main operation the user is performing, the dialog box can also change in appearance and / or functionality. Thus, in such cases, the user does not need to search for other menu items after highlighting the selected text in order to initiate modifications, extractions, etc. of the captured group text fragments. In addition, in certain embodiments, the user interface provided for generating regular expressions can include three highlighting modes: nested automatic, nested manual, and single - layer. In some cases, the default operation mode can be to identify the entire cell as the highlighted area, and the user can also highlight one or more additional subsequences within the highlighted cell. In other modes, the user can be allowed to manually specify two highlights within the data cells of the table data display. In other modes, the user can be allowed to manually specify an outer highlight without an inner highlight. These other modes may be more suitable for "semi - structured" data, such as data columns consisting of tweets or other long strings (e.g., browser "user - agent" strings). "Semi - structured" data refers to data that can be displayed in tabular form within the user interface, but the columns within the table consist of unstructured text.
[0113] In some such embodiments, internal and external selections (e.g., highlighting) made by the user via the user interface can be distinguished by color coding. For example, external highlighting of positive examples can be shown in a first text / background color combination, while internal highlighting of positive examples can be shown in a different contrasting text / background color combination.
[0114] As described above, the user can specify the selection of a capture group by selecting a character subsequence. The GUI can be used to facilitate user selection via highlighting (or other indication). Figure 13 Examples are shown where an example user interface screen with a table data display is shown. In this example, Figure 13 Highlighting within a column value is depicted, for example, caused by the user dragging the mouse over one or more desired elements of the column value. Note that the "cell" in which the user highlighting is performed can exhibit a color change indicating the selected column value. This color change can be interpreted as automatic highlighting in response to the user highlighting.
[0115] Figure 14 and 15 are example user interface screens showing the generation of regular expressions and capture groups based on selecting data from a table display. In these examples, Figure 14 and Figure 15 show additional user interface windows that automatically display the detected user highlighting 1401 within the table data display. The window includes a field 1402 for displaying positive examples, a field for displaying negative examples, and a field for displaying the dynamically (and almost instantaneously) generated regular expression in response to selecting a positive example from the table data display. In these examples, the user highlighting within the column value 1401 can be equivalent to the user highlighting within the automatic highlighting. Thus, highlighting the area code by the user not only results in the highlighted area code 1401 by the user, but also causes the rest of the phone number to be filled in the positive example field 1402.
[0116] However, it should be understood that user highlighting is not limited to the performance within automatic highlighting. For example, alternatively, user highlighting can be performed within other user highlighting. As another example, user highlighting can alternatively be performed without any internal highlighting (e.g., further highlighting within the highlighted text). These alternative examples are particularly suitable for semi-structured data, such as data columns including "Tweets" or other long strings (e.g., browser "user agent" strings).
[0117] In addition, when generating the corresponding regular expression, other column values 1402 that match the regular expression can be identified based on additional automatic highlighting. In Figure 14 and15 In the example shown, additional automatic highlighting indicates elements of these other column values that match the capture groups of the generated regular expression. The additional automatic highlighting can be performed using a color different from the color used for user highlighting.
[0118] As Figure 15 shown, additional user highlighting is shown to indicate that the user selects other examples. The additional user highlighting can be performed in a manner similar to that described above. Thus, Figure 15 the user interface in [FIGURE] shows the filling of other examples in field 1502 for displaying positive examples. This can occur in response to detecting additional user highlighting. Additionally, the generated regular expression 1503 can be updated dynamically and almost instantaneously so that it matches all positive examples 1502. In response to generating the updated regular expression, the automatic highlighting of other column values 1504 that match the updated regular expression can also be updated. In some implementations, dynamic color coding can also be used. For example, the matches can be color-coded using a first color (e.g., blue), the positive examples can be color-coded using a second color (e.g., green), and the negative examples can be color-coded using a third color (e.g., red).
[0119] Figure 16A and 16B are example user interface screens showing the generation of a regular expression based on the selection of positive and negative examples from a table display. In Figure 16A - 16B , individual examples from the positive example field 1602 can be removed from the positive example field 1603 and / or moved to the negative example field 1603. Within the user interface, this can be performed, for example, by the user clicking (e.g., right-clicking) on one of the examples to select it. This selection can cause the user interface to display a menu 1602 that includes a delete option and a change option. Thereafter, clicking on an option can cause the corresponding function to be executed.
[0120] In Figure 16A and 16BIn the example shown, the result of the user selecting to change the option is to move the selected example to the negative example field 1603, such that the regular expression 1601 is updated to the regular expression 1604, which can be generated dynamically and near instantaneously (e.g., between 30 ms and 9000 ms in some embodiments). In response to generating the updated regular expression 1604, the automatic highlighting of other column values that match the updated regular expression can also be updated within the table data display. Additionally, automatic highlighting can be performed on some or all of the negative examples, including any column values corresponding to the negative examples, which can be highlighted using a color different from any of the colors used above, or otherwise distinguished within the user interface using other visual techniques.
[0121] In some embodiments, specifying a negative example via the user interface does not require first specifying the example as a positive example and then converting it to a negative example, as Figure 16A and 16B shown. Instead, negative examples can be specified in a variety of ways. For example, the user can select (e.g., right click) a column value via the user interface (e.g., one of the other column values for which automatic highlighting is performed to indicate its match with the generated regular expression), which can cause a menu to be displayed that includes an option to specify the selected column value as a negative example (e.g., “Create New Counterexample”).
[0122] Thus, using Figure 16A and 16B the examples shown, in response to generating the updated regular expression 1604, the automatic highlighting of other column values that match the updated regular expression can also be updated. In these examples, the updated regular expression specifies a phone number ending with “9”.
[0123] Briefly returning to Figure 14 and 15 , when the user clicks or otherwise selects the “Extract” button, an operation can be initiated to extract the highlighted text fragments within all cells that match the current regular expression 1403 or 1503. Although in Figure 14 and Figure 15Although not shown in the figure, in some embodiments, the user interface may attach to or replace the "Extract" button to provide other selectable buttons. For example, a "Replace" button may be presented as an option to replace the element highlighted by the user with an element specified by the user. Additionally or alternatively, one or more "Delete" buttons may be presented as options to effectively replace the element highlighted by the user with an empty one. For example, one or both of the "Delete Segment" operation and / or the "Delete Line" operation may be implemented, which will respectively delete the text segment or line highlighted by the user. Additional operations that may be implemented in various embodiments may include: the "Keep Line" operation, the "Split" operation (e.g., highlighting a comma and then extracting the comma-separated components into separate new columns), and the "Obfuscate" operation (e.g., replacing the highlighted text / capture group with a "#" sequence or other symbols).
[0124] In this example, in response to selecting the "Extract" button, an extract operation may be added to the list of transformation scripts to be executed by downstream operations. In some embodiments, the list of transformation scripts may be displayed in a part of the user interface for the user to view / modify. Alternatively, the extract operation may be executed on-the-fly to generate a new column (e.g., an element corresponding to the highlighted part of the positive example) that includes the content of the regular expression capture group. In Figure 14 and 15 the example shown, a new column and / or a new table of area codes may be generated in response to the selection of the "Extract" button.
[0125] Figure 17 is another example user interface screen according to one or more embodiments described herein, which shows generating a regular expression and capture groups based on selecting data from a table display.
[0126] VI. Regular Expression Generation Using the Longest Common Subsequence Algorithm over Spans
[0127] Additional aspects described herein relate to regular expression generation based on the LCS algorithm from one or more data input character sequences, but where the regular expression generator 110 can also handle characters that occur only in some examples. To handle characters that occur only in some input examples, spans can be defined, where the minimum and maximum occurrences of the regular expression code are traced. For example, for the character sequence input of "9pm" and "9pm", there is an optional space between the number and the "pm" text. In this case, when a certain span (e.g., a single space between "9" and "pm") may not exist in all given input examples, the minimum occurrence can be set to zero. These minimum and maximum numbers can then be mapped to the regular expression multiplicity syntax. The longest common subsequence (LCS) algorithm can operate on spans of characters obtained from the input examples, including "optional" spans (e.g., minimum length of zero) that do not occur in every input example. As described below, consecutive spans can be merged during the execution of the LCS algorithm. In this case, when the additional optional spans carried are no longer consecutive, the LCS algorithm can also operate recursively on these optional spans. That is, although the operation of the LCS algorithm is recursive in nature, in these cases, the entire LCS algorithm can operate recursively (e.g., recursively operating on a recursive LCS algorithm). Among other technical advantages, this can allow for shorter, neater, and more readable regular expression generation. For example, (am|am) (i.e., with an optional space before am) can be generated without recursively operating the LCS algorithm, while recursively operating the LCS algorithm can result in the regular expression being generated as (?am), which is shorter and neater.
[0128] Figure 18FIG. 18 is a flowchart showing a process of generating a regular expression including an optional span using the Longest Common Subsequence (LCS) algorithm, according to one or more embodiments described herein. In step 1801, a regular expression generator 110 may receive one or more character sequences corresponding to positive regular expression examples as input data. In step 1802, the regular expression generator 110 may convert the character sequences into regular expression codes. Thus, steps 1801 and 1802 may be similar or identical to the corresponding examples discussed above. Then, in step 1803, the regular expression codes may also be converted into span data structures (or spans). As described above, each span may include a data structure storing a character class code (e.g., regex code) and a repetition count range (e.g., a minimum count and / or a maximum count). In step 1804, the regular expression generator 110 may execute the LCS algorithm, providing a set of spans as input to the algorithm. The output of the LCS algorithm in this example may include an output set of spans, including at least one span having a minimum repetition count range equal to zero, which corresponds to an optional span within the output of the LCS algorithm. Finally, in step 1805, the regular expression generator 110 may generate a regular expression based on the output of the LCS algorithm (including the optional span).
[0129] Figure 19 FIG. 19 is an example diagram showing the generation of a regular expression using the Longest Common Subsequence (LCS) algorithm, where the generated regular expression includes an optional span. In this example, the two input data character sequences are "8am" and "9pm". As described above, the input data character sequences are first converted into regular expression codes (step 1802), and then into spans (step 1803). The spans may be provided as input to the LCS algorithm (step 1804), and the LCS output includes an optional span Z <0,1> , indicating that an optional single space may be a text sequence of a digit and two letters. That is, the superscript notation in this example may include two digits, namely, a minimum repetition count range (e.g., 0) and a maximum repetition count range (e.g., 1) applied to the previous code (e.g., Z = space). Finally, a regular expression may be generated based on the output spans of the LCS algorithm, and the optional span may be converted into the corresponding regular expression code "pZ * ".
[0130] In some embodiments, the regular expression generator 110 reproducing and using the optional space during the execution of the LCS algorithm may provide additional technical advantages in terms of performance and readability. For example, when generating a regular expression, it is desirable to be able to handle characters common to all given examples and characters that appear only in some examples in certain cases.
[0131] In some embodiments, for each span data structure, the minimum occurrence count of a category code and the maximum occurrence count of a category code can be tracked. In cases where there is no span at all in one or more given examples, the minimum number is set to zero. As another example, to generate a regular expression to handle the spelled-out months of the year, the minimum and maximum numbers can then be mapped to a regular expression multiplicity syntax that includes curly braces (e.g., [A-Za-z]{3,9}).
[0132] In some embodiments, the regular expression generator 110 can track the minimum and maximum occurrence counts for each span, but can also handle additional implementation details. For example, as a result of processing optional spans and running the LCS over character spans, the regular expression generator 110 can be configured to detect and merge consecutive spans throughout the execution of the LCS algorithm. Additionally, any extra optional spans that are carried sometimes occur consecutively, and the LCS algorithm may also need to run recursively over these spans. For example, in some cases, the regular expression generator 110 modifies and / or extends the LCS algorithm to favor (or weight) fewer transitions between optional and required sequence elements (e.g., spans). For example, grouping optional spans together can minimize the number of grouping parentheses that must be used in the regular expression, which can improve the human readability of the generated regular expression. In some cases, if the resulting lengths are equal even after considering the optional spans, the regular expression generator 110 can exhibit a preference for alternatives with fewer transitions between optional and required spans. For example, in certain cases, a standard LCS algorithm can be implemented to favor selecting the longer sequence at its decision points. However, at decision points where the options have equal length, the configuration preference can be programmed into the regular expression generator 110. For example, one such configuration preference can be to favor the shorter sequence (once the optional spans are taken into account). Thus, the customized LCS within this configuration can be optimized for both longer sequences (of required spans) and shorter sequences (of total required and optional spans).
[0133] In some embodiments, the generated regular expression may be more readable if it starts with a required span (which can also serve as a mental anchor for human readers) instead of starting the regular expression with an optional span. Thus, in some cases, if the resulting options have an equal number of transitions, the option with an earlier non-optional span may be selected. Additionally, in some embodiments, the LCS algorithm performed by the regular expression generator 110 may be configured to push all whitespace (including optional spans corresponding to whitespace) to the right within the regular expression. By pushing all whitespace to the right, the likelihood of merging whitespace spans together may increase, which can simplify the resulting regular expression and improve readability. Thus, during the execution of the LCS algorithm, when determining that two sets of substrings have the same LCS, the set that contributes to improved readability may be selected instead of arbitrarily choosing one of the two sets. Further, in some embodiments, the LCS algorithm may be configured to favor a greater number of required spans and / or fewer optional spans for improved readability.
[0134] As described above, in some cases, negative examples can also be based on optional spans. For example, a user may provide positive examples of "ab" and "a2b" and a negative example of "a3b". In this case, the example implementation may fail because it may attempt to discriminate based only on required spans and the digit "2" is within an optional span. In such a case, the user can be warned of the failure, and options can be provided to the user via a user interface to manually fix the generated regular expression and / or remove some of the negative examples.
[0135] In some embodiments, there may be an isSuccess that is returned as part of the JSON returned from a REST service. In some embodiments, when isSuccess = false, the generated regex can turn a different color (e.g., red).
[0136] VII. Regular Expression Generation Using a Combinatorial Longest Common Subsequence Algorithm
[0137] Other aspects described herein relate to combinatorial search, where the LCS algorithm performed by the regular expression generator 110 may be run multiple times to generate a "correct" regular expression (e.g., a regular expression that properly matches all given positive examples and properly excludes all given negative examples), and / or generate multiple correct regular expressions from which the most desirable or best regular expression can be selected. For example, during combinatorial search, the full LCS algorithm and regular expression generation process may be run multiple times, including different combinations / permutations of text processing directions, different anchoring, and other different characteristics of the LCS algorithm.
[0138] Figure 20 FIG. 2000 is a flowchart showing a process of generating a regular expression based on a combined execution of a longest common subsequence (LCS) algorithm. In step 2001, a regular expression generator 110 may receive an input data character sequence corresponding to a positive example. In step 2002, the regular expression generator 110 may iterate through various different combinations of execution techniques of the LCS algorithm. As shown in this example, during each iteration of step 2002, the regular expression generator 110 may select different combinations of the following LCS algorithm execution parameters (or features): anchoring (i.e., not anchored, anchored to the beginning of the line, anchored to the end of the line), processing direction (i.e., right-to-left order, left-to-right order), pushing spaces (i.e., pushing or not pushing spaces), and collapsing spans (i.e., collapsing or not collapsing spans). In step 2003, the LCS algorithm is run on the input data character sequence (or, if the input character sequence is first transformed, on the regular expression code), where the LCS algorithm is configured based on the parameters / features selected in step 2002. In step 2004, the output of the LCS algorithm may be stored by the regular expression generator 110, including data such as whether the algorithm successfully identified the LCS and the length of the corresponding regular expression. In step 2005, the process may iterate until the LCS algorithm has been run with all possible combinations of the parameters / features of the combined search. Finally, in step 2006, a particular output from one of the LCSs is selected as the best output (e.g., based on success and regular expression length), and a regular expression may be generated based on the selected LCS algorithm output.
[0139] In various embodiments, a combined search such as the one described above may be performed for various different combinations of parameters / features. Figure 20 For example, in some embodiments, the LCS algorithm may anchor the regular expression to the beginning of the text using the caret ^, and / or anchor the regular expression to the end of the text using the dollar sign $. In some cases, such anchoring may result in the generation of a shorter regular expression. Anchors may be particularly useful when the user wishes to find a specific pattern at the beginning and / or end of a string. For example, the user may want a product name at the beginning. To avoid confusing the LCS algorithm with different numbers of words describing the product name, the regex may be anchored to the beginning of the string using the caret, as shown below.
[0140] In addition, in some embodiments, the LCS algorithm may be executed using forward or reverse input data (or, similarly, the LCS algorithm may be configured to receive input data in the normal order and then reverse the order before executing the algorithm). Thus, in some embodiments, the combined search of the LCS algorithm that may be performed on the input character sequence or code pair may be:
[0141] 1. Normal (right-to-left) order, not anchored to the start or end
[0142] 2. Normal (right-to-left) order, using caret ^ to anchor to the start of the line
[0143] 3. Normal (right-to-left) order, using dollar $ to anchor to the end of the line
[0144] 4. Reversed (left-to-right) order, not anchored to the start or end
[0145] 5. Reversed (left-to-right) order, using caret ^ to anchor to the start of the line
[0146] 6. Reversed (left-to-right) order, using dollar $ to anchor to the end of the line
[0147] In this example, among the six executions of the LCS, the shortest resulting regular expression can be selected (step 2006).
[0148] In some embodiments, the combinatorial search of the LCS algorithm can also iterate over the greedy quantifier "?" and the non-greedy quantifier "??". For example, by default, if there is an optional span, a single question mark is emitted. For example, [A-Z]+(?:[A-Z]\.)?[A-Z]+ represents a first name and a last name, with an optional middle initial. If a satisfactory regular expression cannot be found when using the greedy quantifier, the combinatorial search can attempt to replace all question mark quantifiers with double question mark quantifiers (e.g., [A-Z]+(?:[A-Z]\.)??[A-Z]+). The double question mark corresponds to a non-greedy quantifier, which can instruct the downstream regular expression matcher to enter the backtracking mode to find a match.
[0149] Furthermore, in some embodiments, the combinatorial search of the LCS algorithm can also iterate over whether to prefer spaces on the right. For example, as described above, a strategy of pushing spaces to the right can be used in some embodiments. For example, when the LCS algorithm faces an arbitrary choice between otherwise equal options, it is desirable that the space spans can be merged together, resulting in fewer overall spans. This feature adds another option to the combinatorial search, namely either pushing the spaces to the right or performing according to the traditional LCS method, i.e., making a decision arbitrarily.
[0150] In addition, in some embodiments, the combinatorial search of the LCS algorithm can also traverse literal characters that are common in all examples by running the LCS on the original string. In such embodiments, the LCS algorithm can be configured to identify and align common words. As used herein, a "common word" can refer to a word that appears in each positive example. Once the common words are identified, their span types can be converted from letters to words, and then they can be naturally aligned by subsequent runs of the LCS algorithm.
[0151] Thus, in the following example, the combinatorial search can iterate over several parameters / features to perform the full LCS algorithm 96 times. The various parameters / features to iterate over in this example are:
[0152] · Anchoring (3) (values = ^, $, or neither)
[0153] · Push spaces (2) (values = yes or no)
[0154] · Merge low-cardinality spans to wildcards (2) (values = yes or no)
[0155] · Greedy quantifier? (2) (values = yes or no)
[0156] · Align LCS algorithm on generic tokens (2) (values = yes or no)
[0157] · Use "\w" to denote alphanumeric, or treat letters "\pL" and digits "\pN" as separate spans (2) (values = yes or no)
[0158] As described above, in this example, the full LCS algorithm will be executed 96 times (e.g., 3 * 2 * 2 * 2 * 2 * 2 = 96).
[0159] However, in other embodiments, the regular expression generator 110 can provide a performance enhancement by which only the first three features in the above list (anchoring, push spaces, and merging low-cardinality spans to wildcards) can participate in the combinatorial search. This may result in a much lower number of times the full LCS algorithm has to be executed (e.g., 3 * 2 * 2 = 12 times). In such an embodiment, although the last three features in the above list (greedy quantifier, aligning the LCS algorithm to generic tokens, and using "\w" to denote alphanumeric, or treating letters "\pL" and digits "\pN" as separate spans) do not participate in the combinatorial search, these features can be tested individually and sequentially at the end. A technical advantage can be achieved in such an embodiment because partitioning the search space in this way can still find a satisfactory regular expression but with an approximately 8-fold performance acceleration.
[0160] To illustrate this, the following example of combinatorial search can provide a performance advantage over the previous example. In this example, the combinatorial search can be performed based on the following parameters / features to iterate over:
[0161] · Anchoring (3): BEGINNING_OF_LINE_MODE, END_OF_LINE_MODE, NO_EOL_MODE
[0162] · Order / Direction (2): Left-to-Right (Normal) LCS and Right-to-Left (Reverse) LCS
[0163] · Push (2): Whether to attempt to push spaces to the right within the LCS algorithm
[0164] · Compress to Wildcard (2): Whether to attempt to compress long sequences that only occur sometimes down to the wildcard.*?
[0165] The combinations in this example can result in running the full algorithm 3 * 2 * 2 * 2 = 24 times. Then, the regular expression generator 110 can take the best result among the 24 results of the LCS algorithm, where "best" can mean (a) the LCS algorithm was successful, and (b) the shortest regular expression was generated. The regular expression generator 110 can then perform the following three additional tasks:
[0166] 1. Attempt to compress sequences of letters and numbers that are not interrupted by spaces, punctuation, or symbols into a new span type I called alphanumeric, corresponding to the generated regular expression \w. This can be useful for hexadecimal digits found in IPv6 addresses in clickstream logs (see novelty 64 from April 2019).
[0167] 2. Attempt to use the non-greedy quantifier?? instead of the greedy quantifier?
[0168] 3. Attempt to align literally
[0169] VIII. Hardware Overview
[0170] Figure 21 A simplified diagram depicting a distributed system 2100 for implementing an embodiment is shown. In the illustrated embodiment, the distributed system 2100 includes one or more client computing devices 2102, 2104, 2106, and 2108 coupled to a server 2112 via one or more communication networks 2110. The client computing devices 2102, 2104, 2106, and 2108 can be configured to execute one or more applications.
[0171] In various embodiments, the server 2112 can be adapted to run one or more services or software applications that enable the automatic generation of regular expressions as described in this disclosure. For example, in certain embodiments, the server 2112 can receive user input data sent from a client device, where the user input data is received by the client device through a user interface displayed at the client device. The server 2112 can then convert the user input data into a regular expression, which is transmitted to the client device for display through the user interface.
[0172] In some embodiments, server 2112 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as under a software as a service (SaaS) model, to users of client computing devices 2102, 2104, 2106, and / or 2108. Users operating client computing devices 2102, 2104, 2106, and / or 2108 may then utilize one or more client applications to interact with server 2112 to utilize the services provided by these components.
[0173] In Figure 21 the configuration shown, server 2112 may include one or more components 2118, 2120, and 2122 that implement the functions performed by server 2112. These components may include software components, hardware components, or combinations thereof that may be executed by one or more processors. It should be understood that a variety of different system configurations are possible, which may differ from distributed system 2100. Thus, Figure 21 the embodiment shown is an example of a distributed system for implementing the embodiment system and is not intended to be limiting.
[0174] Users may use client computing devices 2102, 2104, 2106, and / or 2108 to execute one or more applications that may generate regular expressions in accordance with the teachings of the present disclosure. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 21 only four client computing devices are depicted, any number of client computing devices may be supported.
[0175] Client devices may include various types of computing systems, such as portable handheld devices, general-purpose computers such as personal computers and laptop computers, workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems (e.g., Microsoft Apple or UNIX-like operating systems, such as LINUX or LINUX-like operating systems such as Google Chrome TM OS), including various mobile operating systems (e.g., Microsoft Windows Windows Android TM , ). Portable handheld devices may include cellular phones, smart phones (e.g., )), tablet computers (e.g., ), personal digital assistants (PDAs), etc. Wearable devices can include Google head-mounted displays and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices (e.g., Microsoft gaming consoles with or without gesture input devices, Sony systems, various gaming systems provided by ), etc.). Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (e.g., email applications, short message service (SMS) applications), and can use various communication protocols.
[0176] Network 2110 can be any type of network familiar to those skilled in the art, which can support data communication using any of a variety of available protocols, including but not limited to, TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internetwork Packet Exchange), , etc. By way of example only, network 2110 can be a local area network (LAN), an Ethernet-based network, Token Ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a network operating under any one of the Institute of Electrical and Electronics Engineers (IEEE) 1002.11 protocol groups, and / or any other wireless protocol).
[0177] Server 2112 can be composed of one or more general-purpose computers, dedicated server computers (e.g., including PC (personal computer) servers, servers, midrange servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or any other suitable arrangement and / or combination. Server 2112 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization, such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server. In various embodiments, server 2112 can be suitable for running one or more services or software applications that provide the functions described in the foregoing disclosure.
[0178] The computing system in server 2112 can run one or more operating systems, including any of the operating systems discussed above, as well as any commercially available server operating system. Server 2112 can also run any of a variety of additional server applications and / or middleware applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, servers, database servers, etc. Exemplary database servers include, but are not limited to, database servers commercially available from (International Business Machines), etc.
[0179] In some implementations, server 2112 can include one or more applications to analyze and combine data feeds and / or event updates received from users of client computing devices 2102, 2104, 2106, and 2108. By way of example, the data feeds and / or event updates can include, but are not limited to feeds, updates, or real-time updates and continuous data streams received from one or more third-party information sources, which can include real-time events related to sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc. Server 2112 can also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 2102, 2104, 2106, and 2108.
[0180] Distributed system 2100 can also include one or more data repositories 2114, 2116. In certain embodiments, these data repositories can be used to store data and other information. For example, one or more of data repositories 2114, 2116 can be used to store information such as new data columns that match regular expressions generated by the system. Data repositories 2114, 2116 can reside in various locations. For example, the data repositories used by server 2112 can be local to server 2112 or can be remote from server 2112 and communicate with server 2112 via a network-based or dedicated connection. Data repositories 2114, 2116 can be of different types. In certain embodiments, the data repositories used by server 2112 can be databases, such as relational databases, such as those provided by Oracle and other vendors. One or more of these databases can be adapted to store, update, and retrieve data to and from the database in response to commands in SQL format.
[0181] In some embodiments, one or more of the data repositories 2114, 2116 may also be used by an application to store application data. The data repositories used by the application may be of different types, such as a key-value repository, an object repository, or a general-purpose repository supported by a file system.
[0182] In some embodiments, the functions described in this disclosure may be provided as services via a cloud environment. Figure 22 is a simplified block diagram of a cloud-based system environment according to some examples, in which various services may be provided as cloud services. In Figure 22 the example shown, the cloud infrastructure system 2202 may provide one or more cloud services that may be requested by users using one or more client computing devices 2204, 2206, and 2208. The cloud infrastructure system 2202 may include one or more computers and / or servers, which may include those computers and / or servers described above for server 2112. The computers in the cloud infrastructure system 2202 may be organized as general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0183] The network 2210 may facilitate data communication and exchange between the clients 2204, 2206, and 2208 and the cloud infrastructure system 2202. The network 2210 may include one or more networks. The networks may be of the same type or different types. The network 2210 may support one or more communication protocols, including wired and / or wireless protocols, to facilitate communication.
[0184] Figure 22 The example depicted in is only one example of a cloud infrastructure system and is not intended to be limiting. It should be understood that in some other examples, the cloud infrastructure system 2202 may have more or fewer components than those Figure 22 depicted, may combine two or more components, or may have a different component configuration or arrangement. For example, although Figure 22 depicts three client computing devices, any number of client computing devices may be supported in alternative examples.
[0185] The term cloud service is generally used to refer to services provided by a service provider's system (e.g., cloud infrastructure system 2202) on demand and via a communication network such as the Internet to users. Generally, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own enterprise servers and systems. The cloud service provider's system is managed by the cloud service provider. Thus, a customer can utilize the cloud services provided by the cloud service provider without having to purchase separate licenses, support, or hardware and software resources for these services. For example, the cloud service provider's system can host applications, and users can order and use the applications on demand via the Internet without having to purchase the infrastructure resources for executing the applications. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Several providers offer cloud services. For example, several cloud services are provided by Oracle in Redwood Shores, California including, for example, middleware services, database services, Java cloud services, etc.
[0186] In some embodiments, the cloud infrastructure system 2202 can use different models such as under a software as a service (SaaS) model, a platform as a service (PaaS) model, an infrastructure as a service (IaaS) model, and other models including hybrid service models to provide one or more cloud services. The cloud infrastructure system 2202 can include a set of applications, middleware, databases, and other resources capable of providing various cloud services.
[0187] The SaaS model enables applications or software to be delivered as a service to customers via a communication network such as the Internet, without the customers having to purchase the hardware or software for the underlying applications. For example, the SaaS model can be used to provide customers with access to on-demand applications hosted by the cloud infrastructure system 2202. Oracle Examples of SaaS services provided by Oracle include, but are not limited to, various services for human resources / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytics services, social applications, etc.
[0188] The IaaS model is generally used to provide infrastructure resources (e.g., servers, storage, hardware, and network resources) to customers in the form of cloud services to provide elastic computing and storage capabilities. Oracle offers various IaaS services.
[0189] The PaaS model is generally used to provide platform and environmental resources as a service, enabling customers to develop, run, and manage applications and services without having to procure, build, or maintain such resources. Oracle Examples of PaaS services provided include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), Data Management Cloud Service, various application development solution services, etc.
[0190] Cloud services are typically provided in an on-demand self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. For example, a customer can order one or more services provided by the cloud infrastructure system 2202 via a subscription order. Then, the cloud infrastructure system 2202 performs processing to provide the services requested in the customer's subscription order. The cloud infrastructure system 2202 can be configured to provide one or more cloud services.
[0191] The cloud infrastructure system 2202 can provide cloud services via different deployment models. In the public cloud model, the cloud infrastructure system 2202 can be owned by a third-party cloud service provider, and the cloud services are provided to any general public customer, where the customer can be an individual or an enterprise. Under the private cloud model, the cloud infrastructure system 2202 can operate within an organization (e.g., within an enterprise organization) and provide services to customers within that organization. For example, the customers can be various departments of an enterprise, such as the human resources department, the payroll department, etc., or even individuals within the enterprise. In the community cloud model, the cloud infrastructure system 2202 and the provided services can be shared by several organizations in a related community. Various other models can also be used, such as a hybrid of the above models.
[0192] The client computing devices 2204, 2206, and 2208 can be of different types (e.g., Figure 21 the devices 2102, 2104, 2106, and 2108 depicted therein), and are capable of operating one or more client applications. A user can use the client device to interact with the cloud infrastructure system 2202, such as requesting services provided by the cloud infrastructure system 2202.
[0193] In some embodiments, the processing performed by the cloud infrastructure system 2202 to provide management-related services can involve big data analytics. Such analytics can involve using, analyzing, and manipulating large datasets to detect and visualize various trends, behaviors, relationships, etc. in the data. The analytics can be performed by one or more processors, possibly processing data in parallel, performing simulations using the data, etc. For example, big data analytics can be performed by the cloud infrastructure system 2202 to automatically determine regular expressions. The data used for this analysis can include structured data (e.g., data stored in a database or structured according to a structured model) and / or unstructured data (e.g., data blobs (binary large objects)).
[0194] As shown Figure 22 in the example of Figure 22 , the cloud infrastructure system 2202 may include infrastructure resources 2230, which are used to assist in the provision of various cloud services provided by the cloud infrastructure system 2202. The infrastructure resources 2230 may include, for example, processing resources, storage or memory resources, networking resources, and the like.
[0195] In some embodiments, to facilitate the efficient provision of these resources to support the various cloud services provided by the cloud infrastructure system 2202 for different customers, these resources may be bundled into resource sets or resource modules (also referred to as "pods"). Each resource module or pod may include a pre-integrated and optimized combination of one or more types of resources. In some embodiments, different pods may be pre-provided for different types of cloud services. For example, a first set of pods may be provisioned for database services, and a second set of pods may be provided for Java services, which may include a different resource combination than the pods in the first set of pods, and so on. For some services, the resources allocated for providing the service may be shared among the services.
[0196] The cloud infrastructure system 2202 itself may internally use services 2232 that are shared by different components of the cloud infrastructure system 2202 and facilitate the cloud infrastructure system 2202 in providing services. These internal shared services may include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelisting services, high availability, backup and recovery services, services for enabling cloud support, email services, notification services, file transfer services, and the like.
[0197] The cloud infrastructure system 2202 may include multiple subsystems. These subsystems may be implemented in software or hardware or a combination thereof. As shown Figure 22As shown, the subsystem may include a user interface subsystem 2212 that enables users or customers of the cloud infrastructure system 2202 to interact with the cloud infrastructure system 2202. The user interface subsystem 2212 may include various different interfaces, such as a web interface 2214, an online store interface 2216 where cloud services provided by the cloud infrastructure system 2202 are advertised and available for purchase by consumers, and other interfaces 2218. For example, a customer may use a client device to request (service request 2234) one or more services provided by the cloud infrastructure system 2202 using one or more of the interfaces 2214, 2216, and 2218. For example, a customer may access an online store, browse the cloud services provided by the cloud infrastructure system 2202, and place a subscription order for one or more services provided by the cloud infrastructure system 2202 that the customer wishes to subscribe to. The service request may include information identifying the customer and one or more services the customer wishes to subscribe to. For example, a customer may subscribe to order services related to automatically generating regular expressions provided by the cloud infrastructure system 2202.
[0198] In certain embodiments, for example Figure 22 In the example depicted in, the cloud infrastructure system 2202 may include an order management subsystem (OMS) 2220 configured to process new orders. As part of this processing, the OMS 2220 may be configured to: create an account for the customer (if not already existing); receive billing and / or account information from the customer, which will be used to charge the customer for the requested services; verify the customer information; after verification, book the order for the customer; and coordinate various workflows to prepare the order for fulfillment.
[0199] Once properly verified, the OMS 2220 may then invoke an order fulfillment subsystem (OPS) 2224, which is configured to fulfill resources for the order, including processing, memory, and networking resources. Fulfillment may include allocating resources for the order and configuring the resources to assist with the services requested by the customer order. The manner in which resources are fulfilled for an order and the type of resources fulfilled may depend on the type of cloud service the customer has ordered. For example, according to one workflow, the OPS 2224 may be configured to determine the specific cloud service requested and identify multiple pods that may have been pre-configured for that specific cloud service. The number of pods allocated for the order may depend on the size / amount / level / scope of the requested service. For example, the number of pods to be allocated may be determined based on the number of users the service is to support, the duration of the service request, etc. Then, the allocated pods may be customized for the specific requesting customer to provide the requested service.
[0200] The cloud infrastructure system 2202 can send a response or notification 2244 to the requesting customer to indicate when the requested service is ready for use. In some cases, information (e.g., a link) that enables the customer to start using and leveraging the benefits of the requested service can be sent to the customer. In certain embodiments, for a customer who requests an automatically generated regular expression related service, the response can include instructions that, when executed, cause a user interface to be displayed.
[0201] The cloud infrastructure system 2202 can provide services to multiple customers. For each customer, the cloud infrastructure system 2202 is responsible for managing information related to one or more subscription orders received from the customer, maintaining customer data related to the orders, and providing the requested services to the customer. The cloud infrastructure system 2202 can also collect usage statistics regarding the customer's use of the subscription services. For example, statistics such as the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and system downtime can be collected. This usage information can be used to bill the customer. For example, billing can be on a monthly cycle.
[0202] The cloud infrastructure system 2202 can provide services to multiple customers in parallel. The cloud infrastructure system 2202 can store information for these customers, which may include proprietary information. In certain embodiments, the cloud infrastructure system 2202 includes an Identity Management Subsystem (IMS) 2228 that is configured to manage customer information and provide separation of the managed information such that information related to one customer cannot be accessed by another customer. The IMS 2228 can be configured to provide various security-related services, such as identity services; information access management, authentication, and authorization services; services for managing customer identities and roles and related capabilities, etc.
[0203] Figure 23 An example of a computer system 2300 is shown. In some embodiments, the computer system 2300 can be used to implement any of the above systems. As Figure 23 shown, the computer system 2300 includes various subsystems, including a processing subsystem 2304 that communicates with a number of other subsystems via a bus subsystem 2302. These other subsystems can include a processing acceleration unit 2306, an I / O subsystem 2308, a storage subsystem 2318, and a communication subsystem 2324. The storage subsystem 2318 can include non-transitory computer-readable storage media, including a storage medium 2322 and a system memory 2310.
[0204] The bus subsystem 2302 provides a mechanism for enabling the various components and subsystems of the computer system 2300 to communicate with each other as expected. Although the bus subsystem 2302 is schematically shown as a single bus, alternative examples of bus subsystems may utilize multiple buses. The bus subsystem 2302 can be any of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, a local bus using any of various bus architectures, and the like. For example, such architectures can include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus, which can be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard, etc.
[0205] The processing subsystem 2304 controls the operation of the computer system 2300 and can include one or more processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The processors can include single-core or multi-core processors. The processing resources of the computer system 2300 can be organized into one or more processing units 2332, 2334, etc. The processing units can include one or more processors, one or more cores from the same or different processors, a combination of cores and processors, or other combinations of cores and processors. In some embodiments, the processing subsystem 2304 can include one or more dedicated coprocessors, such as a graphics processor, a digital signal processor (DSP), etc. In some embodiments, some or all of the processing units of the processing subsystem 2304 can be implemented using custom circuits (such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs)).
[0206] In some embodiments, the processing units in the processing subsystem 2304 can execute instructions stored in the system memory 2310 or the computer-readable storage medium 2322. In various examples, the processing units can execute various programs or code instructions and can maintain multiple concurrently executing programs or processes. At any given time, some or all of the program code to be executed can reside in the system memory 2310 and / or the computer-readable storage medium 2322, including possibly on one or more storage devices. Through appropriate programming, the processing subsystem 2304 can provide the various functions described above. In the case where the computer system 2300 is executing one or more virtual machines, one or more processing units can be assigned to each virtual machine.
[0207] In certain embodiments, the processing acceleration unit 2306 can optionally be used to perform custom processing or to offload some of the processing performed by the processing subsystem 2304, thereby accelerating the overall processing performed by the computer system 2300.
[0208] The I / O subsystem 2308 may include devices and mechanisms for inputting information to the computer system 2300 and / or for outputting information from or via the computer system 2300. Generally, the use of the term input device is intended to include all possible types of devices and mechanisms for inputting information to the computer system 2300. User interface input devices may include, for example, a keyboard, a pointing device such as a mouse or trackball, a touchpad or touchscreen integrated into a display, a scroll wheel, a click wheel, a dial, a button, a switch, a keypad, an audio input device having a voice command recognition system, a microphone, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognition devices, such as Microsoft Kinect motion sensors, Microsoft Xbox 360 game controllers, devices that provide an interface for receiving input using gestures and voice commands. User interface input devices may also include eye gesture recognition devices, such as Google Blink detectors that detect eye activity from a user (e.g., "blinking" when taking a picture and / or making a menu selection) and convert the eye gesture into input for an input device (e.g., Google ). Additionally, user interface input devices may include voice recognition sensing devices that enable a user to interact with a voice recognition system (e.g., Navigator).
[0209] Other examples of user interface input devices include, but are not limited to, three-dimensional (3D) mice, joysticks or trackpoints, gamepads and graphics tablets, and audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers 3D scanners, 3D printers, laser rangefinders, and eye gaze tracking devices. Additionally, user interface input devices may include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasound imaging devices. User interface input devices may also include, for example, audio input devices such as MIDI keyboards, digital musical instruments, and the like.
[0210] In general, the use of the term output device is intended to include all possible types of devices and mechanisms for outputting information from the computer system 2300 to a user or other computer. User interface output devices can include display subsystems, indicator lights, or non-visual displays such as audio output devices. The display subsystem can be a cathode ray tube (CRT), a flat panel device (e.g., a flat panel device using a liquid crystal display (LCD) or a plasma display), a projection device, a touch screen, etc. For example, user interface output devices can include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automotive navigation systems, plotters, voice output devices, and modems.
[0211] The storage subsystem 2318 provides a repository or data store for storing the information and data used by the computer system 2300. The storage subsystem 2318 provides a tangible non-transitory computer-readable storage medium for storing the basic programming and data constructs that provide the functionality of certain examples. The storage subsystem 2318 can store software (e.g., programs, code modules, instructions) that, when executed by the processing subsystem 2304, provides the above-described functionality. This software can be executed by one or more processing units of the processing subsystem 2304. The storage subsystem 2318 can also provide a repository for storing data used in accordance with the teachings of the present disclosure.
[0212] The storage subsystem 2318 can include one or more non-volatile storage devices, including volatile and non-volatile storage devices. As Figure 23 shown, the storage subsystem 2318 includes system memory 2310 and computer-readable storage medium 2322. The system memory 2310 can include multiple memories, including volatile main random access memory (RAM) for storing instructions and data during program execution and non-volatile read-only memory (ROM) or flash memory in which fixed instructions are stored. In some implementations, the basic input / output system (BIOS) is typically stored in the ROM, and the basic input / output system (BIOS) contains basic routines such as those that help transfer information between components within the computer system 2300 during startup. The RAM typically contains the data and / or program modules currently being operated on and executed by the processing subsystem 2304. In some implementations, the system memory 2310 can include various different types of memories, such as static random access memory (SRAM), dynamic random access memory (DRAM), etc.
[0213] By way of example and not limitation, as Figure 23As shown, the system memory 2310 may load the executing application 2312, which may include various applications such as a web browser, a middleware application, a relational database management system (RDBMS), etc., program data 2314, and an operating system 2316. As an example, the operating system 2316 may include various versions of Microsoft Apple and / or Linux operating systems, various commercially available or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google OS, etc.) and / or mobile operating systems such as iOS, Phone, OS, OS, OS operating systems, etc.
[0214] The computer-readable storage medium 2322 may store programming and data constructs that provide some example functionality. The computer-readable medium 2322 may provide storage for computer-readable instructions, data structures, program modules, and other data for the computer system 2300. Software (programs, code modules, instructions) that provides the above functionality when executed by the processing subsystem 2304 may be stored in the storage subsystem 2318. For example, the computer-readable storage medium 2322 may include non-volatile memory such as a hard disk drive, a magnetic disk drive, an optical disk drive (e.g., CD ROM, DVD, disc) or other optical media. The computer-readable storage medium 2322 may include, but is not limited to, drives, flash memory cards, universal serial bus (USB) flash drives, secure digital (SD) cards, DVD discs, digital video tapes, etc. The computer-readable storage medium 2322 may also include solid-state drives (SSDs) based on non-volatile memory (such as flash-based SSDs, enterprise flash drives, solid-state ROMs, etc.), SSDs based on volatile memory (such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs), and hybrid SSDs that use a combination of DRAM and flash-based SSDs.
[0215] In certain embodiments, the storage subsystem 2318 may also include a computer-readable storage medium reader 2320, which may also be connected to the computer-readable storage medium 2322. The reader 2320 may receive data from a storage device such as a disc, a flash drive, etc. and be configured to read data from the storage device.
[0216] In some embodiments, computer system 2300 may support virtualization technologies, including but not limited to virtualization of processing and memory resources. For example, computer system 2300 may provide support for executing one or more virtual machines. In some embodiments, computer system 2300 may execute programs such as hypervisors that facilitate the configuration and management of virtual machines. Each virtual machine may be allocated memory, computing (e.g., processors, cores), I / O, and networking resources. Each virtual machine generally runs independently of other virtual machines. A virtual machine typically runs its own operating system, which may be the same as or different from the operating systems executed by other virtual machines on computer system 2300. Thus, computer system 2300 may run multiple operating systems simultaneously.
[0217] Communication subsystem 2324 provides an interface to other computer systems and networks. Communication subsystem 2324 serves as an interface for receiving data from and sending data to computer system 2300. For example, communication subsystem 2324 may enable computer system 2300 to establish communication channels to one or more client devices via the Internet for receiving information from and sending information to the client devices.
[0218] Communication subsystem 2324 may support wired and / or wireless communication protocols. In some embodiments, communication subsystem 2324 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular phone technologies, advanced data network technologies such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution)), WiFi (IEEE 802.XX series standards, or other mobile communication technologies, or any combination thereof), global positioning system (GPS) receiver components, and / or other components. In some embodiments, in addition to or instead of a wireless interface, communication subsystem 2324 may provide a wired network connection (e.g., Ethernet).
[0219] Communication subsystem 2324 may receive and send data in various forms. In some embodiments, in addition to other forms, communication subsystem 2324 may receive input communications in the form of structured and / or unstructured data feeds 2326, event streams 2328, event updates 2330, etc. For example, communication subsystem 2324 may be configured to receive (or send) data feeds 2326 from users of social media networks and / or other communication services in real time, such as feeds, updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0220] In some embodiments, the communication subsystem 2324 may be configured to receive data in the form of a continuous data stream, which may include an event stream 2328 of real-time events and / or event updates 2330, which may be continuous or unbounded in nature and have no explicit end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, and the like.
[0221] The communication subsystem 2324 may also be configured to transfer data from the computer system 2300 to other computer systems or networks. The data may be transferred in a variety of different forms (e.g., structured and / or unstructured data feeds 2326, event streams 2328, event updates 2330, etc.) to one or more databases that may communicate with one or more stream data source computers coupled to the computer system 2300.
[0222] The computer system 2300 may be one of various types, including handheld portable devices (e.g., cellular phones, computing tablets, PDAs), wearable devices (e.g., Google head-mounted displays), personal computers, workstations, mainframes, kiosks, server racks, or any other data processing system. Due to the ever-changing nature of computers and networks, Figure 23 the description of the computer system 2300 depicted herein is for illustrative purposes only. Many other configurations with more or fewer components than the Figure 23 system shown are possible. Based on the disclosure and teachings provided herein, those of ordinary skill in the art will understand other ways and / or methods of implementing the various examples.
[0223] Although specific examples have been described, various modifications, changes, alternative structures, and equivalents are possible. The examples are not limited to operating in certain specific data processing environments, but may operate freely in multiple data processing environments. Additionally, although some instances have been described using a series of specific transactions and steps, it will be apparent to those skilled in the art that this is not intended to be limiting. Although some flowcharts depict operations as sequential processes, many operations may be performed in parallel or concurrently. Moreover, the order of the operations may be rearranged. The flow may have other steps not included in the figure. The various features and aspects of the above examples may be used alone or in combination.
[0224] In addition, while certain examples have been described using a specific combination of hardware and software, it should be recognized that other combinations of hardware and software are possible. Some examples can be implemented only in hardware, or only in software, or using a combination thereof. The various processes described herein can be implemented in any combination on the same processor or different processors.
[0225] Where a device, system, component, or module is described as being configured to perform certain operations or functions, such configuration can be implemented, for example, by designing an electronic circuit to perform the operations, by programming a programmable electronic circuit (such as a microprocessor) to perform the operations, e.g., by executing computer instructions or code, or a processor or core programmed to execute code or instructions stored on a non-volatile memory medium, or any combination thereof. Processes can communicate using a variety of techniques, including but not limited to conventional techniques for inter-process communication, and different pairs of processes can use different techniques, or the same pair of processes can use different techniques at different times.
[0226] Specific details are given in this disclosure to provide a thorough understanding of the examples. However, the examples can be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary detail to avoid obscuring the examples. This specification provides only examples and is not intended to limit the scope, applicability, or configuration of other examples. Instead, the foregoing description of the examples will provide those skilled in the art with an enabling description for implementing various instances. Various changes can occur in the function and arrangement of the elements.
[0227] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it will be apparent that additions, deletions, omissions, and other modifications and changes can be made without departing from the broader spirit and scope set forth in the claims. Thus, while specific examples have been described, these examples are not intended to be limiting. Various modifications and equivalents are within the scope of the appended claims.
[0228] In the foregoing specification, aspects of the present disclosure have been described with reference to specific instances thereof, but those skilled in the art will recognize that the present disclosure is not limited thereto. The various features and aspects disclosed above can be used alone or in combination. Additionally, the examples can be used in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings are to be regarded as illustrative and not restrictive.
[0229] In the foregoing description, for purposes of illustration, methods were described in a particular order. It should be understood that in alternative examples, these methods may be performed in an order different from the order described. It should also be understood that the above methods may be performed by hardware components or may be implemented in a sequence of machine-executable instructions that may be used to program a machine, such as a general or special-purpose processor or logic circuits, to perform these methods. These machine-executable instructions may be stored on one or more machine-readable media, such as a CD-ROM or other type of optical disc, a floppy disk, a ROM, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, or other types of machine-readable media suitable for storing electronic instructions. Alternatively, these methods may be performed by a combination of hardware and software.
[0230] Where a component is described as being configured to perform certain operations, such configuration may be accomplished, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0231] Although illustrative examples of the present application have been described in detail herein, it should be understood that, except for the limitations of the prior art, the concepts of the present invention may be embodied and used in other ways and that the appended claims are intended to be construed to include such variations.
[0232] Where a component is described as "configured to" perform certain operations, such configuration may be accomplished, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
Claims
1. A method for generating a regular expression, comprising: receiving, by a regular expression generator including one or more processors, first input data including one or more positive character sequences, each of the one or more positive character sequences corresponding to a positive example to be matched with the regular expression generated by the regular expression generator; generating, by the regular expression generator, a first regular expression, wherein the first regular expression matches each of the one or more positive examples; receiving, by the regular expression generator, second input data including one or more negative character sequences, each of the one or more negative character sequences corresponding to a negative example not to be matched with the regular expression generated by the regular expression generator; in response to receiving the second input data, determining whether each of the one or more negative examples matches the first regular expression; and in response to determining that at least one negative example matches the first regular expression: (a) determining a character subsequence at a position within the first regular expression; (b) determining a replacement character sequence, wherein the replacement character sequence at the position within the regular expression differentiates the one or more positive examples from the one or more negative examples; and (c) updating the first regular expression by replacing the determined character subsequence within the first regular expression with the replacement character sequence.
2. The method according to claim 1, wherein determining the character subsequence at the position within the first regular expression comprises: determining the position within the first regular expression; retrieving, from each of the one or more positive examples and from each of the one or more negative examples, a text segment corresponding to the position within the first regular expression; and determining the character subsequence as one or more characters at the position within the first regular expression that differentiates the one or more positive examples from the one or more negative examples.
3. The method according to claim 2, wherein determining the position within the first regular expression comprises: determining a first number of characters at a prefix portion of the first regular expression that differentiates the one or more positive examples from the one or more negative examples; determining a second number of characters at a suffix portion of the first regular expression that differentiates the one or more positive examples from the one or more negative examples; and selecting, at least in part based on whether the first number of characters or the second number of characters is shorter, the prefix portion or the suffix portion as the position within the first regular expression.
4. The method according to claim 3, wherein determining the position within the first regular expression further comprises: executing a formula to determine the position within the first regular expression, wherein the formula weights the prefix portion of the first regular expression relative to the suffix portion.
5. The method according to claim 2, wherein the determined position within the first regular expression is an intermediate span position that does not correspond to a prefix portion or a suffix portion of the first regular expression.
6. The method according to claim 2, wherein determining the replacement character sequence includes determining a plurality of replacement character sequences, and wherein updating the first regular expression includes replacing the determined character subsequence within the first regular expression with the plurality of replacement character sequences.
7. The method according to claim 1, wherein determining the replacement character sequence includes: determining a first number of characters at the position within the first regular expression that can distinguish between the one or more positive examples and the one or more negative examples, and a corresponding first number of replacement character sequences, each replacement character sequence of the corresponding first number of replacement character sequences having the first number of characters; determining a second number of characters at the position within the first regular expression that can distinguish between the one or more positive examples and the one or more negative examples, and a corresponding second number of replacement character sequences, each replacement character sequence of the corresponding second number of replacement character sequences having the second number of characters; and selecting, based on (a) the magnitudes of the first number of characters and the second number of characters and (b) the magnitudes of the corresponding first number of replacement character sequences and the corresponding second number of replacement character sequences, the first number of characters or the second number of characters for the replacement character sequence within the first regular expression.
8. A system for generating a regular expression, the system comprising: a processing unit including one or more processors; and a memory storing instructions that, when executed by the processing unit, cause the system to: receive first input data including one or more positive character sequences, each of the one or more positive character sequences corresponding to a positive example to be matched with the regular expression generated by the regular expression generator; generate a first regular expression, wherein the first regular expression matches each of the one or more positive examples; receive second input data including one or more negative character sequences, each of the one or more negative character sequences corresponding to a negative example that is not to be matched with the regular expression generated by the regular expression generator; in response to receiving the second input data, determine whether each of the one or more negative examples matches the first regular expression; and in response to determining that at least one negative example matches the first regular expression: (a) determine a character subsequence at a position within the first regular expression; (b) determine a replacement character sequence, wherein the replacement character sequence at the position within the regular expression distinguishes between the one or more positive examples and the one or more negative examples; and (c) update the first regular expression by replacing the determined character subsequence within the first regular expression with the replacement character sequence.
9. The system according to claim 8, wherein determining the character subsequence at the position within the first regular expression comprises: determining the position within the first regular expression; retrieving, from each of the one or more positive examples and from each of the one or more negative examples, a text fragment corresponding to the position within the first regular expression; and determining the character subsequence as one or more characters at the position within the first regular expression that can distinguish between the one or more positive examples and the one or more negative examples.
10. The system according to claim 9, wherein determining the position within the first regular expression comprises: determining a first number of characters at a prefix portion of the first regular expression that can distinguish between the one or more positive examples and the one or more negative examples; determining a second number of characters at a suffix portion of the first regular expression that can distinguish between the one or more positive examples and the one or more negative examples; and selecting, at least in part based on whether the first number of characters or the second number of characters is shorter, the prefix portion or the suffix portion as the position within the first regular expression.
11. The system according to claim 10, wherein determining the position within the first regular expression further comprises: executing a formula to determine the position within the first regular expression, wherein the formula weights the prefix portion of the first regular expression relative to the suffix portion.
12. The system according to claim 9, wherein the determined position within the first regular expression is an intermediate-span position that does not correspond to the prefix portion of the first regular expression or the suffix portion of the first regular expression.
13. The system according to claim 9, wherein determining the replacement character sequence comprises determining a plurality of replacement character sequences, and wherein updating the first regular expression comprises replacing the determined character subsequence within the first regular expression with the plurality of replacement character sequences.
14. The system according to claim 8, wherein determining the replacement character sequence comprises: determining a first number of characters at the position within the first regular expression that can distinguish between the one or more positive examples and the one or more negative examples and a corresponding first number of replacement character sequences, each of the corresponding first number of replacement character sequences having the first number of characters; determining a second number of characters at the position within the first regular expression that can distinguish between the one or more positive examples and the one or more negative examples and a corresponding second number of replacement character sequences, each of the corresponding second number of replacement character sequences having the second number of characters; and Select the first number of characters or the second number of characters for the replacement character sequence within the first regular expression based on (a) the sizes of the first number of characters and the second number of characters and (b) the sizes of the corresponding first number of replacement character sequences and the corresponding second number of replacement character sequences.
15. A non-transitory computer-readable medium for generating a regular expression, the computer-readable medium including computer-executable instructions that, when executed on a computer system, cause the computer system to: Receive first input data including one or more positive character sequences, each of the one or more positive character sequences corresponding to a positive example to be matched by a regular expression generated by the regular expression generator; Generate a first regular expression, wherein the first regular expression matches each of the one or more positive examples; Receive second input data including one or more negative character sequences, each of the one or more negative character sequences corresponding to a negative example not to be matched by a regular expression generated by the regular expression generator; In response to receiving the second input data, determine whether each of the one or more negative examples matches the first regular expression; And In response to determining that at least one negative example matches the first regular expression: (a) Determine a character subsequence at a position within the first regular expression; (b) Determine a replacement character sequence, wherein the replacement character sequence at the position within the regular expression differentiates the one or more positive examples from the one or more negative examples; And (c) Update the first regular expression by replacing the determined character subsequence within the first regular expression with the replacement character sequence.
16. The computer-readable medium of claim 15, wherein determining the character subsequence at the position within the first regular expression includes: Determine the position within the first regular expression; Retrieve a text fragment corresponding to the position within the first regular expression from each of the one or more positive examples and from each of the one or more negative examples; And Determine the character subsequence as one or more characters at the position within the first regular expression that differentiates the one or more positive examples and the one or more negative examples.
17. The computer-readable medium of claim 16, wherein determining the position within the first regular expression includes: Determine a first number of characters at a prefix portion of the first regular expression that differentiates the one or more positive examples and the one or more negative examples; Determine a second number of characters at a suffix portion of the first regular expression that differentiates the one or more positive examples and the one or more negative examples; And Select at least in part the prefix portion or the suffix portion as the position within the first regular expression based on whether the first number of characters or the second number of characters is shorter.
18. The computer-readable medium according to claim 17, wherein determining the position within the first regular expression further comprises: executing a formula to determine the position within the first regular expression, wherein the formula weights the prefix portion of the first regular expression relative to the suffix portion.
19. The computer-readable medium according to claim 16, wherein the determined position within the first regular expression is an intermediate-span position that does not correspond to the prefix portion of the first regular expression or the suffix portion of the first regular expression.
20. The computer-readable medium according to claim 16, wherein determining the replacement character sequence includes determining a plurality of replacement character sequences, and wherein updating the first regular expression includes replacing the determined character subsequence within the first regular expression with the plurality of replacement character sequences.
21. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Techniques for similarity analysis and data enrichment using knowledge sources
US10210246B2
Interactive splitting of a column into multiple columns
US20180113894A1
Information extraction and annotation systems and methods for documents
US8856642B1