System and method for automated data schema validation with large language model integration

CA3266172A1Pending Publication Date: 2026-09-21THE TORONTO DOMINION BANK
0 Cites 0 Cited by

Patent Information

Application Number
CA3266172
Authority / Receiving Office
CA · CA
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2026-09-21
Patent Text Reader

Abstract

Systems and methods for automated data schema validation are disclosed. Exemplary implementations may: obtain a data set to be validated with respect to a data validation schema; define a data subset based on a reduced representation of the data set; automatically generate, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema; automatically validate the data subset using the automatically generated set of data validation rules; and output a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules. The data subset may be provided as an input to a large language model using prompt chaining. Resources may be saved compared to known manual approaches, and are easily adaptable to new rule demands.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR AUTOMATED DATA SCHEMA VALIDATION WITH LARGE LANGUAGE MODEL INTEGRATION FIELD

[0001] The present disclosure relates to computing, including but not limited to computing platforms, methods, and storage media for automated data schema validation. BACKGROUND

[0002] In data communications and networking, it is often necessary for servers and applications to send and receive data. Such data can be included in a file or data set that includes an extremely large amount of data, such as tens of thousands of records or a million or more records.

[0003] It can be desirable to validate the data, for example prior to data migration, loading or retrieval. A data schema may define different rules for different columns of data, where validation is to be performed against the data schema. Often rules defining the data schema are hard-coded. Manual validation of data in such large amounts is time-consuming and prohibitive.

[0004] Improvements in approaches for automated data schema validation are desirable. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Embodiments of the present disclosure will now be described, by way of example only, with reference to the attached Figures.

[0006] FIG. 1 illustrates a system configured for automated data schema validation, in accordance with one or more embodiments.

[0007] FIG. 2 illustrates a system configured for automated data schema validation, in accordance with one or more embodiments.

[0008] FIG. 3 illustrates a method for automated data schema validation, in accordance with one or more embodiments.

[0009] FIG. 4 illustrates validation using regular expressions, in accordance with one or more embodiments.

[0010] FIG. 5 illustrates a data validation report, in accordance with one or more2 embodiments.

[0011] FIG. 6 illustrates validation using regular expressions, in accordance with one or more embodiments.

[0012] FIG. 7 illustrates validation using regular expressions from defined instructions, in accordance with one or more embodiments.

[0013] FIG. 8 illustrates a method for automated data schema validation using regular expressions directly from data, in accordance with one or more embodiments.

[0014] FIG. 9 illustrates a method for automated data schema validation using regular expressions directly from data, for example using clustered reduction, in accordance with one or more embodiments.

[0015] FIG. 10 illustrates an example method for automated data schema validation using reduced data, for example using clustered reduction after prompt chaining, in accordance with one or more embodiments.

[0016] FIG. 11 illustrates an example pipeline for automated data schema validation using regular expressions using a hosted solution with APIs, in accordance with one or more embodiments. DETAILED DESCRIPTION

[0017] Computing platforms, methods, and storage media for automated data schema validation are disclosed. Exemplary implementations may: obtain a data set to be validated with respect to a data validation schema; define a data subset of the data set based on a reduced representation of the data set; automatically generate, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema; automatically validate the data subset using the automatically generated set of data validation rules; and output a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[0018] The present disclosure provides a system to ensure data integrity by automatically validating a data schema using regular expressions that employ customizable rules.

[0019] Embodiments of the present disclosure provide a system to automatically validate3 data according to a data schema. In an embodiment, the data schema uses regular expressions that also employ customizable rules, allowing the data schema to be customized. After loading the data, each column may be checked against a customizable column-specific rule. The validation may be performed on a subset of the entire data set. The system may then report the results, for example indicating details on each expression where the validation failed and what the values were.

[0020] There is a technical problem associated with known approaches in that current approaches require hard coding of every single rule for a validation schema. This requires prior knowledge of the validation schema(s), as well as hard coding of associated validation rules prior to performing validation. This limits the capability, and requires recoding for any new variation of the schema. Even attempts at automating such hard coded validation approaches require changes in the source code if there are changes. Embodiments of the present disclosure provide a technical solution of automatically performing data validation on a large amount of data, using a data schema based on a data subset, and in some embodiments based on regular expressions and including customizable rules.

[0021] One aspect of the present disclosure relates to a computing platform configured for automated data schema validation. The computing platform may include a non-transient computer-readable storage medium having executable instructions embodied thereon. The computing platform may include one or more hardware processors configured to execute the instructions. The processor(s) may execute the instructions to obtain a data set to be validated with respect to a data validation schema. The processor(s) may execute the instructions to define a data subset of the data set based on a reduced representation of the data set. The processor(s) may execute the instructions to automatically generate, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. The processor(s) may execute the instructions to automatically validate the data subset using the automatically generated set of data validation rules. The processor(s) may execute the instructions to output a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[0022] Another aspect of the present disclosure relates to a method for automated data4 schema validation. The method may include obtaining a data set to be validated with respect to a data validation schema. The method may include defining a data subset of the data set based on a reduced representation of the data set. The method may include automatically generating, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. The method may include automatically validating the data subset using the automatically generated set of data validation rules. The method may include outputting a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[0023] Yet another aspect of the present disclosure relates to a non-transient computerreadable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method for automated data schema validation. The method may include obtaining a data set to be validated with respect to a data validation schema. The method may include defining a data subset of the data set based on a reduced representation of the data set. The method may include automatically generating, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. The method may include automatically validating the data subset using the automatically generated set of data validation rules. The method may include outputting a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[0024] For the purpose of promoting an understanding of the principles of the disclosure, reference will now be made to the features illustrated in the drawings and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended. Any alterations and further modifications, and any further applications of the principles of the disclosure as described herein are contemplated as would normally occur to one skilled in the art to which the disclosure relates. It will be apparent to those skilled in the relevant art that some features that are not relevant to the present disclosure may not be shown in the drawings for the sake of clarity.

[0025] Certain terms used in this application and their meaning as used in this context are5 set forth in the description below. To the extent a term used herein is not defined, it should be given the broadest definition persons in the pertinent art have given that term as reflected in at least one printed publication or issued patent. Further, the present processes are not limited by the usage of the terms shown below, as all equivalents, synonyms, new developments and terms or processes that serve the same or a similar purpose are considered to be within the scope of the present disclosure.

[0026] FIG. 1 illustrates a system 100 configured for automated data schema validation, in accordance with one or more embodiments. The system 100 comprises a data validator 102 configured to obtain a data set 104 to be validated with respect to a data validation schema. The data set may comprise any set of data for which validation is desirable prior to an operation such as a data transfer, and may be obtained from one or more local or remote storage locations, as one or more files or data stored in a database or other storage. The data validator 102 may comprise a non-transient computer-readable storage medium having executable instructions embodied thereon, and may comprise one or more hardware processors configured to execute the instructions to perform a method as described herein.

[0027] The data validator 102 may be configured to define a data subset of the data set based on a reduced representation of the data set. A goal of the reduced representation of the data set may be to reduce the size of the data set while still preserving the most important information. The reduced representation of the data set may be defined based on a data reduction technique configured to reduce the size of the data set in order to facilitate the automatic generation of data validation rules, preferably without substantially affecting or sacrificing accuracy.

[0028] For example, the data subset and the reduced representation of the data set may be defined using dimensionality reduction, numerosity reduction, data compression, or other data reduction techniques. The data subset and the reduced representation of the data set may be defined using a filter strategy using a selection metric defined on the basis of some clusters or marginal points. After finding the subset of data, the performance of the reduced representation of the data may be checked using a machine learning algorithm. The data subset and the reduced representation of the data set may be defined using a wrapper technique, for example using a classification algorithm for data selection. For example, based on an inference drawn from the6 classification algorithm, data may be added or removed from the subset.

[0029] The data validator 102 may comprise a data validation rule generator 106 configured to automatically generate, with respect to the data subset, a set of data validation rules for the data set. The set of data validation rules may comprise regular expressions defining customizable column-specific rules and defining the data validation schema. A regular expression may be defined as a pattern, text string or sequence of characters for describing a search pattern.

[0030] The data validator 102 may be configured to automatically validate the data subset using the automatically generated set of data validation rules. The data validator may be configured to output a data validation report 108 indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules. Outputting the data validation report 108 assists in identifying instances of failed data validation, so that the failed validation operations may be corrected. The identified failed operations may also be used to improve generation or re-generation of the rules, for example by providing the failed operations as an input to an artificial intelligence system, or machine learning system.

[0031] A data validator 102, or schema validation tool, in accordance with one or more embodiments may be configured to ensure data integrity across multiple formats including CSV, PSV, SQL, Oracle DB etc. By harnessing the power of regular expressions, which are patterns used to match character combinations in strings, embodiments of the present disclosure may provide improved flexibility and precision in schema validation. Whether the scenario comprises migrating data between systems, integrating data from multiple sources or simply ensuring data consistency, the system may provide comprehensive reporting and customizable rules.

[0032] In an embodiment, the system 100 is configured to generate the validation rules automatically or on demand, using artificial intelligence (AI) including a large language model (LLM), based on one or more of several approaches, for example: using the advanced concept of semantic data clustering; or automatic rule generation based on prompt chaining.

[0033] A number of different use cases will be discussed below in relation to the figures. The use cases include automatic generation of regular expressions / rules in various manners.

[0034] Embodiments of the present disclosure may be used in one or more of the following7 use cases: Data Migration: ensuring data integrity and format consistency during database migrations; Data Ingestion: validating incoming data from various sources, to ensure it meets required formats before processing and storage; ETL Processes: validate data at each stage of Extract, Transform, Load processes to ensure consistency and correctness.

[0035] Embodiments of the present disclosure may provide one or more of the following benefits or advantages compared to known approaches: saves resources and / or effort as compared to manual validation, since it is practically impossible to test large data manually; usage of regular expression avoids hardcoding of rules at the code level, making it easily adaptable to any new rule demands; ability to assign any number of rules to a single column; optionally using AI to improve data validation; reduces any prework of configuring complex rules; generating rules without any prompt or human intervention; autonomously flagging suspicious data entries; seamless integration with CI / CD (continuous integration and continuous delivery / deployment).

[0036] FIG. 2 illustrates a system 200 configured for automated data schema validation, in accordance with one or more embodiments. In some embodiments, system 200 may include one or more computing platforms 202. Computing platform(s) 202 may be configured to communicate with one or more remote platforms 204 according to a client / server architecture, a peer-to-peer architecture, and / or other architectures. Remote platform(s) 204 may be configured to communicate with other remote platforms via computing platform(s) 202 and / or according to a client / server architecture, a peer-to-peer architecture, and / or other architectures. Users may access system 200 via remote platform(s) 204.

[0037] Computing platform(s) 202 may be configured by machine-readable instructions 206. Machine-readable instructions 206 may include one or more instruction modules. The instruction modules may include computer program modules. The instruction modules may include one or more of data set obtaining module 208, data subset defining module 210, data validation rule generation module 212, data subset validating module 214, data validation report outputting module 216, optional LLM integration module 218, and / or other instruction modules.

[0038] Data set obtaining module 208 may be configured to obtain a data set to be validated with respect to a data validation schema. The data set may be obtained from a local or remote8 data source, and may be stored as one or more files or as data in one or more databases.

[0039] Data subset defining module 210 may be configured to define a data subset of the data set based on a reduced representation of the data set. Some examples of defining the data subset based on a reduced representation of the data set may be based on data sampling, dimensionality reduction, data compression, data discretization or feature selection.

[0040] Data validation rule generation module 212 may be configured to automatically generate, with respect to the data subset, a set of data validation rules for the data set. The set of data validation rules may include regular expressions defining customizable column-specific rules and defining the data validation schema. The data validation rules may be automatically generated based on analysis of one or more column headers in the data subset. The regular expressions may be generated directly from the data subset. The regular expressions may be automatically generated directly from the data subset independent of any validation instructions or user intervention. The regular expressions may be generated based on defined instructions from one or more of emails, requirement documents, and direct specifications. The data validation rules may include rules automatically generated based on defined instructions from one or more of emails, requirement documents, and direct specifications.

[0041] In an example embodiment, based on an automated review of emails, requirement documents, and / or direct specifications, the data validation rule generation module 212 may be configured to automatically generate the regular expressions and / or the set of data validation rules. For example, if a requirement document specifies that Column F of a data set is to be formatted in a date format of YYYY-MM-DD, the data validation rule generation module 212 may be configured to automatically generate an associated regular expression, and may automatically generate a data validation rule for the expected date format, including the generated regular expression. In another example, if a set of emails is reviewed and determined to include a specification that an email sender is verified with respect to having a domain @example.com, then the data validation rule generation module 212 may be configured to automatically generate a suitable regular expression, and may automatically generate a data validation rule for the expected domain in an email sender field, including the generated regular expression.

[0042] Data subset validating module 214 may be configured to automatically validate the9 data subset using the automatically generated set of data validation rules. The data subset may be automatically validated using correlation matching with in-memory context. Data subset validating module 214 may be configured to automatically validate a remaining portion of the data set in addition to the data subset using the automatically generated set of data validation rules. For example, after data subset validating module 214 completes automatic validation of the data subset, data subset validating module 214 may proceed with automatically validating a remaining portion of the data set, which may comprise the rest of the data set that is not a part of the data subset that was already automatically validated.

[0043] Data validation report outputting module 216 may be configured to output a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules. The data validation report may indicate results from the data subset and the remaining portion of the data set that failed the validation based on the customizable column-specific rules. Contents of the data validation report may be used to identify failed validation results that require correction, so that correction may be completed prior to completion of a data transfer operation or other type of operation.

[0044] In an example embodiment, the system 200 may comprise large language model integration module 218. LLM integration module 218 may be configured to provide the data subset as an input to a large language model using prompt chaining, and to receive, as a large language model output, the automatically generated set of data validation rules. LLM integration module 218 may be configured to transform the data subset to a data format compatible with prompt chaining for the large language model, for example prior to providing the data subset as an input to the LLM. In an implementation, LLM integration module 218 may be configured to enable the data validation report outputting module 216 to provide the data validation report as an input to an LLM, so that the LLM may provide an output with suggested modifications to the automatically generated data validation rules in order to reduce the occurrence of data validation failures, thereby improving system performance and reducing processor and memory usage.

[0045] In some embodiments, computing platform(s) 202, remote platform(s) 204, and / or external resources 220 may be operatively linked via one or more electronic communication links. For example, such electronic communication links may be established, at least in part,10 via a network such as the Internet and / or other networks. It will be appreciated that this is not intended to be limiting, and that the scope of this disclosure includes implementations in which computing platform(s) 202, remote platform(s) 204, and / or external resources 220 may be operatively linked via some other communication media.

[0046] A given remote platform 204 may include one or more processors configured to execute computer program modules. The computer program modules may be configured to enable an expert or user associated with the given remote platform 204 to interface with system 200 and / or external resources 220, and / or provide other functionality attributed herein to remote platform(s) 204. By way of non-limiting example, a given remote platform 204 and / or a given computing platform 202 may include one or more of a server, a desktop computer, a laptop computer, a handheld computer, a tablet computing platform, a NetBook, a Smartphone, a gaming console, and / or other computing platforms.

[0047] External resources 220 may include sources of information outside of system 200, external entities participating with system 200, and / or other resources. In some embodiments, some or all of the functionality attributed herein to external resources 220 may be provided by resources included in system 200.

[0048] Computing platform(s) 202 may include electronic storage 222, one or more processors 224, and / or other components. Computing platform(s) 202 may include communication lines, or ports to enable the exchange of information with a network and / or other computing platforms. Illustration of computing platform(s) 202 in FIG. 2 is not intended to be limiting. Computing platform(s) 202 may include a plurality of hardware, software, and / or firmware components operating together to provide the functionality attributed herein to computing platform(s) 202. For example, computing platform(s) 202 may be implemented by a cloud of computing platforms operating together as computing platform(s) 202.

[0049] Electronic storage 222 may comprise non-transitory storage media that electronically stores information. The electronic storage media of electronic storage 222 may include one or both of system storage that is provided integrally (i.e., substantially nonremovable) with computing platform(s) 202 and / or removable storage that is removably connectable to computing platform(s) 202 via, for example, a port (e.g., a USB port, a firewire port, etc.) or a drive (e.g., a disk drive, etc.). Electronic storage 222 may include one or more11 of optically readable storage media (e.g., optical disks, etc.), magnetically readable storage media (e.g., magnetic tape, magnetic hard drive, floppy drive, etc.), electrical charge-based storage media (e.g., EEPROM, RAM, etc.), solid-state storage media (e.g., flash drive, etc.), and / or other electronically readable storage media. Electronic storage 222 may include one or more virtual storage resources (e.g., cloud storage, a virtual private network, and / or other virtual storage resources). Electronic storage 222 may store software algorithms, information determined by processor(s) 224, information received from computing platform(s) 202, information received from remote platform(s) 204, and / or other information that enables computing platform(s) 202 to function as described herein.

[0050] Processor(s) 224 may be configured to provide information processing capabilities in computing platform(s) 202. As such, processor(s) 224 may include one or more of a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information. Although processor(s) 224 is shown in FIG. 2 as a single entity, this is for illustrative purposes only. In some embodiments, processor(s) 224 may include a plurality of processing units. These processing units may be physically located within the same device, or processor(s) 224 may represent processing functionality of a plurality of devices operating in coordination. Processor(s) 224 may be configured to execute modules 208, 210, 212, 214, 216, and / or 218, and / or other modules. Processor(s) 224 may be configured to execute modules 208, 210, 212, 214, 216, and / or 218, and / or other modules by software; hardware; firmware; some combination of software, hardware, and / or firmware; and / or other mechanisms for configuring processing capabilities on processor(s) 224. As used herein, the term “module” may refer to any component or set of components that perform the functionality attributed to the module. This may include one or more physical processors during execution of processor readable instructions, the processor readable instructions, circuitry, hardware, storage media, or any other components.

[0051] It should be appreciated that although modules 208, 210, 212, 214, 216, and / or 218 are illustrated in FIG. 2 as being implemented within a single processing unit, in embodiments in which processor(s) 224 includes multiple processing units, one or more of modules 208, 210, 212, 214, 216, and / or 218 may be implemented remotely from the other modules. The12 description of the functionality provided by the different modules 208, 210, 212, 214, 216, and / or 218 described below is for illustrative purposes, and is not intended to be limiting, as any of modules 208, 210, 212, 214, 216, and / or 218 may provide more or less functionality than is described. For example, one or more of modules 208, 210, 212, 214, 216, and / or 218 may be eliminated, and some or all of its functionality may be provided by other ones of modules 208, 210, 212, 214, 216, and / or 218. As another example, processor(s) 224 may be configured to execute one or more additional modules that may perform some or all of the functionality attributed below to one of modules 208, 210, 212, 214, 216, and / or 218.

[0052] FIG. 3 illustrates a method 300 for automated data schema validation, in accordance with one or more embodiments. The operations of method 300 presented below are intended to be illustrative. In some embodiments, method 300 may be accomplished with one or more additional operations not described, and / or without one or more of the operations discussed. Additionally, the order in which the operations of method 300 are illustrated in FIG. 3 and described below is not intended to be limiting.

[0053] In some embodiments, method 300 may be implemented in one or more processing devices (e.g., a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information). The one or more processing devices may include one or more devices executing some or all of the operations of method 300 in response to instructions stored electronically on an electronic storage medium. The one or more processing devices may include one or more devices configured through hardware, firmware, and / or software to be specifically designed for execution of one or more of the operations of method 300.

[0054] An operation 302 may include obtaining a data set to be validated with respect to a data validation schema. Operation 302 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to data set obtaining module 208, in accordance with one or more embodiments.

[0055] An operation 304 may include defining a data subset of the data set based on a reduced representation of the data set. Operation 304 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is13 the same as or similar to data subset defining module 210, in accordance with one or more embodiments.

[0056] An operation 306 may include automatically generating, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. Operation 306 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to data validation rule generation module 212, in accordance with one or more embodiments.

[0057] An operation 308 may include automatically validating the data subset using the automatically generated set of data validation rules. Operation 308 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to data subset validating module 214, in accordance with one or more embodiments.

[0058] An operation 310 may include outputting a data validation report indicating results from the subset of the data set that failed the validation based on the customizable columnspecific rules. Operation 310 may be performed by one or more hardware processors configured by machine-readable instructions including a module that is the same as or similar to data validation report outputting module 216, in accordance with one or more embodiments.

[0059] FIG. 4 illustrates validation using regular expressions, in accordance with one or more embodiments. A user interface, template or other indicator may specify a data set or file that includes the data to be validated. For example, the user interface may include a link such as “Get Columns” that, when activated, automatically fetches all of the data to be validated. As shown in FIG. 4, a rule may be specified in the form of regular expressions. Use of regular expressions provides a highly flexible approach, without a need for hard coding. For example: a regular expression for column 1 shows an alphanumeric value with 4 characters, shown as “^[a-zA-Z0-9][4]$”; the character “^” may be used to match the beginning of a string, and “$” may be used to match the end of a string. A “Generate Rules” indicator may, when activated, initiate rule generation, for example if there is an issue with the automatic generation or such rules, or simply to initiate the automatic generation of the rules. A “Validate” indicator may, when activated, initiate validation of the data subset based on the automatically generated rules.14 This may be used, for example, if there is an issue with automatically performing the validation, or simply to initiate the automatic performance of the validation.

[0060] A regular expression in FIG. 4 for “BusinessDate” shows a date in specific format YYYMMDD. A regular expression for Column 2 in FIG. 4 shows that the value can be all numbers and likely be of any length. A rule may be defined to specify that the system is to test an entire row, or only a data subset of the data set, such as the first 10 characters or first 10 rows. A user may specify rules in plain English, and click on Generate Rules button, as shown in FIG. 4, to generate a version of the rules using regular expressions. In an example embodiment, the system may use an LLM to generate regular expressions based on Englishlanguage requirements, as described in relation to FIG. 2.

[0061] FIG. 5 illustrates a data validation report, in accordance with one or more embodiments. A system, such as system 100 of FIG. 1 or system 200 of FIG. 2, may be configured to output a report, for example as shown in FIG. 5, indicating every single expression where it failed, and what the values were. As indicated earlier, the report may be used as a means for correcting entries that failed the validation, and optionally as an input to a machine learning system to generate an output to be used to improve the automatic validation rule generation to reduce the occurrence of validation failures.

[0062] FIG. 6 illustrates validation using regular expressions, in accordance with one or more embodiments. The example in FIG. 6 shows correlation matching in a memory context. Using correlation matching, a system in accordance with one or more embodiments may retain a context of all of the columns, and a plain regular expression may not be desirable, so the system may use a custom algorithm that also employs regular expressions. For example, as shown in FIG. 6, Column2 is defined with respect to, and contains, elements from other columns. The system may have an in-memory context of the entire row and can be validated. The system may automatically learn the run-time value of other values / parameters, for example using correlation matching. The system may use a method to retain the context of the other columns. Based on the context, the rules may be automatically generated. There may be a certain type of request where a regular expression cannot be determined at runtime. Using correlation matching, a certain column must contain first few characters of another column. There may not be a regular expression that can validate this, so the system may be configured15 to create a custom solution that uses a regular expression. The example of FIG. 6 can use single rules or multiple rules, and may match a part of Column1 and a part of Column2 with some regular expression (regex).

[0063] FIG. 7 illustrates validation using regular expressions from defined instructions, in accordance with one or more embodiments. A system in accordance with one or more embodiments may be configured to obtain data or instructions from emails, from requirements documentation, from directly specified instructions, or other mechanisms. The system may be configured to automatically generate regular expressions based on the obtained defined instructions from one or more of emails, requirement documents, and direct specifications, as described earlier in relation to FIG. 2. The system may provide the obtained data to an LLM to automatically generate regular expressions, to leverage the LLM to automate certain features.

[0064] FIG. 8 illustrates a method for automated data schema validation using regular expressions directly from data, in accordance with one or more embodiments. The example in FIG. 8 shows automatic generation of regular expressions / rules directly from data. The method of FIG. 8 may be performed autonomously by a machine, without manual intervention of humans. After retrieving the data, a system performing the method shown in FIG. 8 may be configured to reduce the data, for example reducing from 30,000 rows to the top 5 rows. Then, in an example embodiment using prompt chaining with an LLM, the system may be configured to generalize this into plain language using general classifications, and generate rules, and create final rules for schema validation. The system may use a subset of rows, or a data subset, to generate data / rules for the entire data set.

[0065] Without data reduction, it is too time consuming to generate data validation rules based on all of the rows of data. The method of FIG. 8 may automatically create a data validation schema without even being given any instructions. The system may be configured to obtain a data set, reduce the data set based on one of the strategies, for example defining a data subset, then automatically generate rules for the data, and automatically test the remaining data based on those rules.

[0066] The particular method shown in FIG. 8 is not designed to specify any of the rules or instructions. The system is configured to automatically classify, create rules on its own, and perform validation. For example, the method may be deployed at night, and be configured to16 automatically provide an output in the morning of the results of the applied algorithm and testing. The system may be configured to see a subset of the data, go to the prompt chaining pipeline and automatically generate rules without specifying the column (e.g. column called email). Passing the data through a prompt chaining pipeline enables the method to automatically create the rules.

[0067] FIG. 9 illustrates a method for automated data schema validation using regular expressions directly from data, for example using clustered reduction, in accordance with one or more embodiments. The example workflow of FIG. 9 may be configured to leverage a data clustering algorithm, as part of the system / algorithm to automatically generate rules, in particular prior to prompt chaining steps as described herein.

[0068] FIG. 10 illustrates an example method for automated data schema validation using reduced data, for example using clustered reduction after prompt chaining, in accordance with one or more embodiments. The example in FIG. 10 shows automatic generation of regular expressions directly from data, for example using clustered reduction after the prompt chaining, in contrast to FIG. 9. This enables autonomous execution on the reduced data for immediate validation. It can be desirable, whether with manual or automatic generation, to perform tasks on a smaller set or subset of data, to have a first pass of testing without incurring time delays. In the embodiment of FIG. 10, the final schema validation may be performed on a smaller set of data. Using clustered reduction, the method may be configured to generate the smaller set of data.

[0069] FIG. 11 illustrates an example pipeline for automated data schema validation using regular expressions using a hosted solution with APIs, in accordance with one or more embodiments. For all of the aspects described in relation to FIG. 4 through FIG. 10, these implementations may be performed in a programmable manner, for example as shown in FIG. 11 using a hosted solution with APIs. A schema validation pipeline may receive, via an API, data sources and conditions, perform the schema validation steps, and return the results to the API, which may then send a response including the test subset.

[0070] Embodiments of the present disclosure provide a hosted schema validation system and method with automatic rule generation and integration. The schema validation tool, using a system and method, may use regular expressions to validate data entries. The tool allows for17 the automatic generation of rules based on English language requirements, for example using an LLM. Embodiments of the present disclosure provide a hosted solution with APIs that integrate with the DevOps pipeline and various data sources. The system may automatically generate regular expressions and rules from data without specifying instructions. Reducing the data before generating rules helps ensure accurate validation and prevent overfitting.

[0071] Embodiments of the present disclosure provide for automatic generation of regular expressions directly from data without manual intervention of humans. The approach may include an entire workflow, including random and top classifications, through automatic generation and testing. Data clustering and structured data reduction methods may be used to generate rules for schema validation. Embodiments of the present disclosure may be provided as a hosted solution with APIs and a web interface for automation of schema validation.

[0072] In the preceding description, for purposes of explanation, numerous details are set forth in order to provide a thorough understanding of the embodiments. However, it will be apparent to one skilled in the art that these specific details are not required. In other instances, well-known electrical structures and circuits are shown in block diagram form in order not to obscure the understanding. For example, specific details are not provided as to whether the embodiments described herein are implemented as a software routine, hardware circuit, firmware, or a combination thereof.

[0073] Embodiments of the disclosure can be represented as a computer program product stored in a machine-readable medium (also referred to as a computer-readable medium, a processor-readable medium, or a computer usable medium having a computer-readable program code embodied therein). The machine-readable medium can be any suitable tangible, non-transitory medium, including magnetic, optical, or electrical storage medium including a compact disk read only memory (CD-ROM), digital versatile disk (DVD), Blu-ray Disc Read Only Memory (BD-ROM), memory device (volatile or non-volatile), or similar storage mechanism. The machine-readable medium can contain various sets of instructions, code sequences, configuration information, or other data, which, when executed, cause a processor to perform steps in a method according to an embodiment of the disclosure. Those of ordinary skill in the art will appreciate that other instructions and operations necessary to implement the described implementations can also be stored on the machine-readable medium. The18 instructions stored on the machine-readable medium can be executed by a processor or other suitable processing device, and can interface with circuitry to perform the described tasks.

[0074] The above-described embodiments are intended to be examples only. Alterations, modifications and variations can be effected to the particular embodiments by those of skill in the art without departing from the scope, which is defined solely by the claims appended hereto.

[0075] Embodiments of the disclosure can be described with reference to the following clauses, with specific features laid out in the dependent clauses.

[0076] One aspect of the present disclosure relates to a system configured for automated data schema validation. The system may include one or more hardware processors configured by machine-readable instructions. The processor(s) may be configured to obtain a data set to be validated with respect to a data validation schema. The processor(s) may be configured to define a data subset of the data set based on a reduced representation of the data set. The processor(s) may be configured to automatically generate, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. The processor(s) may be configured to automatically validate the data subset using the automatically generated set of data validation rules. The processor(s) may be configured to output a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[0077] In some implementations of the system, the processor(s) may be configured to provide the data subset as an input to a large language model using prompt chaining. In some implementations of the system, the processor(s) may be configured to receive, as a large language model output, the automatically generated set of data validation rules.

[0078] In some implementations of the system, the processor(s) may be configured to transform the data subset to a data format compatible with prompt chaining for the large language model.

[0079] In some implementations of the system, the processor(s) may be configured to receive the automatically generated set of data validation rules comprising rules automatically generated based on analysis of one or more column headers in the data subset.

[0080] In some implementations of the system, the processor(s) may be configured to19 automatically generate the regular expressions directly from the data subset.

[0081] In some implementations of the system, the processor(s) may be configured to automatically generate the regular expressions directly from the data subset independent of any validation instructions or user intervention.

[0082] In some implementations of the system, the processor(s) may be configured to automatically generate the regular expressions based on defined instructions from one or more of emails, requirement documents, and direct specifications.

[0083] In some implementations of the system, the processor(s) may be configured to automatically validate a remaining portion of the data set in addition to the data subset using the automatically generated set of data validation rules. In some implementations of the system, the processor(s) may be configured to output a data validation report indicating results from the data subset and the remaining portion of the data set that failed the validation based on the customizable column-specific rules.

[0084] In some implementations of the system, the processor(s) may be configured to automatically validate the subset of the data set using correlation matching with in-memory context.

[0085] Another aspect of the present disclosure relates to a processor-implemented method for automated data schema validation. The method may include obtaining a data set to be validated with respect to a data validation schema. The method may include defining a data subset of the data set based on a reduced representation of the data set. The method may include automatically generating, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. The method may include automatically validating the data subset using the automatically generated set of data validation rules. The method may include outputting a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[0086] In some implementations of the method, it may include providing the data subset as an input to a large language model using prompt chaining. In some implementations of the method, it may include receiving, as a large language model output, the automatically generated20 set of data validation rules.

[0087] In some implementations of the method, it may include transforming the data subset to a data format compatible with prompt chaining for the large language model.

[0088] In some implementations of the method, it may include receiving the automatically generated set of data validation rules comprising rules automatically generated based on analysis of one or more column headers in the data subset.

[0089] In some implementations of the method, it may include automatically generating the regular expressions directly from the data subset.

[0090] In some implementations of the method, it may include automatically generating the regular expressions directly from the data subset independent of any validation instructions or user intervention.

[0091] In some implementations of the method, it may include automatically generating the regular expressions based on defined instructions from one or more of emails, requirement documents, and direct specifications.

[0092] In some implementations of the method, it may include automatically validating a remaining portion of the data set in addition to the data subset using the automatically generated set of data validation rules. In some implementations of the method, it may include outputting a data validation report indicating results from the data subset and the remaining portion of the data set that failed the validation based on the customizable column-specific rules.

[0093] In some implementations of the method, it may include automatically validating the subset of the data set using correlation matching with in-memory context.

[0094] Yet another aspect of the present disclosure relates to a non-transient computerreadable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method for automated data schema validation. The method may include obtaining a data set to be validated with respect to a data validation schema. The method may include defining a data subset of the data set based on a reduced representation of the data set. The method may include automatically generating, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. The method may include automatically validating the data subset21 using the automatically generated set of data validation rules. The method may include outputting a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[0095] In some implementations of the computer-readable storage medium, the method may include providing the data subset as an input to a large language model using prompt chaining. In some implementations of the computer-readable storage medium, the method may include receiving, as a large language model output, the automatically generated set of data validation rules.

[0096] In some implementations of the computer-readable storage medium, the method may include transforming the data subset to a data format compatible with prompt chaining for the large language model.

[0097] In some implementations of the computer-readable storage medium, the method may include receiving the automatically generated set of data validation rules comprising rules automatically generated based on analysis of one or more column headers in the data subset.

[0098] In some implementations of the computer-readable storage medium, the method may include automatically generating the regular expressions directly from the data subset.

[0099] In some implementations of the computer-readable storage medium, the method may include automatically generating the regular expressions directly from the data subset independent of any validation instructions or user intervention.

[00100] In some implementations of the computer-readable storage medium, the method may include automatically generating the regular expressions based on defined instructions from one or more of emails, requirement documents, and direct specifications.

[00101] In some implementations of the computer-readable storage medium, the method may include automatically validating a remaining portion of the data set in addition to the data subset using the automatically generated set of data validation rules. In some implementations of the computer-readable storage medium, the method may include outputting a data validation report indicating results from the data subset and the remaining portion of the data set that failed the validation based on the customizable column-specific rules.

[00102] In some implementations of the computer-readable storage medium, the method may include automatically validating the subset of the data set using correlation matching with22 in-memory context.

[00103] Still another aspect of the present disclosure relates to a system configured for automated data schema validation. The system may include means for obtaining a data set to be validated with respect to a data validation schema. The system may include means for defining a data subset of the data set based on a reduced representation of the data set. The system may include means for automatically generating, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. The system may include means for automatically validating the data subset using the automatically generated set of data validation rules. The system may include means for outputting a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[00104] In some implementations of the system, the system may include means for providing the data subset as an input to a large language model using prompt chaining. In some implementations of the system, the system may include means for receiving, as a large language model output, the automatically generated set of data validation rules.

[00105] In some implementations of the system, the system may include means for transforming the data subset to a data format compatible with prompt chaining for the large language model.

[00106] In some implementations of the system, the system may include means for receiving the automatically generated set of data validation rules comprising rules automatically generated based on analysis of one or more column headers in the data subset.

[00107] In some implementations of the system, the system may include means for automatically generating the regular expressions directly from the data subset.

[00108] In some implementations of the system, the system may include means for automatically generating the regular expressions directly from the data subset independent of any validation instructions or user intervention.

[00109] In some implementations of the system, the system may include means for automatically generating the regular expressions based on defined instructions from one or23 more of emails, requirement documents, and direct specifications.

[00110] In some implementations of the system, the system may include means for automatically validating a remaining portion of the data set in addition to the data subset using the automatically generated set of data validation rules. In some implementations of the system, the system may include means for outputting a data validation report indicating results from the data subset and the remaining portion of the data set that failed the validation based on the customizable column-specific rules.

[00111] In some implementations of the system, the system may include means for automatically validating the subset of the data set using correlation matching with in-memory context.

[00112] Even another aspect of the present disclosure relates to a computing platform configured for automated data schema validation. The computing platform may include a nontransient computer-readable storage medium having executable instructions embodied thereon. The computing platform may include one or more hardware processors configured to execute the instructions. The processor(s) may execute the instructions to obtain a data set to be validated with respect to a data validation schema. The processor(s) may execute the instructions to define a data subset of the data set based on a reduced representation of the data set. The processor(s) may execute the instructions to automatically generate, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema. The processor(s) may execute the instructions to automatically validate the data subset using the automatically generated set of data validation rules. The processor(s) may execute the instructions to output a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

[00113] In some implementations of the computing platform, the processor(s) may execute the instructions to provide the data subset as an input to a large language model using prompt chaining. In some implementations of the computing platform, the processor(s) may execute the instructions to receive, as a large language model output, the automatically generated set of data validation rules.

[00114] In some implementations of the computing platform, the processor(s) may execute24 the instructions to transform the data subset to a data format compatible with prompt chaining for the large language model.

[00115] In some implementations of the computing platform, the processor(s) may execute the instructions to receive the automatically generated set of data validation rules comprising rules automatically generated based on analysis of one or more column headers in the data subset.

[00116] In some implementations of the computing platform, the processor(s) may execute the instructions to automatically generate the regular expressions directly from the data subset.

[00117] In some implementations of the computing platform, the processor(s) may execute the instructions to automatically generate the regular expressions directly from the data subset independent of any validation instructions or user intervention.

[00118] In some implementations of the computing platform, the processor(s) may execute the instructions to automatically generate the regular expressions based on defined instructions from one or more of emails, requirement documents, and direct specifications.

[00119] In some implementations of the computing platform, the processor(s) may execute the instructions to automatically validate a remaining portion of the data set in addition to the data subset using the automatically generated set of data validation rules. In some implementations of the computing platform, the processor(s) may execute the instructions to output a data validation report indicating results from the data subset and the remaining portion of the data set that failed the validation based on the customizable column-specific rules.

[00120] In some implementations of the computing platform, the processor(s) may execute the instructions to automatically validate the subset of the data set using correlation matching with in-memory context.

[00121] A further aspect of the present disclosure relates to a system configured for automated data schema validation. The system may include one or more hardware processors configured by machine-readable instructions. The processor(s) may be configured to obtain a data set to be validated with respect to a data validation schema. The processor(s) may be configured to define a data subset of the data set based on a reduced representation of the data set. The processor(s) may be configured to automatically generate, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules defining the25 data validation schema. The processor(s) may be configured to automatically validate the data subset using the automatically generated set of data validation rules. The processor(s) may be configured to output a data validation report indicating results from the subset of the data set that failed the validation based on the set of data validation rules.

Claims

CLAIMS:

1. An apparatus configured for automated data schema validation, the apparatus comprising: a non-transient computer-readable storage medium having executable instructions embodied thereon; and one or more hardware processors configured to execute the instructions to: obtain a data set to be validated with respect to a data validation schema; define a data subset of the data set based on a reduced representation of the data set; automatically generate, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema; automatically validate the data subset using the automatically generated set of data validation rules; and output a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

2. The apparatus of claim 1 wherein the one or more hardware processors are further configured to execute the instructions to: provide the data subset as an input to a large language model using prompt chaining; and receive, as a large language model output, the automatically generated set of data validation rules.

3. The apparatus of claim 2 wherein the one or more hardware processors are further configured to execute the instructions to:27 transform the data subset to a data format compatible with prompt chaining for the large language model.

4. The apparatus of claim 1 wherein the one or more hardware processors are further configured to execute the instructions to: receive the automatically generated set of data validation rules comprising rules automatically generated based on analysis of one or more column headers in the data subset.

5. The apparatus of claim 1 wherein the one or more hardware processors are further configured to execute the instructions to: automatically generate the regular expressions directly from the data subset.

6. The apparatus of claim 1 wherein the one or more hardware processors are further configured to execute the instructions to: automatically generate the regular expressions directly from the data subset independent of any validation instructions or user intervention.

7. The apparatus of claim 1 wherein the one or more hardware processors are further configured to execute the instructions to: automatically generate the regular expressions based on defined instructions from one or more of emails, requirement documents, and direct specifications.

8. The apparatus of claim 1 wherein the one or more hardware processors are further configured to execute the instructions to: automatically validate a remaining portion of the data set in addition to the data subset using the automatically generated set of data validation rules; and28 output a data validation report indicating results from the data subset and the remaining portion of the data set that failed the validation based on the customizable column-specific rules.

9. The apparatus of claim 1 wherein the one or more hardware processors are further configured to execute the instructions to: automatically validate the subset of the data set using correlation matching with in-memory context.

10. A method for automated data schema validation, comprising: obtaining a data set to be validated with respect to a data validation schema; defining a data subset of the data set based on a reduced representation of the data set; automatically generating, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema; automatically validating the data subset using the automatically generated set of data validation rules; and outputting a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

11. The method of claim 10 further comprising: providing the data subset as an input to a large language model using prompt chaining; and receiving, as a large language model output, the automatically generated set of data validation rules.

12. The method of claim 11 further comprising: transforming the data subset to a data format compatible with prompt chaining for the large language model.29 13. The method of claim 10 further comprising: receiving the automatically generated set of data validation rules comprising rules automatically generated based on analysis of one or more column headers in the data subset.

14. The method of claim 10 further comprising: automatically generating the regular expressions directly from the data subset.

15. The method of claim 10 further comprising: automatically generating the regular expressions directly from the data subset independent of any validation instructions or user intervention.

16. The method of claim 10 further comprising: automatically generating the regular expressions based on defined instructions from one or more of emails, requirement documents, and direct specifications.

17. The method of claim 10 further comprising: automatically validating a remaining portion of the data set in addition to the data subset using the automatically generated set of data validation rules; and outputting a data validation report indicating results from the data subset and the remaining portion of the data set that failed the validation based on the customizable column-specific rules.

18. The method of claim 10 further comprising: automatically validating the subset of the data set using correlation matching with inmemory context.

19. A non-transient computer-readable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method for automated data schema validation, comprising:30 obtaining a data set to be validated with respect to a data validation schema; defining a data subset of the data set based on a reduced representation of the data set; automatically generating, with respect to the data subset, a set of data validation rules for the data set, the set of data validation rules comprising regular expressions defining customizable column-specific rules and defining the data validation schema; automatically validating the data subset using the automatically generated set of data validation rules; and outputting a data validation report indicating results from the subset of the data set that failed the validation based on the customizable column-specific rules.

20. The non-transient computer-readable storage medium of claim 19 wherein the method further comprises: providing the data subset as an input to a large language model using prompt chaining; and receiving, as a large language model output, the automatically generated set of data validation rules.