System and method for data processing
By separating source data, transformation, and output definitions in data processing, the complexity and cost of generating financial reports are reduced, enhancing accuracy and efficiency while allowing for software reuse.
Patent Information
- Application Number
- PCT/AU2024/051402
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-26
AI Technical Summary
Current data processing methods for generating financial reports are complex, resource-intensive, and prone to errors, often requiring custom applications that are costly to develop and difficult to modify.
A data processing method that separates source data definitions, transformation definitions, and output definitions, allowing for easy updating and reuse, and utilizing predefined operators to transform data into output data.
This approach simplifies data processing, improves accuracy and efficiency, and allows for the reuse of software across different data types and purposes, reducing the need for costly custom applications.
Smart Images

Figure AU2024051402_26062025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR DATA PROCESSINGTECHNICAL FIELD
[0001] The present disclosure relates to data processing.BACKGROUND ART
[0002] Data processing is an important part of business analytics, not least in relation to financial data. In short, financial data is generally processed to generate financial reports, which are provided to stakeholders periodically, to enable evaluation of a business.
[0003] A problem is that generation of financial reports is a complex procedure, particularly considering that large amounts of data must be considered. In many cases, such data processing is largely a manual process, often based upon spreadsheets, which is very resource intensive. A further problem with such approach is that it is error prone.
[0004] As a result, custom applications and reports have been developed for businesses, which help create such financial reports in a more automated manner. These custom applications are, however, generally customised for a particular business as each business will have different procedures, and business data may be provided in different ways. As a result, these custom applications are costly to develop.
[0005] While custom applications generally work well initially, they are very rigid and difficult to modify, and as requirements change over time, they become less efficient. As an illustrative example, if the way business data is provided changes (e.g. due to a change in how data is captured), the custom application must generally be recoded, which requires specialist expertise.
[0006] As a result, and as a workaround to avoid such re-coding, spreadsheets are often used to pre-process data, and as an interface between the source data and custom application. As this happens over time, the benefit that the custom application once had slowly diminishes, and ultimately disappears.
[0007] Various standards have been developed for reporting of financial data. This greatly simplifies the process of analysing financial reports. Source data used to generate these reports varies, however, greatly and as such, these standards do not remedy the deficiencies of generating the financial reports, outlined above.
[0008] Similar problems exist in other areas of data processing, including processing of non-financial data.
[0009] As such, there is clearly a need for improved systems and methods for data processing.
[0010] It will be clearly understood that, if a prior art publication is referred to herein, this reference does not constitute an admission that the publication forms part of the common general knowledge in the art in Australia or in any other country.SUMMARY
[0011] The present disclosure relates to systems and methods for data processing, which may at least partially overcome at least one of the abovementioned disadvantages or provide the consumer with a useful or commercial choice.
[0012] With the foregoing in view, the present disclosure in one form, resides broadly in a data processing method comprising: receiving, on a data interface, a source data definition, a transformation definition and an output definition; storing, on a data store, the source data definition, the transformation definition and the output definition for subsequent use and reuse with different source data; receiving, on the data interface, source data for processing; retrieving data from the source data using the source data definition, the data associated with one or more source data fields defined by the source data definition; transforming the data into output data using the transformation definition, the output data associated with one or more output data fields defined by the transformation and output definitions; and generating an output using the output data and the output definition, wherein the transformation definition is defined at least in part by one or more operators of a plurality of predefined operators, the one or more operators transforming the data to the output data using predefined rules.
[0013] Advantageously, the method utilises a separate source data definition, transformation definition and output definition, which simplifies updating of the structure of the source data, what transformations are being performed, and the structure of the output. As an illustrative example, if the structure of the source data changes, only the source data definition need be updated leaving the transformation definition and output definition unchanged.
[0014] This enables reliable, repeatable processes to be easily made and updated, whichin turn improves use of associated technology (rather than work arounds), speeds up generation of subsequent outputs, and improves quality and accuracy of processes and outputs.
[0015] Furthermore, by providing a process that separates the source data definitions, transformation definitions and output definitions from the processing of the data, the same software may be used in relation to different types of data and for different purposes simply by changing the definition files.
[0016] Preferably, the data processing method comprises a business data processing method. The source data may comprise data relating to first attributes of a business, and the output data may comprise data relating to second attributes of the business. In other words, the source data and the output data may comprise different data, rather than the same data in a different format or standard.
[0017] Preferably, source data definition defines a structure of source data. The source data definition may define a location of the one or more source data fields.
[0018] Each source data field may include one or more cells of data. In some embodiments, the source data fields include fields comprising multiple cells of data. The source data fields may include rows or columns of data.
[0019] The source data definition may comprise a template, including the one or more source data fields. The template may comprise an array of cells (e.g. a spreadsheet or similar).
[0020] The source data definition may include names of the one or more source data fields.
[0021] The source data definition may characterise the nature of the data, e.g. whether it relates to a single data point, a row, a column, or the like.
[0022] The source data definition may include links to the source data. The links may comprise Uniform Resource Identifiers (URIs), for example.
[0023] Preferably, the transformation definition defines one or more functions for transforming the data into the output data.
[0024] The one or more functions may be defined at least in part by the operators. The one or more functions may comprise a combination of operators. The one or more functions may include a formula.
[0025] The operators may be selected from a set of predefined operators.
[0026] Examples of operators include addition, subtraction, multiplication, division, and average operators. The skilled addressee will readily appreciate that any suitable operator or function may however be used.
[0027] The function may be defined by a plurality of different operators.
[0028] The transformation definition may further define one or more filters, wherein the data is filtered prior to applying one or more functions thereto. The one or more filters may select a subset of the data. The one or more filters may include criteria in relation to which the subset of the data is selected.
[0029] The filters may be applied to the data itself, or associated data. As an illustrative example, the source data may include a date field and a value field, wherein the value field is filtered according to an associated date field.
[0030] The transformation definition may define a transformation of multiple data points to a single output field.
[0031] The multiple data points may relate to a single field, but at different points of time, for example. The multiple data points may relate to different fields, which are combined using the functions.
[0032] Preferably, the output data definition defines a structure of the output data. The output data definition may define a location of the one or more output data fields.
[0033] Each output data field may include one or more cells of data. In some embodiments, the output data fields include fields comprising multiple cells of data. The output data fields may include rows of columns of data.
[0034] The output data definition may comprise a template, including the one or more output data fields. The template may comprise an array of cells (e.g. a spreadsheet or similar).
[0035] The output data definition may include names of the one or more output data fields.
[0036] The output data definition may define types of output data fields. The output data fields may include plain text fields and graphical fields.
[0037] The source data definition may include one or more requirements regarding the structure of the source data.
[0038] The method may include determining whether the source data complies with thesource data definition. In other words, the method may include validating the source data.
[0039] The method may include providing an exception report in case source data does not comply with the source data definition. Alternatively, the method may include issuing an alert upon determining that the source data does not comply with the source data definition.
[0040] The output may comprise a report. The output may comprise a document or file.
[0041] The output may be compliant to one or more standards or requirements.
[0042] The output may be interactive. The output may be in any suitable format.
[0043] The method may include providing a graphical user interface to a user, for defining the source data definition, the transformation definition and / or the output definition.
[0044] The graphical user interface may enable the source data definition, the transformation definition and / or the output definition to be defined without coding.
[0045] The graphical user interface may include one or more cells in which the operators may be provided.
[0046] The graphical user interface may include one or more cells in which source data fields may be provided.
[0047] The graphical user interface may include one or more cells in which output data fields may be provided.
[0048] The source data definition, the transformation definition and / or the output definition may be defined in separate files. Alternatively, the source data definition, the transformation definition and / or the output definition may be defined in different parts of a single file. Regardless, it is envisaged that each of the source data definition, the transformation definition and / or the output definition may be updated independently of each other.
[0049] The method may include subsequently receiving, on the data interface, further source data for processing. Further data may be retrieved from the further source data using the source data definition. The further data may be transformed into further output data using the transformation definition. An output may be generated using the further output data.
[0050] The method may include training an Al or Machine Learning model with the source data definition, the transformation definition and the output definition. The method may further include training the Al or Machine Learning model using the source data.
[0051] In such case, the Al model may obtain semantic structure from the source data definition, the transformation definition, and / or output definition.
[0052] The source data definition, transformation definition and output definition may provide meaning and links between the data as an input to the learning, which enables the Al to focus on the meaning (semantics) of the data at a much deeper and comprehensive level.
[0053] The Al model may suggest adjustments to the definitions (when required) and create new definitions using what it has learned.
[0054] The Al model may be used to create further definitions (e.g. further source data definitions, transformation definitions and / or output definitions). The Al model may be used to update (refine) the definitions.
[0055] The method may automatically document processes and outcomes (i.e. what has been done and based upon what). The output may be associated with the documented processes, e.g. for audit purposes.
[0056] The method may include analysing the source data definition, transformation definition, and / or the output definition to create a semantic structure representing relationships between one or more data elements.
[0057] The semantic structure may then be used to train one or more artificial intelligence (Al) models.
[0058] The Al models may be used to refine at least one of the source data definition, transformation definition, and output definition.
[0059] The Al models may be used to create one or more new source data definitions, transformation definitions, and output definitions.
[0060] The semantic structure may define a knowledge graph, representing semantic relationships between the data elements of the source data definition, transformation definition, and output definition.
[0061] The semantic structure may be used to train artificial intelligence (Al) or machine learning (ML) algorithms to understand the meaning and relationships of the data elements and business concepts.
[0062] The method may include applying machine learning and natural language processing (NLP) techniques to analyse labels and descriptions associated with source dataand / or output data. The analysis may be used to propose modifications to the source data definition, the output definition, and / or the transformation definition.
[0063] The method may comprise refining the source data definition, the transformation definition, and / or the output definition based on feedback from users.
[0064] The method may comprise refining the source data definition, the transformation definition, and / or the output definition based on insights generated by Al analysis of the definitions and processed data.
[0065] The method may comprise refining the source data definition, the transformation definition, and / or the output definition based on patterns identified in the usage of the data processing system.
[0066] According to another embodiment, the disclosure resides broadly in a data processing system comprising: at least one server including a data interface, and associated with a data store, the at least one server configured to: receive, on the data interface, a source data definition, a transformation definition and an output definition; store, on the data store, the source data definition, a transformation definition and an output definition for subsequent use and reuse with different source data; receive, on the data interface, source data for processing; retrieve data from the source data using the source data definition, the data associated with one or more source data fields defined by the source data definition; transform the data into output data using the transformation definition, the output data associated with one or more output data fields defined by the output definition; and generate an output using the output data and the output definition.
[0067] Any of the features described herein can be combined in any combination with any one or more of the other features described herein within the scope of the invention.
[0068] The reference to any prior art in this specification is not, and should not be taken as an acknowledgement or any form of suggestion that the prior art forms part of the common general knowledge.BRIEF DESCRIPTION OF DRAWINGS
[0069] Various embodiments of the invention will be described with reference to the following drawings, in which:
[0070] Figure 1 illustrates a data processing system, according to an embodiment of the present invention.
[0071] Figure 2 illustrates a screenshot of a source data definition screen, according to an embodiment of the present invention.
[0072] Figure 3 illustrates a screenshot of an output data definition screen, according to an embodiment of the present invention.
[0073] Figure 4 illustrates a screenshot of a transformation definition screen, according to an embodiment of the present invention.
[0074] Figure 5 illustrates a screenshot of a source data file, such as a source data file of a data source of the system of Figure 1 , according to an embodiment of the present invention.
[0075] Figure 6 illustrates a screenshot of an output, such as an output of the system of Figure 1 , according to an embodiment of the present invention.
[0076] Figure 7 illustrates a schematic of a system for processing data, according to an embodiment of the present invention.
[0077] Figure 8 illustrates a data processing method, according to an embodiment of the present invention.
[0078] Exemplary features, embodiments and variations of the invention may be discerned from the following Detailed Description which provides sufficient information for those skilled in the art to perform the invention. The Detailed Description is not to be regarded as limiting the scope of the preceding Summary of the Invention in any way.DESCRIPTION OF EMBODIMENTS
[0079] Embodiments of data processing systems and methods are described below that enable reliable, repeatable processes to be easily made and updated, which in turn improves use of associated technology (rather than work arounds), speeds up generation of subsequent outputs, and improves quality and accuracy of processes and outputs.
[0080] Figure 1 illustrates a data processing system 100, according to an embodiment of the present invention.
[0081] The data processing system 100 includes a server 105, with which a user 110 interacts using a computing device 115 to perform various data processing functions andgenerate reports.
[0082] Initially, the user 110 interacts with the server 105 using the computing device 115 to define a source data definition, a transformation definition, and an output (end report) definition, which are together used to process source data in a reusable way. This may be performed using a web interface, a dedicated application installed on the computing device 115, or any other suitable way.
[0083] As outlined in further detail below, the source data definition defines one or more source data fields, which may be used to retrieve data from data sources, and the output definition defines one or more output data fields, which are used to provide processed data into outputs (e.g. reports). The transformation definition defines one or more functions which are used to transform data retrieved from the source data into the data used in the output.
[0084] Figure 2 illustrates a screenshot 200 of a source data definition screen, according to an embodiment of the present invention. The source data definition screen may be similar or identical to a source data definition screen, used to generate source definition data, in the system 100.
[0085] The source data definition screen includes a plurality of cells 205, arranged in a two- dimensional grid. The user 1 10 is able to enter data into the cells 205 to define a layout of the source data to be used by the system 100.
[0086] In Figure 2, the user 1 10 has entered source field names 210 representing columns in the uppermost cells 205, namely the names “Account Code”, “Account Description”, “Balance” and “Date”. The skilled addressee will readily appreciate, however, that the names are merely labels, and thus may take any suitable form. In other words, new names may be used by the system as a business evolves, as new data or types of data becomes available, or for new businesses or functions.
[0087] Once the user 1 10 has finalised the source data definition, they may select a save button 215, which causes the source data definition to be saved in a data store 105a associated with the server 105.
[0088] In some embodiments, in addition to defining a structure of the source data, the source data definition screen may be configured to enable the user 1 10 to enter rules regarding the data. As an illustrative example, the user 110 may enter that the source field “Date” must be of a format DD-MM-YY (Day, Month, Year), or that the “Account Code” field must have four digits, followed by a dash and then two digits. This enables the system 100 to identify errors insource data at an early stage in the process, rather than processing erroneous data.
[0089] Similarly, the source data definition screen may include functionality (e.g. in the form of buttons) to further characterise the nature of the data, e.g. whether it relates to a single data point, a row, column, or the like.
[0090] Finally, the source data definition screen may include links to the source data, e.g. in the form of Uniform Resource Identifiers (URIs) or the like. This further enables source data from multiple locations to be used.
[0091] The source data definition need only be defined once, upon which it may be re-used until the source data changes.
[0092] Figure 3 illustrates a screenshot 300 of an output data definition screen, according to an embodiment of the present invention. The output data definition screen may be similar or identical to an output data definition screen, used to generate output definition data, in the system 100.
[0093] The output data definition screen includes a plurality of cells 305, arranged in a two- dimensional grid. The user 110 is able to enter data into the cells to define a layout of the data output by the system.
[0094] In Figure 3, the user 110 has entered output field names 310 in a two-dimensional manner, namely with the names “Fixed Assets”, “Inventories”, “Trade Receivables”, “Cash and Cash Equivalents” and “Total Assets” as rows, and [Current FY] and [Previous FY] as columns.
[0095] [Current FY] and [Previous FY] are variables that correspond to the current and previous financial year respectively, and will be replaced by the appropriate values (e.g. dates defining the ends of the financial years) when the report is generated. Any number of such variables may be hardcoded in the system 100. The names “Fixed Assets”, “Inventories”, “Trade Receivables”, “Cash and Cash Equivalents” and “Total Assets” are output data fields, which are used in the transformation.
[0096] Once the user 110 has finalised the output data definition, they may select a save button 315, which causes the output data definition to be saved in the data store 105a.
[0097] Figure 4 illustrates a screenshot 400 of a transformation definition screen, according to an embodiment of the present invention. The transformation definition screen may be similar or identical to a transformation definition screen, used to generate output definition data, in the system 100.
[0098] The transformation definition screen includes a plurality of cells 405 arranged in a pre-defined format with headings. The user 1 10 is able to enter data into the cells to define a transformation performed by the system.
[0099] The cells 405 include an end report data point name cell 405a, in which the user may enter a name of an associated end data point, and an “as of’ field 405b, identifying a date of the associated end data point. Together, these identify output data fields of the output (e.g. report), as an intersection between the name and date.
[0100] The skilled addressee will readily appreciate that any suitable type of data may be used. For example, instead of name and date defining a data point, any two labels (e.g. two names) may define data points.
[0101] Similarly, while the example in Figures 2-4 is two dimensional (name and date on different axes), in other embodiments, any suitable number of dimensions may be used. In case more than two dimensions are used, the output (e.g. report) may be interactive to enable the multiple dimensions to be displayed in an efficient manner (e.g. where the user may fix one or more of the dimensions for display).
[0102] The cells 405 further include an operator cell 405c, a value cell 405d, and a formula cell 405e, which define the transformation to be performed on the data. The operator cell 405c may be used in case a simple transformation is used, e.g. a single operator is used, such as addition, as shown in Figure 4.
[0103] The value cell 405d is used in case the transformation comprises setting a value to the data point in the end report. This is useful where the report has as, at least as some elements, fixed values (e.g. an exchange rate used, a number of work hours per week, or the like).
[0104] The formula cell 405e is used in case the transformation comprises a more complex formula, such as a function referencing multiple points in the source data.
[0105] The skilled addressee will readily appreciate that only one of the operator cell 405c, the value cell 405d and the formula cell 405e need be used.
[0106] Finally, the cells 405 include a value field cell 405f, and one or more filter cells 405g, for identifying the source data used in the transformation.
[0107] The value field cell 405f identifies from where the source data is taken, e.g. in the form of a label (e.g. “Balance”) and / or a URL The one or more filter cells 405g define filterrequirements for the source data point, with reference to other fields.
[0108] In the example of Figure 4, the value field is “Balance” meaning that the data being transformed is retrieved from the field “Balance” in the source data, and the filter cells 405g include cells defining that the field “Account Code” in the source data must have a value of “1000- 01” and that the field “Date” in the source data must have a value of “30 / 06 / 2023”.
[0109] As a result, the system 100 selects only the cells from the source data where these criteria are met as input to the transformation. In other words, the system prefilters the source data.
[0110] As the operator is a “+” symbol (i.e. addition), the cells that meet the above filter criteria are added to the variable of the output names “Cash and Cash Equivalents”.[0011 1] While only a single transformation is illustrated, the skilled addressee will readily appreciate that typically several transformations will be used together to define the output (e.g. report). These multiple transformations may relate to different data points of the output, or a data point of another transformation. As an illustrative example, multiple transformations similar to that in Figure 4 may be used to add balances from multiple account codes to a single end report data point.
[0112] In use, the user 110 requests, using the computing device 1 15, that data be processed according to the previously defined source data definition, the transformation definition, and the end report definition.
[0113] The server 105 retrieves data from one or more data sources 120 according to the transformation definition and the source data definition, and filters same according to the transformation definition.
[0114] Figure 5 illustrates a screenshot 500 of a source data file, such as a source data file of a data source 120, according to an embodiment of the present invention. The source data file is illustrated in Figure 5 as a simplified file with only limited entries for illustrative purposes only. The skilled addressee will readily appreciate that source data files may be as complex as necessary.
[0115] The source data file comprises a two-dimensional array of cells 505, with data arranged in rows, with each column corresponding to a field type (e.g. date).
[0116] Now turning back to Figure 4, the value field is “Balance” and the filter cells 405g define that the field “Account Code” in the source data must have a value of “1000-01” and thatthe field “Date” in the source data must have a value of “30 / 06 / 2023”.
[0117] As a result, the server 105 retrieves a single data entry, namely the value ‘600’, which is the data for the field “Balance” in the only row where the Account Code has a value of “1000-01” and the Date has a value of 30 / 06 / 2023.
[0118] As the transformation is “addition” (symbol “+’), and the output name is “Cash and Cash Equivalents”, the value ‘600’ is added to a variable corresponding to the output name “Cash and Cash Equivalents” and for the date 30 June 2023. Initially, the variables are set to 0, or any other suitable (default) value, but multiple transformations may add data to or modify each associated variable.
[0119] Once all transformations are complete, the data from the variables is written to an output file, according to the output definition.
[0120] The output definition defines a layout of the output names, and thereby enables the user to define a layout of the report (or other output) being produced.
[0121] Figure 6 illustrates a screenshot 600 of an output, such as an output of the system 100, according to an embodiment of the present invention.
[0122] The output includes corresponds to the output data definition, and includes a plurality of cells 605, arranged in a two-dimensional grid, with labels corresponding to those of the output data definition. Furthermore, the variables [Current FY] and [Previous FY] are replaced with data based upon the actual date, and the current and previous financial year ends.
[0123] The value 600 is written to cell B6, which corresponds to the field “Cash and Cash Equivalents” and the date “30 June 2023”.
[0124] As outlined above, typically multiple transformations would be used, which may fill in several of the cells of the output. In use, each of the transformations is processed separately to each other, but the end report data points (fields of the output) are treated like variables that are shared across the transforms. As a result, the different transformations may each modify (e.g. add data to) a single field in the output.
[0125] In some embodiments, the output may be generated in the form of a document, such as a PDF file, an excel sheet, or the like, which is sent or otherwise provided to the user 110 and / or other users.
[0126] Figure 7 illustrates a schematic of a system 700 for processing data, according to anembodiment of the present invention. The system 700 may be similar or identical to the system 100.
[0127] The system includes a processor 705, which receives a source data definition 710, a transformation definition 715, and an end report definition 720. The term ‘processor’ here is referred to generally in relation to a unit that performs processing, regardless of its actual physical structure. The processor 705 may comprise software running on one or more computing devices configured to perform said processing.
[0128] The processor 705, together with the source data definition 710, the transformation definition 715, and the end report definition 720, function as a customised data processing system that is able to take data from one or more data sources 725 and generate an end report 730 (or other output therefrom) in a repeatable manner.
[0129] The processor 705 may be configured to integrate the source data definition 710, the transformation definition 715, and the end report definition 720 into a single processing module, which from a user perspective functions like a custom application. In such case, the processor 705, source data definition 710, transformation definition 715, and end report definition 720 may be compiled into a single useable software program.
[0130] The source data definitions are described above with reference to cells, much like a spreadsheet. The skilled addressee will, however, readily appreciate that the source data definition may define a structure of the source data in any suitable way, including with reference to locations of the one or more source data fields.
[0131] The source data definitions may include URIs or any other identifier of location of the data, to enable the system 100 to retrieve the data. Alternatively, the source data may be uploaded when it is available, wherein the source data definitions may simply identify internal locations of data.
[0132] The source data definitions may identify single cells of data. This is particularly relevant for single pieces of data, such as total sales figures, or the like. Alternatively, the source data may identify multiple cells of data, e.g. rows or columns of data. This is particularly useful for data over sales or timesheet data, having a plurality of entries.
[0133] The source data definition illustrated above comprises a template, including the one or more source data fields. The skilled addressee will, however, readily appreciate that the source data definition may be defined in any suitable manner, including in XML, Extensible Business Reporting Language (XBRL), Mathematical Markup Language (MathML), Structuredquery language (SQL) or the like.
[0134] The transformation definitions described above are similarly described with reference to cells, e.g. in a spreadsheet. The skilled addressee will, however, readily appreciate that the transformation definition may be defined in any suitable manner, including in XML, Extensible Business Reporting Language (XBRL), Mathematical Markup Language (MathML), Structured query language (SQL) or the like.
[0135] The transformation definition may define one or more functions for transforming the data into the output data. The functions may include multiple input variables. For example, the function may generate an average of a number of cells. The functions may alternatively operate on a single input variable.
[0136] The functions may be defined using one or more of a plurality of predefined operators. The operators may include addition, subtraction, multiplication, division, average. The skilled addressee will readily appreciate that any suitable function may however be used.
[0137] As an illustrative example, a function may retrieve hours worked from one column, hourly rate from another column, and a penalty multiplier from a third column, and multiply these to generate an employee salary entry.
[0138] While the above describes basic filtering of data in the transformation definition, the skilled addressee will readily appreciate that any suitable filtering may be used. For example, the filtering may be based upon data in associated fields, or data of the field itself.
[0139] When the source data relates to multiple entries, e.g. rows of data, the transformation may be performed on a row-by-row basis, such that rows of data are provided to the output. Alternatively, the transformation may be provided on groups of data to generate a single output. The transformation definition may include one or more flags to indicate whether data is to be processed on a row-by-row basis or as a group.
[0140] The output data definition described above is similarly described with reference to cells, e.g. in a spreadsheet. The skilled addressee will, however, readily appreciate that the output definition may be defined in any suitable manner, including in XML, Extensible Business Reporting Language (XBRL), Mathematical Markup Language (MathML), Structured query language (SQL) or the like.
[0141] The output may take any suitable form, including a spreadsheet, a document, a text file, e.g. with comma separated values (CSV), or any suitable format. In some embodiments, the output is interactive, to enable a user to be able to interactively view data in the report. Thisis particularly useful when multiple dimensions of data are provided.
[0142] The output may be compliant to one or more standards or requirements. For example, the output may comply with one or more accounting standards regarding reporting of financial data. In some embodiments, the output is defined according to a universal information structure.
[0143] Each output data field may include one or more cells of data. In some embodiments, the output data fields include fields comprising multiple cells of data (e.g. a row or column). This is particularly useful when the source data comprises a row of data, and the output comprises a transformation of the row.
[0144] In addition to defining a structure of the source data for the purpose of locating relevant data, the source data definition may define characteristics of the source data, e.g. for the purpose of verifying the data.
[0145] As an illustrative example, the source data definition may define that data of a particular field must be within certain pre-defined boundaries. For example, hours worked may have the requirement to be a number between 0 and 12 hours, and date may have the format YY-MM-DD with a date in the last 12 months.
[0146] The source data definition may also include any suitable requirements regarding the structure of the source data, e.g. a minimum number of rows.
[0147] The server 105 may determine whether the source data complies with the source data definition. The source data definition may define validation rules, business rules or the like. An alert, notification or exception report may be provided in case source data does not comply with the source data definition.
[0148] The source data definition may comprise a taxonomy. The transformation definition and / or output definition may comprise a taxonomy.
[0149] As outlined above, graphical user interfaces are provided to the user 1 10, for defining the source data definition, the transformation definition and / or the output definition. The graphical user interface enables the source data definition, the transformation definition and / or the output definition to be defined without coding.
[0150] In particular, the graphical user interfaces of Figures 2-4 include cells in which source and output field names and operators may be provided, simply by typing these names in desired locations. This may then be translated to XML, XBRL, MathML, SQL or any other format, butimportantly, the user interface enables the source data definition, the transformation definition and / or the output definition to be defined and / or updated without coding.
[0151] The skilled addressee will, readily appreciate that the above are simplified examples of graphical user interfaces, for the sake of clarity. Further detail regarding use of the user interfaces to generate the definition, and functionality of the user interfaces, is provided below.
[0152] Initially, the output definition, which defines the ‘end report’ is generally created. This term ‘end report’ is, however, used broadly, and can include any type of report, workpaper, dashboard or other output of the system.
[0153] To do this, the user initially gives the output definition a unique name, to enable easy access, identification and reuse of the output definition. The user may also define one or more sections within the end report.
[0154] One or more End Report Data Points are then added, and the End Report Data Points may be separated into sections. Each End Report Data Point comprises a Data Point Name and, optionally, one or more Type / Type Value combinations. In case Type / Type Value combinations are used, each combination comprises an entry in the end report. One or more descriptions may be provided in association with each Data Point Name and Type / Type Value combination.
[0155] In some cases, the Type Values are known when the output definition is defined. For example, if the Type is “Countries”, the Type Values may be Afghanistan, Albania, Algeria, etc. In other cases, the Type Values may be unknown when the output definition is defined. For example, if Type is “Client ID” the Type Values = list of clients, the list of clients may not be known, and may even change over time. If Type Values are unknown, the user may be prompted to provide a location to get those values from when a report instance is created. This location could, for example, comprise an identifier to a field called Client ID in a client database table or a Client ID column in a clients list spreadsheet.
[0156] The user may optionally define how to visualise an instance of the report in the output definition. In such case, a type of visualisation container, such as a graph, tab, sub-tab, column or row within a tab or sub-tab may be defined, for each report sections, and / or each End Report Data Point (i.e. data point name and each type / type value combination).
[0157] The order or layout of sections, data point names and types in the corresponding visualisation container may be also defined.
[0158] Attributes of interest are then identified in the source data and a location of thesource data is defined in a source definition. The source data and location may be in the form of a URL and access credentials, as needed, be provided by uploading a file, or by manually entering the data.
[0159] Finally, a transformation definition is created, defining the transformations used to transform the source data to the output data.
[0160] For each transformation, an End Report Data Point is selected as a target for the transformation. The End Report Data Point is treated as a variable which is initialised to zero and / or an empty string.
[0161] The data point(s) on which the transformation are performed are selected. This may include identifying one or more attributes in the source data, to filter / select relevant source data points, or identification of a single data point.
[0162] Finally, a transformation of the data points is defined using an operator or formula. In case of an operator (e.g. +) and multiple source data points, each of the source data points is transformed using the same operator (e.g. added).
[0163] In some embodiments, the transformations may define values or constants. This is useful as it enables fixed values to be added to a report.
[0164] Once the source, output and transformation definitions are defined, they may be used multiple times.
[0165] The source data definition, the transformation definition and / or the output definition may be defined in separate files. Alternatively, the source data definition, the transformation definition and / or the output definition may be defined in different parts of a single file. Regardless, it is envisaged that each of the source data definition, the transformation definition and / or the output definition may be updated independently of each other.
[0166] For example, if the reporting requirements change, the user 110 is able to load the output definition and make these changes, without risking inadvertently changing the transformations.
[0167] The above examples describe processing of source data once. The source data definition, the transformation definition and / or the output definition are, however, defined in a manner such that they are easily reusable.
[0168] As such, further source data may be received (or retrieved) for further processing.In such case, the further source data may be transformed into further output data using the source data definition, the transformation definition and / or the output definition, as previously defined, and an output may be generated using the further output data.
[0169] In the case of accounting data, new outputs (reports) may be generated based on new source data as it becomes available, e.g. monthly or quarterly. This then enables consistent reporting to be provided across different periods of time, which enables easy and consistent comparison across different time periods.
[0170] While the above specification describes a server, the skilled addressee will readily appreciate that any computing device or number of computing devices may be used.
[0171] As an illustrative example, the system may operate entirely locally, wherein the processing is performed on a local PC (rather than a remote server), without deviating from the scope of the invention.
[0172] Similarly, the processing may be cloud based, wherein a plurality of computing devices work together to perform the functions of the server 105, in a transparent manner. In particular, the different computing devices may function in a similar manner to a single server, and it may not be visible externally exactly which computing device is performing which function.
[0173] As the source, output and transformation definitions are defined independently of the software which uses them, the software may be generic, such that it is able to work on any type of data. In such case, the source, output and transformation definitions may simply be updated for different types of data, while the software remains the same.
[0174] Furthermore, while the Figures illustrate a single user 110, the skilled addressee will readily appreciate that multiple users can and will use the system 100. For example, one user 110 may initially set up the source data definition, transformation definition, and end report definition. Another user 110 may then process data using the definitions. Furthermore, different users 1 10 may use the definitions over time.
[0175] Figure 8 illustrates a data processing method 800, according to an embodiment of the present invention. The method 800 may be similar or identical to the method performed by the system 100, described above.
[0176] At step 805, a source data definition, a transformation definition and an output definition is received, on a data interface. These may be defined using the user interfaces of Figures 2-4 illustrated above, or any suitable means. Preferably, however, the source data definition, transformation definition and output definition are defined without needing to code.
[0177] At step 810, the source data definition, transformation definition and output definition are stored, on a data store, for subsequent use and reuse with different source data. This enables the source data definition, transformation definition and output definition to be defined once, and used consistently with changing data (e.g. as data changes over time).
[0178] At step 815, source data is received, on the data interface, for processing. The source data may be uploaded by a user, or be retrieved from a data source.
[0179] At step 820, data is retrieved from the source data using the source data definition, the data associated with one or more source data fields defined by the source data definition. The source data definition defines a structure of the data, enabling this step. This step may also include validation of the data.
[0180] At step 825, the data is transformed into output data using the transformation definition. The output data is associated with one or more output data fields defined by the output definition.
[0181] The transformation step will generally include operators (of a plurality of operators), forming a function.
[0182] At step 830, an output is generated using the output data and the output definition. The output may comprise a report, a spreadsheet, a document, a file or any other suitable output.
[0183] As different outputs may be generated based upon different source data, with the same source data definition, transformation definition and output definition, outputs may be easily compared with each other.
[0184] In some embodiments, the outputs include graphics or imagery generated according to one or more transformed values. The graphics may be defined in the output definition in a similar manner to data fields, but with a tag indicating that the field relates to graphics. An example of graphics includes a bar chart, a pie chart, for example.
[0185] The graphics may be colour coded, according to values. For example, data below a certain threshold may be coloured red.
[0186] While the above examples show creation of the source data definition, transformation definition and output definition without the need for coding, the skilled addressee will readily appreciate that code defining the relationships between data may be generated.
[0187] In some embodiments, the source data definition, transformation definition andoutput definition are defined in a semantics-based computer language. The source data definition, transformation definition and output definition may be defined in XBRL, XML, MathML, SQL or any suitable language.
[0188] While the above examples illustrate quantitative data, the skilled addressee will readily appreciate that any suitable type of data can be used. The data may be quantitative, qualitative or a combination of quantitative and qualitative. Natural Language Processing (NLP) may be used as part of the processing of the data in case natural language is used, for example.
[0189] In some embodiments, the source data definition may be used at data capture. In such case, validation of the data may be performed as the data is entered, upon which a user may be alerted when data is outside of pre-defined limits.
[0190] In some embodiments, the data is continually retrieved, e.g. from third party sources, and the output is continually updated. Furthermore, the transformation of data may include triggers. As an illustrative example, the source data could include stock price, stock trade volume, and announcements and media news about a company. This data, which is publicly available, can be continually and automatically monitored, and the transformation definition may identify spikes or dips in the stock price, which when detected, could trigger an analysis of the announcements made in a period before and after the spike in terms of positive or negative language used which may have caused the spike.
[0191] In other embodiments, the source data is retrieved from a blockchain. In such case, each new entry on the blockchain may comprise further source data which is automatically processed and updated. Similarly, source definitions, transformation definitions, output definitions, and outputs may all be stored on blockchain.
[0192] The source data definition, transformation definition and output definition provide semantic meaning to each data point in the output, which is particularly useful in training an artificial intelligence (Al) model. In short, the source data definition, transformation definition and output definition together with source data, function much like a knowledge graph, and by understanding relationships, it is easier to understand meaning from the data.
[0193] In contrast to traditional Al training (unsupervised learning), where data is simply provided as input without meaning, the source data definition, transformation definition and output definition provide meaning and links between the data as an input to the learning. In other words, the Al is able to focus on the meaning (semantics) of the data at a much deeper and comprehensive level, which meaning is provided rather than having to be “learned”, as it is provided upfront in a consumable format with the injection of rich semantics and meta-dataspecific to the relevant business domain. The result is a new generation of Al and machine learning algorithms that can leverage previous knowledge and contextual information in a way that is very similar to what the human brain does.
[0194] In particular, the definitions created, as outlined above, provide an ideal knowledge graph relating to, and specific to, the data. As a result, the definitions may be provided as data to the Al / machine learning algorithms, so that the machine learning not only understands the meaning of the data (source data and end report) but also understands the relationships between the data (e.g. how the end report is derived from the source data).
[0195] Because the definitions are used and validated in real day-to-day work, they enable a very focused, continuous, and inexpensive training and validation for Al / machine learning.
[0196] NLP may further be applied to the labels and descriptions in an output definition for each data point (and the labels and descriptions of other end report data points connected to it) to identify likely source data point matches in source data, and thereby propose source data definitions based thereon. As NLP uses understanding of language, context, and relationships between different words and concepts, it is able to find relevant information, and propose source data, even when it is not explicitly stated. This also enables different labels to be added over time, e.g. in the source data, without requiring any change in the system.
[0197] When definitions are proposed, e.g. based on NLP and / or Al or machine learning, the proposed definitions may be provided to the user in a manner that is easily editable. As such, the system provides a tool for simplifying the process for generation definitions, while enabling a user to approve same before use.
[0198] This in turn enables Al to be used to improve the definitions and outputs, generate further insights into the relationships among the data, and apply / adapt definitions to other scenarios.
[0199] In some embodiments, one or more of the source data definition, transformation definition and output definition are generated substantially automatically and from a data base of definitions.
[0200] In some embodiments, the Al utilises at least partially supervised learning. In such case, the Al may propose links / semantics between data points, which may be confirmed by a user. Upon confirmation of each link / semantic, the data which the Al may use grows, which in turn improves the AL
[0201] In some embodiments, the system is used for knowledge harvesting, e.g. in thecontext of an Al system.
[0202] As outlined above, experts may create a source data definition, and transformation definition and an output definition, using a graphical user interface or similar tool. These definitions are based on valuable knowledge regarding where relevant data can be found, how it should be processed, and how the output should be presented.
[0203] This information can be used by the system to create a knowledge base, because it captures real business practices used by the business. This knowledge base may be associated with written procedures that explain what transactions go into each account, accounting standards that govern each financial statement line item, and / or internal policies about financial reporting requirements.
[0204] This information can be used for a variety of purposes by the system including automatic or semi-automatic error identification and updates / repairs when changes are made.
[0205] For example, when input or output data relating to a business is changed, the system may automatically update, or propose updates, to one or more of the definitions.
[0206] For example, when a new bank account is added in a source data definition, the system may suggest modifications to the transformation definition based on the modifications to the source data definition.
[0207] The system may identify similar accounts, e.g. based on account code, data associated therewith, or other factors, and suggests transformations from those similar accounts. These suggestions may be proposed for approval, prior to the transformation definition being modified.
[0208] Similarly, if an output definition is generated or modified to create a new cash flow report, relevant accounts may be identified based on other reports, and proposed for addition to the source data definition. Transformations may be suggested for the transformation definition based on the source data and transformations used in other transformation definition. These suggestions may similarly be proposed for approval, prior to the appropriate definition being created or modified.
[0209] Finally, when an account structure changes in a source data definition, the system may propose changes to transformation definitions and / or output definitions. Furthermore, the system may identify other transformation definitions and / or output definitions that are used together with a source data definition, and propose similar updates to these. These suggestions may similarly be proposed for approval, prior to the appropriate definition being created ormodified.
[0210] In addition to performing calculations and generating the output, the systems and methods may save data associated with the transformation, e.g. for audit purposes. In particular, copies of source data may be saved, as well as copies of the definitions used, in case these were to be updated at a later time.[0021 1] In addition to generating an output, such as a report, in some embodiments the output definition may define an interface to one or more other systems. As an illustrative example, systems may be linked to enable reporting across a supply chain.
[0212] While above descriptions primarily relate to financial data, the skilled addressee will readily appreciate that the systems and methods described have applicability with any type of data, in any jurisdiction, sector (government and private, for profit and non-profit), industry, organization, and information system, and in diverse domains such as business data, health care, sustainability, etc. Furthermore, the systems and methods may provide data consolidation / aggregation, reporting, auditing, monitoring, or business intelligence, as illustrative non-limiting examples.
[0213] An example of an application of the systems and methods described above is in relation to Procure-to-Pay cycles. In such case, the business documents and entries typically involved include a request for quotation; a quotation; an order; goods / services receipt; an invoice; payment and related accounting entries.
[0214] This end-to-end cycle can be supported by a secure distributed public ledger, such as that on a Blockchain, the payload of which at each step comprises a source, transformation or output definition. This allows creation of reusable software applications that support the cycle, and determine significant benefits for all the participants in the network. As an example, buyers would be able to automate all business, accounting and compliance processes, and would be able to gain better conditions from suppliers by allowing them access to easier and more affordable revenues financing. Similarly, suppliers would be able to automate business, accounting and compliance processes, and also leverage the transparency of the system to more easily and affordably mobilize financial resources through revenue financing.
[0215] Furthermore, banks could leverage the transparency of the system to reduce their risk in credit operations with suppliers and increase their volumes, and governments, through tax agencies, nationally and internationally, would achieve greater data quality in compliance processes and an increase in tax revenues determined by the transparency and pervasiveness of the system.
[0216] The same methodology and benefits described for the Procure-to-Pay cycle can be applied to partial cycles derived from it, such as the simple issue and payment of an invoice or a purchase receipt, or completely different cycles where interrelated business documents and entries are at play in an end-to-end process. The end-to-end process can occur within one organisation or span across different organisations that participate in the same business reporting supply chain.
[0217] Another example of an end-to-end process which could consolidate Universal Information Structures is carbon reporting. As businesses strive toward net zero emissions, they also need to account for contributions from their supply chains in terms of goods purchased and / or supply company footprints. Each supply company can have source data definitions, transformation definitions and output definitions which automatically supply the required data together with the links of how it interfaces with the purchasing company.
[0218] This system could continually provide updates and could also utilise blockchain principles to maintain authenticity of the data. Again, these principles apply in the same way to both quantitative and qualitative information. Currently, qualitative information is predominant in environmental, social, and governance (ESG) reporting. The role of qualitative information will become more and more important due to the global ongoing efforts to formalise transparent and reliable indicators for ESG reporting, similar to those available for financial reporting. Universal Information Structures offer a way to verify consistency of qualitative and quantitative ESG information released by one organisation, and to verify its consistency with complementary information released by other organisations that participate in the same supply chain.
[0219] As such, the systems and methods not only provide an efficient means of standardising and automating reporting for both quantitative and qualitative data, but also for consolidation of reporting across supply chains. Consolidation of reporting makes it easier to track things like total carbon footprint or total social good across a supply chain.
[0220] Advantageously, the systems and methods described above utilise a separate source data definition, transformation definition and output definition, which simplifies updating of the structure of the source data, what transformations are being performed, and the structure of the output. For example, if the structure of the source data changes, only the source data definition need be updated leaving the transformation definition and output definition unchanged.
[0221] Furthermore, by separating the source data definition, transformation definition and output definition from the software performing the processing, the software is able to be reused across a wide range of scenarios and applications can be built without writing software code by constructing a source definition, a transformation definition and output definition.
[0222] The source data definition, transformation definition and output definition define an end-to-end taxonomy, from the source data to the output. This enables reliable, repeatable processes to be easily made and updated, which in turn improves use of associated technology (rather than work arounds), speeds up generation of subsequent outputs, and improves quality and accuracy of processes and outputs.
[0223] Furthermore, by providing a process that separates the source data definitions, transformation definitions and output definitions from the processing of the data, the same software may be used in relation to different types of data and for different purposes simply by changing the definition files.
[0224] The systems and methods enable processes to be updated without programming. In short, no specific technical skills are required, only a knowledge of the data and the desired outcomes. This in turn puts power back to users, enabling updating of just an output definition if only an output should change, rather than be exposed to code defining transformations, etc.
[0225] In contrast, in traditional software-based solutions, knowledge about completing each task is locked in the software code. It is not accessible, and it is not usable outside of the task itself.
[0226] In the present specification and claims (if any), the word ‘comprising’ and its derivatives including ‘comprises’ and ‘comprise’ include each of the stated integers but does not exclude the inclusion of one or more further integers.
[0227] Reference throughout this specification to ‘one embodiment’ or ‘an embodiment’ means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, the appearance of the phrases ‘in one embodiment’ or ‘in an embodiment’ in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more combinations.
[0228] In compliance with the statute, the disclosure has been described in language more or less specific to structural or methodical features. It is to be understood that the disclosure is not limited to specific features shown or described since the means herein described comprises preferred forms of putting the disclosure into effect. The disclosure is, therefore, claimed in any of its forms or modifications within the proper scope of the appended claims (if any) appropriately interpreted by those skilled in the art.
Claims
CLAIMS1 . A data processing method comprising: receiving, on a data interface, a source data definition, a transformation definition and an output definition; storing, on a data store, the source data definition, the transformation definition and the output definition for subsequent use and reuse with different source data; receiving, on the data interface, source data for processing; retrieving data from the source data using the source data definition, the data associated with one or more source data fields defined by the source data definition; transforming the data into output data using the transformation definition, the output data associated with one or more output data fields defined by the transformation and output definitions; and generating an output using the output data and the output definition, wherein the transformation definition is defined at least in part by one or more operators of a plurality of predefined operators, the one or more operators transforming the data to the output data using predefined rules.
2. The data processing method of claim 1 , wherein the source data definition defines a structure of source data.
3. The data processing method of claim 2, wherein the source data definition defines a location of the one or more source data fields.
4. The data processing method of claim 1 , wherein each source data field includes one or more cells of data.
5. The data processing method of claim 4, wherein the source data fields include fields comprising multiple cells of data.
6. The data processing method of claim 4, wherein the source data fields include rows or columns of data.
7. The data processing method of claim 1 , wherein the source data definition comprises a template, including the one or more source data fields.
8. The data processing method of claim 7, wherein the template comprises an array of cells (e.g. a spreadsheet or similar).
9. The data processing method of claim 1 , wherein the source data definition includesnames of the one or more source data fields.
10. The data processing method of claim 1 , wherein the source data definition characterises the nature of the data, e.g. whether it relates to a single data point, a row, column, or the like.11 . The data processing method of claim 1 , wherein the source data definition includes links to the source data, e.g. in the form of Uniform Resource Identifiers (URIs).
12. The data processing method of claim 1 , wherein the transformation definition defines one or more functions for transforming the data into the output data, the one or more functions defined at least in part according to the operators.
13. The data processing method of claim 12, wherein the one or more functions comprise a combination of operators defining a formula.
14. The data processing method of claim 12, wherein the transformation definition further defines one or more filters, wherein the data is filtered prior to applying one or more functions thereto.
15. The data processing method of claim 14, wherein the one or more filters select a subset of the data and wherein the filters include criteria in relation to which the subset of the data is selected.
16. The data processing method of claim 14, wherein the filters are applied to associated data (e.g. associated date data).
17. The data processing method of claim 1 , wherein the transformation definition defines a transformation of multiple data points to a single output field.
18. The data processing method of claim 17, wherein the multiple data points relate to a single field, but at different points of time, or to different fields, which are combined using functions.
19. The data processing method of claim 1 , wherein the output data definition defines a structure of the output data.
20. The data processing method of claim 1 , wherein the output data definition defines a location of the one or more output data fields.
21. The data processing method of claim 1 , wherein each output data field include one or more cells of data.
22. The data processing method of claim 1 , wherein the output data definition comprise a template, including the one or more output data fields.
23. The data processing method of claim 1 , wherein the output data definition includes names of the one or more output data fields.
24. The data processing method of claim 1 , wherein the output data definition defines types of output data fields.
25. The data processing method of claim 1 , wherein the source data definition include one or more requirements regarding the structure of the source data.
26. The data processing method of claim 25, wherein the method includes determining whether the source data complies with the source data definition.
27. The data processing method of claim 1 , wherein the method includes providing an exception report or issuing an alert in case source data does not comply with the source data definition.
28. The data processing method of claim 1 , wherein the output comprises a report.
29. The data processing method of claim 1 , wherein the output is compliant to one or more standards or requirements.
30. The data processing method of claim 1 , wherein the output is interactive.
31. The data processing method of claim 1 , wherein the method includes providing a graphical user interface to a user, for defining the source data definition, the transformation definition and / or the output definition.
32. The data processing method of claim 31 , wherein the graphical user interface enables the source data definition, the transformation definition and / or the output definition to be defined without coding.
33. The data processing method of claim 31 , wherein the graphical user interface includes one or more cells in which operators may be provided.
34. The data processing method of claim 31 , wherein the graphical user interface includes one or more cells in which source data fields and / or output data fields are provided.
35. The data processing method of claim 31 , wherein the source data definition, the transformation definition and / or the output definition are defined in separate files.
36. The data processing method of claim 1 , wherein the method further includes: subsequently receiving, on the data interface, further source data for processing.
37. The data processing method of claim 36, wherein the further data is retrieved from the further source data using the source data definition, is transformed into further output data using the transformation definition, and an output is generated using the further output data.
38. The data processing method of claim 1 , further including training an Al or Machine Learning model with the source data definition, the transformation definition and the output definition.
39. The data processing method of claim 38, further including training the Al or Machine Learning model using the source data.
40. The data processing method of claim 38, wherein the Al or Machine Learning model is configured to create further source data definitions, transformation definitions and / or output definitions, or refine the source data definition, the transformation definition and / or the output definition.41 . The data processing method of claim 38, further including analysing the source data definition, transformation definition, and / or the output definition to identify a semantic structure representing relationships between one or more data elements thereof, and training the Al or Machine Learning model using the semantic structure.
42. The data processing method of claim 41 , wherein the semantic structure defines a knowledge graph, representing semantic relationships between the data elements of the source data definition, transformation definition, and output definition.
43. The data processing method of claim 1 , further comprising applying machine learning and / or natural language processing (NLP) techniques to analyse labels and / or descriptions associated with the source data and / or the output data, and proposing modification to the source data definition, the output definition, and / or the transformation definition according thereto.
44. The data processing method of claim 1 , further comprising refining the source data definition, the transformation definition, and / or the output definition based on patterns identified in the usage of the data processing system.
45. A data processing system comprising: at least one server including a data interface, and associated with a data store, the atleast one server configured to: receive, on the data interface, a source data definition, a transformation definition and an output definition; store, on the data store, the source data definition, a transformation definition and an output definition for subsequent use and reuse with different source data; receive, on the data interface, source data for processing; retrieve data from the source data using the source data definition, the data associated with one or more source data fields defined by the source data definition; transform the data into output data using the transformation definition, the output data associated with one or more output data fields defined by the output definition; and generate an output using the output data and the output definition, wherein the transformation definition is defined at least in part by one or more operators of a plurality of predefined operators, the one or more operators transforming the data to the output data using predefined rules.
Citation Information
Patent Citations
XML streaming transformer
US20040034830A1
System, method and computer program for defining a data mapping between two or more data structures
US20040122852A1
Evaluation device, evaluation method, and recording medium
US20190095412A1
Systems and Methods for Unifying Formats and Adaptively Automating Processing of Business Records Data
US20210326350A1
Template mapping system for data translation
US5909570A