Data analytics system and method for machine learning solutions with zero coding
The zero-coding data analytics system addresses the accessibility and efficiency limitations of traditional machine learning by providing a user-friendly interface for developing machine learning solutions, automating variable selection, and enabling interactive data visualization, thus enhancing model development and deployment.
Patent Information
- Application Number
- PCT/IN2024/052232
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-22
AI Technical Summary
Traditional machine learning solutions require extensive coding expertise, mathematical knowledge, and manual effort, limiting their accessibility and efficiency for users without programming skills.
A zero-coding data analytics system and method that uses innovative user interfaces, automated variable selection, and intuitive data handling to empower users in developing machine learning solutions without coding, enabling interactive data visualization and controlled model development.
The system streamlines the development and deployment of machine learning models, making them more accessible and efficient, while maintaining transparency and reducing the need for multiple tools and systems.
Smart Images

Figure IN2024052232_22052025_PF_FP_ABST
Abstract
Description
DATA ANALYTICS SYSTEM AND METHOD FOR MACHINE LEARNING SOLUTIONS WITH ZERO CODINGFIELD OF THE INVENTION
[0001] The present invention relates to data analytics system and method for various machine learning solution with zero coding, particularly to highly evolved interactive dashboard for data visualization getting generated instantly without manual efforts as well as controlled and guided machine learning model development based on best practices of industry.BACKGROUND OF THE INVENTION
[0002] Machine learning has emerged as a transformative technology with applications spanning various industries, from healthcare to finance and beyond. In machine learning / data science applications, there are many challenges that data science model developer faces, due to which developing machine learning solutions becomes a difficult / time consuming job as well as at times, it is less evolved / sub optimal solution due to lack of awareness of mathematical / industrial practices as well as too much manual work required in repetitive coding and execution. However, a significant barrier to its widespread adoption is the requirement for individuals and organisations to possess coding expertise and mathematical / statistical knowledge. Traditional machine learning models often necessitate extensive programming, data preprocessing and algorithm selection, limiting their ability to develop optimal model for data science model developers.
[0003] Traditional machine learning tools does not produce interactive dashboard for data visualization easily and model development with these tools are not simultaneously transparent, zero coding based and highly evolved. There exists a growing demand for machine learning solutions that democratize access to this technology. Such solutions should enable individuals with domain-specific expertise but limited coding and mathematical / statistical knowledge to harness the power of machine learning for their unique applications. These solutions should also streamline the development process, making it more efficient and less relianton traditional coding paradigms such as drag and drop approach which has longer learning curve.
[0004] Most of the system and method requires running of various procedure manually after change in input parameters like introduction of new fields, removal of some fields, application of filter etc. and then see the impact. However, transparency is also a major issue as many systems and method act like black box, where you do not know what is happening and why it is doing certain tasks.
[0005] Traditional machine learning system and methods creates too many datasets as it requires to save data every now and then and one model development itself creates too many datasets causing space crunch. Further there are too many separate system and methods for visualization, supervised model building, text mining etc, as a reason of which users can not conduct all the activities in a single system and method which lead them to use so many systems and methods.
[0006] For visualization, Different people have different need to see two / three things together. Everyone wants customized dashboard as per his need, which is impossible to provide if a developer is setting it up for them. Many a times, developer ends up making many dashboards which is copy of same stuff. The tool allows end users to get the required dashboard by just selection (no drag n drop / parameter creation required).
[0007] Various machine learning frameworks and platforms are available, but they typically demand coding proficiency and a comprehensive understanding of machine learning concepts. Users often face a steep learning curve when attempting to create, train, and deploy-machine learning models, which could be time-consuming and error-prone.
[0008] Therefore, there is a need for a novel approach that empowers users to build and deploy machine learning solutions without manual coding or programming skills. Such a solution would not only accelerate the development process but also broaden the reach of machine learning to a wider audience, allowing experts in diverse fields to leverage its capabilities.SUMMARY OF THE INVENTION
[0009] The present invention proposes a novel approach to developing machine learning solutions with zero coding. By combining innovative user interfaces, automated variable selection, and intuitive data handing, this invention aims to democratize machine learning, making it accessible to individuals and organisations that lack extensive coding experience. This approach has the potential to transform various industries by streamlining the development and deployment of machine learning models, ultimately driving innovation and solving complex problems more efficiently.
[0010] This system and method empower users to explore various machine learning scenarios by allowing them to upload data and immediately get metadata. It facilitates the automatic creation of numerous derived fields based on just opting yes for the choice and offers interactive visualization, crosstab analysis, and dashboard views. Graphical user interface for filters and analysis, with an audit trail, enhances data manipulation and variables relationship understanding. Users can further customize their data with additional numeric or character fields binning. The system and method provide prompts to assist on developing models, which adapts interactive graphical user interface suitable to the user selections of method / technique. Users have control over model development, including choices related to supervised machine learning, such as the number of independent features, target encoding, Max VIF, and score bins. The system and method leverage data and output table analysis to apply various procedures, ultimately generating a final model and associated results and visualization.
[0011] This system and method with zero-coding empowers users to create machine learning solutions without the need for in-depth programming skills. Zero-coding machine learning platforms provide a user-friendly interface where users can upload their data, choose from a range machine learning model, and easily train models for tasks such as predictive modeling (supervised machine learning), clustering (unsupervised machine learning), time series analysis, collaborative filtering, association rules mining, sentiment / text analysis. Theseplatforms also offer automated feature engineering and model tuning, simplifying the entire process.
[0012] Furthermore, this platform bridge the gap between no-code simplicity but less transparency / blackbox approach and traditional coding which offers transparency but requires more coding and statistical knowledge. This allows data science developers to develop highly evolved model more efficiently.BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings are included to provide a further understanding of the invention. These drawings are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention. In the drawings:FIG. 1 is a flow diagram of one embodiment of a system and method which shows development of machine learning solutions with highly evolved interactive dashboard for data visualization with zero coding.FIG. 2 is a flow diagram of one embodiment of a system and method in which supervised, unsupervised, time series analysis, association rules mining, collaborative filtering and text mining machine learning technique is used.FIG. 3 is a flow diagram of one embodiment of a system and method which validate model on external data.FIG. 4 illustrate exemplary user interfaces that facilitate a user's selection of information pertaining to machine learning solution with zero coding.FIG. 5A and 5B illustrates exemplary user interface to get data and treat data that facilitates a user selection and input of information from a data set which can be generated or uploaded.FIG. 5C illustrates exemplary user interface that gives information about the data that has been loaded into system (called Meta Data). It further provides variable name treatment details, outlier detection details, flooring and capping related detail of the data.FIG. 5D illustrates exemplary user interface that gives details of variables which will be made available for visualization and which will be ignored as the system finds them unsuitable (as that might lead to over fitting) for model development purpose as well provides automatically created new variable details (such as indicator / dummy variables, classified numeric variables, derived fields from date fields) along with audit trail of column count.FIG. 5E illustrates exemplary user interface that facilitates a user selection and inputs for custom variables for numeric variables binning without coding. This interface also shows accuracy check of custom numeric field created immediately.FIG. 5F illustrates exemplary user interface that facilitates a user selection and inputs for custom variables for character variables binning without coding. This interface also shows accuracy check of custom character field created immediately.FIG. 6 illustrate exemplary user interfaces for filtering data that facilitate a user's selection of information for applying cascading filter - next filter value based on the current selection. This also provides audit trail for record count in every stage.FIG. 7 illustrate exemplary user interfaces for visual analysis of data, trends and relationship. It produces graphical as well as tabular output based on user selection.FIG. 8 illustrate exemplary user interfaces for cross tab analysis. It facilitates a user's selection of information that produces cross tab result. It also gives ten options to visualize or anomaly detection for cross tab analysis in multiple row / column-based analysis. It also displays the graph of data of cross tab results, which ensures that different groups are coloured differently for easy understanding of groups and group wise details.FIG. 9 illustrate exemplary user interfaces to visualize data. This facilitates a user's selection of information for dashboard view - Multiple variables same time - on the fly modification.FIG. 10 illustrate exemplary user interfaces that facilitate a user's selection of information for setting supervised machine learning model parameters through GUI which only shows applicable algorithms.FIG. 11 illustrate exemplary user interfaces to visualize and understand automated variable selection details. The system applies first bi-variate selection of variables and then applies step wise regression to select variables on the basis of multi variate analysis in case of supervised machine learning, where these techniques are applicable.FIG. 12A, 12B, 12C and 12D illustrate exemplary user interfaces to visualize and understand model development process, final model details, final model strength details and model usage recommendation.FIG. 13 illustrate exemplary user interfaces to validate supervised model on external data that facilitate a user's selection of information for scoring unseen data using saved Supervised Machine Learning model.FIG. 14 illustrate exemplary user interfaces to validate supervised model on external data that facilitate a user's selection of information for setting model validation or score generation parameters.FIG. 15 illustrate exemplary user interfaces that displays model validation result.FIG. 16 illustrates Kolmogorov- Smirnov (KS) chart patterns within an exemplary model validation results.FIG. 17 illustrate exemplary user interfaces that facilitate a user's selection of information for setting unsupervised Machine Learning - model parameters.FIG. 18 illustrate exemplary user interfaces that facilitate a user's selection of information for variable selection for cluster analysis.FIG. 19 illustrate exemplary user interfaces that displays information about variable clustering.FIG. 20 illustrate exemplary user interfaces displays information about homogeneity within clusters with various number of clusters and then displays mathematically chosen optimal number of clusters and information about clusters for observation clustering.FIG. 21 illustrate exemplary user interfaces that is aimed to make it easy to visualize clustering analysis - result.FIG. 22 illustrate exemplary user interfaces that facilitate a user's selection of information for setting parameters for time series analysis.FIG. 23 illustrates analysis variable patterns with an exemplary diagnostic for time series data.FIG. 24 illustrate exemplary user interfaces that displays information about best model selection for forecasting, then best model-based forecast and visualization of actual and forecasted results in continuation to make it easy for users to judge the forecasted figures.FIG. 25 illustrate exemplary user interfaces that facilitate a user's selection of information for association rules mining.FIG. 26 illustrate exemplary user interfaces that facilitate a user's selection of information for setting association rules mining parameters.FIG. 27 illustrate exemplary user interfaces that displays information for association rules mining results.FIG. 28 illustrates exemplary user interface that facilitates interactive 3D visualization of rules of association rules mining result.FIG. 29 illustrate exemplary user interfaces that facilitate a user's selection of information for collaborative filtering.FIG. 30 illustrate exemplary user interfaces that facilitate a user's selection of information for setting collaborative filtering parameters.FIG. 31 illustrate exemplary user interfaces that facilitate a user's selection of information for collaborative filtering result demonstration.FIG. 32 illustrate exemplary user interfaces that facilitate a user's selection of information for taking dump of collaborative filtering model outcome.FIG. 33 illustrate exemplary user interfaces that facilitate a user's selection of information with parameters to decide for derived field creation for text mining - supervised (sentiment analysis) / unsupervised.FIG. 34 illustrate exemplary user interfaces that facilitate a user's selection of information for setting text mining model parameters and updating.FIG. 35 illustrate exemplary user interfaces that displays information for step-by- step text treatment details.FIG. 36 illustrate exemplary user interfaces that displays for text Mining result.FIG. 37 illustrate exemplary user interfaces that displays information for text mining model and accuracy.FIG. 38 illustrate exemplary user interfaces that facilitate a user's selection of information for interactive custom investigation.FIG. 39 illustrate exemplary user interfaces that displays custom investigation result.FIG. 40 illustrate exemplary user interfaces that facilitate a user's selection of information for saving text mining data (with treatment) and / or model.DETAILED DESCRIPTION OF THE INVENTION
[0014] The following disclosure is provided in order to enable a person having ordinary skill in the art to practice the invention. Exemplary embodiments are provided only for illustrative purposes and various modifications will be readily apparent to persons skilled in the art. The general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the invention. Also, the terminology and phraseology used is for the purpose of describing exemplary embodiments and should not be considered limiting. Thus, the present invention is to be accorded the widest scope encompassing numerous alternatives, modifications and equivalents consistent with the principles and features disclosed. For the purpose of clarity, details relating to technical material that is known in the technical fields related to the inventionhave not been described in detail so as not to unnecessarily obscure the present invention.
[0015] The present disclosure describes methods and systems for the present invention that proposes a novel approach for developing machine learning solutions with zero coding and creating highly evolved interactive dashboard for data visualization - getting generated instantly without manual effort. Further, the invention relates to a system and method for dynamically creating and modifying graphical user interface control elements used to display data in a graphical user interface as a digital dashboard. Application extracts data, performs data processing on the extracted data, and displays the crosstab / dashboard etc for visualization purpose. The tool also has save data and model wizards in to download model / data with derived fields / recommendations / scores etc. The system and method instantly work on lots of data preparation job using Artificial Intelligence to create a list of variables for easy usage. No need for drag n drop, simply select the method from a list of applicable method, dependent variable and model gets ready. Additionally, it gives options to remove those variables, which should not be considered as independent variable due to any reason to allow users to control the model development.
[0016] This system and method is entirely non-coding, eliminating the need for coding expertise. It maintains transparency and states the reason why it adopts a particular strategy on a given data in every stage of the data. Additionally, it adopts the concepts of volatility, so, there is no need for running various procedures manually after introducing new fields, bins, filter and removal of fields because tool runs in auto mode. It regenerates all the subsequent statistics / graphs / models on its own. All intermediate data is generated in memory but permits the end data to be saved as a pkl or csv file, preventing excessive intermediate data generation. Only one system and method is required for performing various functions such as visualization, supervised model building, text mining, etc. Its user interface is one of a kind, which make it super simple for users. It brings only relevant options in front of users, which is applicable in the given situation / task along with default values / options. Hence user does not need to learn where to click (or how to writecode) to apply those options. Additionally, the end users do not need to have expertise in statistics and machine learning, they can leave prefilled options as they are, and the model will be optimized based on industry best practices.
[0017] The details of various applicable machine learning procedure of the tool are given below:• Supervised Machine Learning (Classification tree, Regression Tree, Linear Regression, logistic regression, Random Forest classifier, Random Forest regressor, XG Boost classifier, SG Boost regressor) -1. It produces tabular / textual and graphical output of every stage (variable strength, model’s separation power, correlations etc.) based on applicability.2. It does lot of preprocessing of data like a. It renames variables, if initial variable names have some issues b. It does missing value treatment, if opted for c. It does flooring / capping, if opted for d. In generates dummy variables, if missing percentage is too huge3. It also automatically a) generates the dummy variables, b) does classification of numeric variables, c) generates target encoded variables etc. d) When opted, it generates the details of target encoding at every stage.4. It allows control of final variables in the model and VIF while ensuring that no way user can put irrelevant value of modelling parameters like VIF threshold etc.5. In case of tree models, it allows control of decision tree growth,6. It produces model summary on it’s own.7. It allows user to run scoring on out of time data with ease - just upload data, upload older model and set it for running.• Unsupervised Machine Learning / Cluster analysis -1. It gives options of variable selection in the context of subjective segmentation, decide number of segments.2. It conducts variable selection (if opted for), generates segments and3. Also shows statistically chosen elbow point based number of segments and statistics along with user defined number of segments and statistics.• Time Series Analysis - once user uploads the data and selects the time period field, metrics field and number of forecast periods needed, the tool1. Automatically generates the multiplicative as well as addition based Time Series diagnostic.2. The tool highlights those data points in the table, which are too far based on standardized value of residual of the time series analysis based decomposition.3. And then it goes ahead and finds best set of ARIMA parameters. Then it produces best ARIMA model summary statistics and generates forecast based on that.4. Next it produces graph of actual and forecast in continuation so that user can get comfort on forecasted value.• Association Rule Mining - based on transaction history, the tool1. Finds the items which are bought together and2. Gives the rules based on chosen threshold.3. It also generates 3D interactive visualization to understand the rules generated.• Collaborative Filtering - again once the user uploads the data and inform the system about the user field, rating field and item field, the tool takes over.1. It produces the cosine similarity of users and2. Then recommends new options for any given user based on what similar users have used / seen and liked.3. It also allows bulk download of recommendations for all users of the dataset, so that if communication (like SMS, whatsapp, email etc.) through some other system is required, then that can happen.• Text Mining - when user uploads the data and lets the tool know, which is the comment field and which one indicates sentiment, and the system runs text mining stuff.1. The tool also gives control in the hand of user to apply various procedure of text mining, modify stop words and see the result.2. Then it generates the most frequent list of words and graph for each class of outcome (or each type of sentiment) as well3. produces words and associated odds for making a comment positive / negative if it is present.4. It allows users to conduct interactive analysis on specific word and shows results containing that word.• Graphical Regular Analysis -graphical regular analysis is available in all of the above components. However, when user selects the Regular Analysis only in the main page, it relaxes some of the threshold, which usually makes machine learning model poor due to over fitting (like character variable having too many classes) and let’s user run normal analysis.1. The tool allows user to see different kind of charts of metrics based on classes of different variables without effort.2. Allows users to generate cross tab based on choice.3. Provides various options to analyse / visualize / detect anomaly in the cross tab analysis result like a) Values based colour coding (all data wise, row wise, column wise) b) Z-score based colour coding (all data wise, row wise, column wise) c) Z statistics of regression residual based colour coding (row wise, column wise etc.)4. Dashboard view - Additionally, in business, different people have different need to see two / three things together. Many wants customized dashboard as per his need, which is impossible to provide if a developer is setting it up (which happens in most of the popular visualization tools). Many a times, developer ends up making many dashboards on the same data for different person. This tool allows configuration of dashboard as per choice by just selection.
[0018] The method for development of various machine learning solution with zero coding basically comprises of following steps. Please note - users mandatorily have to interact with applications only at Get Data Screen and Model Parameter screen. Rest of the screen are allowing users to get good insights about the data processing / discover trends in data etc but is not mandatory.STEP 1: The user is required to input data through the system main page interface i.e., Home Screen. User inputs are further taken for indicator dummy variable creation, numeric variable classifications, missing value treatment, flooring / capping and user decides which machine learning situation to try. This screen also has uploads data wizard.STEP 2: The system and method after allowing the selection of the required fields by the users reads the file selected in file upload wizard by the user and starts further processing which consists of treatment of the variable names, missing value treatment, for all kind of data and generate indicator variables to indicate any huge missing. It also shows flooring and capping details if it was opted for.STEP 3: The data treatment process selects variables, which can be used in visualizations. It also creates dummy variables for eligible set of character variables. Further, it classifies numeric variables and generate derive fields from the date fields. The data treatment details display insights about the data such as details of variables that will be used / ignored for visualization, details of variables which were used to generate dummy variables, details of numeric variables which got classified, details of derived fields from date fields and audit of column count.STEP 4: The system generates GUI to allow users to do custom numeric variable binning. It is an optional step in the model development. The user is required to input choice through the numeric variable binning interface for the number of variables for which custom numeric binning is needed. Then application generates GUI to allow users to pass custom numeric binning requirement for each of the numeric field individually and it displays accuracy check metrics for the added fields.STEP 5: The system provides custom character variable binning GUI to allow users to do custom character variable binning. The user is required to input number of character variables for which custom binning is required. Then system provides graphical interface to allow user to get customized character binning. The system immediately shows the accuracy of newly created customized character variables.STEP 6: The user is required to input option / choice / value for application of filters. The system provides interface for the categorical variable based filters, numeric variable based filters, date field-based filters. Each filter works on data that comes through previous stage of filtration thus next filters GUI is based on remaining data (always shows relevant filter options only). Filter process generates audit of rows in the data, and it has option for saving the data as well.SPEP 7: The system generates GUI for visual analysis. The user is required to choose option in the visual analysis interface for classification parameter, metrics 1 and 2, method of aggregating metrics, chart type, metrics 1 and 2 interaction method and the tool generates graphical and tabular output immediately. Additionally, it provides other details for both the metrics to provide closer insight on chosen metrics.STEP 8: The system generates GUI to allow users to do cross tab analysis and generates GUI to allow users to perform cross tab analysis. The user is required to choose options on cross tab analysis interface for multi-level rows / multi-level columns / single metric and metrics aggregation. The tool generates table and graph and provides 10 different method of colour coding table for easy visualization and / or anomaly detection.STEP 9: The system also provides dashboard view, where user can see multiple metrics for a given comparison parameter. The system generates dashboard automatically but allows users to modify dashboard as per his requirement by just selection. User has option to choose 1 classification parameter, 4 metrics - 1, 4 metrics 2, 4 Method of aggregating metrics, 4 chart type, 4 methods of metrics 1 and 2 interaction on dashboard interface and get modified dashboard.STEP 10: The model parameter interface is generated depending on machine learning method chosen. The user is required to choose / enter options on model parameter interface that shows suitable options only based on the machine learning method and characteristics of dependent variable.STEP 11: This step is only required in case of supervised machine learning. The user needs to upload data and previously developed model. The user is then required to input data for dependent variable, model type, out of time validation is needed / only scoring is needed. User inputs are further taken for the file name for saving scored data (if needed). The system and method read data files and produce basic statistics and starts further processing which consists of sample data viewing, numeric variable details, high level data analysis. The validation result display out of time model validation results, graphical and tabular display of model strength. If user opts for saving data than it saves the data in two different format (csv and pkl) and provides confirmation on data created with score.
[0019] An embodiment of the present invention illustrates that the system and method is a Streamlit and Python based platform which comprises:• Inputs - Streamlit has been used to develop GUI to collect user inputs. Streamlit components like selection box, radio buttons, input box, slider, file selector etc. is making user inputs very easy.• Core Python is working like a middleware which takes user inputs, processes it and then produces output and lets Streamlit display it.• Output - Streamlit has been used to display information, charts, tables, mark up colourful texts, warnings etc.. Additionally wizard to save / download data and model has been developed using Streamlit.
[0020] The embodiment of the present invention provides an interactive system for visualizing input data organized for a machine learning model or regular analysis, in which system comprises means for input of data by the user from multidimensional data source, means for converting received input data into a graphical form, and means for presenting the received input data to a user as a graphical visualization allowing the user for combining and present the data in form of graphical data visualization elements. Visualizations according to the invention are charts, and tables which are further generating a number of candidate visualizations is based on selected data, and numerically evaluating candidate input variables for their fitness as independent variable in the model, then using them for model development, determination of model strength and confirmation on data created with score. The figures in cross tab tables are highlighted to help user understand the scale of the figures or to detect anomaly.
[0021] The embodiment of the present invention as illustrated in figure 1 the user will enter the inputs in the Interface 101 displayed on the main page. The interface 101 comprises of various required fields such as visualization, indicator dummy variable creation, numeric variable classifications, missing value treatment, flooring / capping treatment, machine learning needed and allowing users to select the required file for selection of information and data content. Based on the selection made on the interface 101 of the user, the internal operation will be executed to read the file, do treatment of variable names, do missing value treatment for all kind of data, do flooring / capping (if opted for) and generate indicator variables to indicate huge missing.
[0022] The meta data 102 display initial data comprise of missing value treatment, viewing sample data, numeric variable details, outlier possibility, flooring / capping applied and high-level data analysis. After the display of meta data 102 procedure, the internal operation starts executing to select variables which can be used in visualizations, create dummy variables for eligible set of character variables, classify numeric variables and generate derive fields from the date fields.
[0023] The data treatment 103 display the details of variables that will be used / ignored for visualization, details of variables which were used to generate dummyvariables, details of numeric variables which got classified, detail of date fields which will be used to generate derived fields and audit of column count. Once the data treatment displays, the internal operation generates GUI to allow users to do custom numeric variable binning.
[0024] The user will enter the inputs in the numeric variable binning interface 104 and it comprises of number of variables for which custom numeric binning is needed, generate GUI to allow users to pass custom numeric binning requirement for each of the numeric field and display accuracy check metrics for added fields. Based on the selection made on the numeric variable binning interface 104 the internal operations generate GUI to allow users to do custom character variable binning.
[0025] The user will enter the inputs in the character variable binning interface 105 and it comprises of number of variables for which custom character binning is needed, generate GUI to allow users to pass custom character binning requirement for each of the character field and display accuracy check metrics. Based on the selection made on the numeric variable binning interface 105 the internal operations generate interface for filtration.
[0026] The user will enter the inputs in the filter interface 106 and it comprises of categorical variable based filters, numeric variable based filters, date field based filters. It always applies filters / provides next filter GUI based on remaining data only (so always displays relevant filter only), generate audit trails of rows count in the data and save the data wizard. Based on the selection made on the filter interface 106 the internal operations generate GUI for visual analysis.
[0027] The user will enter the inputs in the visual analysis interface 107 and it comprises classification parameter, metrics 1 and 2, method of aggregating metrics, chart type, metrics 1 and 2 interaction method and generates graphical and tabular output. The internal operations then generate GUI to allow users to do cross tab analysis (interface 108) and generate GUI to allow users to get dashboard interface (interface 109).
[0028] The user will enter the inputs in the cross tab analysis interface 108 and it comprises multi-level rows / multi-level columns / single metric and metricsaggregation, generate table and graph and 10 different method of colour coding table for easy visualization / anomaly detection.
[0029] The user will enter the inputs in the dashboard interface 109 and it comprises 1 x classification parameter, 4 x metrics - 1 and 2, 4 x Method of aggregating metrics, 4 x chart type, 4 x metrics 1 and 2 interaction method and generate dashboard. Based on the selection made on the dashboard interface 109 the internal operations generate the interface depending on the machine learning method required. The user will enter the inputs in the model parameter interface 110 based on the machine learning method required.
[0030] In an embodiment of the present invention as illustrated in figure 2 for supervised machine learning, the user will enter the inputs in the model parameters interface 201 and it comprises dependent variable, variables which should not be considered, selection of machine learning method, target encoding and encoding process details (if opted for), number of variables from step wise and decision tree and lift table design. Based on the selection made on the model parameters interface 201 the internal operations detect bi-variate strength of the model and perform step wise regression (if applicable). The variable strength / selection 202 display bivariate strength of variables, graphical display of bi-variate strength of variables and step wise variable selection details. After the display of variable strength / selection 202 procedure, the internal operation run subsequent machine learning procedure. The model statistics 203 display optimized model development based on variable significance, variable bi-variate strength, multi collinearity, model strength, model usage recommendations, model visualization and save data and model wizard.
[0031] In an embodiment of the present invention as illustrated in figure 2 for unsupervised machine learning, the user will enter the inputs in the model parameters interface 301 and it comprises variable selection need, eigen value threshold for variable selection and how many clusters are needed. Based on the selection made on the model parameters interface 301 the internal operations detect correlation of variables and perform variable clustering based on user inputs. The variable strength / selection 302 display variables first and second moment, variablestandardization, correlation of variables (tables and visualization), variable clustering & best subset of unrelated variable selection. After the display of variable strength / selection 302 procedure, the internal operations run cluster analysis on all / selected set of variables. The model statistics 303 display cluster analysis and elbow point detection, number of segments and sum of squared errors within clusters visualization, display of clustering solution based on chosen number of clusters as well as mathematically chosen optimal number of clusters.
[0032] In another embodiment of the present invention as illustrated in figure 2 for time series analysis, the user will enter the inputs in the model parameter interface 401 and it comprises variable which represent date, variable which needs to be forecasted and number of forecasts needed. Based on the selection made on the model parameter interface 401, the internal operations run time series decomposition on the data. The variable strength / selection 402 display time series chart, multiplicative and additive decomposition of the time series data, graphical display of visualization and standard deviation-based colour coding of residual in table. After the display of variable strength / selection 402 procedure, the internal operations find best set of parameters for ARIMA model. The model statistics 403 display iteration details of parameter search for ARIMA model, select best parameters for ARIMA model, display forecasted values based on best ARIMA model, visualize actual and forecasted value to get a feel of forecasting accuracy.
[0033] In an embodiment of the present invention as illustrated in figure 2 for association rules mining, the user will enter the inputs in the model parameter interface 501 and it comprises variables which indicate transaction and item. Additionally, it allows user to enter support, confidence and lift threshold needed. Based on the selection made on the model parameter interface 501, the internal operations run procedure to discover association rules. The model statistics 503 display details of association rules generated, and 3D interactive visualization of association rules generated.
[0034] In an embodiment of the present invention as illustrated in figure 2 for collaborative filtering, the user will enter the inputs in the model parameter interface 601 and it comprises variables which indicates user ID, item ID and rating.Based on the selection made on the model parameter interface 601, the internal operations run procedure to discover similar users and recommendations. The model statistics 603 display interactive demo of collaborative filtering for any User ID, shows, similar users and recommendations, wizard to bulk download, similar user details and recommendations for each user.
[0035] In an embodiment of the present invention as illustrated in figure 2 for text mining, the user will enter the inputs in the model parameter interface 701 and it comprises variables which indicate text / comment and (optional) sentiment / class, number of important features required for display of data preparation and texts for stop word additions. Based on the selection made on the model parameter interface 701, the internal operations run data preparation steps. The variable strength / selection 702 display step by step data preparation details for chosen number of records. After the display of variable strength / selection 702 procedure, the internal operations run procedure to find most frequent words and words associated with particular sentiment / class. The model statistics interface 703 display of original and modified comments, display of most frequent words (in all data or for each class - if class field is available in data), treated words for classification and associated odds, model accuracy report on test data and wizard to display actual record along with modified text pertaining to a particular word and sentiment.
[0036] In an embodiment of the present invention as illustrated in figure 3 for out of time model validation / scoring engine, the user will enter the inputs in the load model and data interface 801 and it comprises dependent variable / model type, out of time validation is needed / only scoring is needed and the file name for serving scored data (if needed). Based on the selection made on the load model and data interface 801, the internal operations read data files and produce basic statistics. The meta data 802 display viewing sample data, numeric variable details and high level data analysis. After the display of meta data 802 procedure, the internal operations run scoring / model validation. The validation results 803 displays out of time model / validation results, graphical display of model strength and confirmation on data created with score.
[0037] In an embodiment of the present invention as illustrated in figure 4, the user will enter the inputs in the interface 101 displayed on the main page as it illustrates exemplary user interface 101 that facilitate a user's selection of information pertaining to machine learning solution with zero coding.
[0038] In an embodiment of the present invention as illustrated in figure 5A and 5B, the user will input information from the data set as it illustrates exemplary user interface 101 to get data and treat data that facilitates a user selection and input of information from a data set which can be generated or uploaded. Based on the selection made on the interface 101 of the user, the internal operation will be executed to read the file, do treatment of variable names, do missing value treatment for all kind of data, do flooring / capping treatment and generate indicator variables to indicate huge missing.
[0039] In an embodiment of the present invention as illustrated in figure 5C, the meta data 102 display initial data which comprise of missing value treatment, viewing sample data, numeric variable details, flooring / capping details and high- level data analysis as it illustrates exemplary user interface to get data and treat data that facilitates the preparation of meta data 102. After the display of meta data 102 procedure, the internal operation starts executing to select variables which can be used in visualizations, create dummy variables for eligible set of character variables, classify numeric variables and generate derive fields from the date fields.
[0040] In an embodiment of the present invention as illustrated in figure 5D, the data treatment 103 display the details of variables that will be used / ignored for visualization, details of variables which were used to generate dummy variables, details of numeric variables which got classified and audit of column count as it illustrates exemplary user interface to get data and treat data that facilitates the automatically created data treatment 103 details. Once the data treatment 103 displays, the internal operation generates GUI to allow users to do custom numeric variable binning.
[0041] In an embodiment of the present invention as illustrated in figure 5E, the user will enter the inputs in the numeric variable binning interface 104 and it comprises of number of variables for which custom numeric binning is needed,generate GUI to allow users to pass custom numeric binning requirement for each of the numeric field and display accuracy check metrics for added fields. Based on the selection made on the numeric variable binning interface 104 the internal operations generate GUI to allow users to do custom character variable binning.
[0042] In an embodiment of the present invention as illustrated in figure 5F, the user will enter the inputs in the character variable binning interface 105 and it comprises of number of variables for which custom character binning is needed, generate GUI to allow users to pass custom character binning requirement for each of the character field and display accuracy check metrics. The internal operations generate interface for filtration.
[0043] In an embodiment of the present invention as illustrated in figure 6, the user will enter the inputs in the filter interface 106 and it comprises of categorical variable based filters, numeric variable based filters, date filed based filters, apply filters / next filter GUI based on remaining data (always relevant filter only), generate audit of rows in the data and save the wizard as it illustrate exemplary user interfaces to visualize and develop data that facilitate a user's selection of information for applying cascading filter - next filter value based on the current selection. Based on the selection made on the filter interface 106 the internal operations generate GUI for visual analysis.
[0044] In an embodiment of the present invention as illustrated in figure 7, the user will enter the inputs in the visual analysis interface 107 and it comprises classification parameter, metrics 1 and 2, method of aggregating metrics, chart type, metrics 1 and 2 interaction method and generate graphical and tabular output as it illustrate exemplary user interfaces to visualize and develop data that facilitate a user's selection of information for visual analysis interface 107 of automatically generated interface. The internal operations generate GUI to allow users to do cross tab analysis and generate GUI to allow users to get dashboard interface.
[0045] In an embodiment of the present invention as illustrated in figure 8, the user will enter the inputs in the cross tab analysis interface 108 and it comprises multilevel rows / multi-level columns / single metric and metrics aggregation, generate table and graph and 10 different method of colour coding table for easyvisualization / anomaly detection as it illustrate exemplary user interfaces to visualize and develop data that facilitate a user's selection of information with colour scale or anomaly detention for cross tab analysis in multiple row / columnbased analysis.
[0046] In an embodiment of the present invention as illustrated in figure 9, the user will enter the inputs in the dashboard interface 109 and it comprises 1 x classification parameter, 4 x metrics - 1 and 2, 4 x Method of aggregating metrics, 4 x chart type, 4 x metrics 1 and 2 interaction method and generate dashboard. Based on the selection made on the dashboard interface, graphs gets updated immediately.
[0047] Interface 110 varies as per the selection of machine learning requirement in interface 101. In an embodiment of the present invention as illustrated in figure 10, for supervised machine learning situation, the user will enter the inputs in the model parameters interface 201 and it comprises dependent variable / variables which should not be considered, selection of machine learning method, target encoding and requirement of detail encoding process, number of variables from step wise and decision tree and lift table design as it illustrate exemplary user interfaces to visualize and develop model that facilitate a user's selection of information for setting supervised machine learning model parameters interface 201 through GUI which only shows applicable algorithms. Based on the selection made on the model parameters interface 201 the internal operations detect bi-variate strength of the variables for model and perform step wise regression (multi variate variable selection) (if applicable).
[0048] In an embodiment of the present invention after preparing structured data along with a response variable, there are several steps of model building and steps depends on choice of algorithm like classification tree / random forest / xgboost etc. will not require lots of preprocessing of data because they are tolerant of outliers / missing values of data. Classification tree is usually coarse and random forest / xgboost is usually black box because it has too many trees inside.
[0049] In an embodiment of the present invention logistic regression usually produces parsimony model, which means ideally it should have few variables,which can still predict the class of the record (prospect / customer) with high degree of accuracy. Hence it involves Flooring / capping of variables - to remove outlier values of variables in the data, which distorts the model, missing value treatment - Either the missing value is replaced with some other values or at times the records containing missing values are dropped. Dummy variable creation - if there is character variable with few classes.
[0050] In an embodiment of the present invention logistic regression model building steps comprises:• Dummy variable creation — say a variable Type has value A, B and C. Then one can create Type A, Type B and Type C, all of which will be 0 / 1 kind variable. Where 0 means absence and 1 means presence. It will get populated like• When type field will have value A, Type_A will be 1 otherwise 0.• Target encoding - this is all about creating a numerical field (usually based on response rate) from a character field, because character fields can not be used in a model.• When distinct values of a character field is few, one go for dummy variable creation. When the distinct values are finite but above a threshold, then one go for either target encoding or ignores those character variables.• Most often than not - no one uses those character fields, which has too many distinct values in the model (like Name etc.).• Bi-variate Variable selection - finding those variables which should become part of the multi variate model development.• Multi variate variable selection - find those sets of variables, which together appears making good classification model.• Multi collinearity removal - there might be several independent variables still in the model, which is indicating the same stuff. Hence, they do not necessarily add value to the model. Additionally, this causes instability of coefficients of independent variables of the model. Hence multi collinearity needs to be removed from the parsimony model.• Multi collinearity removal is a repetitive process where one finds the variable which has maximum VIF (variation inflation factor) (let’s call is II) and then finds which is the other variable which is associated with this variable (let’s call it 12). Then check bi-variate strength of II and 12 and keep the one which has more bi-variate strength and drop other.• Then the Multi collinearity removal process is repeated. Again the next variable with maximum VIF is found and we repeat the process of getting rid of weaker variable (one at a time). The process continues till VIF of the model is less than a threshold (usually < 2).• Visual trend analysis - many a times developers wishes to see if the trend is according to common sense or not. For this he usually plots the distribution of variables and overlay dependent variables on top of this to see if it makes sense.• Removal of variables, which is coming insignificant in the model.• Strength of classification model (applicable to all classification techniques not just logistic regression) - once the model is ready, industrial practices revolves around checking how good is model for practical usage. It involves measurement of cumulative response rate against score bins and cumulative non response rate against score bins.
[0051] In an embodiment of the present invention as illustrated in figure 11, the variable strength / selection 202 display bi-variate strength of variables, graphical display of bi-variate strength of variables and step wise variable selection details. After the display of variable strength / selection 202 procedure, the internal operation run subsequent machine learning procedure as it illustrates exemplary user interfaces to visualize and develop data that facilitate a user's selection of information for automated variable strength / sel ection 202 details.
[0052] In an embodiment of the present invention as illustrated in figure 12 A, 12B, 12C and 12D the model statistics 203 display optimized model development based on variable significance, variable bi-variate strength, multi collinearity, model strength, model usage recommendations, model visualization and save data andmodel wizard as it illustrates exemplary user interfaces to visualize and develop model that facilitate model development and visualization.
[0053] In an embodiment of the present invention as illustrated in figure 17, the user will enter the inputs in the model parameters interface 301 and it comprises variable selection need, eigen value threshold for variable selection and how many clusters are needed. Based on the selection made on the model parameters interface 301 the internal operations detect correlation of variables and perform variable clustering based on user inputs as it illustrates exemplary user interfaces that facilitate a user's selection of information for setting unsupervised Machine Learning - model parameters.
[0054] In an embodiment of the present invention as illustrated in figure 18, 19, 20 and 21 the variable strength / selection 302 display variables first and second moment, variable standardization, correlation of variables (tables and visualization), variable clustering & best subset of unrelated variable selection. After the display of variable strength / selection 302 procedure, the internal operations run cluster analysis on all / selected set of variables. The model statistics 303 display cluster analysis and elbow point detection, number of segments and sum of squared errors within clusters visualization, display of clustering solution based on chosen number of clusters as well as mathematically chosen optimal number of clusters as it illustrate exemplary user interfaces that facilitate a user's selection of information for variable selection for clustering analysis, variable clustering summary, observation clustering for clustering analysis, and clustering analysis - result.
[0055] In an embodiment of the present invention as illustrated in figure 22, the user will enter the inputs in the model parameter interface 401 and it comprises variable which represent date, variable which needs to be forecasted and number of forecasts needed as it illustrate exemplary user interfaces that facilitate a user's selection of information for setting parameters for time series analysis. Based on the selection made on the model parameter interface 401, the internal operations run time series decomposition on the data.
[0056] In an embodiment of the present invention as illustrated in figure 23, the variable strength / selection 402 display time series chart, multiplicative and additive decomposition of the time series data, graphical display of visualization and standard deviation-based colour coding of residual in table. After the display of variable strength / selection 402 procedure, the internal operations find best set of parameters for ARIMA model as it illustrates sales patterns within an exemplary diagnostic for time series data.
[0057] In an embodiment of the present invention as illustrated in figure 24, the model statistics 403 display iteration details of parameter search for ARIMA model, select best parameters for ARIMA model, display forecasted values based on best ARIMA model, visualize actual and forecasted value to get a feel of forecasting accuracy as it illustrates exemplary user interfaces that facilitate a user's input of information for best model-based forecast.
[0058] In an embodiment of the present invention as illustrated in figure 25, 26, 27 and 28 the user will enter the inputs in the model parameter interface 501 and it comprises variables which indicate transaction, item, support, confidence and lift threshold. Based on the selection made on the model parameter interface 501, the internal operations run procedure to discover association rules. The model statistics 503 display details of association rules generated, and 3D interactive visualization of association rules generated as it illustrates exemplary user interfaces that facilitate a user's selection of information for association rules mining, setting association rules mining parameters, association rules mining results and data contents with interactive 3D visualization of rules for association rules mining result visualization.
[0059] In an embodiment of the present invention as illustrated in figure 29, 30, 31 and 32 the user will enter the inputs in the model parameter interface 601 and it comprises variables which indicates user ID, item ID and rating. Based on the selection made on the model parameter interface 601, the internal operations run procedure to discover similar users and recommendations. The model statistics 603 display interactive demo of collaborative filtering for any User ID, which shows similar users and recommendations, wizard to bulk download, similar user detailsand recommendations for each user as it illustrate exemplary user interfaces that facilitate a user's selection of information for collaborative filtering, setting collaborative filtering parameters, collaborative filtering result demonstration and taking dump of collaborative filtering model outcome.
[0060] In an embodiment of the present invention as illustrated in figure 33 and 34 the user will enter the inputs in the model parameter interface 701 and it comprises variables which indicate text / comment and (optional) sentiment / class, number of important features required for display of data preparation and texts for stop word additions. Based on the selection made on the model parameter interface 701, the internal operations run data preparation steps. The variable strength / selection 702 display step by step data preparation details for chosen number of records. After the display of variable strength / selection 702 procedure, the internal operations run procedure to find most frequent words and words associated with particular sentiment / class. The model statistics interface 703 display of original and modified comments, display of most frequent words (in all data or for each class - if class field is available in data), treated words for classification and associated odds, model accuracy report on test data and wizard to display actual record along with modified text pertaining to a particular word and sentiment as it illustrate exemplary user interfaces that facilitate a user's selection of information with parameters to decide for text mining - supervised (sentiment analysis) / unsupervised and setting text mining model parameters and updating.
[0061] In an embodiment of the present invention the system and method uses plenty of machine learning libraries available as open source and it also uses many graphics libraries available as open source.
[0062] In an embodiment of the present invention as illustrated in an example, an educational firm targeted 100 thousand prospects for a particular course and distributed brochure and CD about the course. Each brochure was costing 200 INR. Some 1000 finally took the course. That makes the cost of acquisition per student = 200* 100,000 / 1000 = 20,000 per student. However, the management of the educational firm feels that the course is much less profitable to them due to such a high cost of acquisition per student. They wonder if they can target much lessernumber of prospects and still can get almost same number of final students (like by targeting say 20,000 prospects, if they can get say 900 students, which will make their cost of acquisition per student = 200* 20,000 / 900 = 40000 / 9. Supervised machine learning comes into picture in such situation. The firm creates structured data of all the prospects (like their age, gender, address, educational background etc.) and tags each records where they converted or not (Response column if they took course, then Yes otherwise No). Then supervised machine learning is all about using above data and finding those sets of profile of prospects, who has high chance of taking the course.
[0063] In accordance with an example embodiment a cancer research institutes collects data (such as their demographic, food habits, past treatment history etc.) about many patients who are undergoing a particular kind of treatment and tags each patient as recovered / not recovered at the end of treatment period. Now they can analyse this data using supervised machine learning to detect the profile of patients who has high chance of recovery with this kind of treatment and then direct patients with less chance of survival to other methods of treatment.
[0064] In another embodiment of the present invention as illustrated in an example, a resort can find profile of prospects, which has high chance of taking their holiday package.
[0065] In accordance with an example embodiment a bank can find profile of prospective loan applicants who have high chance of converting in to bad loans - means they would not pay back the loans taken.
[0066] In accordance with an example embodiment a credit card company can find profile of customers, who are more suitable for an additional credit card which provides high reward on travel.
[0067] The system and method described herein offer several advantages over conventional machine learning solutions by producing highly evolved interactive dashboard for data visualization instantly without any effort also the change of graph type / metrics only requires another selection. The dashboard is even better in visualization than best visualization tools used for modelling purpose (please note - creating dashboard in most tools is not instant and it requires certain level ofunderstanding and aptitude to be able to generate dashboard, which not all users are able to do but, in this tool, it happens automatically). It applies artificial intelligence to create the list of variables, which can be considered as dependent variable and then allows user to select one of them as dependent variable. The moment user selects the dependent variable, it gives relevant applicable machine learning technique based on dependent variable statistics using Al. One cannot select any method in this tool, which is not applicable. It adopts blank canvas based approach, where it brings those prompts for user inputs / selections that is applicable in a given scenario. In a way, tool drags and brings required prompts in front of user rather than user dragging / going to the menu to select required options (which requires higher skills and has long learning curve). Or in other words tool serves the user by bringing relevant option / choice in front of user.
[0068] Most of the tool require coding expertise however the user in present method selects the dependent column, the column that user wants to ignore and the technique from option box using GUI. Tool instantly generates the required code on the fly and creates the output. It works on the produced statistical summary output on its own and then go on to build model after applying industry standard best practices on different results. Also be the binning of numeric / character field, applying filter on data, reducing VIF of model, deciding best set of variables etc. nothing requires coding. Additionally, when creating bins etc, it immediately generates the summary output to check the accuracy of process. While it allows user to make an entry to decide VIF / Number of independent variable etc., if user is novice (or he is a business user), he / she can leave prefilled option as it is and the model will be optimal based on evolved industry practices. Additionally for business users, it produces model recommendations in English texts. Maintaining the transparency is one of the major advantages, even though the tool runs in auto mode for most of the part, it writes clearly the reason that why did it adopt a particular strategy on a given data in every stage of the data.
[0069] None of the procedure requires manual execution after change in input parameters like introduction of new fields, removal of some fields, application of filter because, it applies the concept of volatility to the fullest. So, the moment userapplies filter, adds some fields and or introduces new binning, removes certain fields, changes value of VIF threshold or number of independent variables through a particular stage all subsequent stages run on its own and produces the end result on its own. It also generates all the intermediate data in the memory only but allows end data to be saved as pkl file or csv file. Hence it does not generate too many datasets in between. It is quite versatile and single tool allows users to conduct Supervised machine learning - for development and score generation, Cluster analysis, Time series analysis, Collaborative filtering, Association rule mining, Text mining and Graphical regular analysis which makes it much simpler to develop models due to zero coding, Blank canvas and Al based approach to get required inputs, inbuilt procedure for subsequent stages, transparency and above all speed of execution. Here it allows end user to configure dashboard as per his need, by just selection (no drag n drop). Then by just making different selections in the auto generated dashboard, the dashboard refreshes immediately.
[0070] It runs perfectly in multiuser situation, where different users are submitting different requirement on different browsing instances. Even one user can submit different job on different instances of browser. It can serve the entire organization and there is no need to different environment like - development, scoring, visualization server etc. Less outside help needed - a number of places, it writes the details, which are good enough that user can understand calculations or what is going on himself / herself easily, without looking for reference material. The design is very intuitive which makes segregation of menu in logical way so that anyone can understand the same within very less time. Industry practice-based threshold is one of the biggest advantages as the user cannot enter irrelevant values as each of the input has threshold defined. All the relevant procedure based on the provided inputs are applied automatically, which will be time saving and will enable even less technical person to complete the task in a fairly quick time.
[0071] It provides GUI for users to let them provide relevant inputs. It gives optional control in the hand of users to control development of models. It ensures that user inputs are not absurd (unacceptable as per mathematical / industrial practices). It only asks for relevant options based on data and analysis objective. Itdoes not ask for drag and drop, rather brings relevant options in front of users for selection. Users does not need to know or call specific machine learning algorithm / packages or graphical packages. It does in between transformation job to make sure that output of one machine learning packages can act as smooth input for some machine learning and / or graphics packages and overall, the system can produce useful and concise (and interactive if applicable) / display. It runs several logic / iterations / checks inside to analyse the output of machine learning procedure and run some other / repeat some procedure based on situation. It allows end user to create customized fields with zero coding. It demonstrates step by step details, which is concise yet good enough for transparency. Fool proofing - except the input data upload, nowhere user can select anything that makes the system crash. It produces industry practices based models. It produces industry practices / Basel norms based output like KS / GINI. It provides textual interpretation of output tables for helping users. At times, it enhances the quality of output by running some procedure on top of the output of a machine learning package (like colour coding irregular component based on z statistics) to make it easy for user to understand / interpret it. Overall, it makes super simple for end users to develop some specific types of models (supervised, unsupervised, time series etc.) and analysis (cross tab analysis, dashboards etc.).
[0072] While the preferred embodiments of the invention have been illustrated and described, it will be clear that the invention is not limited to these embodiments only. Numerous modifications, changes, variations, substitutions and equivalents will be apparent to those skilled in the art without departing from the spirit and scope of the invention, as described in the claims.
Claims
We claim:
1. A data analytics system involving creating and deploying machine learning solutions with zero coding, said system comprising: an input module in a GUI to get and treat data comprising a main page interface for taking input, displaying input data, meta data, and treated data wherein user inputs are taken for numeric variable binning interface and character data binning input is through character variable binning interface; a processing module to visualize and develop model comprising: a apply filter interface; a visual analysis interface; a cross tab analysis interface; a dashboard interface; a model parameter interface for taking user input and selection of machine learning method; an output module for displaying variable strength / selection, steps of model development, model statistics, recommendations for model usage and validation of selected machine learning model through a validation module. wherein the data analytics system involving creating and deploying machine learning solutions WITH zero coding.
2. The data analytics system as claimed in claim 1, wherein a validation module to validate selected machine learning model comprises a load model and data interface for taking user inputs; a meta data display interface for scoring and validation;a validation result display interface for out of time model validation results, graphical display of model strength and confirmation on data created with score.
3. The data analytics system as claimed in claim 1, wherein the system is zero coding and allow users to create customized fields.
4. An interactive system for visualizing input data organized for a machine learning model, said system comprising: a) means for input / upload of data and model (for validation of supervised machine learning solution on out of time data) in GUI from users; b) means for converting received input data into a graphical form; c) means for presenting the received input data to a user as a graphical visualization in form of charts and tables; d) means for prompting a user to select dependent variable feature with an out of time model validation / scoring engine; e) a filter being added to said visualization for filtering said data; f) a colour coding being added to said visualization for anomaly detection.
5. The interactive system for visualizing input data as claimed in claim 4, wherein generating a number of candidate visualizations is based on selected data, and numerically evaluating candidate input variables for their fitness as independent variable in the model, then using them for model development, determination of model strength and confirmation on data created with score.
6. The interactive system for visualizing input data as claimed in claim 4, wherein the said filter being selected from the group comprising of categorical variable based filters, numeric variable based filters, date filed based filters, apply filters / next filter GUI based on remaining data (always relevant filter only) and generate audit of rows in the data.
7. The interactive system for visualizing input data as claimed in claim 4, wherein the colour coding in the said visualization for easy visualization or detection of anomaly is applied automatically.
8. A data analytics method involving creating and deploying machine learning solutions with zero coding, said method comprising: an input module in the GUI to get and treat data which comprises: taking user inputs on the main page interface; displaying meta data from initial data; displaying data treatment details; taking user inputs on the numeric variable binning interface; taking user inputs on the character variable binning interface; a processing module in the GUI to visualize and develop model which comprises: taking user inputs on the apply filter interface; taking user inputs on the visual analysis interface; taking user inputs on the cross tab analysis interface; taking user inputs on the dashboard interface; taking user inputs on the model parameter interface for selection of machine learning method; an output module for displaying variable strength / selection, steps of model development, model statistics, recommendations for model usage and validation of selected machine learning model through a validation module. taking user inputs on load model and data interface; displaying meta data for scoring and validation;displaying validation result for out of time model validation results; displaying graphical model strength and confirmation on data created with score; wherein the data analytics method involving creating and deploying machine learning solutions is zero coding.
9. The data analytics method as claimed in claim 8, wherein after display of meta data step, the internal operation starts executing to select variables for visualization.
10. The data analytics method as claimed in claim 8, wherein the data treatment details is displaying insights about the data for visualization.
11. The data analytics method as claimed in claim 8, wherein the system is zero coding and allow users to create customized fields by selection of fields.
12. An interactive method for visualizing input data organized for a machine learning model, said system comprising: a) taking input of data in GUI from users; b) converting received input data into a graphical form; c) presenting the received input data to a user as a graphical visualization in form of charts and tables; d) prompting a user to select data and layout features with an out of time model validation / scoring engine; e) adding a filter to said visualization for filtering said data; f) adding a colour coding to said visualization for anomaly detection;13. The interactive method as claimed in claim 12, wherein generating a curated set of applicable options is automated and is based on the data and analysis objective, providing only relevant options in hand of user to control development of models.
Citation Information
Patent Citations
End-to-end machine learning pipelines for data integration and analytics
WO2022197669A1
Methods and systems for identification and visualization of bias and fairness for machine learning models
WO2022240860A1