A METHOD AND SYSTEM FOR RAPID, EXPLAINABLE, AND FIELD-AWARENESS AUTOMATED FEATURE ENGINEERING.

TR202608054TPending Publication Date: 2026-06-22QNB BANK ANONİM ŞİRKETİ
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
QNB BANK ANONİM ŞİRKETİ
Filing Date
2023-12-19
Publication Date
2026-06-22

Smart Images

  • Figure 00000039_0000
    Figure 00000039_0000
  • Figure 00000040_0000
    Figure 00000040_0000
Patent Text Reader

Abstract

The described invention relates to an automated feature engineering process based on generator-type feature generation, integrating domain knowledge through generation and elimination rules. This invention is located in the field of artificial intelligence, and particularly machine learning, and aims to improve usability, enhance prediction performance, reduce the use of processor, memory, storage, and similar resources, and decrease computation time.
Need to check novelty before this filing date? Find Prior Art

Description

36405.01 1 TARIFF FAST, EXPLAINABLE, AND FIELD-AWARENESS. A METHOD FOR AUTOMATIC FEATURE ENGINEERING AND THE SYSTEM The present invention, in particular, focuses on Deep Feature Synthesis (DFS). automatic generation based on constructor type attribute generation. Feature engineering approaches to the field of artificial intelligence and more specifically This relates to the field of machine learning and aims to improve ease of use. Improving prediction performance and the processor, memory, storage unit and 10 It aims to reduce calculation time with similar resource consumption. An attribute is a phenomenon that is generally observed but can be individually measured or It is defined as a computable property. These attributes are one of... Bringing them together provides information about the relevant phenomenon. Feature engineering, It can be defined as the creation or selection of attributes. 15 Creating interpretable, explainable, and effective attributes for the machine This determines the success of learning processes. Experts use feature creation. Although they are quite successful in this regard, the feature generation process is generally machine-based. One of the learning techniques is the longest-running predictive modeling processes. This constitutes the stage. Attribute creation or selection processes are performed by machine 20 It is estimated that this covers approximately seventy percent of the learning process. With digitalization and the widespread use of the Internet, the amount of data and The diversity has increased considerably in recent years. This diversity is reflected in data analytics and This has made information extraction processes more complex and time-consuming. Advanced computing resources (such as CPU and GPU) are required to process the data. (processor units, distributed servers), high-capacity data storage systems and parallel processing infrastructures are needed for these datasets. Because they can reach very large sizes, processing times can be longer, and Obtaining results can take hours or even days. This situation, organizations can make real-time decisions or respond effectively to events 30 36405.01 2 This can limit their capabilities. Increased processing times, data analysis... by causing delays in projects and thus missing opportunities This can result in the above-mentioned technical problems. Addressing this issue is of great importance because the data is dynamic in nature. Existing attributes, changes in distributions, or events like pandemics 5 As a result, they may lose performance over time. Also, time New data sources or new products may emerge within this context. In addition, experts attribute some important data characteristics to personal biases. This can be overlooked by machine learning models. This can lead to a decrease in performance. Therefore, the 10 mentioned above... Addressing performance losses is also important. In addition, sectors such as banking, finance, medicine, and the like are heavily involved. It is subject to regulations. Regulatory frameworks and many national regulatory bodies, The attributes used in the models should be "reasonable," and the models should be designed to be human-made. It is necessary that it be subject to review and that the final decisions be explainable. In addition, business and operations units also gain insights regarding customers. In order to obtain and build trust in the models, the models must be explainable. It demands that these requirements be met. These requirements make businesses complex and difficult to interpret. This limits the use of powerful models; instead, it uses simpler models. 20 complex but information-rich attributes that can function effectively This leads to relying on field expertise in its development. In this context The selected attributes must be interpretable and explainable. Therefore, addressing the explainability problem mentioned above is important. Specifically, 25 that were found to be significant when examined by a field expert. Attributes are defined as domain-interpretable attributes. Attribute creation or attribute selection are generally manual processes. However, this enables the automatic creation of attributes from the data. There are various methods. There are four main types of methods: (a) model (b) feature selection based on pre-trained neural networks middleware outputs 30 36405.01 3 (c) use of new attributes from input data or existing attributes (d) genetics Feature generation using evolutionary computation methods through programming. creation. Model-based approaches improve model performance. Effective for a single model or multiple models through feature selection. 5 It aims to create working features. Neural networks take their inputs linearly. It converts into outputs through a series of non-transformations, and the said Each stage in the sequence is generally considered a layer. Neural the parameters of the networks will be adapted to the desired outputs It is being updated. The outputs of the middleware layers are 10 learned features. It can be interpreted. The advantages of this approach include the manual interpretation of attributes. there is no need to design it as such and the number of attributes is kept under control This includes preventing feature explosion by controlling the system. However, Uncertainty regarding the transferability of learned attributes to new problems and the fact that the attributes in question are not explainable, that is, 15 produced by this method The fact that attributes cannot be interpreted intuitively by humans is important. These are among the disadvantages. Constructivist approaches, on the other hand, involve a model. to generate attributes regardless of whether they are dependent on predefined basic operations It applies this process iteratively to the data. These approaches are specific to a particular area. It is not model-dependent and can be calculated using appropriate basic operations and rules in area 20. It allows for the creation of attributes that can be interpreted in terms of depth. Deep Feature Synthesis (DFS) and AutoFeat methods, It is among the feature synthesis methods of the generator type. Deep Feature Synthesis Synthesis (DFS) involves multiple tables typically encountered in enterprise applications. It is designed to work. AutoFeat, on the other hand, was developed for scientific data. 25 This method prevents the creation of physically meaningless attributes. the ability to specify the units of input to the user or data scientist for this purpose It provides. However, it is a suitable solution in terms of corporate applications. It does not offer. Genetic programming and evolutionary computation approaches, genetic programming 30 and combines existing features by utilizing evolutionary computation principles. 36405.01 4 These approaches range from basic operations (to constructive type approaches) similarly) or by using genetic methods to create new features the creation of a suitability of the performance of the attributes in question. its evaluation through its function (e.g., attribute importance) and related This involves sampling a new population. The new features include 5... Crossover and mutation processes can also be used in its creation. However, attributes created in this way have limited interpretability. It has the potential to lose. In the known state of the art, U.S. patent document number US11392607B2, a Online scoring in a computing environment using 10 or more processors. It describes automated feature engineering during the process. In the known state of the art, US patent application number US20200175314A1 the document describes predictive data analytics with automated feature extraction. It explains that the structures of the present invention are based on low-order feature extraction. Methods, devices, systems, and computing for predictive data analytics. 15 It provides devices, computing assets and / or similar structures. In the known state of the art, U.S. patent document number US11042145B2, Predictive maintenance and, more specifically, reinforcement learning for predictive maintenance. It explains how to learn health indicators using this method. One of the aims of the current invention is to find area 20 for successful machine learning models. the area necessary for generating new attributes that can be interpreted in this sense The aim is to reduce the need for expertise, especially for new fields. eliminating the need for expert feature engineers who may not be available The aim is to remove it. Another objective of the present invention is to be interpretable in terms of field and / or field 25 processes of creating and / or selecting attributes that have awareness to reduce memory and storage requirements during automation and The aim is to reduce resource consumption. 36405.01 Another objective of the present invention is to enable multiple users to utilize the invention. If used by the system, feature generation experiments The goal is to ensure that resources are managed effectively. Another objective of the present invention is to identify statistically significant features. the suggestion and the explainability and predictive power of the model based on attributes 5 The aim is to identify the attribute quality that affects performance. Another objective of the present invention is to enable an expert to analyze the generated and selected attributes. The goal is to offer a system that can be implemented without the need for a programmer. Another objective of the present invention is to meet both internal and external requirements. When using machine learning methods, it is understandable to humans and 10 automatic generation and / or selection of interpretable attributes Another aim of the present invention is to provide solutions to machine learning problems. while working, domain-interpretable attributes are automatically generated. during the creation and / or selection process to users and / or data scientists The goal is to provide ease of use. 15 Another aim of the current invention is to use it in machine learning problems. attributes processed from a database and interpretable in terms of field The aim is to make it easy to create. Drawings illustrating the organizational structure aimed at achieving the purpose of the present invention. It is below: 20 Figure 1 - A conceptual architecture suitable for applications of the current invention. It is an example. All parts in the figures are numbered, and their corresponding reference numbers are provided. shown below: 1. User 25 2. Front-End Server 3. Experiment Timer 4. Backend Orchestra 5. Run Timer 36405.01 6 6. Area Manager 7. Experiment Manager 8. Data Quality Test Module 9. Target Data Reader 10. Objective Tables 5 11. Raw Data Reader 12. Raw Tables 13. Sequential Data Reader 14. Ordered Data 15. Numeric Feature Seeker 10 16. Time Window Attribute Seeker 17. Categorical Attribute Seeker 18. Subject Knowledge Test Module 19. Candidate Digital Feature Generator 20. Candidate Time Window Attribute Generator 15 21. Candidate Categorical Attribute Generator 22. Sequential Filter Module 23. Model Filter 24. Candidates Module 25. Operations Manager 20 26. Production Digital Attribute Generator 27. Production Time Window Attribute Generator 28. Production Categorical Attribute Generator 30. Field and Experiment Database 31. Source Database 25 32. Target Database 33. Field Knowledge Injection Module 34. Resource Optimized Candidate Attribute Generation Module 35. Module for Creating Domain-Aware Production Attributes It is a computer-implemented method for automated feature synthesis, 30 the method in question: 36405.01 7 - To generate attributes, communicate with at least one user interface. At least one user on at least one server in the state (1) connecting, - for the user (1) to have access to at least one user interface authorization, 5 - the user (1) must have at least one attribute to create an attribute. select the data stored in the database, - Logical elimination rules by user (1) its definition, - Machine learning problem entered by user (1) 10 In accordance with the data regarding its information, it is in communication with the server. Candidate attributes are determined using a database. research, - through the logical elimination rules entered by user (1) elimination of unnecessary attributes, 15 - the remaining attributes, the statistical properties of the filtering and the machine provided that the learning model is based on performance, with the server filtered by another module that is in communication with it, - filtered attributes and / or attribute creation templates, 20 to another database that is in communication with the server to be recorded, - filtered attributes and / or attribute creation templates, with at least one command from the user (1) through the user interface the creation and - generated filtered attributes and / or attribute creation 25 templates for use in machine learning problems This involves saving it to another database. In the computer-applied method implemented according to the current invention, firstly, the most A small number of users (1) communicate with at least one server to create attributes 30 36405.01 8 It connects to at least one user interface that is currently in use. The connection After its installation, the user (1) can access at least one user interface. is authorized and user (1) to find a solution to a machine learning problem For this purpose, it selects data stored in at least one database. Then Logical elimination rules are defined by user (1). Then 5 Logical elimination rules are defined by the user. The following applies: The purpose of the rules is to calculate new values ​​using one or more input data. Interpreting variables is difficult, complex, and computationally time-consuming. The goal is to prevent this from happening. In this context, depth, fundamental processing characteristics, and conversion are discussed. Types and similar rules are defined by user (1). Then, 10 data provided by user (1) regarding the machine learning problem using a server that communicates with a database containing the candidate Attribute generation templates are being explored. Candidate attribute generation. Templates are computational plans for generating attributes, and these At this stage, the attributes have not yet been created from the data. By user (1) 15 Unnecessary candidate attributes are identified based on defined logical elimination rules. Attribute creation templates are being eliminated. The remaining candidate attribute creation templates are: data in batches, in a way that is suitable for the available computing resources. It is used to create attributes. Then, the remaining attributes... It is filtered by another module that communicates with the server and 20 the filtering process in question is based on statistical properties and a machine learning model It is based on performance. Filtered attributes and / or attribute creation. The templates are saved by a server to another database. This then filtered attributes and / or attribute creation templates, user 25 based on at least one command given by user (1) through the interface are created. Finally, the generated and filtered attributes and / or Feature generation templates are used in machine learning problems. It is saved to another database. This is a preferred method. In its structure, these steps are performed by multiple different servers. is being brought. 30 36405.01 9 In a construction of the present invention, the method also requires the user (1) to have at least one Preliminary Connecting to the End Server (2) and the user (1) through the user screen This involves authorizing users to access a limited user interface. User screen; fields user screen, experiments user screen and runs It is selected from a group consisting of user screens. 5 on the Front End Server (2) The areas user screen is an interface for the Area Manager (6) users can add new logical rules and / or existing logical rules It allows you to remove and / or view the rules. Front End The experiments on the server (2) user screen, directed to the Experiment Manager (7) It is an interface that allows users to configure and initiate new experiments and / or 10 It allows viewing the characteristics and results of existing experiments. It provides. The runs on the Front End Server (2) user screen, Run It is an interface for the manager (25) and users can start a new run configuration and startup and / or the status of existing runs It allows for listing. 15 In a structure of the present invention, the method also includes the information received from the user (1). the command(s) and / or requests are processed by the Front End Server (2) The step of passing the data to a Backend Orchestrator (4) in the REST API structure It includes. In a structure of the present invention, the method also includes 20 received from the user (1). command and / or commands and / or requests by Back End Orchestra (4) This involves the step of passing the task to a subtask scheduler, and that subtask is... Experiment Timer (3) which plans and manages experiment requests. Operations that plans and manages the production of experimental requests It is selected from a group of (5) timers. 25 In a structural development of the present invention, the method also involves feature generation. templates by the Domain Administrator (6) to the user / users (1) stored in a Field and Experiments Database (30) in a way that it can access and / or This includes the step of hiding it. 36405.01 In a structure of the present invention, the method also includes the user (1) user experiments involve the user submitting an experiment request via their screen and the experiment in question request to Experiment Timer (3) via Back End Orchestra (4) It includes the transmission step. In a structure of the present invention, the method also includes experiment 5 by the user (1). Estimating the resource requirements of the given command(s). for the purpose of previous command settings and / or metadata relating to source data training and / or experimenting with a resource utilization regression model using if the estimated resource requirements for the command are meetable It includes the configuration step of the Experiment Manager (7). 10 The resource usage regression model uses information about previous instructions and By utilizing the results, the execution times and resource requirements of new instructions can be determined. a machine learning that makes predictions about parameters such as these It can be explained as a model. Resource utilization is the output of the regression model. In this way, resource planning is carried out more efficiently. 15 For example, an experiment conducted using 100 raw data sources lasted one hour and If it only used a maximum of 5 GB of RAM, try a new experiment with similar settings; it should yield similar results. It is anticipated that this will create a level of resource utilization. In a structure of the present invention, the method also involves the estimation of the experimental command. If the resource requirement can be met, the Experiment Manager (7) 20 Following its structuring, data analysis and machine learning related to the experiment. This includes the step of being managed by the Experiment Manager (7). In a structured approach to the present invention, the method also includes the machinery related to the experiment. at least an example or examples of the learning problem(s) and / or problems. 25 includes the step of reading a Target Table (10) from a Source Database. where user (1), user / users (1) machine learning to identify the problem and to carry out the experimental request / command. To select the relevant attributes, the Target Table must contain at least one Target Data. He chooses through his reader (9). 36405.01 11 In a structure of the present invention, the method also includes the experimental request / command. In order to achieve this, depending on the selected machine learning problems, including examples of the machine learning problem and / or related data of the examples at least one Raw Table (12), through at least one Raw Data Reader (11) This includes the step of reading from the source database. 5 In a structure of the present invention, the method also includes the experimental request / command. For the purpose of implementation, at least one Sequential number representing time and event series. Data (14) is arranged chronologically through at least one Sequential Data Reader (13). This includes the step of reading from the Source Database. In a structure of the present invention, the method also involves processing large amounts of data 10 Target Data to enable processing with limited resources The aggregate of Reader (9), Raw Data Reader (11) and Sequential Data Reader (13) It involves using a transactional approach. In a structure of the present invention, the method also includes the Target Data Reader (9), 15 read via Raw Data Reader (11) and Sequential Data Reader (13) data quality includes data type checking, missing data completion, and outlier handling. A Data Quality Tester that performs its operations using standard techniques. It includes monitoring through Module (8). Data Quality Test Module (8), In addition to standard techniques for processing outliers, clustering is used. It also utilizes methods such as k-Nearest 20. The values ​​of the generated features are calculated using k-Nearest 20 methods. Using methods such as k-Nearest Neighbors and hierarchical clustering They are clustered together. This clustering process groups similar values ​​together in the same clusters. based on the principle of collecting and assigning distant values ​​to different sets. It is based on. Outliers that fall outside these sets are considered outliers. It is accepted and replaced with values ​​from the closest clusters. Thus, 25 Outliers that emerge during the production phase are processed effectively. In a structured approach to the present invention, the method also involves the experimental results. viewing by user / users (1) and / or experimental results Field and Experiments by Experiment Manager (7) for future use Saving to the database (30) and / or Target Data Reader (9), Raw Data 30 36405.01 12 Data read by Reader (11) and Sequential Data Reader (13) and Data Quality Data controlled by Test Module (8) by Experiment Manager (7) a Field Knowledge Injection Module (33) to be sent and / or Experiment Manager (7) includes the steps for creating a search space. The search space can be created from all 5 options based on the settings specified in the experiment request. Each test request includes feature probabilities (combinations of raw data). There is a search space for this. The search space contains a large number of meaningful and meaningless objects. It is usually quite large due to the potential features it may contain. Therefore, the search Space needs to be explored intelligently. In a structure of the present invention, the method also involves the experimental request / command, 10 Subspaces within the search space through the Domain Knowledge Test Module (18) One or more sub-sections of the Field Information Injection Module (33) that detects This involves sending the data to the module, and that sub-module is the experiment. Field Information Injection settings are based on raw data type and conversion functions. Module (33) by Digital Attribute Seeker (15), Time Window 15 A group consisting of Feature Seeker (16) and Categorical Feature Seeker (17) It is selected from among them. Field Knowledge Test Module (18) allows you to access the experiment settings, raw data type and conversion. It can be predicted and explained according to its functions, in other words, its valuable sub-functions. It is structured to detect spaces. These functions are located in Area 20. Information is provided by the Injection Module (33). In a structure of the present invention, the method also includes metadata of the source data. Using, Numeric Feature Seeker (15), Time Window Feature Seeker from the group formed by (16) and the Categorical Attribute Seeker (17) 25 attribute creation templates using at least one submodule It includes the production step. In a structural development of the present invention, the method also involves feature generation. the creation of templates by the Domain Knowledge Injection Module (33) then the aforementioned feature generation templates are Resource Optimized 36405.01 13 This includes the step of submitting the Candidate Attribute Generation Module (34), where Attribute creation templates depend on the attribute's data type (Candidate Digital). Attribute Generator (19), Candidate Time Window Attribute Generator (20) from within the group consisting of and / or Candidate Categorical Attribute Generator (21) The system is redirected to the selected relevant creation module. The relevant attributes are in Source 5. It is created using data read from the database. Attribute creation. The steps are executed in parallel and in batches, and computational resources Optimizing batch processing size using resource utilization regression model. is being done. In a structure of the present invention, the method also involves processing large amounts of data 10 Target Data to enable processing with limited resources The aggregate of Reader (9), Raw Data Reader (11) and Sequential Data Reader (13) It involves using a transactional approach. In a structuring of the present invention, the method also includes candidate attributes, Candidate 15 Digital Feature Generator (19), Candidate Time Window Feature Generator by the group consisting of (20) and / or Candidate Categorical Attribute Generator (21) After creation, the selected candidate attributes are Resource Optimized. Sequential Filter Module (22) within the Candidate Attribute Creation Module (34) It includes the transfer step. The Sequential Filter Module (22) contains 20 candidate attributes. filtering by statistical features such as correlation, variance, and Gini scores. It is carried out based on this. In a structure of the present invention, the method also involves sequential filtering of attributes. from (22) statistical filtering methods / steps in the module After passing through, Resource Optimized Candidate Attribute Generation 25 Transferring the module (34) to the Model Filter Module (23) It includes the Model Filter Module (23), which is determined by the experiment command and the user. (1) a number of attributes defined by or an acceptable machine machine learning until the learning performance level is reached 36405.01 14 counting features using feature selection algorithms based on performance It reduces. In a structure of the present invention, the method also includes Domain Knowledge Injection. Module (33) and Resource Optimized Candidate Attribute Generation The module (34) performs sequential and collective searches in a loop until the search space is scanned. It includes the step of executing the process in a batch manner. All batch operations result in... The obtained attributes are temporarily stored on a hard disk in the Candidates Module (24). It is stored as such. In a framework for the current invention, the method also involves scanning the experimental space. After completion, all attributes obtained as a result of the batch processes will be 10 Determined in accordance with the experiment request / command given by user (1) depending on the number of features and / or the desired machine learning performance In order to select the generated attributes, the Sequential Filter Module (22) and This includes the step of passing the Model Filter Module (23) one last time. The generated attributes and their templates are in the Field and Experiment Database (30) 15 is stored. Following the elimination, creation, and selection steps, the selected Attributes and / or results of elimination steps with alternative attributes Front End It is presented on the server (2) and can be observed by the user (1). Also, For each attribute generated and selected by the Data Quality Test Module (8) Identified missing data completion, outlier processing, and statistical elimination 20 The methods / steps are also stored in the Field and Experiment Database (30), and Preliminary By user (1) so that it is open to inspection via End Server (2) It can be observed. Production / running command by user (1) via Front End Server (2) and / or when commands and / or requests are given, all of them are a Field and Experiment 25 Missing data completion, outlier processing and stored in the database (30) Using statistical screening methods / steps and selected features, the user (1) from the data to provide a solution to a machine learning problem Attribute creation is being performed. 36405.01 In a structure of the present invention, the method also includes the Front End by the user (1). The results of the experiment and / or the proposed attributes and / or on the server (2) Observation of the outputs of the screening steps and alternative features then, a production request is made by the user (1) via the user screen. Operation 5 via Back End Orchestra (4) to provide and / or production request It includes the step of transmitting it to its timer (5). In a structured approach to the present invention, the method also incorporates the selected experimental results. A Works Manager (25) to manage production demand using configuration and / or Field and Experiment Database (30) and / or Resource Reading the raw data in the database (31) and / or Data Quality of the raw data 10 This includes the step of editing by Test Module (8). In a preferred configuration of the present invention, in addition to the expression of production demand; The expression "request for employment" can also be used. In a structured approach to the present invention, the method also involves Data Quality Testing of raw data. After being edited by Module (8), the raw data in question will be in Area 15 and previously in the Experiment Database (30) and / or the Source Database (31) discovered and saved attributes and / or attribute generation templates Domain-Aware Manufacturing, which manages one or more sub-modules that produce To one or more sub-modules of the Attribute Creation Module (35) This includes the sending step, and the sub-module in question is Production Digital 20. Attribute Generator (26), Production Time Window Attribute Generator (27) from a group consisting of and / or Production Categorical Attribute Generator (28) is selected. In a structure of the present invention, the method also includes the attributes associated with them. Production Digital Attribute Generator (26), which produces in parallel, Production Time 25 Window Attribute Generator (27) and Production Categorical Attribute This includes the step of running the generator (28). According to the invention, the structured instructions for carrying out the steps of the method configured to provide the processor, at least one processor and that minimum 36405.01 16 a system that includes at least one memory module connected to a processor and is designed for automatic feature synthesis. a system. When executed by a computer, the computer's method according to the invention A computer program containing instructions that enable it to perform certain steps. When executed by a computer, the computer, according to the invention, performs 5 of the method. by the computer containing instructions that enable it to perform the steps a readable storage medium. The term "automatic feature synthesis" refers to feature generation and / or discovery and / or feature generation. creating and / or filtering and / or generating and / or selecting and / or It is used to encompass the meanings of elimination. 10 The machine that the user (1) and / or data scientist is trying to solve Examples of learning problems include credit scoring tasks (credit, credit-related tasks). deposit account, credit card), risk modeling (lending, limit management, delay management and collection management) and marketing (pricing, customer acquisition retention, behavioral segmentation, customer profitability, product trend, and campaign 15 Banking problems such as management can be given. Possible attributes include: the number of payment delays of 1-31 days in the last 24 months. Credit card installment usage rate over the last 3 months, available credit card limit. ratio between credit limit, total debt-to-asset ratio, overdrafts in the last 6 months. account usage changes, customer's other financial institutions 20 all previous loans granted and reported to the Credit Bureau, monthly balance of the customer's previous loans held at the Credit Bureau information such as the number of consumer loans used in the last year, annual income-debt ratio The difference lies in the sector where the customer shops the most, and the annual change in the number of products. average profit per campaign and campaign acceptance rate over the last 6 months, and similar factors. They can be described as attributes. The system performs feature extraction and generation through the use of a distributed server network. by optimizing processes, thereby increasing processing speed and calculation efficiency. This helps reduce costs and improve scalability. 36405.01 17 Traditional feature engineering methods, especially with large datasets and many When working with a large number of requests, in terms of processing speed and scalability... It encounters various limitations. The feature engineering software described... the system, for feature engineering purposes, involves numerous aspects of large-scale data. 5 designed to address the challenges associated with processing the request It incorporates a distributed architecture. Incoming requests are divided into smaller subsets. It is divided, and these subsets form a network of interconnected servers. It is distributed across the system. This distribution allows each server to manage requests in a manageable manner. Optimize computational efficiency by enabling it to work on a subset of the data. Distributed architecture allows feature engineering tasks to be performed in parallel. This allows for processing. Each server has its own allocated request. Thanks to its ability to work independently on a subset, multiple attributes The engineering task is being carried out simultaneously, thus allowing for preprocessing. The time required is significantly reduced. Thanks to the distributed architecture, the system, the work distributing the workload across multiple servers to handle numerous requests and big data 15 It can process sets seamlessly. Parallel processing and multiple The use of a server, compared to traditional single-server approaches, increases processing power. It provides significant reductions in processing times. Optimizing computing resources. Thanks to this, distributed architecture enables high-performance single-server processing. This reduces the need to invest in hardware. This reduces costs for organizations by 20. It provides cost savings. The modular nature of the distributed architecture, new attributes. If engineering techniques and improvements emerge, they will be added to the system. This allows for easy integration, thus enabling changing data preprocessing. Adaptability to their needs is ensured. The system, each component (Experiment Timer (3), Run Timer (5), 25 Work Managers (25) ensure consistent working environments, separate for the purpose of isolating dependencies and enabling easy distribution It uses a microservice architecture where the components are encapsulated within a container. These containers are part of a server cluster managed by Kubernetes. It is working on the Kubernetes cluster, distributing the containers, 30 It manages the scaling and orchestration. Cluster load balancing, 36405.01 18 Features such as automatic scaling, self-improvement, and resource management. This ensures that raw data is never transferred between microservices. and each microservice that needs raw data receives that data from the Source. It retrieves from the database (31). The system efficiently distributes experiment requests to distributed Experiment Managers (7) 5 It includes an Experiment Timer (3) that directs. Experiment Managers (7) are separate We are conducting feature engineering experiments on servers, and this As a result, data preprocessing optimized for machine learning applications and Improved scalability is provided. Experiment Timer (3), attribute a central body responsible for managing the distribution of engineering experiment requests 10 It is a component. It includes experimental requests containing experimental parameters and dataset information. The timer handles these demands through load balancing and resource allocation. Depending on its availability, it can be used by Experience Managers (7) with a smart It allocates resources in this way. The demand for more experience from available resources... If it arrives, the Experiment Timer (3) will have a 15 pending experiments as the queue is maintained and Experiment Managers (7) become available It assigns the experiments to the relevant Experiment Supervisors. Multiple Experiment Managers (7) are deployed within a distributed server architecture. It can be obtained. Each manager has the necessary resources and attribute engineering. It is equipped with capabilities. Users can submit their feature engineering experience requests to 20 It sends requests to the Experiment Timer (3). These requests are sent to the dataset, to the feature engineering techniques to be applied and any specific parameters It contains information relating to. Multiple Experiment Managers (7), different experiments It processes requests simultaneously on separate server nodes. This Parallel processing significantly reduces the total time required for feature engineering experiments by 25%. It reduces the results to a certain extent. Experiment Managers (7) complete their tasks. It transmits to the Experiment Timer (3). The Experiment Timer (3) transmits the said bringing together the results, relating to the findings from various experiments It provides a consolidated view and displays these on the Experiments user screen. It offers. 30 36405.01 19 Feature engineering experiments from the experimental phase to the production phase. such as data transfer, data consistency, resource optimization, and scalability. It can be a complex process involving challenges. This invention, a successful attribute... engineering configurations seamlessly integrated into the production environment To enable its transfer, Operation 5 is implemented along with a distributed architecture. By using its Timer (5) and Run Managers (25) It resolves the problems. The completion of the experiments and the user's feedback on the results... After being satisfied, the user selected the feature engineering experiment. It can decide to transfer it to the production environment. The Run Timer (5) It facilitates the transition. The Run Timer (5) is for specific experiments or 10 User requests for transferring configurations to the production environment Users (1) are producing which experimental results and configurations. It sends execution requests indicating that it will be transferred to the environment. Execution Its timer (5) handles incoming run requests based on urgency, data availability and It prioritizes based on factors such as computational resources. Experiment 15 Transition from the production phase to the production phase, Run Timer (5) and Run This is facilitated through the managers (25) and successful configurations This enables quick activation. Operation Timer factors such as server availability, load balancing, and resource requirements Taking into account the specific operation requests, the appropriate Operation 20 will process them. It appoints its managers (25). Operations Managers (25) attribute engineering at the production level They are responsible for the execution and maintenance of the processes. Operation Managers receive instructions from the Run Timer (5) and select configurations consistent with incoming data in the production environment 25 It enables the implementation of feature engineering applied during the experiments. configurations for the production environment to ensure consistent data preprocessing. It is transferred without any problems. Operation Managers (25) server network distributed throughout, feature engineering in the production environment It is responsible for the execution of the configurations. Multiple Runs 30 The manager (25) displays the production data simultaneously on separate server nodes. 36405.01 It is able to process data in parallel. Parallel execution increases data preprocessing speed and overall efficiency. The system increases the parallel processing capabilities of Operation Managers (25). thanks to this, as data volumes and production demands increase, effectively It is scalable.

Claims

36405.01 21 REQUESTS 1. A computer-based method for automated feature synthesis. and the method in question is: - To generate attributes, communicate with at least one user interface. at least one user on at least one server in the state (1) 5 connecting, - for the user (1) to have access to at least one user interface authorization, - the user (1) must have at least one attribute to create an attribute. select the data stored in the database, 10 - Logical elimination rules by user (1) its definition, - machine learning problem entered by user (1) In accordance with the data regarding its information, it is in communication with the server. 15 candidate attributes were selected using a database. research, - through the logical elimination rules entered by user (1) eliminating unnecessary attributes, - the remaining attributes, the statistical properties of the filtering and the machine Provided that the learning model is based on performance, the server and 20 filtered by another module that is in communication with it, - filtered attributes and / or attribute creation templates, to another database that is in communication with the server to be recorded, - filtered attributes and / or attribute creation templates, 25 with at least one command from the user (1) through the user interface the creation and - generated filtered attributes and / or attribute creation templates for use in machine learning problems This involves saving it to another database. 30 36405.01 22 2. According to any of the previous requests, the method also requires the user (1) to... connect to a Front End Server (2) and the user's (1) user screen the step of authorizing access to at least one user interface via It includes the following user screen: Fields user screen, 5 The Experiments user screen consists of an Experiments user screen and a Runs user screen. They are selected from within the group.

3. According to the method in the request number 2, Fields (2) on the Front End Server The user screen is the interface for the Domain Manager (6), and users 10 to add and / or remove new logical rules and / or existing ones It allows for the visualization of logical rules.

4. According to requests 2 and 3, the method is in the Front End Server (2) The Experiments user screen is the interface to the Experiment Manager (7), 15 enabling users to configure and launch new experiments and / or It allows viewing the characteristics and results of existing experiments. It provides.

5. According to requests 2, 3 and 4, the method is as follows: Front End Server (2) 20 Runs user screen, interface to Run Manager (25) It allows users to configure and start a new run. and / or allow listing the status of existing runs. It provides.

6. According to any of the previous requests, the method also requires the user (1) the received command(s) and / or requests, Front End Server (2) transmitted by a Backend Orchestrator (4) in a RestAPI structure It includes the step. 36405.01 23 7. According to any of the previous requests, the method also requires the user (1) the received command(s) and / or requests, Backend The step of passing the orchestrator (4) to a sub-task scheduler It includes the aforementioned sub-task scheduler; experiment requests 5 Experiment requests with an Experiment Scheduler (3) that plans and manages a Run Scheduler that plans and manages the start of production (5) are selected from a group consisting of.

8. The method also creates attributes according to any of the previous requirements. templates in Area 10 in a way that the user / users (1) can access Storage in a Field and Experiment Database (30) by the Manager (6) and / or includes the storage step.

9. According to any of the previous requests, the method also requires the user (1), The experiments involve the user submitting an experiment request via their screen and the 15 in question. Experiment request via Back End Orchestra (4) Experiment It includes the step of transmitting it to its timer (3).

10. In accordance with any of the previous requirements, the method is also experimental. as source 20 of the command(s) given by user (1). In order to estimate the requirement, previous command settings and / or A resource utilization regression using metadata from source data. Estimated resource for training the model and / or for the experimental command. If the requirement can be met, the Experiment Manager (7) It includes the steps for structuring. 25 11. Depending on any of the previous requests, the method also applies to the experiment command. if the estimated resource requirements related thereto can be met After configuring the Experiment Manager (7), data relating to the experiment analysis and machine learning steps by Experiment Manager (7) 30 It includes the management step. 36405.01 24 12. The method also applies according to any of the previous requirements, the user / users (1) to identify the machine learning problem and relevant attributes for executing the experiment request / command In order to select, the user (1) must have at least one Target Data Reader (9) 5 through which it can select the machine learning problem related to the experiment and / or At least one Objective Table containing examples and / or examples of the problems (10) includes the step of reading from a Source Database (31).

13. The method also applies according to any of the previous requests, experiment 10 the selected machine in order to execute the request / command depending on the learning problems, related to the machine learning problem At least one Raw Table containing the sample and / or related data of the samples (12), Source via at least one Raw Data Reader (11). It includes the step of reading from the database (31). 15 14. The method also applies to experiments, according to any of the previous requirements. time and event series in order to fulfill the request / command at least one Sequential Data representing (14), at least one Sequential Data Reader (13) 20 from the Source Database (31) in chronological order. It includes the reading step.

15. According to any of the previous requests, the method also applies to large quantities. to enable data to be processed with limited resources Target Reader (9), Raw Data Reader (11) and Sequential Data 25 The reader (13) should use the batch processing approach and / or Target Reader (9), Raw Data Reader (11) and Sequential Data Reader (13) The quality of the data read through this method, data type checking, missing data Methods such as hierarchical clustering with completion and k-Nearest Neighbors Data Quality 30 performs outlier processing operations using It includes steps involving auditing via Test Module (8). 36405.01 16. In accordance with any of the previous requirements, the method also applies to the experimental results. observation and / or experiment by user / users (1) The results are to be used in the future by the Experiment Manager (7). Recording to the Field and Experiments Database (30) and / or Target Reader 5 (9), Raw Data Reader (11) and Sequential Data Reader (13) data read and checked by Data Quality Test Module (8) A Field Knowledge Injection Module (33) by the Experiment Manager (7) sending and / or a search space by the Experiment Manager (7) It includes the steps for its creation. 10 17. In addition to any of the previous requirements, the method also involves experimentation. the request / command is searched via the Domain Knowledge Test Module (18) Domain Information Injection, which identifies subspaces within a given space. The step of sending the module (33) to one or more sub-modules is 15 It includes the sub-module in question, experiment settings, raw data type, and Field Information Injection Module (33) according to transformation functions by Digital Attribute Seeker (15), Time Window Attribute A group consisting of the Seeker (16) and the Categorical Attribute Seeker (17) It is selected from among them. 20 According to the method in accordance with request no. 17, the Field Knowledge Test Module (18) is mentioned. The subject functions are by the Domain Knowledge Injection Module (33) If provided, the experiment settings, raw data type, and conversion It can be estimated and explained according to its functions, in other words, it is valuable 25 It is structured in such a way as to detect subspaces.

19. The method also includes attribute creation according to any of the previous requirements. Digital templates are created using metadata from the source data. Attribute Seeker (15), Time Window Attribute Seeker (16) and 30 36405.01 26 At least one submodule selected from the Categorical Attribute Seeker (17) It includes the step of creating it through.

20. The method also includes attribute creation according to any of the previous requirements. templates by Field Knowledge Injection Module (33) 5 After their creation, one of the aforementioned attribute generation templates Source Optimized Candidate Attribute Generation Module (34) It includes the transmission step, where attribute creation templates are located. depending on the attribute's data type, to the relevant creation modules It is being directed and the creation module in question is Candidate Digital 10 Attribute Generator (19), Candidate Time Window Attribute Generator (20) and / or a group consisting of Candidate Categorical Attribute Generator (21) It is selected from among them.

21. According to request number 19, the relevant attributes in the method are from Source 15 It is created from the data read from the database (31) and the attribute The creation steps are performed in parallel and in batches. and computing resources and batch processing size resource usage. It is optimized using a regression model.

22. The method also requires, according to any of the previous requirements, that the candidate attributes, Candidate Digital Feature Generator (19), Candidate Time Window Feature Generator (20) and / or Candidate Categorical Attribute Generator (21) After being formed by a group, the selected candidate Attributes Source Optimized Candidate Attribute Generation Module 25 (34) enters a Sequential Filter Module (22) and then Sequential The Filter Module (22) correlates, variances of the candidate attributes in question. and filtering based on statistical features such as Gini scores It includes. 36405.01 27 23. The method also applies to the sequential arrangement of attributes, according to any of the previous requirements. Statistical screening methods / steps (22) in the Filter Module After passing through, Resource Optimized Candidate Attribute Generation Steps to enter a Model Filter Module (23) within Module (34) It includes 5.

24. According to the request number 22, in the method, Model Filter Module (23), attribute number, feature selection based on machine learning performance using algorithms, by the user (1) with the experiment command to reduce to a predetermined and given number or an acceptable 10 It is configured to reduce performance based on machine learning.

25. The method also includes Domain Information according to any of the previous requirements. Injection Module (33) and Source Optimized Candidate Attribute The Creation Module (34) will run a loop until the search space is scanned. the step of running them in a sequential and batch manner It includes.

26. According to any of the preceding requirements, the method also involves the experimental space. After scanning, the obtained attributes are in the Candidates Module (24) 20 This involves temporarily storing the data on a hard drive.

27. According to any of the preceding requirements, the method also involves the experimental space. After scanning, all attributes obtained from the batch processes, 25 in accordance with the experiment request / command given by user (1) specified number of features and / or desired machine learning This is the last time we'll select attributes based on performance. Passing through the Sequential Filter Module (22) and the Model Filter Module (23) It includes the steps. 36405.01 28 28. According to any of the previous requests, the method also applies to the user (1) experimental results and / or on the Front End Server (2) by outputs of proposed attributes and / or screening steps and alternatives After observing the attributes, the user (1) by the user a production request is submitted via the screen and / or the said production 5 its request to the Run Timer via the Back End Orchestra (4) (5) includes the steps of transmission.

29. The method also applies to the chosen experiment, according to any of the previous requirements. Run 10 to manage production demand using the results. Configure of the Manager (25) and / or Field and Experiment Database (30) and / or reading the raw data in the Source Database (31) and / or The raw data is processed by the Data Quality Test Module (8). It includes the steps.

30. The method also processes raw data according to any of the previous requirements. After being regulated by the Quality Test Module (8), the said raw data, previously discovered and in the Field and Experiment Database (30) and / or attributes stored in the Source Database (31) and / or a 20 that manages the submodules and / or submodules that generate attribute templates. A Field Aware Production Attribute Creation Module (35) This includes the step of sending it to one or more submodules, the word The subject of the sub-module is Production Digital Attribute Generator (26), Production Time Attribute Generator Window (27) and / or Production Categorical Attribute It is selected from a group of creators (28). 25 31. In accordance with any of the preceding requirements, the method also relates to them. Production Digital Feature Generator (26), which generates features in parallel, Production Time Window Attribute Generator (27) and Production Categorical The steps for creating attributes by the Attribute Generator (28) are described in 30 It includes. 36405.01 29 32. To perform the steps of the method in request number 1. configured to provide structured instructions to the processor, en a system containing at least one processor and at least one memory connected to that processor, 5 A system for automatic feature synthesis.

33. When executed by a computer, the computer's response to prompt number 1 a method containing instructions that enable it to perform the steps computer program. 10 34. When executed by a computer, the computer's response to prompt number 1 a computer containing instructions that enable the method to perform its steps a readable storage medium. 36405.01 18 CLAIMS 1. A computer implemented method for synthesizing automatic feature comprising: - connecting at least a user (1) to at least a user interface in communication with at least a server to generate features, 5 - authorizing the user (1) to access at least one user interface, - selecting the data by user (1) stored in at least a database to generate features, - defining logic elimination rules by the user (1), - exploring candidate features by using a database in communication 10 with the server in accordance with data about the machine learning problem information entered by the user (1), - eliminating the unnecessary features via the user (1) entered logic elimination rules, - filtering the remaining features by another module in 15 communication with the server, wherein the filtering is based on statistical properties and machine learning model performance, - saving filtered features and / or feature generation templates to another database in communication with the server, - generating the filtered features and / or the feature generation 20 templates by at least a command of the user (1) through the user interface, - saving generated the filtered features and / or the feature generation templates to some other database for use in the machine learning problems. 25 2. The method according to any one of the preceding claims further comprising the step of: connecting the user (1) to at least a frontend server (2) and authorizing the user (1) to access at least one user interface through the user 30 36405.01 19 screen, wherein the user screen selected from a group consisting of a domains user screen, an experiments user screen and a runs user screen.

3. The method according to claim 2, wherein the domains user screen in the Frontend Server (2), is the interface to the Domain Manager (6) and lets 5 users add / remove new logical rules and / or view existing ones.

4. The method according to claim 2 and 3, wherein the experiments user screen in the Frontend Server (2), is the interface to the Experiment Manager (7) and lets users configure and start new experiments, and / or view the 10 properties and results of existing ones.

5. The method according to claim 2, 3 and 4 wherein runs user screen in the Frontend Server (2), is the interface to the Run Manager (25) and lets users configure and start a new run and / or list the status of existing ones. 15 6. The method according to any one of the preceding claims, further comprising the step of: the command and / or commands and / or orders received from the user (1) transmitting to a Backend Orchestrator (4) in a RestAPI structure by the Frontend Server (2). 20 7. The method according to any one of the preceding claims, further comprising the step of: the command and / or commands and / or orders received from the user (1) transmitting to a subtask schedular by the Backend Orchestrator (4), wherein the subtask schedular selected from a 25 group of consisting an experiment schedular (3) which plans and manages experiment orders and a run schedular (5) plans and manages producing the experiment orders.

8. The method according to any one of the preceding claims, further 30 comprising the step of: storing and / or saving the feature generation 36405.01 templates in a Domain and Experiments DB (30) by the Domain manager (6) in a way that the user / users (1) can access.

9. The method according to any one of the preceding claims, further comprising the step of: giving the experiment order by the user (1) through 5 the user experiments user screen which transmits the experiment order to Experiment Schedular (3) through Backend Orchestrator (4).

10. The method according to any one of the preceding claims, further comprising the steps of: training a resource usage regression model by using 10 the previous command settings and / or the meta information of source data to estimate resource need of command and / or commands given by the user (1) for the experiment and / or setting up the Experiment Manager (7) if the estimated resource need of the experiment command is provided.

11. The method according to any one of the preceding claims, further comprising the step of: after setting up the Experiment Manager (7) if the estimated resource need of the experiment command is provided, managing data analysis and machine learning steps of the experiment by the Experiment Manager (7). 20 12. The method according to any one of the preceding claims, further comprising the step of: reading of at least a Target Table (10) which is a sample and / or samples of the machine learning problem and / or problems of the experiment from a Source DB wherein the user (1) can select via at least 25 a Target Reader (9) to identify the machine learning problem of the user / users (1) to select related features to make the experiment order / command happen.

13. The method according to any one of the preceding claims, further 30 comprising the step of: reading of at least a Raw Table (12) which is 36405.01 21 subjective data of the sample and / or samples of the machine learning problem from the Source DB (31) based on the selected machine learning problems via at least a Raw Reader (11) to make the experiment order / command happen.

14. The method according to any one of the preceding claims, further comprising the step of: reading of at least a Sequential Data (14) which represents time and event series from the Source DB (31) in time order via at least a Sequential Reader (13) to make the experiment order / command happen. 10 15. The method according to any one of the preceding claims, further comprising the steps of: utilizing of the Target Reader (9), Raw Reader (11) and the Sequential Reader (13) steps in batches to allow large amounts of data to be processed with limited resources and / or controlling the quality of 15 data read with the Target Reader (9), the Raw Reader (11) and the Sequential Reader (13) via a Data Quality Tester (8) which performs data type checking, missing data imputation and outlier handling with methods such as k-Nearest Neighbors and hierarchical clustering.

16. The method according to any one of the preceding claims, further comprising the steps of: observing experiment results by the user / users (1), and / or saving experiment results to Domain and Experiment DB (30) to future usage of the experiment results by the Experiment Manager (7), and / or sending the data read by Target Reader (9), the Raw Reader (11) and 25 the Sequential Reader (13) and controlled by the Data Quality Tester (8) to a Domain Knowledge Injection module (33) by the Experiment Manager (7) and / or creating a search space by the Experiment Manager (7).

17. The method according to any one of the preceding claims, further 30 comprising the step of: sending the experiment order / command to a 36405.01 22 submodule and / or submodules of the Domain Knowledge Injection module (33) which detects subspaces in the search space through a Domain Knowledge Tester (18), wherein the submodule selected from a group of consisting a Numerical Feature Searcher (15), a Time Window Feature Searcher (16) and a Categorical Feature Searcher (17) according to 5 experiment settings, type of raw data and transformation functions by the Domain Knowledge Injection Module (33).

18. The method according to Claim 17, wherein the Domain Knowledge Tester (18) is configured to detect predictable and explainable, in other words 10 valuable, subspaces according to the experiment settings, type of raw data and transformation functions, where these functions are given by the Domain Knowledge Injection Module (33).

19. The method according to any one of the preceding claims, further 15 comprising the step of: generating the feature generation templates by using at least one submodule which are the Numerical Feature Searcher (15), the Time Window Feature Searcher (16) and the Categorical Feature Searcher (17) from the meta information of source data.

20. The method according to any one of the preceding claims, further comprising the step of: after generating the feature generation templates by the Domain Knowledge Injection Module (33), forwarding to the feature generation templates to a Resource Optimized Candidate Feature Creation module (34) where the feature generation templates are routed to 25 corresponding generation modules wherein the generation module selected from a group of consisting a Candidate Numerical Feature Generator (19), a Candidate Time Window Feature Generator (20) and / or a Candidate Categorical Feature Generator (21) based on the feature’s data type. 36405.01 23 21. The method according to Claim 19, wherein the related features are generated from data read from the Source DB (31) and the features generation steps are performed in parallel and in batches, and the computational resources and batch size are optimized using the resource usage regression model. 5 22. The method according to any one of the preceding claims, further comprising the step of: after generating candidate features by a group of consisting a Candidate Numerical Feature Generator (19), a Candidate Time Window Feature Generator (20) and / or a Candidate Categorical Feature 10 Generator (21), the selected candidate features are entered into a Sequential Filter module (22) within the Resource Optimized Candidate Feature Creation module (34) and then the Sequential Filter module (22) filters the candidate features based on statistical properties such as correlation, variance, and Gini scores. 15 23. The method according to any one of the preceding claims, further comprising the steps of: after passing the statistical elimination methods / steps in the Sequential Filter module (22), the features enter a Model Filter module (23) within the Resource Optimized Candidate Feature 20 Creation module (34).

24. The method according to Claim 22, wherein the Model Filter module (23) is configured to reduce the number of features by using feature selection algorithms based on machine learning performance, either to a user (1) 25 defined number which is given with the experiment command or based on acceptable machine learning performance.

25. The method according to any one of the preceding claims, further comprising the step of: running the Domain Knowledge Injection module 30 36405.01 24 (33) and the Resource Optimized Candidate Feature Creation module (34) sequentially and in batches in a loop until the search space is scanned.

26. The method according to any one of the preceding claims, further comprising the step of: after the experimental space is scanned resulting 5 features temporary saved in a hard disk in a Candidates module (24).

27. The method according to any one of the preceding claims, further comprising the steps of: after the experimental space is scanned, all the features from the batches are passing through the Sequential Filter module 10 (22) and the Model Filter module (23) for the last time to select the generated features based on the number features determined in accordance with the experiment order / command given by the user (1) and / or the desired machine learning performance.

28. The method according to any one of the preceding claims, further comprising the steps of: after observing the experiment results and / or suggested features and / or outputs of the elimination steps and alternative features on Frontend Server (2) by the user (1), a production order given by the user (1) through the user screen, and / or transmitting the production order 20 to Run Schedular (5) through Backend Orchestrator (4).

29. The method according to any one of the preceding claims, further comprising the steps of: setting up a Run Manager (25) to manage production order using selected experiment results and / or reading the raw 25 data in the Domain and Experiment DB (30) and / or Source DB (31) and / or arranging raw data by the Data Quality Tester (8).

30. The method according to any one of the preceding claims, further comprising the steps of: after arranging raw data by the Data Quality Tester 30 (8), sending the raw data to a submodule and / or submodules of a Domain 36405.01 Aware Production Feature Creation module (35) which manages a submodule and / or submodules that produce previously discovered and saved features and / or feature templates in the Domain Experiment DB (30) and / or Source DB (31) wherein the submodule selected from a group of consisting a Production Numerical Feature Generator (26), a Production 5 Time Window Feature Generator (27) and / or a Production Categorical Feature Generator (28).

31. The method according to any one of the preceding claims, further comprising the steps of: after the Production Numerical Feature Generator 10 (26), the Production Time Window Feature Generator (27) and the Production Categorical Feature Generator (28) that parallelly produce features related to themselves.

32. A system for synthesizing automatic feature comprising: at least a processor and at least a memory coupled to at least the processor and configured to provide the processor with instructions configured to perform the steps of the method of claim 1.

33. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the steps of the method of claim 1.

34. A computer-readable storage medium comprising instructions which, when 25 executed by a computer, cause the computer to carry out the steps of the method of claim 1.