Automatic Configuration Processing Method and System Developed for Risk Control Index Calculation Based on FlinkSQL
By adopting FlinkSQL-based automatic configuration processing method in Flink application development, the problem of data source configuration redundancy and coupling is solved, the development efficiency is improved and the cost is reduced, and the analysis of data systems that do not support SQL is supported.
Patent Information
- Application Number
- CN202211108450.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-09-09
AI Technical Summary
During the Flink application development process, the configuration and code of multiple data sources, multiple structural types of data and multiple data implementation systems are redundant and messy, with frequent updates and iterations, and high coupling, resulting in reduced development efficiency and increased development costs.
An automatic configuration and processing method based on FlinkSQL is adopted for risk control indicator calculation and development. Through three steps of reading data process, data processing conversion and data implementation, the automatic configuration and processing of data sources are realized, and the development process is simplified.
It effectively solves the problem of configuration and code redundancy in Flink application development, improves development efficiency, reduces development costs, and provides SQL script support for data systems that do not support SQL.
Smart Images

Figure CN115469941B_ABST
Abstract
Description
Technical Field
[0001] An automatic configuration processing method and system for risk control index calculation and development based on FlinkSQL, which is used to customize source data, the configuration and call redundancy problems of data landing, and the distribution and processing of different structured data during the Flink application development process, belonging to the technical field of Flink real-time batch data processing. Background Art
[0002] Apache Flink (Flink) is a framework and distributed processing engine for stateful computing on unbounded and bounded data streams. Flink can run in all common cluster environments and can perform calculations at in-memory speed and any scale.
[0003] Flink is a distributed processing engine for both streaming data and batch data. It is mainly implemented in Java code. Currently, it mainly relies on the contributions of the open source community for development. For Flink, the main scenarios it has to handle are streaming data, and batch data is just an extreme special case of streaming data. In other words, Flink will treat all tasks as streams, which is also its greatest feature. Flink can support local rapid iteration and some circular iteration tasks.
[0004] Flink is powerful and supports the development and operation of various different types of applications. Its main features include: integration of batch and stream, precise state management, event time support, and exactly-once state consistency guarantee, etc. Flink can not only run on various resource management frameworks including YARN, Mesos, and Kubernetes, but also supports independent deployment on bare-metal clusters. In the case of enabling the high-availability option, it has no single point of failure problem. It has been proven that Flink can be scaled to thousands of cores, its state can reach the TB level, and it can still maintain the characteristics of high throughput and low latency. There are many demanding stream processing applications around the world running on Flink.
[0005] Currently, Flink is widely used in the field of big data, especially in real-time computing. However, during the Flink application development process, the following technical problems exist:
[0006] The problems of configuration and code redundancy and clutter of multiple data sources (such as Kafka, relational databases), multiple structural types of data, and multiple data landing systems (such as Es, Kafka, relational databases), frequent update and iteration, and high coupling degree, resulting in reduced development efficiency and increased development costs. Summary of the Invention
[0007] In view of the problems in the above research, the purpose of the present invention is to provide an automatic configuration processing method and system for risk control index calculation and development based on FlinkSQL, so as to solve the problems of configuration and code redundancy and clutter of multiple data sources, multiple structural types of data, and multiple data landing systems in the process of Flink application development, frequent update and iteration, and high coupling degree, resulting in reduced development efficiency and increased development costs.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] An automatic configuration processing method for risk control index calculation and development based on FlinkSQL includes the following steps:
[0010] Step 1. Read the data flow:
[0011] According to the data source, the data source type parameter is passed in, and the configuration parameters are read from the configuration reference data of the data source type parameter or an external configuration file. Based on the consumption interface of the data source that has been inherited by Flink, the data source connection logic is obtained, and a consumer instance object is obtained according to the configuration parameters to read the data, and a Flink data set is returned. Among them, the data source is Kafka or a relational database, and the relational database includes mysql and oracle. The Flink data set is the intermediate data result in the Flink application program, including the DataStream data set or the DataSet data set, and the initial data set is generated after reading the data source;
[0012] Step 2. Data processing and transformation:
[0013] Based on the risk control engine system, the attributes in the risk control JSON message are obtained, and the attributes are mapped to the logical field names and stored in the configuration table. Based on the configuration table, the sql script, and the given pseudo-sql script, the java parsing code of JSON is parsed. And based on the java parsing code of JSON and the Flink data set containing the topic engine fields obtained by reading the topic engine data of the data source in the risk control engine system in Step 1, it is registered as a Flink table and corresponding to the table name in the pseudo-sql script, and a Flink data set is returned;
[0014] Step 3. Data landing:
[0015] Based on the Flink data set obtained in Step 2, the configuration reference data is passed in for data landing.
[0016] Further, the specific steps of Step 1 are:
[0017] Step 1.1 Define a data source distribution class and method for the data source reading interface:
[0018] Define a data source distribution class to distinguish between batch processing and stream processing. The data source distribution class includes batch processing BatchSourceData and stream processing StreamSourceDatas;
[0019] The data source distribution method getData() is used to call the init() method of the data source template class and parse the configuration file based on the incoming data source type parameter of batch processing BatchSourceData or stream processing StreamSourceDatas. The specific process is as follows: According to the incoming data source type parameter of batch processing BatchSourceData and stream processing StreamSourceDatas, call the built-in interface of Flink to connect to the data source to call the corresponding init() method of the data source template class. The built-in interface of Flink to connect to the data source means that Flink inherits the consumption interface of the data source by default. The consumption interface refers to the connector that integrates the data source with Flink. Obtain the consumer instance object through the configuration parameters;
[0020] Step 1.2 Define the data source startup method for development:
[0021] Since Flink inherits the consumption interface of the data source by default, define the data source startup method onStart() to obtain the data source connection logic without calling the consumption interface of the data source, that is, there is no need to rewrite the operation logic of reading data. Specifically: The data source has defined SQL to obtain the tables and fields to be read in the logic of establishing a connection and preparing for reading the database driver parameters in the early stage. In the data source startup method onStart(), the defined data reading logic is passed into the execution method ResultSet, and a Flink data set is returned;
[0022] Step 1.3 Define and develop a data source template class and methods for data source configuration based on Steps 1.1 - 1.2:
[0023] Define a data source template class to distinguish between batch processing and stream processing. The data source template class includes the data stream of batch processing BatchMysqlSourceData and the data source of stream processing StreamKafkaSourceData;
[0024] The method init() of the data source template class is used to receive the data source type parameter of the batch data stream BatchMysqlSourceData or the stream processing data source StreamKafkaSourceData passed in by the data source distribution method getdata() to call the method init() of the data source template class. The method init() of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameter data of the data source type parameter in the data source type parameter or the external configuration file to obtain the parsed configuration parameters, and call the built-in interface of flink to connect to the data source to return the Flink data set for batch processing or stream processing;
[0025] Step 1.4 Define and develop the specific data source class and method:
[0026] Based on the specific data source class in Steps 1.1 - 1.3, that is, based on the Flink application calling Data Source A, through the data source type parameter of Data Source A passed in by the data source distribution method getData(), call the method init() of the data source template class based on the data source type parameter of Data Source A. The method init() of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameter data of the data source type parameter of Data Source A in the data source type parameter or the external configuration file to obtain the parsed configuration parameters, pass the configuration parameters into the method init() of the data source template class, then pass them into the built-in interface in flink for reading Data Source A to create a consumer instance object of Data Source A, and then call the core method addSource of flink to pass the consumer instance object as a parameter to read the data. After the data is read, the result set is returned. Among them, the result is a data set of database queries, and Data Source A is a Kafka, mysql, or oracle data source.
[0027] Furthermore, the data source startup method onStart() in Step 1.2 also defines a data source exception handling and post-processing method:
[0028] Data source exception handling method: During the process of obtaining and reading the data source, use try catch to capture exceptions, including connection timeouts, configuration exceptions, or data source exceptions;
[0029] Post-processing method:
[0030] After the data source connection is successful and the data reading is completed, call the database driver to create a prepared statement PrepareStatement and the close() method of the connection interface name Connection to close the prepared statement and the connection to recycle resources.
[0031] Further, the specific steps of step 2 are:
[0032] Step 2.1 Define the mapping between source fields and logical fields:
[0033] Obtain the attributes in the risk control JSON message through the risk control engine system configuration, define the logical field name according to the business logic, obtain the corresponding channels and paths at all levels, correspond the attributes in the JSON message to the logical field name, and obtain the hierarchical structure, JSON object name and JSON object array name according to the complete path. Among them, the JSON attribute can correspond to the json object of the A level, the json object of the A level can correspond to the json object of the B level, and so on. According to the complete path, the multi-level structure can be obtained. The attributes in the risk control JSON message include the data source type and shcema information. The shcema information is the metadata information of the source data;
[0034] Step 2.2 configures the mapping relationship into the database, that is, configures the mapping relationship between each level, JSON object name and JSON object array name obtained in step 2.1 into a configuration table in the database for storing the JSON data structure;
[0035] Step 2.3 Develop an automated script to encapsulate the parsing logic:
[0036] Develop the calculation conversion and aggregation logic to form a SQL script and encapsulate it into a function for storing SQL scripts. Based on the given pseudo SQL script (the given data of business personnel analyzing business indicators, such as students and teachers), use the encapsulated SQL script to parse the required logical fields in the configuration table to obtain the JSON Java parsing code as the result output returned by the function. The encapsulated SQL script and Java parsing code are stored in the Flink operator.
[0037] Step 2.4 Implement Flink custom risk control engine indicators based on the pseudo SQL script:
[0038] Based on step 1, read the topic engine data of the data source in the risk control engine system, and process the topic engine data based on the json parsing tool fastjson in the Flink conversion operator and the json parsing code, and finally return it to the Flink data set. The json parsing tool fastjson in the Flink conversion operator is a javaassist or asm dynamic generation method
[0039] Step 2.5 Register the Flink data table and calculate the risk control engine indicators:
[0040] Register the Flink dataset as a Flink table through the built-in method fromDataStream() of Flink, and map it to the table name in the pseudo SQL script. Then, call the sqlQuery() method built into Flink Sql to return the Flink dataset.
[0041] Furthermore, the specific steps of step 3 are as follows:
[0042] Step 3.1 Define the development data source landing distribution class and method:
[0043] Define the data landing encapsulation class, including batch processing and stream processing. Since in the batch processing and stream processing modes, the landing target system is essentially materialization and persistence, and there is no essential difference between the two modes. Therefore, define the data source landing distribution class SinkAssigner of the landing target system, and define the static method sinkAssign() of the data source landing distribution class SinkAssigner to pass in the configuration parameters based on the Flink dataset obtained in step 2.5. Among them, SinkAssigner is the parameter passed in during landing, and sink represents the data landing target system;
[0044] Step 3.2 Define the development data landing template class and method:
[0045] The name of the data landing template class is named after the landing target system + Sink. Define the data landing template class method handleSink() corresponding to the data landing template class. When the static method sinkAssign() of the data source landing distribution class SinkAssigner passes in the parameters and determines that the data source is the landing target, at this time, the data source type parameter or the configuration file parameter is called by the data source template class method init() to call the data source startup method onStart() to call the execution method ResultSet to parse the configuration. After parsing the configuration, call the data landing template class method handleSink() to process the data landing, and finally, the sinkPost() method processes the data landing exception and the subsequent operations after the data is landed in the database;
[0046] Step 3.3 Define the development data landing class and method:
[0047] Develop specific data landing classes and methods based on the data source floor classes and methods. Specifically: land data to data source A, pass in the parameter data source based on the Flink dataset, that is, configure and pass in the url, username, and password of the data source. At this time, distribute to the data landing template class and method to call the init() method of the data source template class to call the data source startup method onStart(). Through the data source startup method onStart(), call the execution method ResultSet to configure the environment, establish a database connection object and a prepared statement. Then define the sql statement for writing to the database in data source A in the handleSink() method called by the data landing template class, and then execute the sql statement to land the data in the database.
[0048] Furthermore, the specific steps for the sinkPost() method in step 3.2 to handle data landing exceptions and subsequent operations after data landing are as follows:
[0049] Exception handling method for data landing target: During the process of obtaining and reading the data source, use trycatch to capture exceptions, including connection timeouts, configuration exceptions, or data source exceptions;
[0050] Post-processing method:
[0051] After the data source connection is successful and the data reading is completed, call the database driver to create a prepared statement PrepareStatement and the close() method of the Connection interface of the connection to close the prepared statement and the connection to recycle resources.
[0052] An automatic configuration processing system for risk control index calculation and development based on FlinkSQL, including:
[0053] Data reading process module:
[0054] According to the data source type parameter passed in by the data source, read the configuration parameters from the configuration parameter data of the data source type parameter or the external configuration file, and obtain the data source connection logic based on the consumption interface of the data source already inherited by Flink to obtain the consumer instance object according to the configuration parameters to read the data, and return the Flink dataset. Among them, the data source is Kafka or a relational database, and the relational database includes mysql and oracle. The Flink dataset is the intermediate data result in the Flink application program, including the DataStream dataset or the DataSet dataset. The initial dataset is generated after reading the data source;
[0055] Data processing module:
[0056] Obtain the attributes in the risk control JSON message based on the risk control engine system, map the attributes to the logical field names and store them in the configuration table, parse the JSON Java parsing code based on the configuration table, SQL script and the given pseudo-SQL script, and register the Flink dataset containing the topic engine fields obtained from the topic engine data of the data source in the risk control engine system based on the JSON Java parsing code and step 1 as a Flink table, corresponding to the table name in the pseudo-SQL script, and return the Flink dataset;
[0057] Data landing module:
[0058] The Flink dataset obtained in step 2 is passed into the configuration parameter data for data landing.
[0059] Furthermore, the specific implementation steps of the data reading process module are as follows:
[0060] Step 1.1 Define and develop a data source distribution class and method for the data source reading interface:
[0061] Define a data source distribution class to distinguish between batch processing and stream processing. The data source distribution class includes batch processing BatchSourceData and stream processing StreamSourceDatas;
[0062] The data source distribution method getData() is used to call the init() method of the data source template class based on the incoming data source type parameter of batch processing BatchSourceData or stream processing StreamSourceDatas, and parse the configuration file. The specific process is as follows: According to the incoming data source type parameter of batch processing BatchSourceData and stream processing StreamSourceDatas, call the built-in interface of flink to connect to the data source to call the corresponding init() method of the data source template class. The built-in interface of flink to connect to the data source means that flink inherits the consumption interface of the data source by default. The consumption interface refers to the connector for integrating the data source with Flink, and the consumer instance object is obtained through the configuration parameters;
[0063] Step 1.2 Define and develop a data source startup method:
[0064] Since Flink itself inherits the consumption interface of the data source, a data source startup method onStart() is defined to obtain the data source connection logic without calling the consumption interface of the data source, that is, there is no need to rewrite the operation logic of reading data. Specifically: the data source has defined SQL to obtain the tables and fields to be read in the logic of establishing a connection and preparing for reading the database driver parameters in the early stage. In the data source startup method onStart(), the defined data reading logic is passed into the execution method ResultSet, and a Flink data set is returned;
[0065] Step 1.3 Define and develop a data source template class and method for data source configuration based on Steps 1.1 - 1.2:
[0066] Define a data source template class to distinguish between batch processing and stream processing. The data source template class includes the batch data stream BatchMysqlSourceData and the stream processing data source StreamKafkaSourceData;
[0067] The method init() of the data source template class is used to receive the data source type parameter of the batch data stream BatchMysqlSourceData or the stream processing data source StreamKafkaSourceData passed in by the data source distribution method getdata() to call the method init() of the data source template class. The method init() of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameters of the data source type parameter or the external configuration file to obtain the parsed configuration parameters, and call the built-in interface for Flink to connect to the data source to return a batch processing or stream processing Flink data set;
[0068] Step 1.4 Define and develop specific data source classes and methods:
[0069] Based on the specific data source classes in Steps 1.1 - 1.3, that is, based on the Flink application calling Data Source A, the data source type parameter of Data Source A passed in through the data source distribution method getData(). Based on the data source type parameter of Data Source A, the data source template class method init() is called. The data source template class method init() is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameters of the data source type parameter of Data Source A or in the external configuration file to obtain the parsed configuration parameters and pass the configuration parameters into the data source template class method init(). Then, an interface built into Flink for reading Data Source A is used to create a consumer instance object of Data Source A. Then, the core method addSource of Flink is called to pass the consumer instance object as a parameter to read the data. After the data is read, a result set is returned. Among them, the result is a data set of database queries. Data Source A is a Kafka, mysql, or oracle data source.
[0070] Further, the specific implementation steps of the data processing module are as follows:
[0071] Step 2.1 Define the mapping between source fields and logical fields:
[0072] Obtain the attributes in the risk control JSON message through the risk control engine system configuration, define the logical field names according to the business logic, obtain the corresponding channels and each level of paths, map the attributes in the JSON message to the logical field names, and obtain each hierarchical structure, JSON object name, and JSON object array name according to the complete path. Among them, the JSON attribute can correspond to the json object at level A, and the json object at level A can correspond to the json object at level B, and so on. The multi - level structure can be obtained according to the complete path. The attributes in the risk control JSON message include the data source type and shcema information. The shcema information is the metadata information of the source data;
[0073] Step 2.2 Configure the mapping relationship into the database, that is, configure the mapping relationships of each level, JSON object name, and JSON object array name obtained in Step 2.1 into the configuration table in the database used to store the JSON data structure;
[0074] Step 2.3 Develop an automated script to encapsulate the parsing logic:
[0075] Develop the calculation conversion and aggregation logic to form an SQL script and encapsulate it into a function for storing the SQL script. Based on the given pseudo-SQL script (data given by business personnel for analyzing business metrics, such as students, teachers, etc.), use the encapsulated SQL script to parse the required logical fields in the configuration table to obtain the Java parsing code of JSON, and output it as the result returned by the function. Among them, the encapsulated SQL script and the Java parsing code are stored in the Flink operator;
[0076] Step 2.4 Implement the Flink custom risk control engine metrics according to the pseudo-SQL script:
[0077] Based on the topic engine data of the data source in the risk control engine system read in Step 1, and use the JSON parsing tool fastjson in the Flink transformation operator to cooperate with the JSON parsing code to process the topic engine data, and finally return it to the Flink dataset. Among them, the JSON parsing tool fastjson in the Flink transformation operator is a method dynamically generated by javaassist or asm;
[0078] Step 2.5 Register the Flink data table and calculate the risk control engine metrics:
[0079] Register the Flink dataset as a Flink table through the Flink built-in method fromDataStream(), corresponding to the table name in the pseudo-SQL script, and then call the sqlQuery() method built in Flink Sql to return the Flink dataset.
[0080] Furthermore, the specific implementation steps of the data landing module are as follows:
[0081] Step 3.1 Define and develop the data source landing distribution class and method:
[0082] Define the data landing encapsulation class, including batch processing and stream processing. Since the landing target system is essentially materialized and persisted in the batch processing and stream processing modes, and there is no essential difference between the two modes. Therefore, define the data source landing distribution class SinkAssigner of the landing target system, and define the static method sinkAssign() of the data source landing distribution class SinkAssigner to pass in the configuration parameter data based on the Flink dataset obtained in Step 2.5. Among them, SinkAssigner is the parameter passed in during landing, and sink represents the data landing target system;
[0083] Step 3.2 Define and develop the data landing template class and method:
[0084] The class name of the data landing template is named after the landing target system + Sink. Define the data landing template class method handleSink() for the corresponding data landing template class. When the static method sinkAssign() of the data source landing distribution class SinkAssigner passes in parameters and determines that the data source is the landing target, at this time, the data source type parameter or the configuration file parameter is called by the data source template class method init() to call the data source startup method onStart() to call the execution method ResultSet to parse the configuration. After parsing the configuration, call the data landing template class method handleSink() to process the data landing. Finally, the sinkPost() method processes data landing exceptions and subsequent operations after the data is stored in the database;
[0085] Step 3.3 Define the development of the data landing class and method:
[0086] Develop specific data landing classes and methods according to the data source landing board classes and methods. Specifically: to land data to the A data source, based on the Flink dataset passing in parameters of the data source, that is, passing in the url, username, and password configurations of the data source. At this time, it is distributed to the data landing template class and method to call the data source template class method init() to call the data source startup method onStart(). Through the data source startup method onStart(), call the execution method ResultSet to configure the environment, establish a database connection object and a prepared statement, and then define the sql statement for writing to the database in the data landing template class method handleSink() for the A data source. Then execute the sql statement to store the data in the database.
[0087] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0088] First, the present invention perfectly decouples and encapsulates the three essential steps in the flink application, namely Source, Tranformation, and Sink, to simplify the development process;
[0089] Second, based on the data process, data processing transformation, and data landing in this case, the present invention solves the problems of configuration and code redundancy and clutter of multiple data sources, multiple structural types of data, and multiple data landing systems in the Flink application development process, frequent update and iteration, and high coupling degree, resulting in reduced development efficiency and increased development costs;
[0090] Third, the present invention provides support for sql scripts for data systems that do not support sql, facilitating business personnel to analyze data for various systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 It is a schematic diagram of the process framework of the present invention;
[0092] Figure 2 It is a schematic structural diagram of the configuration table in the present invention;
[0093] Figure 3 It is a schematic diagram of the Java parsing code that parses JSON from the encapsulated SQL script in the present invention. Specific implementation manners
[0094] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners.
[0095] An automatic configuration processing method for risk control index calculation and development based on FlinkSQL includes the following steps:
[0096] Step 1. Data reading process:
[0097] According to the data source, the data source type parameter is passed in, and the configuration parameters are read from the configuration parameter data of the data source type parameter or an external configuration file. Based on the consumption interface of the data source that has been inherited by Flink itself, the data source connection logic is obtained to obtain a consumer instance object according to the configuration parameters to read the data, and a Flink data set is returned. Among them, the data source is Kafka or a relational database, and the relational database includes mysql and oracle. The Flink data set is an intermediate data result in the Flink application program, including a DataStream data set or a DataSet data set. The initial data set is generated after reading the data source;
[0098] The specific steps are as follows:
[0099] Step 1.1 Define a data source distribution class and method for the data source reading interface:
[0100] Define a data source distribution class to distinguish batch processing and stream processing. The data source distribution class includes batch processing BatchSourceData and stream processing StreamSourceDatas;
[0101] The data source distribution method getData() is used to call the init() method of the data source template class and parse the configuration file based on the incoming data source type parameters of batch processing BatchSourceData or stream processing StreamSourceDatas. The specific process is as follows: According to the incoming data source type parameters of batch processing BatchSourceData and stream processing StreamSourceDatas, call the built-in interface of Flink to connect to the data source to call the corresponding init() method of the data source template class. The built-in interface for Flink to connect to the data source (such as the built-in Kafka interface) means that Flink inherits the consumption interface of the data source by default. The consumption interface refers to the connector for integrating the data source with Flink, and obtain the consumer instance object through the configuration parameters;
[0102] Step 1.2 Define the data source startup method for development:
[0103] Since Flink inherits the consumption interface of the data source by default, define the data source startup method onStart() to obtain the data source connection logic without calling the consumption interface of the data source, that is, there is no need to rewrite the operation logic of reading data. Specifically: The data source has defined the SQL to obtain the tables and fields to be read in the logic of establishing the connection and preparing for reading data by obtaining the database driver parameters in the early stage. In the data source startup method onStart(), the defined data reading logic is passed into the execution method ResultSet, and a Flink dataset is returned;
[0104] The data source startup method onStart() also defines the data source exception handling and post-processing methods:
[0105] Data source exception handling method: In the process of obtaining and reading the data source, use try catch to wrap and catch exceptions, including connection timeouts, configuration exceptions, or data source exceptions;
[0106] Post-processing method:
[0107] After the data source connection is successful and the data reading is completed, call the close() method of the PreparedStatement and Connection interfaces created by the database driver to close the PreparedStatement and the connection to recycle resources.
[0108] Step 1.3 Define and develop the data source template class and methods for data source configuration based on Steps 1.1 - 1.2:
[0109] Define data source template classes to distinguish between batch processing and stream processing. The data source template classes include the batch data stream BatchMysqlSourceData and the stream processing data source StreamKafkaSourceData;
[0110] The method init() of the data source template class is used to receive the data source type parameter of the batch data stream BatchMysqlSourceData or the stream processing data source StreamKafkaSourceData passed in by the data source distribution method getdata() to call the init() method of the data source template class. The init() method of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameter data of the data source type parameter in the external configuration file to obtain the parsed configuration parameters, and call the built-in interface of Flink to connect to the data source to return a Flink data set for batch processing or stream processing;
[0111] Step 1.4 Define the development of specific data source classes and methods:
[0112] Based on the specific data source classes in Steps 1.1 - 1.3, that is, based on the Flink application calling Data Source A, the data source type parameter of the type of Data Source A passed in by the data source distribution method getData() (that is, the data source distribution method getData() determines which specific data source class to call to read data according to the transmitted data source type parameter (for example, when Data Source A is mysql, call the data source template class of mysql, and when Data Source A is kafka, call the data source template class of kafka)). Based on the data source type parameter of the type of Data Source A, call the init() method of the data source template class. The init() method of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet (the execution method ResultSet here is the execution method for Data Source A, that is, to find the data to be queried in Data Source A) to parse the configuration parameter data of the data source type parameter of Data Source A in the external configuration file to obtain the parsed configuration parameters, pass the configuration parameters into the init() method of the data source template class, then pass them into the built-in interface in Flink for reading Data Source A to create a consumer instance object of Data Source A, and then call the core method addSource of Flink to pass the consumer instance object as a parameter to read data. After the data is read, the result set is returned. Among them, the result is a data set of database queries. Data Source A is a Kafka, mysql, or oracle data source.
[0113] Step 2. Data processing and transformation:
[0114] Based on the risk control engine system, the attributes in the risk control JSON message are obtained, and the attributes are mapped to the logical field names and stored in the configuration table. Based on the configuration table, SQL script and the given pseudo SQL script, the JSON java parsing code is obtained. Based on the JSON java parsing code and the topic engine data of the data source in the risk control engine system read in step 1, the Flink data set containing the data of the topic engine field is obtained and registered as a Flink table, and corresponds to the table name in the pseudo SQL script, and the Flink data set is returned;
[0115] The specific steps are:
[0116] Step 2.1 Define the mapping between source fields and logical fields:
[0117] Obtain the attributes in the risk control JSON message through the risk control engine system configuration, define the logical field name according to the business logic, obtain the corresponding channels and paths at all levels, correspond the attributes in the JSON message to the logical field name, and obtain the hierarchical structure, JSON object name and JSON object array name according to the complete path. Among them, the JSON attribute can correspond to the json object of the A level, the json object of the A level can correspond to the json object of the B level, and so on. According to the complete path, the multi-level structure can be obtained. The attributes in the risk control JSON message include the data source type and shcema information. The shcema information is the metadata information of the source data, such as the field name, field type and length of the table in the relational data;
[0118] Step 2.2 configures the mapping relationship into the database, that is, configures the mapping relationship between each level, JSON object name and JSON object array name obtained in step 2.1 into a configuration table in the database for storing the JSON data structure;
[0119] Step 2.3 Develop an automated script to encapsulate the parsing logic:
[0120] Develop the calculation conversion and aggregation logic to form an SQL script and encapsulate it into a function for storing SQL scripts. Based on the given pseudo-SQL script (data given by business personnel for analyzing business metrics, such as students, teachers, etc.), use the encapsulated SQL script to parse the required logical fields in the configuration table to obtain the Java parsing code of JSON, which is output as the result returned by the function. Among them, the pseudo-SQL script is the given field, and the SQL script is to query and obtain the corresponding logical fields in the configuration table from the complete path and structure of the configuration table to the JSON message to obtain the Java parsing code (for example, to query the user ID, the query SQL statement: select user_id from user; then we can obtain the logical field we need to query, user_id, through the from keyword in the SQL script by Java code, and then pass user_id as a parameter into the custom function to obtain the Java parsing code of the logical field (the configuration table has been configured in step 2.2 above)). Among them, the encapsulated SQL script and the Java parsing code are stored in the Flink operator. The Flink transformation operator includes flatMap and Map;
[0121] Step 2.4 Implement the Flink custom risk control engine metrics according to the pseudo-SQL script:
[0122] Based on step 1, read the topic engine data of the data source in the risk control engine system, and use the JSON parsing tool fastjson in the Flink transformation operator in combination with the JSON parsing code to process the topic engine data, and finally return it to the Flink dataset. Among them, the JSON parsing tool fastjson in the Flink transformation operator is a method dynamically generated by javaassist or asm;
[0123] Specifically: There are three types in Json, namely Json object, array, and property, and the data to be obtained is the value corresponding to the property (equivalent to obtaining the value through the key). There is one or more layers of Json objects and arrays above the property. The configuration table configures the complete path of the property, that is, to obtain which Json objects and which arrays exist in each layer of this property. Then, according to this rule, the Java parsing code is based on the existing Json parsing tool fastjson (that is, calling database functions, including dynamically generating methods using javaassist or asm) (if it is other message, other parsing tools are used): If it is a Json array, the Json array is obtained through the method getJSONArray(json array name), and loop traversal code is generated for the Json array object; if it is a Json object, the Json object is obtained through the method getJSONObject(json object name), and then continue to process the next layer of Json object or Json array until the last one is the property. Among them, the property needs to configure the data type to generate code logic. For example, if the type is a string and the type in the configuration table is String, then get concatenated with String(property name) is generated, and the Java logic getString(property name) is obtained to obtain the property value. Among them, the Json object and Json array are identified as hierarchical fields, and the Json object or Json array is distinguished by the first two digits of the field value (the Json object starts with {}, and the Json array starts with [], which is convenient for parsing). Finally, through path-level grouping, the code of the same group is merged. Among them, the process of the configuration table and the encapsulated sql script parsing and returning the Java parsing code is as Figure 2 and Figure 3 shown.
[0124] Step 2.5 Register the Flink data table and calculate the risk control engine metrics:
[0125] Register the Flink data set obtained in Step 2.4 as a Flink table through the Flink built-in method fromDataStream(), and correspond it to the table name in the pseudo sql script. Then call the sqlQuery() method built in Flink Sql to return the Flink data set.
[0126] Step 3. Data landing:
[0127] Based on the Flink data set obtained in Step 2, pass in the configuration parameter data for data landing.
[0128] The specific steps are as follows:
[0129] Step 3.1 Define the development data source landing distribution class and method:
[0130] Define the data landing encapsulation class, including batch processing and stream processing. Due to the batch processing and stream processing modes, the landing target system is essentially materialized and persisted. Since there is no essential difference between the two modes, define the data source landing distribution class SinkAssigner for the data landing target system, and define the static method sinkAssign() of the data source landing distribution class SinkAssigner to pass in the configuration parameters based on the Flink data set obtained in step 2.5. Among them, SinkAssigner is the parameter passed in during landing, and sink represents the data landing target system;
[0131] Step 3.2 Define the development data landing template class and method:
[0132] The name of the data landing template class is named after the landing target system + Sink. Define the data landing template class method handleSink() of the corresponding data landing template class. When the static method sinkAssign() of the data source landing distribution class SinkAssigner passes in parameters and determines that the data source is the landing target, at this time, the data source type parameter or the configuration file parameter is parsed by the data source template class method init() calling the data source startup method onStart(). After parsing the configuration, call the data landing template class method handleSink() to process the data landing. Finally, the sinkPost() method processes the data landing exception and the subsequent operations after the data is stored in the database;
[0133] The specific steps for the sinkPost() method to process the data landing exception and the subsequent operations after the data is stored in the database are as follows:
[0134] Exception handling method for the data landing target: During the process of obtaining and reading the data source, use trycatch to capture exceptions, including connection timeouts, configuration exceptions, or data source exceptions;
[0135] Post-processing method:
[0136] After the data source connection is successful and the data reading is completed, call the database driver to create a prepared statement PrepareStatement and the close() method of the connection interface name Connection to close the prepared statement and the connection to recycle resources.
[0137] Step 3.3 Define the development data landing class and method:
[0138] Develop specific data landing classes and methods according to the data source floor class and method. Specifically: land data to data source A, and pass in the parameter data source based on the Flink dataset, that is, configure and pass in the url, username, and password of the data source. At this time, it is distributed to the data landing template class and method to call the method init() of the data source template class to call the data source startup method onStart(). Through the data source startup method onStart(), call the execution method ResultSet to configure the environment, establish a database connection object and a prepared statement, and then define and call the sql statement (the sql statement is the specific writing logic, that is, query the required table through the filtering condition, write which fields, that is, to the required table, such as insert into table name 1 Select field 1, field 2 from table name 2 where filtering condition) for writing to the database in the method handleSink() of the data landing template class for data source A. Then execute the sql statement to land the data to the database.
[0139] Embodiment
[0140] 1. Business personnel analyze the data of the risk control engine system, calculate indicators, provide and encapsulate sql scripts.
[0141] 2. Developers prepare the relevant configuration parameters of kakfa in the risk control engine system and encapsulate them into the data source distribution class.
[0142] 3. The data source distribution class calls the data source startup method init() of the data source class to call the execution method ResultSet, parses the configuration parameters, reads the corresponding topic data in the data source kafka of the risk control engine system, and returns the Flink dataset (DataStream, DataSet).
[0143] 4. Based on the given pseudo sql script and the encapsulated sql script, find the logical fields in the configuration table and obtain the java parsing code from the JSON message.
[0144] 5. Read the topic engine data of the data source in the risk control engine system, and process the topic engine data based on the json parsing tool fastjson in the Flink transformation operator in cooperation with the json parsing code, and finally return it to the Flink dataset. Among them, the json parsing tool fastjson in the Flink transformation operator is a method dynamically generated by javaassist or asm.
[0145] 6. The distribution method of the data landing configuration parameters of the developer is passed into the data landing distribution class. Based on the Flink data set, the parameter data source is passed in, and the method init() of the data source template class of mysql is called (that is, the database for data landing is different from the database for data reading) to call the execution method ResultSet to configure the environment, establish a database connection object and a prepared statement. Then, the sql statement for writing to the database of the mysql data source in the method handleSink() of the data landing template class is defined, and then the sql statement is executed to land the data in the database.
[0146] The above are only representative embodiments among the numerous specific application scopes of the present invention, and do not constitute any limitation to the protection scope of the present invention. Any technical solutions formed by transformation or equivalent replacement fall within the scope of the protection of the rights of the present invention.
Claims
1. An automatic configuration processing method for risk control index calculation and development based on FlinkSQL, characterized in that, It includes the following steps: Step 1. Read the data flow: According to the data source, the data source type parameter is passed in. The configuration parameters are read from the configuration parameter data of the data source type parameter or the external configuration file. Based on the consumption interface of the data source that Flink already inherits, the data source connection logic is obtained, and the consumer instance object is obtained according to the configuration parameters to read the data, and the Flink dataset is returned. Among them, the data source is Kafka or a relational database, and the relational database includes mysql and oracle. The Flink dataset is the intermediate data result in the Flink application, including the DataStream dataset or the DataSet dataset. The initial dataset is generated after reading the data source; Step 2. Data processing and transformation: Based on the risk control engine system, the attributes in the risk control JSON message are obtained, and the attributes are mapped to the logical field names and stored in the configuration table. Based on the configuration table, the sql script and the given pseudo sql script, the java parsing code of JSON is parsed. And based on the java parsing code of JSON and the Flink dataset containing the topic engine field data obtained by reading the topic engine data of the data source in the risk control engine system in Step 1, it is registered as a Flink table and corresponding to the table name in the pseudo sql script, and the Flink dataset is returned; Step 3. Data landing: Based on the Flink dataset obtained in Step 2, the configuration parameter data is passed in for data landing.
2. The automatic configuration processing method developed based on the calculation of risk control indicators using FlinkSQL according to claim 1, wherein, The specific steps of Step 1 are as follows: Step 1.1 Define and develop a data source distribution class and method for the data source reading interface: Define a data source distribution class to distinguish batch processing and stream processing. The data source distribution class includes batch processing BatchSourceData and stream processing StreamSourceDatas; The data source distribution method getData() is used to call the data source template class method init() based on the incoming data source type parameter of batch processing BatchSourceData or stream processing StreamSourceDatas, and parse the configuration file. The specific process is: according to the incoming data source type parameter of batch processing BatchSourceData and stream processing StreamSourceDatas, based on the data source type parameter, call the built-in interface of Flink to connect to the data source to call the corresponding data source template class method init(). The built-in interface of Flink to connect to the data source means that Flink already inherits the consumption interface of the data source. The consumption interface refers to the connector for integrating the data source with Flink, and the consumer instance object is obtained through the configuration parameters; Step 1.2 Define and develop a data source startup method: Since Flink comes with a consumption interface that inherits from the data source, the data source startup method onStart() is defined to obtain the data source connection logic without calling the consumption interface of the data source, that is, there is no need to rewrite the operation logic of reading data. Specifically: The data source has defined SQL to obtain the tables and fields to be read in the logic of establishing a connection and preparing for reading the database driver parameters in the early stage. In the data source startup method onStart(), the defined data reading logic is passed into the execution method ResultSet, and a Flink data set is returned; Step 1.3 Define and develop a data source template class and methods for data source configuration based on Step 1.1 - Step 1.2: Define a data source template class to distinguish between batch processing and stream processing. The data source template class includes the batch data stream BatchMysqlSourceData and the stream processing data source StreamKafkaSourceData; The method init() of the data source template class is used to receive the data source type parameter of the batch data stream BatchMysqlSourceData or the stream processing data source StreamKafkaSourceData passed in by the data source distribution method getdata() to call the method init() of the data source template class. The method init() of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameters of the data source type parameter in the data source type parameter or the external configuration file to obtain the parsed configuration parameters, and call the built-in interface in Flink to connect to the data source to return a batch processing or stream processing Flink data set; Step 1.4 Define and develop specific data source classes and methods: Based on the specific data source class in Step 1.1 - Step 1.3, that is, based on the Flink application calling data source A, through the data source type parameter of data source A passed in by the data source distribution method getData(), call the method init() of the data source template class based on the data source type parameter of data source A. The method init() of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameters of the data source type parameter of data source A in the data source type parameter or the external configuration file to obtain the parsed configuration parameters and pass the configuration parameters into the method init() of the data source template class. Then, pass it into the built-in interface in Flink for reading data source A to create a consumer instance object of data source A, and then call the core method addSource of Flink to pass the consumer instance object as a parameter to read data. After the data is read, the result set is returned. Among them, the result is a data set of database queries. Data source A is a Kafka, mysql, or oracle data source.
3. The automatic configuration processing method developed based on the risk control index calculation of FlinkSQL according to claim 2, characterized in that, The data source startup method onStart() in Step 1.2 also defines data source exception handling and post-processing methods: Data source exception handling method: When obtaining and reading data sources, use try catch to capture exceptions, including connection timeout, configuration exception, or data source exception. Post-processing method: After the data source is connected successfully and the data is read, the database driver is called to create the prepared statement PrepareStatement and the connection interface name Connection's close() method to close the prepared statement and connection to recycle resources.
4. An automatic configuration processing method developed based on the calculation of risk control indicators using FlinkSQL according to claim 3, characterized in that The specific steps of step 2 are: Step 2.1 Define the mapping between source fields and logical fields: Obtain the attributes in the risk control JSON message through the risk control engine system configuration, define the logical field name according to the business logic, obtain the corresponding channels and paths at all levels, correspond the attributes in the JSON message to the logical field name, and obtain the hierarchical structure, JSON object name and JSON object array name according to the complete path. Among them, the JSON attribute can correspond to the json object of the A level, the json object of the A level can correspond to the json object of the B level, and so on. According to the complete path, the multi-level structure can be obtained. The attributes in the risk control JSON message include the data source type and shcema information. The shcema information is the metadata information of the source data; Step 2.2 configures the mapping relationship into the database, that is, configures the mapping relationship between each level, JSON object name and JSON object array name obtained in step 2.1 into the configuration table in the database for storing the JSON data structure; Step 2.3 Develop an automated script to encapsulate the parsing logic: Develop the calculation conversion and aggregation logic to form a SQL script and encapsulate it into a function for storing SQL scripts. Based on the given pseudo SQL script, use the encapsulated SQL script to parse the required logical fields in the configuration table to obtain the JSON Java parsing code as the result output returned by the function. The encapsulated SQL script and Java parsing code are stored in the Flink operator. Step 2.4 Implement Flink custom risk control engine indicators based on the pseudo SQL script: Based on step 1, read the topic engine data of the data source in the risk control engine system, and process the topic engine data based on the json parsing tool fastjson in the Flink conversion operator and the json parsing code, and finally return it to the Flink data set. The json parsing tool fastjson in the Flink conversion operator is a javaassist or asm dynamic generation method; Step 2.5 Register the Flink data table and calculate the risk control engine indicators: The Flink dataset is registered as a Flink table through the Flink built-in method fromDataStream(), and corresponds to the table name in the pseudo SQL script. Then the Flink Sql built-in sqlQuery() method is called to return the Flink dataset.
5. The automatic configuration processing method developed based on the risk control index calculation of FlinkSQL according to claim 4, characterized in that, The specific steps of step 3 are: Step 3.1 Define and develop data source distribution classes and methods: Define the data landing encapsulation class, including batch processing and stream processing. Due to the batch processing and stream processing modes, the landing target system is essentially materialized and persisted. Since there is no essential difference between the two modes, define the data source landing distribution class SinkAssigner for the data landing target system, and define the static method sinkAssign() of the data source landing distribution class SinkAssigner to pass in the configuration parameters based on the Flink dataset obtained in step 2.
5. Among them, SinkAssigner is the parameter passed in during landing, and sink represents the data landing target system; Step 3.2 Define the development data landing template class and method: The name of the data landing template class is named after the landing target system + Sink. Define the data landing template class method handleSink() corresponding to the data landing template class. When the static method sinkAssign() of the data source landing distribution class SinkAssigner passes in parameters and determines that the data source is the landing target, at this time, the data source type parameter or the configuration file parameter is called by the data source template class method init() to call the data source startup method onStart() to call the execution method ResultSet to parse the configuration. After parsing the configuration, call the data landing template class method handleSink() to process the data landing. Finally, the sinkPost() method processes the data landing exception and the subsequent operations after the data is stored in the database; Step 3.3 Define the development data landing class and method: Develop specific data landing classes and methods according to the data source landing board classes and methods. Specifically: to land data to the A data source, pass in the parameters of the data source based on the Flink dataset, that is, pass in the url, username, and password configurations of the data source. At this time, distribute to the data landing template class and method to call the data source template class method init() to call the data source startup method onStart(). Through the data source startup method onStart(), call the execution method ResultSet to configure the environment, establish a database connection object and a prepared statement, and then define the sql statement for writing to the database in the A data source in the data landing template class method handleSink(), and then execute the sql statement to store the data in the database.
6. The automatic configuration processing method developed based on the risk control index calculation of FlinkSQL according to claim 5, characterized in that, The specific steps of the sinkPost() method in step 3.2 to process the data landing exception and the subsequent operations after the data is stored in the database are as follows: Exception handling method for the data landing target: During the process of obtaining and reading the data source, use try catch to capture exceptions, including connection timeouts, configuration exceptions, or data source exceptions; Post-processing method: After the data source connection is successful and the data reading is completed, call the database driver to create a prepared statement PrepareStatement and the close() method of the connection interface name Connection to close the prepared statement and the connection to recycle resources.
7. An automatic configuration processing system developed for risk control index calculation based on FlinkSQL, characterized in that, Include: Data reading process module: According to the data source type parameter passed in by the data source, read the configuration parameters from the configuration parameter data of the data source type parameter or the external configuration file, and obtain the data source connection logic based on the consumption interface of the data source that has been inherited by Flink itself. Then, obtain the consumer instance object according to the configuration parameters to read the data, and return the Flink dataset. Among them, the data source is Kafka or a relational database, and the relational database includes mysql and oracle. The Flink dataset is the intermediate data result in the Flink application, including the DataStream dataset or the DataSet dataset. The initial dataset is generated after reading the data source; Data processing module: Based on the risk control engine system, obtain the attributes in the risk control JSON message, map the attributes to the logical field names and store them in the configuration table. Based on the configuration table, sql script and the given pseudo-sql script, parse the java parsing code of JSON, and register the Flink dataset containing the data of the topic engine fields obtained based on the java parsing code of JSON and the topic engine data of the data source read from the risk control engine system as a Flink table, and correspond to the table name in the pseudo-sql script, and return the Flink dataset; Data landing module: Based on the obtained Flink dataset, pass in the configuration parameter data for data landing.
8. An automatic configuration processing system developed based on the calculation of risk control indicators by FlinkSQL according to claim 7, characterized in that The specific implementation steps of the data reading process module are: Step 1.1 Define and develop a data source distribution class and method for the data source reading interface: Define a data source distribution class to distinguish batch processing and stream processing. The data source distribution class includes batch processing BatchSourceData and stream processing StreamSourceDatas; The data source distribution method getData() is used to call the init() method of the data source template class based on the incoming data source type parameter of batch processing BatchSourceData or stream processing StreamSourceDatas, and parse the configuration file. The specific process is: according to the incoming data source type parameter of batch processing BatchSourceData and stream processing StreamSourceDatas, call the corresponding data source template class method init() based on the data source type parameter through the built-in interface for Flink to connect to the data source. The built-in interface for Flink to connect to the data source refers to the consumption interface of the data source inherited by Flink itself. The consumption interface refers to the connector for the integration of the data source and Flink, and obtain the consumer instance object through the configuration parameters; Step 1.2 Define and develop a data source startup method: Since Flink comes with a consumption interface that inherits from the data source, a data source startup method onStart() is defined to obtain the data source connection logic without calling the consumption interface of the data source, that is, there is no need to rewrite the operation logic for reading data. Specifically: The data source has defined SQL to obtain the tables and fields to be read in the logic of establishing a connection and preparing for reading the database driver parameters in the early stage. In the data source startup method onStart(), the defined data reading logic is passed into the execution method ResultSet, and a Flink dataset is returned; Step 1.3 Define and develop a data source template class and methods for data source configuration based on Step 1.1 - Step 1.2: Define a data source template class to distinguish between batch processing and stream processing. The data source template class includes the batch data stream BatchMysqlSourceData and the stream processing data source StreamKafkaSourceData; The method init() of the data source template class is used to receive the data source type parameter of the batch data stream BatchMysqlSourceData or the stream processing data source StreamKafkaSourceData passed in by the data source distribution method getdata() to call the method init() of the data source template class. The method init() of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameters of the data source type parameter in the data source type parameter or the external configuration file to obtain the parsed configuration parameters, and call the built-in interface for Flink to connect to the data source to return a batch processing or stream processing Flink dataset; Step 1.4 Define and develop specific data source classes and methods: Based on the specific data source class in Step 1.1 - Step 1.3, that is, based on the Flink application calling data source A, through the data source type parameter of data source A passed in by the data source distribution method getData(), call the method init() of the data source template class based on the data source type parameter of data source A. The method init() of the data source template class is responsible for calling the data source startup method onStart() to call the execution method ResultSet to parse the configuration parameters of the data source type parameter of data source A in the data source type parameter or the external configuration file to obtain the parsed configuration parameters and pass the configuration parameters into the method init() of the data source template class. Then, pass it into the built-in interface in Flink for reading data source A to create a consumer instance object of data source A, and then call the core method addSource of Flink to pass the consumer instance object as a parameter to read data. After the data is read, a result set is returned. Among them, the result is a data set of a database query. Data source A is a Kafka, mysql, or oracle data source.
9. An automatic configuration processing system developed based on the calculation of risk control indicators using FlinkSQL according to claim 8, characterized in that, The specific implementation steps of the data processing module are as follows: Step 2.1 Define the mapping between source fields and logical fields: Obtain the attributes in the risk control JSON message through the risk control engine system configuration, define the logical field name according to the business logic, obtain the corresponding channels and paths at all levels, correspond the attributes in the JSON message to the logical field name, and obtain the hierarchical structure, JSON object name and JSON object array name according to the complete path. Among them, the JSON attribute can correspond to the json object of the A level, the json object of the A level can correspond to the json object of the B level, and so on. According to the complete path, the multi-level structure can be obtained. The attributes in the risk control JSON message include the data source type and shcema information. The shcema information is the metadata information of the source data; Step 2.2 configures the mapping relationship into the database, that is, configures the mapping relationship between each level, JSON object name and JSON object array name obtained in step 2.1 into the configuration table in the database for storing the JSON data structure; Step 2.3 Develop an automated script to encapsulate the parsing logic: Develop the calculation conversion and aggregation logic to form a SQL script and encapsulate it into a function for storing SQL scripts. Based on the given pseudo SQL script, use the encapsulated SQL script to parse the required logical fields in the configuration table to obtain the JSON Java parsing code as the result output returned by the function. The encapsulated SQL script and Java parsing code are stored in the Flink operator. Step 2.4 Implement Flink custom risk control engine indicators based on the pseudo SQL script: Based on reading the topic engine data of the data source in the risk control engine system, the topic engine data is processed based on the json parsing tool fastjson in the Flink conversion operator and the json parsing code, and finally returned to the Flink data set. Among them, the json parsing tool fastjson in the Flink conversion operator is a javaassist or asm dynamic generation method; Step 2.5 Register the Flink data table and calculate the risk control engine indicators: The Flink dataset is registered as a Flink table through the Flink built-in method fromDataStream(), and corresponds to the table name in the pseudo SQL script. Then the Flink Sql built-in sqlQuery() method is called to return the Flink dataset.
10. An automatic configuration processing system developed based on the calculation of risk control indicators by FlinkSQL according to claim 9, characterized in that, The specific implementation steps of the data landing module are: Step 3.1 Define and develop data source distribution classes and methods: Define the data landing encapsulation class, including batch processing and stream processing. Due to the batch processing and stream processing modes, the landing target system is essentially materialized and persisted. Since there is no essential difference between the two modes, define the data source landing distribution class SinkAssigner for the data landing target system, and define the static method sinkAssign() of the data source landing distribution class SinkAssigner to pass in the configuration parameters based on the Flink dataset obtained in Step 2.
5. Among them, SinkAssigner is the parameter passed in during landing, and sink represents the data landing target system; Step 3.2 Define the development data landing template class and method: The name of the data landing template class is named after the landing target system + Sink. Define the data landing template class method handleSink() corresponding to the data landing template class. When the static method sinkAssign() of the data source landing distribution class SinkAssigner passes in parameters and determines that the data source is the landing target, at this time, the data source type parameter or the configuration file parameter is called by the data source template class method init() to call the data source startup method onStart() to call the execution method ResultSet to parse the configuration. After parsing the configuration, call the data landing template class method handleSink() to process the data landing. Finally, the sinkPost() method processes the data landing exception and the subsequent operations after the data is landed in the database; Step 3.3 Define the development data landing class and method: Develop specific data landing classes and methods according to the data source landing board classes and methods. Specifically: to land data to the A data source, pass in the parameter data source based on the Flink dataset, that is, pass in the url, username, and password configuration of the data source. At this time, distribute it to the data landing template class and method to call the data source template class method init() to call the data source startup method onStart(). Through the data source startup method onStart(), call the execution method ResultSet to configure the environment, establish a database connection object and a prepared statement, and then define the sql statement for writing to the database in the A data source in the data landing template class method handleSink(). Then execute the sql statement to land the data in the database.
Citation Information
Patent Citations
Service data processing method and device based on Flink engine
CN110704518A
Real-time data processing method, platform and device based on Flink
CN113010512A