Bank customer manager performance real-time evaluation and early warning method based on machine learning

Through machine learning and real-time data processing technology, combined with the random forest model and the Flink platform, real-time evaluation and early warning of bank account manager performance is achieved, and the lag and inaccuracy problems in traditional methods are solved, and the accuracy of evaluation and data processing efficiency are improved.

CN120146647APending Publication Date: 2025-06-13JIANGSU SUNING BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510068583.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing bank account manager performance evaluation methods have lag and inaccuracy, which cannot meet the needs of real-time and efficient processing. Traditional systems have shortcomings in real-time data, personalized analysis and intelligent decision-making support.

Method used

Using machine learning methods, data is integrated through offline big data platforms and distributed file systems, random forest models are trained for performance evaluation, and real-time data processing is performed in combination with Flink computing platform, rolling windows are defined for performance summary calculations, and dynamic warning rules are configured for real-time warning.

Benefits of technology

Real-time evaluation and early warning of account manager performance, improve the accuracy and scalability of the model, support diversified data sources, solve data silos, and provide real-time data analysis and decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146647A_ABST
    Figure CN120146647A_ABST
Patent Text Reader

Abstract

The invention discloses a machine learning-based bank customer manager performance real-time evaluation and early warning method. The method comprises the steps of collecting data, and generating a detail data table and a summary data table; integrating data in the detail data table and the summary data table, and outputting the data to a distributed file system for storage; the offline performance evaluation model is trained, and after the offline performance evaluation model is qualified, the offline performance evaluation model is serialized and stored in a specified directory; summarizing and calculating the performance condition of the day in real time; and performing grouping and aggregation processing on the integrated data and the performance condition of the current day, inputting the grouped and aggregated data into an offline performance evaluation model to obtain a performance scoring result of each customer manager, and performing real-time early warning according to the performance scoring result and early warning rule configuration. The system can collect and analyze interaction data and market dynamics of customer managers in real time, can provide instant insight and decision support for a management layer, and can provide accurate performance analysis and improvement suggestions for the management layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of real-time performance evaluation and early warning methods for bank customer managers based on machine learning, and particularly relates to a real-time performance evaluation and early warning method for bank customer managers based on machine learning. Background Art

[0002] In the past few decades, the fintech (financial technology) field has undergone tremendous changes. Banks and financial institutions have gradually shifted from traditional paper transactions and manual processing to digital and automated operations. With the popularization of Internet and mobile technologies, customers' expectations for bank services have also been continuously increasing. They hope to access financial services anytime and anywhere and enjoy fast and secure transaction experiences. However, this transformation has also brought new challenges, especially in data processing and management.

[0003] The banking industry is facing fierce market competition and increasingly complex customer needs. As the bridge between banks and customers, the performance of customer managers directly affects customer loyalty and the profitability of banks. Traditional performance evaluation methods often rely on historical data and subjective evaluations, suffering from problems of lag and inaccuracy. With the rapid development of big data and artificial intelligence technologies, banks have started to seek more intelligent solutions to achieve real-time evaluation and early warning of customer managers. There are also the following solutions in the industry:

[0004] Solution 1) Offline batch processing system

[0005] With the progress of data processing technologies, banks have gradually introduced batch processing systems to improve the efficiency and accuracy of evaluations. Batch processing systems can process large-scale data without affecting daily operations through automated data collection, integration, and analysis. Such systems support the calculation of complex performance indicators, the generation of regular performance reports, and in-depth analysis of historical data. In addition, batch processing systems also have anomaly detection and early warning functions to help banks promptly identify and solve potential problems. By optimizing resource allocation and compliance checks, batch processing systems provide an efficient and reliable performance evaluation solution for banks. However, despite the wide application of these technical means, there is still room for improvement in data real-time, personalized analysis, and intelligent decision support.

[0006] Solution 2) Data warehouse architecture

[0007] A data warehouse is a system specifically designed to store and manage large amounts of historical data, supporting complex query and analysis tasks. By integrating data from multiple sources, a data warehouse provides a unified view for enterprises, helping decision-makers conduct strategic analysis. The architecture of a data warehouse usually includes data extraction, transformation, and loading (ETL) processes, which extract data from various operating systems, clean and transform it, and then load it into the warehouse. The design of a data warehouse is usually based on a star or snowflake architecture, supporting multidimensional data analysis (OLAP). This architecture allows users to analyze data through different dimensions (such as time, geographical location, product category, etc.) to discover hidden patterns and trends. A data warehouse also supports complex query and report generation, helping enterprises gain an advantage in the competition.

[0008] Solution 3) Traditional API interface

[0009] An API (Application Programming Interface) interface is a standardized method for connecting different software systems and applications, allowing data exchange and function calls between them. Traditional API interfaces are usually based on the HTTP protocol and use standards such as REST or SOAP for communication. API interfaces play a key role in modern software development, supporting system integration and interoperability. The wide application of API interfaces enables enterprises to quickly integrate third-party services, expanding their functions and market coverage. However, traditional API interfaces also face some challenges. As the complexity of systems increases, the management and version control of APIs become more difficult. In addition, the performance and security issues of API interfaces may affect the overall reliability of the system. To address these issues, many enterprises have started to adopt API gateways and microservices architectures to improve the management efficiency and security of APIs.

[0010] The above solutions 1, 2, and 3 cannot meet the real-time requirements of banks. Due to its inherent latency, a batch processing system cannot achieve real-time data processing. Although the data warehouse architecture is suitable for in-depth analysis, it performs poorly in terms of real-time and high-concurrency processing. Moreover, problems such as a large data redundancy and low disk application performance will occur. Although a distributed database has scalability, the complex management and debugging processes affect the real-time response ability. The performance bottleneck of traditional API interfaces is obvious in high-concurrency situations, making it difficult to support the real-time data interaction requirements of the banking industry. Therefore, these solutions cannot fully meet the strict requirements of banks for real-time and efficient processing. Summary of the Invention

[0011] The object of the present invention is to provide a method for real-time evaluation and early warning of the performance of bank customer managers based on machine learning in view of the deficiencies of the existing technology.

[0012] To achieve the above object, the present invention provides a method for real-time evaluation and early warning of the performance of bank account managers based on machine learning, including:

[0013] Collect data from the data source, perform preliminary cleaning on the collected data, then convert the preliminarily cleaned data into a unified format using business rules, store it in the detail data table, and further process it based on the detail data table to obtain a summary data table with multiple cycle dimensions;

[0014] Call hive through the offline big data platform to integrate the data in the detail data table and the summary data table, output the data integrated by the offline big data platform to the distributed file system for storage, and use the data date as the partition;

[0015] Train the offline performance evaluation model based on the data saved in the distributed file system. After the offline performance evaluation model is trained successfully, save it to the specified directory through serialization;

[0016] For the real-time data of the current day, after the business system processes and converts each piece of data, encapsulate it into a kafka message, and use the flink computing platform as the consumer of kafka to process the performance data of the account manager on the current day. Define a rolling window in units of days to calculate the performance situation on the current day in real time;

[0017] Load the saved offline performance evaluation model, store the offline performance evaluation model object in memory through deserialization, group and aggregate the integrated data and the performance situation on the current day of different account managers according to the unique key key, and input the grouped and aggregated data into the offline performance evaluation model to obtain the performance score results of each account manager, and perform real-time early warning according to the performance score results and the early warning rule configuration.

[0018] Further, the preliminary cleaning includes duplicate removal, format conversion, and basic verification.

[0019] Further, the field names of the detail data table include data date, account manager work number, account manager name, dimension name, dimension value, and occurrence time.

[0020] Further, the dimension names include the number of communications with customers, the number of applications, the credit amount, the loan amount, the financial product sales amount, the bad debt recovery amount, and customer evaluation and feedback.

[0021] Further, the data integrated by the offline big data platform is saved in the distributed file system in CSV format. When training the offline performance evaluation model using the data saved in the distributed file system, the toolkits provided by Weka are called to create a process to load the CSV format data files stored in the distributed file system into the memory, and then they are parsed into ARFF format data sets through an ArffSaver object.

[0022] Further, the offline performance evaluation model is a random forest model, and its training method is specifically as follows:

[0023] (1) Load the ARFF format data set, and set the last field in the ARFF format data set as the label;

[0024] (2) Instantiate a random forest classifier;

[0025] (3) Set the maximum depth of the tree to 0, and at the same time set the number of features considered during each tree split to 0, and then set the random seed;

[0026] (4) Use ten-fold cross-validation to evaluate the performance of the model.

[0027] The specific method of real-time warning is as follows:

[0028] Further, the specific steps of real-time warning are as follows:

[0029] 1) Read all rule configurations in the database;

[0030] 2) Receive the performance scoring results of account managers sent, and compare the performance scoring results of each account manager with the current corresponding rules. When the rules are met, it indicates that a warning is required, and read the action field in the rule table to indicate what kind of warning method to use, and write the data into the warning table;

[0031] 3) Start a scheduled task to regularly read the data in the warning table, and select the warning method according to the value of the action field;

[0032] 4) Save the warning data to the detailed data table, and display the warning data as a visual report through the BI report platform;

[0033] 5) Configure the service interface for calling by the http interface or other client systems.

[0034] Beneficial effects: 1) The present invention solves the problem of bloated design of the offline data model, and can easily expand the performance sample data fields of account managers in the existing model without changing the underlying model architecture design; it solves data sparsity and improves space utilization efficiency;

[0035] 2) The present invention solves the problem of the single source of offline performance data sources, integrates the multi-dimensional data model of the business system, provides richer learning samples for the machine learning model, and makes the model accuracy higher; at the same time, it supports the iteration of daily training sample data, and the export and application of the machine learning model become more general;

[0036] 3) The present invention solves the problem that the real-time data source is strongly dependent on the business system, decouples from multiple business systems, and meets the diversity of the real-time data source of the customer manager in the way of data synchronization, breaking the data island;

[0037] 4) The present invention solves the problem that the performance evaluation and early warning of the customer manager are relatively lagging. By combining the flink real-time computing platform with the machine learning model, it uses the streaming data processing method instead of batch processing to achieve the real-time performance evaluation and early warning. By dynamically defining the early warning threshold and configuring multiple early warning mechanisms, the visualization display platform helps the bank management and customer managers to conduct real-time data analysis and decision support. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic structural diagram of the method for real-time evaluation and early warning of the performance of bank customer managers based on machine learning according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The present invention will be further clarified below with reference to the drawings and specific embodiments. These embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0040] As Figure 1 shown, the embodiment of the present invention provides a method for real-time evaluation and early warning of the performance of bank customer managers based on machine learning, including:

[0041] Collect data from data sources, perform preliminary cleaning on the collected data, then apply business rules to transform the preliminarily cleaned data into a unified format, store it in a detailed data table, and further process it based on the detailed data table to obtain a summary data table with multiple cycle dimensions. Specifically, the specific implementation steps of offline data collection and processing include the following stages: First, in the data collection stage (ODS layer), it is necessary to identify and connect all relevant data sources, such as the core banking system, CRM system, and marketing platform, etc. During the extraction process, perform preliminary cleaning on the data, including duplicate removal, format conversion, basic verification (such as date format, numerical range), etc., to ensure the integrity and consistency of the data. Next, in the data processing and storage stage (DWD layer), further transform and standardize the data in the ODS layer, and apply business rules to convert the data into a unified format. By integrating the data from different systems, comprehensively and multi-dimensionally perform fine-grained processing on the detailed information of customer managers. Since the data dimensions of customer managers may be very numerous, the detailed data table is specifically shown in Table 1:

[0042] Table 1: Detailed Data Table

[0043] Field Name Field Type Field Comment etl_date varchar Data Date manager_no varchar Customer Manager's Employee ID manager_name varchar Customer Manager's Name key_name varchar Dimension Name value_name varchar Dimension Value occur_time varchar Occurrence Time

[0044] These dimension names can be: the number of communications with customers, the number of applications, the credit amount, the loan amount, the wealth management sales amount, the bad debt recovery amount, customer evaluations and feedback, etc., and each business attribute has the time of its value at that time point. For example, if customer manager A completed a credit situation at a certain time point, then the occurrence time at this time is the time point value of the credit occurrence time. At the same time, the offline data has already been tagged with a score for this piece of data, which is for the convenience of subsequent model training. Considering that the dimension fields of the business attributes of customer managers will be very numerous, and it is impossible for each customer manager to have all the field attributes. In most cases, there will only be partial field attributes, which will cause sparsity of the table and is not conducive to subsequent data analysis. Therefore, the advantage of this design is that it can make full use of the data space, with a clear structure and complete time dimension data, which is convenient for subsequent analysis using models. Furthermore, based on the detailed table, further process to obtain a summary data table model with multiple cycle dimensions. Among them, there are many dimension modules in the performance dimension. For the convenience of expansion, the present invention defines multiple extensible fields, and normalization processing needs to be added during dimension summarization. The summary data table is specifically shown in Table 2:

[0045] Table 2: Summary Data Table

[0046] Field Name Field Type Field Comment etl_date varchar Data Date performancePeriod varchar Dimension Summary 1: Day 2: Week 3: Month 4: Quarter 5: Year manager_no varchar <![CDATA[Customer manager Employee ID of Science and Engineering > manager_name varchar Customer Manager's Name performance1 varchar Summary Value of Performance Dimension 1 performance2 varchar Summary Value of Performance Dimension 2 performance_ext varchar Performance Dimension Summary Extended Field (Multiple) score varchar Performance Evaluation (Normalized Processing)

[0047] In the process of performance evaluation of bank account managers, the above-obtained basic data is the starting point for further analysis. This data usually includes the historical performance records, sales data, customer feedback, and market activity participation records of account managers. Moreover, the detailed data table and the summary data table have corresponding normalized scoring data, which can help us train the offline performance evaluation model well. The data in the detailed data table and the summary data table can be integrated by calling hive through the offline big data platform, and the data integrated by the offline big data platform is output to the distributed file system (HDFS) for storage. The storage format is preferably CSV, and the data date is used as the partition, so that there is one data file per day, such as the directory: / hdfs / manager / performace / 20241130.

[0048] Based on the data saved in the distributed file system, the offline performance evaluation model is trained. After the offline performance evaluation model is trained successfully, it is saved in serialized form to the specified directory. The offline performance evaluation model in the embodiment of the present invention is preferably a random forest model, because the random forest improves the accuracy and stability of the model by integrating the prediction results of multiple decision trees. Each tree learns different feature combinations of the data during training, thereby reducing the risk of overfitting and being able to effectively process data sets with a large number of features. Its randomness in feature selection helps it find the most informative feature combinations and perform well even when the number of features is more than the number of samples. The specific steps are as follows:

[0049] 1) After the offline data is ready, the model will be notified to read the data in the hdfs daily folder for training. Since the file is of the csv type and the weka framework requires the input format to be ARFF, first, the data format conversion is performed. The tool package provided by weka is called to load the CSV file into memory, and then it is parsed into the ARFF format through the ArffSaver object for subsequent model loading and training.

[0050] 2) Model training: Use the random forest algorithm in the weka library to train the model. After training, it is saved in serialized form to the specified directory, such as on disk or in SDOSS, to lay the foundation for subsequent real-time evaluation. The method for defining model training and saving is as follows:

[0051]

[0052]

[0053] The specific step description is as follows:

[0054] (1) Load the dataset in ARFF format and set the last field in the ARFF-format dataset as the label, which represents the variable to be predicted, i.e., the performance score of the account manager.

[0055] (2) Instantiate a random forest classifier. Generally, the higher the number of trees, the higher the accuracy of the model, but at the same time, the computational complexity will also be higher. Considering comprehensively, we set the number of trees to 100, and at this time, the accuracy and timeliness are better.

[0056] (3) Set the maximum depth of the tree to 0, and at the same time set the number of features considered when splitting the tree each time to 0. Then set the random seed to ensure the repeatability of the results.

[0057] (4) Adopt ten-fold cross-validation to evaluate the performance of the model. After passing the qualification, output the model to the specified directory for subsequent real-time calls.

[0058] For the real-time data of the day, after the business system processes and converts each piece of data, it packages it into a kafka message, and uses the flink computing platform as the consumer of kafka to process the performance data of the account manager on the same day. Define a rolling window in units of days to summarize and calculate the performance situation of the day in real time.

[0059] General message entity class design:

[0060]

[0061] Among them, this entity class design is very general and can extend most performance-related information. The message data source of flink can be multiple business systems. When an account manager completes a piece of performance data, kafka messages are sent synchronously, or through a data real-time synchronization tool, in the master-slave architecture mode of MySQL, subscribing to the binlog in the slave mode, and synchronizing a system to parse this binlog and send kafka messages. Considering the characteristics of different systems, generally we will use them in combination. The specific code is as follows:

[0062]

[0063]

[0064] The random forest model of Weka requires the input classes to be converted into the Instance object type. It also needs to further transform the Kafka messages, extract the features in the Kafka messages, and obtain each input feature value through the provided jsonObject.getkey() method. By traversing and parsing each value in sequence, a complete Instance object can be obtained. Once this object is obtained, it can be put into our model to predict the evaluation results and generate label data for display.

[0065] The tumbling window divides the continuous data stream into chunks of a fixed length according to time. We use a daily tumbling window, that is, the length of each window is one day. The purpose of the tumbling window is to timely observe the performance evaluation of the account manager from the start to the current point in time every day. For example, if account manager A completes an intake at 8 am and then completes a loan disbursement at 9 am, then what we evaluate is the performance of the account manager in an entire interval, rather than the data of a single transaction. After the calculation is completed, the real-time data can be imported into the database and can be conveniently displayed through a dashboard or visualization tool.

[0066] Tumbling window design:

[0067]

[0068] Define a tumbling window with a date of one day through TumblingProcessingTimeWindows.of(Time.days(1)). Within each tumbling window, Flink will collect all the data that enters this time period and then batch process this data. This processing method ensures that each piece of data passes through the model evaluation, which is suitable for application scenarios that require real-time analysis of each data point. And the non-overlapping feature of the tumbling window ensures that each piece of data is only processed once, avoiding the problem of duplicate calculation. We will group by account manager to form data for each account manager in the dimension of time sequence, and then the artificial intelligence model can analyze today's performance based on historical performance evaluations, helping account managers and management better identify the current situation.

[0069] Load the saved offline performance evaluation model, store the offline performance evaluation model object in memory through deserialization, group and aggregate the integrated data of different account managers and the daily performance according to the unique key "key", and input the grouped and aggregated data into the offline performance evaluation model to obtain the performance score results of each account manager, and perform real-time warnings according to the performance score results and the configured warning rules. In the offline stage, we have integrated the performance data from different account managers and trained a prediction model through the random forest algorithm. It can be understood that we have obtained the performance details and summary data of each account manager in multiple dimensions, and evaluated the performance data by adding labels to the model. For example, according to the specific performance of the account manager, it is evaluated according to the grades from A to E. In real-time evaluation, these offline data are our sample data. Based on this data, combined with the data within a one-day window that occurs in real-time on the same day, we dynamically and timely give performance evaluation information data so that they can quickly understand their own performance. The code for loading the saved offline performance evaluation model is as follows:

[0070]

[0071] After the offline performance evaluation model is trained, the structure and parameters of the model (such as the splitting conditions of each tree, the output of leaf nodes, etc.) will be serialized and saved to a file. This "modelPath" is the path of the model. The "loadRandomForestModel" method is responsible for reading the stored model file, and then the random forest model object will be stored in memory through deserialization.

[0072] In this instance, that is, the employee number corresponding to the account manager is used as the unique key, and different account managers will be assigned to different slots for grouping and aggregation processing. After the aggregation processing is completed, it will be immediately input into the random forest model, and then a data result returned by the model will be obtained. The specific code is as follows:

[0073]

[0074]

[0075] In the callback apply method, we obtain specific context information. The TimeWindow is the time range of the current window, which is from 00:00 of the current day to the current time point value. The input is all elements within the window. Since the customer managers were grouped by key earlier, all multi-dimensional performance information of this customer manager received through Kafka messages on this day will be obtained. After summarizing the detail data table and the summary data table, it enters the random forest model to predict the performance score result of this customer manager and return it. The returned prediction information can be written into the database or sent out as a Kafka message for further analysis later.

[0076] During real-time warning, each piece of performance evaluation data received is saved. Define the performance threshold for each customer manager, which can be dynamically modified and imported daily through the operation module. Because banking business is complex and involves many modules, different performance target designs can be made according to the different business lines where each customer manager is located. The specific rule table is shown in Table 3:

[0077] Table 3: Rule Table

[0078] Field Name Field Type Field Comment biz_no varchar Applied Business Rule condition varchar Execution Condition action varchar Execution Action priority varchar Priority, the higher the value, the higher the priority isActive varchar Is Active

[0079] The event data of the customer manager can be encapsulated using the Event class. The specific code is as follows:

[0080]

[0081] The specific steps of real-time warning are as follows:

[0082] 1) Read all rule configurations in the database. There will be multiple different business rule attributes for biz_no, and the relationship between one biz_no and manager is 1:N. Save the rules in the JVM memory, such as the caching tool provided by Google, and set the expiration time. Pull once every ten minutes to ensure the latestness of the rules and at the same time ensure business performance, because when the quantity is large, the query efficiency of each database query is low. When there is no hit in the cache, then perform a database query operation and then save it in the cache;

[0083] 2) Receive the performance score results of the customer managers sent. Compare the performance score result of each customer manager with the corresponding current rule. When the rule is satisfied, it indicates that a warning needs to be issued, and read the action field of the rule table to indicate what warning method to perform, and write the data into the warning table;

[0084] 3) Start the scheduled task to read the data in the early warning table regularly. At the same time, pay attention to which information has been warned and do not need to be warned repeatedly. Select the warning method according to the value of the action field. The conventional warning methods include mobile app warning, SMS warning, email warning, etc.;

[0085] 4) At the same time, the early warning data today will also be saved as detailed data in the detailed data table. Such early warning data can also be displayed as a visual report through the BI report platform for better analysis. And the real-time data of the day will also be used as sample data for offline training. Through the feedback of dynamic sample data, the offline performance evaluation model is constantly enriched and improved;

[0086] 5) Configure the service interface for http interface or other client systems to call, such as the upgrade communication system, cockpit system, etc.

[0087] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, the other parts not specifically described belong to the prior art or common general knowledge. Without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A real-time evaluation and early warning method for bank account manager performance based on machine learning, characterized in that: include: Collect data from the data source and perform preliminary cleaning on the collected data. Then, apply business rules to convert the preliminarily cleaned data into a unified format and store it in a detailed data table. The detailed data table is further processed to obtain a summary data table of multiple period dimensions. Integrate the data in the detailed data table and the summary data table by calling hive through the offline big data platform, and output the data integrated by the offline big data platform to the distributed file system for storage, and use the data date as the partition; Training the offline performance evaluation model based on the data stored in the distributed file system, and saving the offline performance evaluation model to a designated directory through serialization after the training is qualified; For the real-time data of the day, the business system processes and converts each piece of data, encapsulates it into a Kafka message, and uses the Flink computing platform as the consumer of Kafka to process the performance data of the account manager on the day, defines a rolling window with days as the unit, and summarizes and calculates the performance of the day in real time; Load the saved offline performance evaluation model, store the offline performance evaluation model object in memory by deserialization, group and aggregate the integrated data and daily performance of different account managers according to the unique key, and input the grouped and aggregated data into the offline performance evaluation model to obtain the performance score result of each account manager, and make real-time warnings based on the performance score result and warning rule configuration.

2. According to claim 1, a real-time evaluation and early warning method for bank account manager performance based on machine learning is characterized in that: The preliminary cleaning includes deduplication, format conversion and basic verification.

3. The real-time evaluation and early warning method for bank account manager performance based on machine learning according to claim 1 is characterized in that: The field names of the detailed data table include data date, account manager ID, account manager name, dimension name, dimension value and occurrence time.

4. A real-time evaluation and early warning method for bank account manager performance based on machine learning according to claim 3, characterized in that: The dimension names include the number of communications with customers, the number of applications, the credit amount, the loan amount, the wealth management sales amount, the bad debt collection amount and the customer evaluation and feedback.

5. The real-time evaluation and early warning method for bank account manager performance based on machine learning according to claim 1 is characterized in that: The data integrated by the offline big data platform is saved in the distributed file system in CSV format. When the offline performance evaluation model is trained using the data saved in the distributed file system, the CSV format data file stored in the distributed file system is loaded into the memory by calling the toolkit provided by weka, and then parsed into an ARFF format data set through the ArffSaver object.

6. A real-time evaluation and early warning method for bank account manager performance based on machine learning according to claim 5, characterized in that: The offline performance evaluation model is a random forest model, and its training method is as follows: (1) loading the dataset in the ARFF format, and setting the last field in the dataset in the ARFF format as a label; (2) Instantiate a random forest classifier; (3) Set the maximum depth of the tree to 0, set the number of features considered each time the tree splits to 0, and then set the random seed; (4) Ten-fold cross validation was used to evaluate the performance of the model.

7. The real-time evaluation and early warning method for bank account manager performance based on machine learning according to claim 1 is characterized in that: The specific steps for real-time warning are as follows: 1) Read all rule configurations in the database; 2) Receive the performance score results of the account managers sent, and compare the performance score results of each account manager with the current corresponding rules. When the rules are met, it indicates that an early warning is needed, and reads the action field of the table rule table to indicate the type of early warning method, and writes the data into the early warning table; 3) Start the scheduled task, read the data in the warning table regularly, and select the warning method according to the value of the action field; 4) Save the warning data to the detailed data table and display the warning data as a visual report through the BI report platform; 5) Configure the service interface for calling by http interface or other client systems.