Method for switching data sources when calling the springboot service interface based on the xxl-job timed task
By obtaining performance data in xxl-job timing tasks, dynamically adjusting connection pools and visualizing them, the problems of dynamicity and visualization during data source switching are solved, efficient and stable data source switching is achieved, and system performance and user experience are improved.
Patent Information
- Application Number
- CN202510548874.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The prior art lacks dynamicity and visualization in the data source switching process in xxl-job timing tasks, resulting in cumbersome operations and difficulty in adapting to changing data environments, making it difficult for users to intuitively perceive the switching state and results.
Generate JSON structured logs by obtaining performance data, dynamically adjust the connection pool size, record the connection pool status, generate real-time feedback data streams and visualize displays, automatically calibrate the configuration when the response time exceeds the threshold, set the timeout automatic rollback mechanism, and dynamically adjust the configuration refresh frequency to meet business requirements.
It realizes efficient and stable switching in multiple data source scenarios, improves system performance and reliability, and enhances user transparency and system stability.
Smart Images

Figure CN120066747B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method for switching data sources when an xxl-job timed task calls a springboot service interface. Background Art
[0002] The timed task system plays a crucial role in modern software architectures. Its core value lies in ensuring the efficient scheduling and stable execution of tasks, and it is widely used in fields such as distributed systems, data processing, and business process automation. As an important tool for enterprise-level task scheduling, xxl-job has become the industry standard with its high performance and ease of use. However, with the increase in business complexity and the need for diverse data sources, how to flexibly switch data sources during task scheduling and visually present this process has become a key issue in enhancing system adaptability and user experience.
[0003] Currently, there are still significant limitations in the implementation of data source switching in many timed task systems. Traditional methods often rely on hard coding or static configuration, lacking dynamism and flexibility, resulting in cumbersome switching operations and difficulty in adapting to changing data environments. In addition, existing solutions usually only focus on functional implementation and ignore the visual display of the switching process. It is difficult for users to intuitively perceive the execution status and results of the switching. This lack of information presentation not only increases the complexity of operations but also reduces the maintainability and transparency of the system. In this field, the core challenges mainly focus on two technical factors: the dynamic implementation of data source switching and the visualization of the process. First, dynamically switching data sources requires configuration adjustments according to the call specification document during task execution, which requires the system to have a high degree of flexibility and real-time performance, and existing mechanisms often struggle to meet this requirement. Second, the visual display of the switching process involves the organic combination of status tracking, result feedback, and interface design, but currently, there is a lack of effective technical means to integrate these aspects into an intuitive experience. Due to these two technical factors not being fully addressed, the system is prone to unique problems such as configuration chaos or user understanding difficulties when facing multi-data source scenarios. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for switching data sources when an xxl-job timed task calls a springboot service interface, so as to solve the problems existing in the prior art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for switching data sources when an xxl-job timed task calls a springboot service interface, including:
[0006] S1. Obtain performance data collected once per second, generate a JSON structured log file, and obtain the target data source identifier and connection information;
[0007] S2. Load the target data source identifier through the preset scheduling engine interface, dynamically adjust the connection pool size to 50 connections, and confirm that the new data source connection has been activated;
[0008] S3. Record the change in the number of connections during the connection pool status monitoring, obtain the timestamp of the switching operation and the INFO-level event tracking log, and generate a real-time feedback data stream containing the switching status;
[0009] S4. Map the real-time feedback data stream to the visualization component, generate a dynamic update progress bar and result prompt for the number of connections, and obtain a user-perceivable switching process display;
[0010] S5. If it is detected in the switching status feedback that the response time threshold exceeds 500 milliseconds, extract the identifier and parameters of the abnormal data source from the multi-data source scenario, and determine whether the calibrated data source meets the expectations;
[0011] S6. Obtain the connection information of the calibrated data source and reload it into the scheduling engine during task execution, generate an updated adjustment record in CSV format, and confirm that the switching process has returned to normal;
[0012] S7. Generate a state change trend chart for the switching results in the multi-data source scenario;
[0013] S8. Collect performance metric type data, analyze whether the response time of the switching operation is lower than 300 milliseconds and the result of the success rate calculation formula, and determine whether the dynamic switching mechanism meets the requirements of the business scenario;
[0014] S9. Adjust the configuration refresh frequency of the scheduling engine to update once every 10 seconds, and automatically roll back the configuration when the threshold is exceeded, to obtain a stable task execution environment suitable for the multi-data source scenario.
[0015] Preferably, the S1 includes:
[0016] Obtain the performance data collected per second from the task execution through the RESTful API, store it as an initial data set. For the initial data set, parse the collected data within the sliding window of the most recent 5 minutes to obtain a time series data set. Use the time series data set to calculate the response time of each data point to obtain a response time series. If a certain data point in the response time series exceeds the preset time threshold, mark it as abnormal to obtain an abnormal marking series. Through the abnormal marking series and the success rate formula, that is, the success rate is the total number of data points minus the number of abnormal points and then divided by the total number of data points, calculate the success rate within the sliding window to obtain a success rate series. Obtain the success rate series and the data source identifier, generate a JSON structured log file containing the response time threshold and the success rate to obtain a log output file. According to the log output file, extract the data source identifier and connection information to obtain a complete description of the target data source.
[0017] Preferably, S2 includes:
[0018] Obtain response time data by parsing the log file, perform threshold parsing on the response time using a preset threshold to obtain a parsing result. If the parsing result exceeds the preset threshold, load the data source identifier through the scheduling engine, determine the target data source, adjust the connection pool size according to the identification information of the target data source to obtain a new connection number configuration, activate the new data source through the new connection number configuration, determine whether the activation status is completed, obtain the confirmation information of the activation status, load and verify the availability of the new data source through the interface, determine that the connection pool adjustment takes effect, extract the change trend of the response time from the log file, analyze the correlation between the change trend and the connection pool size using the random forest algorithm to obtain optimized connection parameters, update the preset threshold of the scheduling engine according to the optimized connection parameters, and judge the stability of the data source connection.
[0019] Preferably, S3 includes:
[0020] Capture the change in the number of active connections in the connection pool through the status tracking module, record the time point of the change to obtain preliminary data, extract the timestamp corresponding to the switching operation from the preliminary data, match the INFO-level event logs to generate a log correlation set, analyze the corresponding relationship between the switching operation and the switching state using the log correlation set to determine the triggering condition for the state switch. If the triggering condition meets the preset threshold, generate a data flow containing the switching state through the real-time feedback mechanism, obtain the change trend of the quantity in the data flow, judge the dynamic adjustment requirement of the active connections, update the connection pool status according to the dynamic adjustment requirement to obtain the output of the optimized monitoring module. For the output of the optimized monitoring module, generate a real-time feedback data stream to determine the final switching state record.
[0021] Preferably, S4 includes:
[0022] Obtain the real-time data stream from the data source, use stream processing technology to parse the number of active connections to obtain structured data, convert the structured data into the input format of the visualization component through data mapping technology to generate component rendering parameters, use interface integration technology to render the dynamic progress bar and result prompt to display the change in the number of active connections, determine the visual output. If the number of active connections exceeds the preset threshold, adjust the length of the progress bar and the content of the result prompt through dynamic update technology to obtain a real-time feedback display. According to the real-time feedback display, detect the state change during the switching process, use the state machine model to judge the transition of the connection state to determine the switching completion flag, present the switching completion flag through the refresh mechanism of the visualization component, update the dynamic progress bar and the result prompt to obtain the switching effect perceived by the user. If the refresh frequency of the switching effect is lower than the preset threshold, adjust the sampling rate of the data mapping to optimize the processing efficiency of the real-time data stream and generate a smooth feedback display.
[0023] Preferably, the S5 includes:
[0024] Obtain response time data from the state feedback. If the response time exceeds the preset threshold, determine the abnormal data source through time series analysis to obtain the identifier of the abnormal data source. According to the identifier of the abnormal data source, extract the corresponding data parameters from the multi-data source scenario, group the parameters using clustering analysis to obtain the feature set of the abnormal data parameters, calculate the parameter deviation value by comparing the feature set of the abnormal data parameters with the average value of the adjustment basis, determine whether the deviation value exceeds the preset range. If the deviation value exceeds the preset range, adjust the data source configuration according to the deviation value, predict the adjusted parameter value using linear regression analysis to obtain the calibrated configuration parameters, extract the key performance indicators from the calibrated configuration parameters, judge whether the calibrated data source meets the expectation by comparing with the expected value to obtain the judgment result. According to the judgment result, obtain the identifier of the data source that does not meet the expectation, further detect the change trend of its response time using time series analysis to obtain the optimization direction, and adjust the configuration parameters of the data source through the optimization direction, and loop through the calibration and judgment processes to obtain the finally data source configuration that meets the expectation.
[0025] Preferably, the S6 includes:
[0026] Obtain the calibrated connection information from the data source, verify the information integrity through the verification algorithm to obtain the verified connection data set, use the priority round-robin scheduling algorithm to extract the priority tags from the verified connection data set and load them into the task scheduling engine to obtain the initialized scheduling configuration. Run the dynamic adjustment mechanism through the task scheduling engine, adjust the connection pool size according to the priority tags to generate the adjusted connection pool parameters, extract the scale change data from the adjusted connection pool parameters, use the CSV format conversion tool to generate a structured adjustment record file, obtain the key fields in the adjustment record file, detect the status of the switching process through the log analysis tool to judge whether the switching process returns to normal. If the switching process returns to normal, extract the performance indicators from the operation log of the scheduling engine, use the statistical analysis method to obtain the evaluation data of the system stability, and update the priority rules of the task scheduling engine according to the evaluation data of the system stability to generate the optimized scheduling configuration.
[0027] Preferably, the S7 includes:
[0028] Obtain adjustment record data by parsing a CSV - formatted file, extract multi - data - source information according to preset fields, separate switching scenario data from the multi - data - source information, use time - series analysis to determine the state changes of CPU occupancy and memory usage, generate a trend chart for the state change data, calculate the coordinate points of dynamic switching using a line - chart algorithm, obtain the trend - chart coordinate - point data, render the content on the interface through a visualization component. If the switching scenario changes, update the state changes of CPU occupancy and memory usage according to the new scenario data, recalculate the trend - chart coordinate points based on the updated state changes, and refresh the dynamic - switching effect shown on the interface. Adjust the parameters of the visualization component according to the user - interaction records to obtain the optimized trend - chart display content.
[0029] Preferably, S8 includes:
[0030] Obtain CPU occupancy and memory - usage performance - metric data through the RESTful API during task operation, store it as a time - series data set, pre - process the time - series data set, calculate the response time of the dynamic - switching operation to generate a response - time series. If any value in the response - time series exceeds a preset threshold, mark it as an exception, count the number of exceptions to obtain the exception ratio, calculate the success rate through the exception ratio and the total number of switching operations. The success - rate formula is: S=(N - E) / N, where S represents the success rate, N represents the total number of switching operations, and E represents the number of exceptions, to obtain the success - rate value. Use the support - vector - machine algorithm to classify the response - time series and the performance - metric data to judge the stability of the dynamic - switching mechanism, obtain the classification result, and determine whether the dynamic - switching mechanism meets the business - scenario requirements based on the classification result and the success - rate value to generate a judgment result. If the judgment result is not satisfied, optimize the performance - metric data by adjusting the resource - allocation strategy during task operation, and repeat the above steps until a judgment result that meets the requirements is obtained.
[0031] Preferably, S9 includes:
[0032] Divide time windows by collecting time, calculate the average value within the window to obtain the data - fluctuation benchmark, adjust the refresh frequency of the scheduling engine according to the average value to determine the configuration refresh period, use a timeout - handling mechanism to monitor the refresh frequency, and trigger an automatic rollback if it exceeds the threshold to restore to the previous stable configuration. Obtain the input data of multiple data sources, allocate resources through the task - execution module, judge the execution efficiency, evaluate the stable environment for the task - execution state, adjust the time - window range if the fluctuation exceeds the expectation to obtain a new average value, update the configuration refresh frequency with the new average value to determine the operating parameters of the scheduling engine, obtain the operating log of the scheduling engine, and judge the overall stability of the system by analyzing the timeout records and rollback times in the log.
[0033] As can be seen from the above technical solutions, the present invention has the following beneficial effects:
[0034] This method collects performance data in real time through RESTful APIs, analyzes the response time threshold and success rate within a sliding window, dynamically adjusts the connection pool size, uses a status tracking module to record changes in the connection pool, generates a real-time feedback data stream and visualizes it. When the response time exceeds the threshold, it automatically extracts the abnormal source from multiple data sources, recalibrates the configuration and loads it into the scheduling engine, adopts priority-based round-robin scheduling, generates an adjustment record in CSV format, and displays the switching effect through a visualization component. The present invention also analyzes whether the switching operation meets the business requirements according to performance indicators, dynamically adjusts the configuration refresh frequency, and sets up a timeout automatic rollback mechanism, thereby realizing efficient and stable switching in a multi-data source scenario, and improving the overall performance and reliability of the system. Brief Description of the Drawings
[0035] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0037] As Figure 1 shown, the present invention provides a technical solution: a method for switching data sources when calling the springboot service interface based on the xxl-job timed task, including:
[0038] S1. Obtain performance data collected once per second, generate a JSON structured log file, and obtain the target data source identifier and connection information;
[0039] S2. Load the target data source identifier through a preset scheduling engine interface, dynamically adjust the connection pool size to 50 connections, and determine that the new data source connection has been activated;
[0040] S3. Record the change in the number of connections in the connection pool status monitoring, obtain the switching operation timestamp and INFO-level event tracking log, and generate a real-time feedback data stream containing the switching status;
[0041] S4. Map the real-time feedback data stream to a visualization component, generate a dynamic update progress bar and result prompt for the number of connections, and obtain a display of the switching process that can be perceived by the user;
[0042] S5. If it is detected in the switching state feedback that the response time threshold exceeds 500 milliseconds, extract the identifiers and parameters of the abnormal data sources from the multi-data source scenario, and determine whether the calibrated data sources meet the expectations;
[0043] S6. Obtain the connection information of the calibrated data sources and reload it into the scheduling engine during task execution, generate an updated adjustment record in CSV format, and determine that the switching process has returned to normal;
[0044] S7. Generate a state change trend chart for the switching results in the multi-data source scenario;
[0045] S8. Collect data of performance metric types, analyze whether the response time of the switching operation is lower than 300 milliseconds and the result of the success rate calculation formula, and determine whether the dynamic switching mechanism meets the requirements of the business scenario;
[0046] S9. Adjust the configuration refresh frequency of the scheduling engine to be updated every 10 seconds, and automatically roll back the configuration when the threshold is exceeded to obtain a stable task execution environment suitable for the multi-data source scenario.
[0047] This method is based on the xxl-job platform for timed task management and combines with the springboot service interface to achieve dynamic switching of data sources. In step S1, through the performance monitoring means with a sampling frequency of 1 second, the system operation status data is captured and the logs are structured and recorded in JSON format to facilitate the extraction of the target data source identifier and its connection configuration. In S2, the target data source information is dynamically loaded through the interface exposed by the scheduling engine, and the connection pool configuration is automatically adjusted at runtime, increasing the number of connections to 50 to meet the data source switching requirements under high-concurrency requests, and verifying whether the connection is correctly activated. In S3, with the help of the status monitoring tool, the changes in the connection pool are continuously tracked, the key operation timestamps and INFO logs are obtained, and a feedback data stream that can be used to trace the execution effect of the switching is generated. In S4, the feedback data is mapped to the visualization interface in real time to generate a dynamic progress bar and result prompt, improving the user's perception and judgment ability of the switching status. In S5, when the feedback data indicates that the response time exceeds 500 milliseconds, the abnormal data source is identified through the backtracking mechanism and recalibrated according to the standard parameters. In S6, the calibrated connection information is reloaded into the scheduling engine, and the adjustment process is persistently saved in the CSV record format. In S7, a state change trend graph can be generated to reflect the adaptation of the switching mechanism in a multi-data source environment. In S8, the switching mechanism is evaluated by collecting performance indicators such as response time and success rate to determine whether it meets the business objectives. In S9, by setting the configuration refresh frequency and triggering the threshold control, an automatic rollback configuration strategy is implemented when an exception occurs during the switching, thus constructing a dynamically schedulable environment that can be stably executed. This method has a high degree of automation and real-time response capabilities, and can achieve accurate and controllable data source switching in a dynamic business environment. On the one hand, through the dynamic adjustment of the connection pool and the visualization feedback mechanism, the system stability and user transparency are effectively improved. On the other hand, by introducing the exception detection and rollback mechanism, the fault tolerance and recovery capabilities of the system in high-concurrency or abnormal scenarios are enhanced. At the same time, the introduction of structured logs and trend graphs helps with long-term monitoring and optimization of the switching strategy. The performance analysis mechanism ensures that this method meets the high requirements of enterprise-level scenarios in terms of response time, stability, and success rate. The overall solution has good scalability and adaptability, and is suitable for complex data call logic scenarios in multi-source data platforms or microservice architectures.
[0048] In the product recommendation service of a large e-commerce platform "Platform X", to improve the query response speed of cross-border users, the present invention is adopted for data source switching. The platform is configured with data sources in multiple countries / regions, such as datasource_us, datasource_cn, datasource_uk, etc. Whenever a user requests access, the system dynamically switches to the corresponding regional data source according to the IP mapping rule, in combination with the xxl-job scheduled task and the method described in the present invention. After actual deployment, the average response time is reduced to 210 milliseconds during the Double Eleven period, and the success rate reaches 98.6%. The system administrator found that the European nodes fluctuated frequently through the trend chart, and timely optimized the data source configuration to avoid system crashes, significantly improving the stability of the business system and customer satisfaction.
[0049] S1 includes obtaining the performance data collected per second from the task runtime through the RESTful API, storing it as an initial data set. For the initial data set, parsing the collected data within the sliding window of the most recent 5 minutes to obtain a time series data set. Using the time series data set, calculating the response time of each data point to obtain a response time series. If a certain data point in the response time series exceeds the preset time threshold, it is marked as an anomaly to obtain an anomaly marking sequence. Through the anomaly marking sequence and the success rate formula, that is, the success rate is the total number of data points minus the number of anomaly points and then divided by the total number of data points, calculating the success rate within the sliding window to obtain a success rate sequence. Obtaining the success rate sequence and the data source identifier, generating a JSON structured log file containing the response time threshold and the success rate to obtain a log output file. According to the log output file, extracting the data source identifier and connection information to obtain a complete description of the target data source.
[0050] In this embodiment, the key to step S1 is to collect performance data from the task runtime environment at a frequency of once per second through the RESTful API interface, and store these data in JSON format to form an initial data set. The data content usually includes timestamp, response time, data source identifier, execution status, etc. During the data processing, the system performs a sliding window analysis on the initial data set. Specifically, a time series data set is constructed from the data collected in the last 5 minutes. Since 60 data points are collected per minute, each window contains 300 data points. Each data point contains a corresponding response time, forming a response time series. For example, the i-th response time is denoted as Ri. The system sets a response time threshold, such as 500 milliseconds. For each data point in the response time series, if its response time Ri exceeds the threshold, the point is marked as an "abnormal point", otherwise it is marked as a "normal point". Through this rule, an anomaly marking sequence can be generated, where the value at each position is 1 indicating that the point is abnormal and 0 indicating normal. Based on this anomaly marking sequence, the success rate within the sliding window can be further calculated. The definition of the success rate is: the total number of data points in the current sliding window minus the number of abnormal points, and then divided by the total number of data points. In other words, success rate = (total number of data points - number of abnormal points) ÷ total number of data points. For example, in a 5-minute window, 300 data points are collected, and 12 abnormal points are detected. Then the success rate for this time period is (300 - 12) divided by 300, which is 0.96. The system repeats the sliding window calculation, sliding forward one or more data points each time, and continuously calculates the success rates for multiple time periods, forming a success rate time series. Subsequently, the system combines the success rate value for each time period with the corresponding response time threshold and data source identifier to generate structured JSON log data for subsequent analysis and as a basis for judging the stability of data source switching. Finally, by parsing these JSON logs, the identifier and connection information of the target data source can be quickly extracted, preparing for the dynamic selection and switching of the data source. This processing logic provides the basic capabilities of real-time monitoring and dynamic adjustment for the task scheduling system, with high stability and high response performance. This method uses the sliding window mechanism and mathematical model to evaluate the task running stability in real time, introduces a quantifiable and continuous success rate calculation index, and improves the sensitivity and accuracy of anomaly recognition. Compared with traditional methods based on static thresholds or log matching, this solution has the following advantages: High-sensitivity anomaly detection ability: It can quickly identify performance degradation under high concurrency or critical load; Data-driven decision-making mechanism: The system performance trend can be judged through the success rate sequence to assist in optimizing the scheduling strategy; Standardized output interface: JSON format logs can be directly connected to log collection systems such as ELK and Fluentd to improve compatibility; Enhanced system adaptability and stability: Avoid misjudgment based on single-point response time and improve the overall fault tolerance.
[0051] In the equipment control center of a certain intelligent manufacturing platform, the system schedules multiple device status collection service tasks through xxl-job. To avoid misjudgment caused by network jitter, the platform introduces the method described in the present invention to monitor the response situation within a 5-minute window in real time during the task execution process. One day at noon, due to the core route switch, the response time of some data sources increased abnormally to 800 ms. The system automatically identified the abnormal data points and marked the tenant_012 data source as "unstable in performance", and the success rate dropped to 92.3%. The administrator checked the JSON log and the trend chart, quickly locked the problem data source and implemented manual intervention. Thanks to the sliding window evaluation and the JSON-structured log, the system realized an efficient, stable, and traceable dynamic data source switching mechanism without affecting the main process of equipment control.
[0052] S2 includes obtaining response time data by parsing the log file, performing threshold parsing on the response time using a preset threshold to obtain a parsing result. If the parsing result exceeds the preset threshold, loading the data source identifier through the scheduling engine to determine the target data source, adjusting the connection pool size according to the identification information of the target data source to obtain a new connection number configuration, activating the new data source through the new connection number configuration, judging whether the activation status is completed, obtaining the confirmation information of the activation status, loading and verifying the availability of the new data source through the interface to determine that the connection pool adjustment takes effect, extracting the change trend of the response time from the log file, analyzing the correlation between the change trend and the connection pool size using the random forest algorithm to obtain optimized connection parameters, updating the preset threshold of the scheduling engine according to the optimized connection parameters, and judging the stability of the data source connection.
[0053] In step S2 of this embodiment, the system completes the intelligent switching and performance optimization of the data source through the log analysis and connection pool parameter adjustment mechanism. The operation process is as follows: First, the system regularly extracts response time data from the log files generated during the task running process to construct a time series, where each data point represents the response time of an interface request. The system sets a preset response time threshold, such as 500 milliseconds, and then judges each response time data point against this threshold. If the response time of a certain data point exceeds this threshold, it is regarded as a performance anomaly. When the system detects one or more abnormal response times, it will trigger the data source switching mechanism. The scheduling engine will load the identification information of the target data source (such as data source name, database type, connection string, etc.) according to the business context of the current task, and determine the new data source to be connected. According to the configuration of the new data source, the system will readjust the connection pool size. For example, if the original maximum connection number of the connection pool is 30, it can be adjusted to 50 according to the load pressure. The system calls the connection pool parameter modification interface for adjustment and restarts the connection pool after adjustment to activate the new data source configuration. After the connection pool is adjusted and the new data source is activated, the system will confirm the availability of the new connection through the health check interface. Common confirmation methods include requesting the database status interface or executing a basic SQL statement (such as SELECT 1) to verify the connection validity. If the connection is successful, it means that the connection pool adjustment has taken effect. Next, the system will analyze the change trend of the response time from the previously extracted response time logs, that is, the difference in response time for each time point. For example, subtract the response time of the i-th time point from the response time of the i-1-th time point to obtain a set of continuous change values. The system uses these change values together with the connection pool size at the corresponding time points as training data and inputs them into the machine learning model for analysis. In this embodiment, the random forest algorithm is selected as the modeling tool. The system uses historical operation data to train the model so that it can learn the relationship between the connection pool size and the change in response time. When the model training is completed, the system can predict a more optimal connection pool parameter configuration according to the current response trend. Once the optimized connection pool parameters are obtained, the system will write these parameters into the configuration of the scheduling engine and synchronously update the response time threshold. The new threshold will no longer be a static setting but the average value calculated based on recent real response data. For example, if 300 response time data points are collected in the past 5 minutes and the total is 135000 milliseconds, the new average threshold will be 135000 divided by 300, resulting in 450 milliseconds. Finally, the system will continuously monitor the stability of the connection, such as detecting whether there are new abnormal responses within 10 minutes or whether the database health check remains passed. If the system performs normally during this period, it can be confirmed that the current configuration has good stability.
[0054] In this embodiment, by introducing a response time analysis and connection pool parameter optimization mechanism based on the random forest algorithm, the intelligence level of the data source switching process and the stability of system operation are significantly improved. Compared with the traditional solution that relies on fixed connection pool configurations, this method can real-time sense the fluctuations in connection performance and dynamically adjust the connection pool size according to the actual response data, avoiding resource waste or connection bottlenecks caused by improper parameter settings. Through the machine learning model to model and predict historical data, the transformation of connection optimization from "experience-driven" to "data-driven" is realized, enhancing the system's adaptability to complex operation scenarios such as sudden traffic and high-concurrency access. At the same time, a closed-loop control system is constructed through the scheduling engine and health check mechanism, enabling the connection adjustment process not only to have immediate response capabilities but also to be effectively verified and automatically rolled back after adjustment, ensuring that the task scheduling environment is always in the optimal operating state. This solution is particularly suitable for distributed systems with high stability requirements and frequently changing dynamic access patterns, effectively improving the configuration flexibility, fault self-healing ability, and overall operation and maintenance efficiency of the task scheduling platform in a multi-data source environment.
[0055] In the distributed accounting processing platform of a certain bank, to improve the data access stability of off-site disaster recovery nodes, the embodiment of the present invention is introduced into the system. The initial configuration of the connection pool for each data source node is 30. During the processing peak on a certain day, the response time of the Beijing node fluctuated continuously. The system judged the anomaly through the log and triggered the automatic reconfiguration process. The system collected the recent response time data, analyzed it using the trained random forest model, and output the recommended connection pool size of 48. The system immediately adjusted the pool size and updated the scheduling engine threshold. Subsequently, the response time tended to be stable, and the success rate remained above 98% within 10 minutes. This effectively avoided manual intervention and large-scale service jitter, significantly improving the system availability and fault self-healing ability.
[0056] S3 includes capturing the change in the number of active connections in the connection pool through the status tracking module, recording the time points of the changes to obtain preliminary data, extracting the time stamps corresponding to the switching operations from the preliminary data, matching the INFO-level event logs to generate a log correlation set, analyzing the corresponding relationship between the switching operations and the switching states using the log correlation set, determining the trigger conditions for state switching. If the trigger conditions meet the preset thresholds, a data flow direction containing the switching state is generated through the real-time feedback mechanism, obtaining the trend of the number change in the data flow direction, judging the dynamic adjustment requirements of the active connections, updating the connection pool state according to the dynamic adjustment requirements to obtain the output of the optimized monitoring module. For the output of the optimized monitoring module, a real-time feedback data stream is generated to determine the final switching state record.
[0057] In step S3 of this embodiment, the system continuously monitors the change in the number of active connections in the connection pool through the built-in status tracking module and records the specific time point when each change occurs. These data together constitute a preliminary status data set, which is used to analyze the change trend of the data source connection status during task execution. When the scheduling task triggers the data source switching operation, the system searches for the corresponding INFO-level event logs in the log file. These logs usually contain text markers such as "Start switching data source to XXX" or "Data source switching completed". The system matches these texts with the timestamps recorded in the preliminary status data set to form a set of "log association sets". This association set establishes a temporal association between "connection pool changes" and "switching operations". Subsequently, the system identifies from the log association set whether the switching operation has actually caused a change in the connection pool status. The judgment criteria are as follows: If, after the data source switching event occurs, the system monitors that the change value of the number of active connections is greater than or equal to the set threshold (e.g., 10 connections) within the set time window (e.g., within 30 seconds), it is considered that the switching has indeed caused a status change. Expressing this calculation logic in plain text: The number of active connections before switching is denoted as A1, and the number of active connections after switching is denoted as A2. If the absolute value of the difference between A2 and A1 is greater than or equal to 10, that is, |A2 - A1| ≥ 10, then it is determined that the switching effectively triggers the status change. After the system identifies the status trigger, it organizes the connection change trend collected in the subsequent period of time into a set of data streams. This data stream records the number of active connections per second in chronological order and is used to analyze whether the connection pool is currently in a stable state. Then, the system makes a trend judgment on this data stream by accumulating the change values of the connection numbers at adjacent time points. For example, the connection number differences per second are D1, D2... Dn. If the sum of these differences is greater than zero, it indicates that the usage of the connection pool shows a continuous upward trend and the system load may increase; conversely, if the sum is less than zero, it indicates that the usage pressure of the connection pool is reduced. According to this trend, the system can judge whether it is necessary to adjust the connection pool parameters, such as increasing the maximum number of connections or reducing the minimum number of idle connections. The adjusted results are output through the monitoring module and a feedback data stream is generated again to evaluate whether the configuration adjustment has produced the expected effect. Finally, the system confirms whether to complete the status switching record based on the feedback data, including whether the switching is successful, whether the adjusted parameters take effect, etc., to support subsequent system policy optimization and problem backtracking.
[0058] By introducing a state tracking mechanism based on the changing trend of active connections, this method realizes the dynamic identification, accurate judgment, and real-time feedback of the data source switching process, significantly enhancing the state awareness ability of the task scheduling system and the intelligent level of connection management. Different from the traditional method of indirectly judging the success or failure of switching through task logs, the present invention can quickly determine whether the switching actually takes effect based on specific connection data, and realize adaptive parameter adjustment in combination with the usage trend of the connection pool, effectively avoiding resource waste or task failure caused by unreasonable connection pool configuration. At the same time, this method visualizes the operation and maintenance behavior through the log correlation set and the trend analysis algorithm, enabling the system administrator to grasp the connection changes in real time on the graphical interface, improving the system transparency, operation and maintenance efficiency, and automation level. This mechanism is particularly applicable to high-concurrency task scenarios, can ensure the stable operation of tasks and reduce manual intervention, and has extremely strong engineering practicability and promotion value.
[0059] In the business system of an Internet insurance company, a certain regular insurance push service uses xxl-job for task scheduling. This service needs to frequently switch the customer database to complete concurrent writing during the peak period (10:00-10:30 every day). After introducing the present invention, the system can collect the number of active connections in real time and judge whether each switch is effectively triggered in combination with the event log. A certain monitoring record shows that after the data source is switched, the number of connections jumps from 12 to 34 and continues to increase. The system dynamically increases the maximum number of connections to 60 according to the change trend and stabilizes the switching state within 5 minutes. The administrator can intuitively view the state evolution diagram through the Grafana interface, greatly improving the fault location efficiency and the system operation transparency.
[0060] S4 includes obtaining real-time data streams from the data source, using stream processing technology to parse the number of active connections to obtain structured data, converting the structured data into the input format of the visualization component through data mapping technology to generate component rendering parameters, using interface integration technology to render the dynamic progress bar and result prompt to display the change of the number of active connections, determining the visual output, if the number of active connections exceeds the preset threshold, then adjusting the length of the progress bar and the content of the result prompt through dynamic update technology to obtain the real-time feedback display, detecting the state change of the switching process according to the real-time feedback display, using the state machine model to judge the transition of the connection state, determining the switching completion flag, presenting the switching completion flag through the refresh mechanism of the visualization component, updating the dynamic progress bar and the result prompt to obtain the switching effect perceived by the user, if the refresh frequency of the switching effect is lower than the preset threshold, then adjusting the sampling rate of the data mapping to optimize the processing efficiency of the real-time data stream and generating a smooth feedback display.
[0061] In this embodiment, the key to step S4 is to visually present the data source connection status involved in the task execution process to the user in real time, so as to enhance the observability and interaction experience of the system. The system first continuously obtains the real-time connection status data stream from the target data source or the connection pool module. This data contains key metrics such as the current active connection count and the maximum connection count. After being parsed by the stream processing module (such as Spring WebFlux stream, Kafka Streams stream), the "active connection count" field is extracted from these raw data. This field is converted into a standard structured data format, such as JSON structure, for use as the data input for the front-end component. Next, according to the preset mapping logic, the system converts the active connection count into the display ratio of the progress bar. The conversion relationship is: the current active connection count divided by the maximum connection count, resulting in a percentage value, which directly determines the filling ratio of the progress bar. For example, when the current active connection count is 25 and the maximum connection count is 50, the calculated ratio is 25 divided by 50, and the result is 0.5, that is, 50%. This value is passed to the front-end component to render a dynamic progress bar filled with 50%. On this basis, the system also generates prompt information according to the connection status. If the active connection count exceeds the preset threshold, for example, set to 40, the prompt text will be automatically changed to "Connection count approaching the upper limit" or "Running under high load". Such prompts are displayed side by side with the progress bar through the front-end component, facilitating the user to quickly understand the current system status. In addition, the system introduces a state machine model to judge the change process of the connection status. The state machine includes at least three states, namely "Initialization", "Switching", and "Switching Completed". If the system continuously detects that the connection count changes stably within a period of time (for example, the change amplitude is less than 2 connections for five consecutive times), it will automatically transfer from the "Switching" state to the "Switching Completed" state, and update the prompt content to "Switching Completed". The front-end display part obtains the latest status data through WebSocket push or timed refresh mechanism to ensure real-time information update. If the system detects that the refresh frequency of the visualization component is lower than the set value (for example, once per second), it means that there is a delay or lag in the display. The system will reduce the data sampling frequency to relieve the transmission and processing pressure. For example, sampling once per second is adjusted to sampling once every two seconds. The adjustment strategy is: the sampling frequency is directly proportional to the refresh frequency. When the actual refresh frequency decreases, the sampling frequency is synchronously reduced, thereby improving the data processing efficiency and page fluency.
[0062] By combining the running status of the connection pool with the visualization component, this method constructs a user interface system with real-time, interactive, and intelligent adjustment capabilities, effectively enhancing the visibility of the data source switching process and the operation and maintenance efficiency. Users can use the dynamic progress bar and prompt messages to keep track of the changes in the connection status in real time, so as to quickly identify the current switching stage or the running load situation of the system, greatly enhancing the trust in the system and the operation transparency. The introduction of the state machine model enables the system to automatically identify whether the switching is completed and accurately reflect the state changes, reducing the possibility of human intervention and misjudgment. At the same time, through the dynamic sampling rate adjustment mechanism, the system can reasonably control the resource usage while maintaining the smoothness of the display, avoiding information congestion and front-end performance bottlenecks. This display mechanism centered on user perception is particularly applicable to large-scale data source scheduling or multi-tenant task environments, and has extremely strong practical value and engineering promotion prospects.
[0063] In an intelligent operation and maintenance platform, users batch update multiple customer bills through the task scheduling module. The system is configured to automatically switch to the exclusive database in the customer's location during task execution. During the switching process, the platform uses the visualization component in the present invention to display the active connection number of the connection pool in the form of a dynamic progress bar on the operation and maintenance console, and the status prompt message automatically changes. Users can see prompts such as "switching" and "connection completed" in real time. When the background detects the switching success flag, the page is synchronously updated to "switching success" and enters the data processing state. If the data refresh is too slow due to concurrent requests, the system automatically adjusts the sampling rate from once per second to once every two seconds to ensure smooth front-end display without jamming, significantly enhancing the user's operation confidence and operation and maintenance response efficiency.
[0064] S5 includes obtaining response time data from the state feedback. If the response time exceeds the preset threshold, then determine the abnormal data source through time series analysis to obtain the identifier of the abnormal data source. According to the identifier of the abnormal data source, extract the corresponding data parameters from the multi-data source scenario, group the parameters using clustering analysis to obtain the feature set of the abnormal data parameters. By comparing the feature set of the abnormal data parameters with the average value of the adjustment basis, calculate the parameter deviation value, and determine whether the deviation value exceeds the preset range. If the deviation value exceeds the preset range, then adjust the data source configuration according to the deviation value, use linear regression analysis to predict the adjusted parameter value to obtain the calibrated configuration parameters, extract the key performance indicators from the calibrated configuration parameters, and by comparing with the expected value, judge whether the calibrated data source meets the expectation to obtain the judgment result. According to the judgment result, obtain the identifier of the data source that does not meet the expectation, use time series analysis to further detect the change trend of its response time to obtain the optimization direction, and adjust the configuration parameters of the data source through the optimization direction, and loop through the calibration and judgment processes to obtain the final data source configuration that meets the expectation.
[0065] The system first extracts the response time series from the task status feedback, and the response time at each time point is denoted as Ri. A preset threshold T is set, for example, 500 milliseconds. If the response time Ri at a certain time point is greater than T, the current response is determined to be abnormal, and the corresponding data source identifier Ds is recorded. Next, the system extracts relevant parameters from the connection pool configuration associated with Ds, such as the maximum connection number, the minimum idle connection number, the connection timeout, the verification query statement, etc. These parameters form a feature vector for subsequent analysis. To identify the commonalities of abnormal parameters, the system uses a clustering analysis method (such as the K-Means algorithm) to divide the parameter sets of multiple data sources into several groups according to similarity. The goal is to identify the abnormal feature set to which Ds belongs, that is, the parameters with significant deviation from the mean within its group. Specifically, the system calculates the deviation of each parameter from its clustering center (i.e., the mean). If the deviation value exceeds the set range Δ, the parameter is marked as "to be adjusted". For example, the clustering center of the maximum connection number is 50, and a certain data source is 80. If Δ is set to 20%, the deviation is (80 - 50) / 50 = 0.6, which exceeds Δ = 0.2, so adjustment is required. To ensure a reasonable adjustment direction, the system introduces a linear regression model. The historical connection parameters are used as independent variables, and the response time is used as the dependent variable for training to obtain a prediction model. Then, the adjusted parameters are used as input to predict the response time to determine whether it is expected to meet the desired range (for example, less than 300 milliseconds). If the predicted value still does not meet the expectation, the system continues to adjust the parameters to form a cyclic calibration mechanism. At the same time, the system continuously extracts key performance indicators from the calibrated configuration, such as the average response time corresponding to the maximum connection number, the connection establishment success rate, etc., and compares them with the system's preset performance expectation indicators. If it still does not meet the requirements, the system detects the change trend of the subsequent response time of this data source through time series analysis. The trend analysis method is: fitting the most recent n response times. If the change rate continues to rise or the fluctuation intensifies, it is considered still unstable. Finally, through the above closed-loop process of calibration - judgment - optimization, the system obtains a data source configuration that meets the performance expectation and updates it to the scheduling engine as the official operating parameters.
[0066] This method realizes the automatic identification, precise positioning and intelligent optimization of abnormal data sources in a multi-data source environment by introducing multi-stage processing mechanisms such as response time analysis, clustering recognition, parameter deviation detection, linear regression prediction and trend tracking. It significantly improves the connection stability and performance recovery ability of the system in high-concurrency and complex distributed scenarios. Different from the traditional method that relies on manual parameter tuning or static configuration, the present invention uses a data-driven parameter optimization strategy, enabling the system to have an adaptive ability. It can automatically adjust connection parameters according to performance feedback during real-time operation, shorten the fault recovery time, reduce the frequency of manual intervention, enhance the maintainability and reliability of the system, and is particularly suitable for connection strategy optimization tasks in multi-tenant or dynamic switching scenarios in large-scale microservice systems, with broad engineering application value.
[0067] In the inventory management system of a national chain supermarket, the system uses xxl-job to schedule multiple inventory databases by region. One day, the response time of the South China node continued to be higher than 700 milliseconds, and the system automatically identified it as abnormal. Through cluster analysis, it was found that the maximum connection number of this node was set to 80, which was much higher than the average level. The linear regression model predicted that the corresponding response time for this setting was about 620 milliseconds, which did not meet the standard. The system automatically adjusted the maximum connection number to 55 and continuously monitored the response time trend, which dropped to 280 milliseconds within 5 minutes. Finally, the adjusted parameter was locked and updated to the scheduling configuration, and the entire optimization process did not require manual intervention, significantly improving the connection performance and system response speed.
[0068] S6 includes obtaining calibrated connection information from the data source, verifying the information integrity through a verification algorithm to obtain a verified connection data set, using the priority round-robin scheduling algorithm to extract priority tags from the verified connection data set and load them into the task scheduling engine to obtain an initialized scheduling configuration, running a dynamic adjustment mechanism through the task scheduling engine to adjust the connection pool size according to the priority tags to generate adjusted connection pool parameters, extracting scale change data from the adjusted connection pool parameters, using a CSV format conversion tool to generate a structured adjustment record file, obtaining the key fields in the adjustment record file, detecting the switching process status through a log analysis tool to determine whether the switching process has returned to normal. If the switching process has returned to normal, extracting performance metrics from the scheduling engine running log, using statistical analysis methods to obtain evaluation data on system stability, and updating the priority rules of the task scheduling engine according to the evaluation data on system stability to generate an optimized scheduling configuration.
[0069] In step S6 of this embodiment, the system first extracts connection configuration information from the data source that has been optimized, including key fields such as the maximum number of connections, the minimum number of idle connections, the connection timeout, and the connection verification SQL statement. The system performs integrity verification operations on these fields, such as determining whether a field is missing, whether the data types match, and whether the values are within the set range. For example, the maximum number of connections must be a positive integer, and the timeout should be between 1000 milliseconds and 10000 milliseconds. If all fields meet the conditions, the connection dataset is marked as "valid". Next, the system assigns a priority label to each valid connection dataset. There are usually three types of labels: high priority, medium priority, and low priority. The determination of the label can be generated based on multiple indicators such as the historical performance of the task, such as success rate, average response time, and number of exceptions. For example, tasks with a success rate higher than 95% and a response time lower than 300 milliseconds are given a high priority. When the scheduling engine executes the task to load the connection configuration, it adjusts the parameters of the connection pool, especially the maximum number of connections, according to the priority label. For example, the default connection pool size is 30 connections. If the task is of high priority, the maximum connection number of the connection pool is expanded to 48 connections according to the weight setting. The calculation logic is to multiply the default number of connections by the priority weight coefficient. For example, if the high-priority weight is 1.6, then 30 multiplied by 1.6 gives 48. The above adjustment information is recorded as a structured file and saved in CSV format. The content includes fields such as timestamp, data source identifier, priority label, number of connections before and after adjustment, and reason for adjustment. These record files are convenient for subsequent operation and maintenance personnel to audit and also provide a basis for the continuous optimization of the system. During the task execution process, the system uses the log analysis module to monitor in real time whether the task is running normally, including determining whether there are connection failures, response delays, error messages, etc. If the system does not detect any abnormalities within the set monitoring period and the average response time remains within the ideal range (e.g., lower than 300 milliseconds), it indicates that the switching process has returned to normal. Subsequently, the system extracts performance data from the scheduling log, such as average response time, error rate, task completion time, etc., and then uses statistical methods to calculate the system stability score. The calculation process of this score includes the proportion of successful tasks and the degree of response time fluctuation. For example, if the total number of tasks is 100 and the number of successful tasks is 98, the success rate is 98%. If the response time fluctuation is small, it indicates that the system is running stably. Finally, the stability score is comprehensively obtained to evaluate the rationality of the current scheduling configuration. If the score reaches the expected standard (such as above 0.85), the system will update the priority policy table in the scheduling engine and generate an optimized scheduling configuration for subsequent task scheduling.
[0070] The present invention realizes the full - process intelligent control and scheduling optimization of the data source switching process by introducing a priority scheduling tag and a connection pool dynamic adjustment mechanism, combined with connection information integrity verification, log feedback monitoring, and system stability evaluation and analysis. Its core advantage lies in forming a closed - loop mechanism from connection configuration optimization, task scheduling execution, connection adjustment record, performance monitoring, policy evaluation to rule update, with high automation, self - adaptation, and sustainable optimization capabilities. This method effectively avoids the performance bottleneck and recovery lag problems caused by static configuration in traditional task scheduling, ensures the stable operation of critical tasks in high - concurrency scenarios through a priority weighting strategy, and at the same time improves the overall system resource utilization rate and maintenance efficiency. It is particularly suitable for large - scale multi - tenant task platforms and cross - source distributed scheduling systems, with significant engineering practical value and broad promotion prospects.
[0071] In a provincial - level medical data platform, it is necessary to batch - synchronize the city - and - county - level HIS databases to the provincial - level analysis center every night. This process involves switching the connection configurations of hundreds of data sources. During system operation, 5 municipal nodes are automatically identified with a response time exceeding 600 milliseconds, which is determined to be abnormal. The optimized connection parameters are loaded into the task engine, with the priority set to High, and the maximum number of connections in the connection pool is adjusted from the original 30 to 48. After the scheduling engine runs, the system is stably restored, and the response time drops to 270 milliseconds. The adjustment record is generated and archived in CSV format. Through log analysis, the stability score is obtained as 0.89, and the system updates the priority rules, and the optimized configuration is automatically inherited by the next task. The entire process realizes the closed - loop control from anomaly identification, scheduling optimization, parameter write - back to stability evaluation, greatly improving the robustness and automation level of the cross - database scheduling system.
[0072] S7 includes obtaining adjustment record data by parsing a CSV - format file, extracting multi - data - source information according to preset fields, separating switching - scenario data from the multi - data - source information, using time - series analysis to determine the state changes of CPU occupancy and memory usage, generating a trend graph for the state - change data, calculating the coordinate points of dynamic switching using a line - graph algorithm, obtaining the trend - graph coordinate - point data, rendering the content on the interface through a visualization component. If the switching scenario changes, then update the state changes of CPU occupancy and memory usage according to the new scenario data, recalculate the trend - graph coordinate points based on the updated state changes, refresh the dynamic - switching effect shown on the interface, and adjust the parameters of the visualization component according to the user - interaction record to obtain the optimized trend - graph display content.
[0073] In this embodiment, step S7 mainly demonstrates the changing trends of CPU and memory usage during the data source switching process through the dynamic construction and real-time refreshing mechanism of the trend chart. The system first parses the previously generated CSV format adjustment record file, which contains the resource usage status data during the execution of multi-data source tasks. By setting field filtering conditions, such as "data source identifier", "switching time", "CPU usage rate", "memory usage", etc., the system extracts the records directly related to the switching operation from it to construct a switching scenario data set. Then, the system organizes the CPU and memory usage data in a time series manner. For example, data is collected once per second, and two time series can be formed within a continuous time period: one series recording the CPU usage rate and one series recording the memory usage. Each data point corresponds to a timestamp and a value. The system uses the timestamp as the abscissa and the usage rate or usage as the ordinate to construct the corresponding trend chart. In terms of coordinate point calculation, the abscissa of each point is the specific time point of data collection (for example, the i-th second), and the ordinate is the resource usage value at this time point (such as 75% for CPU and 3.2GB for memory). Connecting all the points to form a broken line can display the resource change trend curve. This line chart can reflect the changes in the system resource load before and after the switch in real time. When the system detects a change in the switching scenario, such as switching from one data source node to another, it will automatically clear the original trend chart data, re-collect the new CPU and memory series according to the new scenario, re-calculate and render the new trend chart. This update process is achieved by refreshing the chart component. At the same time, the system records the user's interaction operations on the chart, such as zooming the coordinate axes, switching the display area, or modifying the chart style, etc. According to these interaction records, the system dynamically adjusts the chart parameters, such as automatically adapting the coordinate scale range, display time window length, chart color scheme, etc., so as to enhance the visualization experience.
[0074] The present invention realizes the intuitive and real-time display of the changes in system resource usage during the data source switching process by constructing a dynamic trend chart based on time series, enhancing the perceivability and control ability of the operation and maintenance personnel on the system operation status. Through continuous monitoring and visual expression of CPU and memory data, users can quickly identify the system load fluctuations caused by the switch, effectively assisting in abnormal judgment and parameter adjustment decisions. At the same time, the trend chart refreshing mechanism ensures that the displayed content is always consistent with the actual state, and the chart optimization design driven by user interaction records makes the display interface more adaptable and interactive. The overall solution improves the operation transparency and performance tuning ability of the task scheduling platform in a multi-data source high-concurrency environment, especially applicable to complex distributed systems that require continuous monitoring and adjustment of connection performance, and has strong engineering practical value and application promotion prospects.
[0075] In a cross-border e-commerce platform, the night system performs scheduled archiving operations on user order data in multiple countries through xxl-job tasks. During a certain execution, the system switches the processing task from the "Singapore node" to the "Australia node", and at the same time, the background records the change trends of CPU and memory usage before and after the node switch. The system generates a dynamic trend chart through the present invention and displays the resource fluctuation curve in real time on the front-end operation and maintenance platform. The administrator observes that the memory usage continues to rise after the switch. After analysis, it is determined that the concurrent access volume of the new node is too large. Then, the administrator manually increases the maximum connection number and simultaneously observes the effect of parameter adjustment on the trend chart to intuitively perceive the performance recovery process. This mechanism not only reduces the time for analyzing switch failures but also improves the response efficiency of operation and maintenance personnel and the system stability.
[0076] S8 includes obtaining CPU occupancy rate and memory usage performance metric data from the task runtime through the RESTful API, storing it as a time series data set, preprocessing the time series data set, calculating the response time of the dynamic switch operation to generate a response time series. If any value in the response time series exceeds a preset threshold, it is marked as an anomaly. The number of anomalies is counted to obtain the anomaly ratio. Through the anomaly ratio and the total number of switch operations, the success rate is calculated. The success rate formula is: S = (N - E) / N, where S represents the success rate, N represents the total number of switch operations, and E represents the number of anomalies. The success rate value is obtained. The support vector machine algorithm is used to classify the response time series and the performance metric data to judge the stability of the dynamic switch mechanism and obtain the classification result. According to the classification result and the success rate value, it is determined whether the dynamic switch mechanism meets the business scenario requirements to generate a judgment result. If the judgment result is not satisfied, the resource allocation strategy during task runtime is adjusted to optimize the performance metric data, and the above steps are repeated until a judgment result that meets the requirements is obtained.
[0077] In step S8 of this embodiment, the system evaluates and optimizes the stability of the dynamic data source switching mechanism by regularly collecting performance indicators and combining response time analysis. The system first obtains real-time performance indicator data from the task runtime environment through the RESTful API interface, including CPU occupancy and memory usage, and records it once per second to form a time series data set. Each piece of data contains fields such as timestamp, CPU usage, and memory usage, and the system organizes these data into a continuous performance sequence. At the same time, in each data source switching operation, the system records the start time and completion time of the operation, and calculates the response time of the switch based on this. All response time data constitute a response time sequence. For example, if the start time of the i-th switch is Tstart and the end time of the i-th switch is Tend, then the response time is Tend minus Tstart, usually in milliseconds. The system sets a preset response time threshold (such as 500 milliseconds) and compares each data point in the response time sequence with the threshold. If a response time exceeds the threshold, it is marked as an "abnormal switch". In the statistical period, let the total number of switching operations be N and the number of abnormalities be E. Then the abnormal ratio is E divided by N, and the success rate is N minus E divided by N. The calculation formula can be expressed as: the success rate is equal to (the total number of switching operations minus the number of abnormalities) divided by the total number of switching operations. In order to further determine whether the switching mechanism is stable, the system combines the above response time data with the CPU and memory performance indicators of the corresponding period to form a feature vector, and uses the support vector machine algorithm for classification analysis. The classification goal is to determine whether the switching behavior belongs to the "stable" or "unstable" category. The support vector machine maps the input features to a high-dimensional space through a training model and divides the two categories of results at the optimal interval. Finally, the system evaluates whether the switching mechanism meets the current business scenario requirements based on the classification results and the above success rate value. If the success rate is lower than the set qualified threshold (such as 95%), or there are more "unstable" marks in the classification results, it is determined that the current mechanism does not meet the business requirements. At this time, the system automatically starts the resource allocation optimization strategy, including increasing CPU resources, expanding memory capacity, or adjusting connection pool parameters. After the optimization is completed, the system re-collects data and executes the above evaluation process to form a closed-loop control. The switching mechanism can only be considered stable and usable until the success rate and classification results reach the set standards.
[0078] Through the construction of a dynamic evaluation system integrating performance monitoring, response time analysis, and machine learning classification, the present invention has successfully realized the stability judgment and parameter optimization functions of the data source switching mechanism during the operation period. Compared with the traditional method relying on static rules or manual monitoring, this method can automatically identify performance bottlenecks and response anomalies, calculate the switching success rate in real time, and intelligently classify the stability of switching behaviors based on the support vector machine algorithm, greatly improving the adaptive ability and response sensitivity of the switching mechanism. When the judgment mechanism does not meet the business requirements, the system can independently execute resource configuration adjustment and form a repeated evaluation closed-loop after optimization, making the whole process highly automated, data-driven, and capable of real-time feedback. This solution is particularly suitable for microservice architectures or distributed task systems with large-scale multiple data sources, significantly reducing the operation and maintenance costs and improving the reliability, efficiency, and intelligence level of the system.
[0079] In a certain intelligent medical data platform, the system needs to synchronize electronic medical record data between different regional hospitals. To improve efficiency, the platform uses xxl-job to execute cross-database switching tasks for different time windows. During operation on a certain day, the system found that when switching from the "South China Medical Database" to the "Beijing-Tianjin Node", the response time frequently exceeded 500 ms, which was determined to be abnormal. Through analysis, it was found that the CPU usage rate was as high as 92%, and the memory was also close to the upper limit. The SVM model judged that the switching process was "unstable", and the success rate dropped to 91%. The system automatically upgraded the task Pod from 2 cores to 4 cores, expanded the memory from 2 GB to 3 GB, and increased the upper limit of the connection pool. After rescheduling, the response time dropped to 310 ms, and the success rate rebounded to 98.4%. The model automatically updated the training data to form a closed-loop optimization. This solution restored service stability without manual intervention, demonstrating the excellent performance of this patent solution in intelligent scheduling and dynamic self-healing.
[0080] S9 includes dividing time windows by collection time, calculating the average value within the window to obtain the data fluctuation benchmark, adjusting the refresh frequency of the scheduling engine according to the average value, determining the configuration refresh period, using a timeout processing mechanism to monitor the refresh frequency, triggering an automatic rollback if the threshold is exceeded, restoring to the previous stable configuration, obtaining the input data of multiple data sources, allocating resources through the task execution module, judging the execution efficiency, evaluating the stable environment for the task execution status, adjusting the time window range if the fluctuation exceeds the expectation, obtaining a new average value, updating the configuration refresh frequency through the new average value, determining the operating parameters of the scheduling engine, obtaining the operating logs of the scheduling engine, and judging the overall stability of the system by analyzing the timeout records and rollback times in the logs.
[0081] In this embodiment, in step S9, by setting a time window, the performance data collected during the task execution process is periodically statistically analyzed, the refresh frequency of the scheduling engine is dynamically adjusted, and the stable operation of the scheduling system is ensured through the timeout monitoring and rollback mechanism. The system first collects the operation data at the set time interval and divides it into several time windows of a fixed length. Within each time window, the system statistically analyzes the average values of key metrics such as response time, CPU usage rate, and memory usage. For example, within a 60-second window, if the response time is recorded as several data points, the system sums these data points and divides by the number of data points to obtain the average response time for this time period. This average value is used as the "data fluctuation benchmark". Based on this benchmark, the system dynamically calculates the refresh period of the scheduling engine. Specifically, if the current average response time is greater than the set ideal response time, the system proportionally extends the refresh period; if it is lower than the ideal value, the system shortens the refresh period, thereby achieving the optimization of the scheduling frequency to adapt to the task pressure. The system is equipped with a refresh timeout monitoring mechanism. If it is detected that the actual refresh time for two consecutive times exceeds the preset time threshold, a rollback operation is automatically executed to restore the current configuration to the stable parameter values of the previous round. In addition, the system combines the task execution module to monitor the resource allocation status and task completion efficiency during the processing of multiple data sources, such as processing time consumption, concurrency, and failure rate. If the task execution status fluctuates significantly beyond expectations within the current time window, for example, the standard deviation of the response time is greater than the set value, the system will automatically adjust the length of the analysis window, recalculate the average value, and update a more reasonable refresh period. To continuously evaluate the operation stability of the scheduling engine, the system also extracts the entries related to "refresh timeout events" and "configuration rollback records" from the operation logs and calculates their proportions in the total number of operations. For example, if the total number of refreshes is 100 times, and there are 6 timeouts and 2 rollbacks, the timeout rate is 6% and the rollback rate is 2%. If such indicators continuously remain at a low level, it can be determined that the current scheduling strategy is stable and reliable.
[0082] Through the dynamic analysis method based on time window and the adaptive refresh frequency control mechanism, the present invention realizes the intelligent configuration update and anomaly self-healing capabilities of the task scheduling system in a multi-data source environment. Compared with the traditional fixed-period refresh or manual adjustment methods, this method can flexibly adjust the refresh period of the scheduling engine according to the real-time fluctuations of the system operation data, and automatically trigger the configuration rollback in case of anomalies, thus effectively reducing the instability risk caused by frequent configuration changes in the system. At the same time, the system improves the task execution efficiency and resource utilization rate by adjusting the time window length and optimizing the refresh rhythm. During continuous operation, the system also timely captures performance bottlenecks and anomaly indicators through log data analysis, providing a solid basis for system stability evaluation and strategy evolution. The overall method features high real-time performance, high stability and high self-adaptability, and is particularly suitable for distributed microservice platforms that need to stably support high-frequency scheduling and dynamic task switching, with significant engineering feasibility and promotion value.
[0083] In a certain traffic big data platform, xxl-job is used to schedule the timed update of hundreds of urban traffic event databases. The system evaluates the refresh frequency of the scheduling engine in real time according to the change of night data synchronization traffic. At 3 am one time, the system detected a significant increase in the response time in the southwest region, and the refresh period was adjusted to twice the original, while triggering the rollback mechanism to avoid configuration instability. Subsequently, the analysis time window was automatically expanded from the original 60 seconds to 180 seconds, and the fluctuation average value was recalculated and the scheduling parameters were updated. This operation effectively alleviated the performance risk brought by configuration drift during peak periods, maintained the stable operation of the entire scheduling system in a high-concurrency environment, and improved the overall system availability and automatic operation and maintenance level.
[0084] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for switching data sources when calling the springboot service interface based on the xxl-job timed task, characterized in that Including: S1. Obtain the performance data collected once per second, generate a JSON structured log file, and obtain the target data source identifier and connection information; S2. Load the target data source identifier through a preset scheduling engine interface, dynamically adjust the connection pool size to 50 connections, and determine that the new data source connection has been activated; S3. Record the change in the number of connections in the connection pool status monitoring, obtain the switching operation timestamp and INFO-level event tracking log, and generate a real-time feedback data stream containing the switching status; S4. Map the real-time feedback data stream to a visualization component, generate a dynamically updated progress bar and result prompt for the number of connections, and obtain a user-perceivable switching process display; S5. If it is detected in the switching status feedback that the response time threshold exceeds 500 milliseconds, extract the identifier and parameters of the abnormal data source from the multi-data source scenario, and determine whether the calibrated data source meets the expectations; S6. Obtain the connection information of the calibrated data source and reload it into the scheduling engine during task runtime, generate an updated adjustment record in CSV format, and determine that the switching process has returned to normal; S7. Generate a status change trend chart for the switching results in the multi-data source scenario; S8. Collect performance metric type data, analyze whether the response time of the switching operation is lower than 300 milliseconds and the result of the success rate calculation formula, and determine whether the dynamic switching mechanism meets the requirements of the business scenario; S9. Adjust the configuration refresh frequency of the scheduling engine to update once every 10 seconds, and automatically roll back the configuration when the threshold is exceeded, to obtain a stable task execution environment suitable for the multi-data source scenario.
2. The method for switching data sources when invoking the springboot service interface based on the xxl-job timed task according to claim 1, wherein: The S1 includes: Obtain the performance data collected per second from the task runtime through the RESTful API, store it as an initial data set. For the initial data set, parse the collected data within the sliding window of the most recent 5 minutes to obtain a time series data set. Use the time series data set to calculate the response time of each data point to obtain a response time series. If a certain data point in the response time series exceeds the preset time threshold, mark it as abnormal to obtain an abnormal marking series. Through the abnormal marking series and the success rate formula, that is, the success rate is the total number of data points minus the number of abnormal points and then divided by the total number of data points, calculate the success rate within the sliding window to obtain a success rate series. Obtain the success rate series and the data source identifier, generate a JSON structured log file containing the response time threshold and the success rate to obtain a log output file. According to the log output file, extract the data source identifier and connection information to obtain a complete description of the target data source.
3. The method for switching data sources when invoking the springboot service interface based on the xxl-job timed task according to claim 1, wherein: The S2 includes: Obtain response time data by parsing the log file, perform threshold parsing on the response time using a preset threshold to obtain a parsing result. If the parsing result exceeds the preset threshold, load the data source identifier through the scheduling engine to determine the target data source. According to the identifier information of the target data source, adjust the connection pool size to obtain a new connection number configuration. Activate the new data source through the new connection number configuration, determine whether the activation status is completed, obtain the confirmation information of the activation status, load and verify the availability of the new data source through the interface to determine that the connection pool adjustment takes effect. Extract the change trend of the response time from the log file, analyze the correlation between the change trend and the connection pool size using the random forest algorithm to obtain optimized connection parameters. Update the preset threshold of the scheduling engine according to the optimized connection parameters, and judge the stability of the data source connection.
4. The method for switching data sources when invoking the springboot service interface based on the xxl-job timed task according to claim 1, characterized in that: The S3 includes: Capture the change in the number of active connections in the connection pool through the status tracking module, record the time point of the change to obtain preliminary data, extract the timestamp corresponding to the switching operation from the preliminary data, match the INFO-level event logs to generate a log correlation set, analyze the corresponding relationship between the switching operation and the switching state using the log correlation set to determine the trigger condition for the state switch. If the trigger condition meets the preset threshold, generate a data flow direction containing the switching state through the real-time feedback mechanism, obtain the change trend of the quantity in the data flow direction, judge the dynamic adjustment requirement of the active connections, update the connection pool status according to the dynamic adjustment requirement to obtain the optimized monitoring module output. For the optimized monitoring module output, generate a real-time feedback data stream to determine the final switching state record.
5. The method for switching data sources when the xxl-job timed task calls the springboot service interface according to claim 1, characterized in that: The S4 includes: Obtain the real-time data stream from the data source, parse the number of active connections using stream processing technology to obtain structured data, convert the structured data into the input format of the visualization component through data mapping technology to generate component rendering parameters, use interface integration technology to render the dynamic progress bar and result prompt to display the change in the number of active connections, determine the visual output. If the number of active connections exceeds the preset threshold, adjust the length of the progress bar and the content of the result prompt through dynamic update technology to obtain the real-time feedback display. According to the real-time feedback display, detect the state change during the switching process, use the state machine model to judge the transition of the connection state to determine the switching completion flag, present the switching completion flag through the refresh mechanism of the visualization component, update the dynamic progress bar and the result prompt to obtain the switching effect perceived by the user. If the refresh frequency of the switching effect is lower than the preset threshold, adjust the sampling rate of the data mapping to optimize the processing efficiency of the real-time data stream and generate a smooth feedback display.
6. The method for switching data sources when the xxl-job timed task calls the springboot service interface according to claim 1, characterized in that: The S5 includes: Obtain response time data from the state feedback. If the response time exceeds the preset threshold, determine the abnormal data source through time series analysis, obtain the identifier of the abnormal data source, extract the corresponding data parameters from the multi-data source scenario according to the identifier of the abnormal data source, group the parameters using clustering analysis to obtain the feature set of abnormal data parameters, calculate the parameter deviation value by comparing the feature set of abnormal data parameters with the average value of the adjustment basis, determine whether the deviation value exceeds the preset range. If the deviation value exceeds the preset range, adjust the data source configuration according to the deviation value, use linear regression analysis to predict the adjusted parameter value to obtain the calibrated configuration parameters, extract the key performance indicators from the calibrated configuration parameters, judge whether the calibrated data source meets the expectation by comparing with the expected value to obtain the judgment result. According to the judgment result, obtain the identifier of the data source that does not meet the expectation, use time series analysis to further detect the change trend of its response time to obtain the optimization direction, and adjust the configuration parameters of the data source through the optimization direction, and loop through the calibration and judgment processes to obtain the final data source configuration that meets the expectation.
7. The method for switching data sources when the xxl-job timed task calls the springboot service interface according to claim 1, characterized in that: The S6 includes: Obtain the calibrated connection information from the data source, verify the information integrity through the verification algorithm to obtain the verified connection data set, use the priority round-robin scheduling algorithm to extract the priority tags from the verified connection data set and load them into the task scheduling engine to obtain the initialized scheduling configuration. Run the dynamic adjustment mechanism through the task scheduling engine, adjust the connection pool size according to the priority tags to generate the adjusted connection pool parameters, extract the scale change data from the adjusted connection pool parameters, use the CSV format conversion tool to generate a structured adjustment record file, obtain the key fields in the adjustment record file, detect the status of the switching process through the log analysis tool, and judge whether the switching process returns to normal. If the switching process returns to normal, extract the performance indicators from the operation log of the scheduling engine, use statistical analysis methods to obtain the evaluation data of system stability, and update the priority rules of the task scheduling engine according to the evaluation data of system stability to generate the optimized scheduling configuration.
8. The method for switching data sources when invoking the springboot service interface based on the xxl-job timed task according to claim 1, wherein: The S7 includes: Obtain the adjustment record data by parsing the CSV format file, extract the multi-data source information according to the preset fields, separate the switching scenario data from the multi-data source information, use time series analysis to determine the state changes of the CPU occupancy rate and memory usage, generate a trend chart for the state change data, calculate the coordinate points of the dynamic switching using the line chart algorithm, obtain the trend chart coordinate point data, and display the content through the visualization component rendering interface. If the switching scenario changes, update the state changes of the CPU occupancy rate and memory usage according to the new scenario data, recalculate the trend chart coordinate points through the updated state changes, and refresh the dynamic switching effect displayed on the interface. Adjust the parameters of the visualization component according to the user interaction record to obtain the optimized trend chart display content.
9. The method for switching data sources when the xxl-job timed task calls the springboot service interface according to claim 1, characterized in that: The S8 includes: Obtain the CPU occupancy rate and memory usage performance metric data from the task runtime through the RESTful API, store it as a time series data set, preprocess the time series data set, calculate the response time of the dynamic switching operation, generate a response time series. If any value in the response time series exceeds the preset threshold, mark it as an anomaly, count the number of anomalies to obtain the anomaly ratio, and calculate the success rate through the anomaly ratio and the total number of switching operations. The success rate formula is: S = (N - E) / N, where S represents the success rate, N represents the total number of switching operations, and E represents the number of anomalies. Obtain the success rate value, use the support vector machine algorithm to classify the response time series and the performance metric data, judge the stability of the dynamic switching mechanism to obtain the classification result, determine whether the dynamic switching mechanism meets the business scenario requirements based on the classification result and the success rate value, generate a judgment result. If the judgment result is not satisfied, optimize the performance metric data by adjusting the resource allocation strategy during the task runtime, and repeat the above steps until a judgment result that meets the requirements is obtained.
10. The method for switching data sources when calling the springboot service interface based on the xxl-job timed task according to claim 1, characterized in that: The S9 includes: Divide the time window by collecting time, calculate the average value within the window to obtain the data fluctuation benchmark, adjust the refresh frequency of the scheduling engine according to the average value to determine the configuration refresh period, use the timeout handling mechanism to monitor the refresh frequency, and trigger an automatic rollback if the threshold is exceeded to restore to the previous stable configuration. Obtain the input data of multiple data sources, allocate resources through the task execution module, judge the execution efficiency, evaluate the stable environment for the task execution status, and adjust the time window range if the fluctuation exceeds the expectation to obtain a new average value. Update the configuration refresh frequency through the new average value to determine the operating parameters of the scheduling engine, obtain the operating log of the scheduling engine, and judge the overall stability of the system by analyzing the timeout records and rollback times in the log.
Citation Information
Patent Citations
Connection pool management method and device, equipment and storage medium
CN119814554A
Multi-task data analysis method and device and storage medium
CN119847752A