ECG anomaly prediction method based on distributed improved whale optimization algorithm
By adopting the distributed improved whale optimization algorithm in ECG abnormality prediction and using the Spark platform and Levy flight and Brownian motion to optimize the random forest model parameters, the accuracy and efficiency problems of traditional ECG prediction methods were solved, and early diagnosis and real-time monitoring of cardiovascular diseases were achieved.
Patent Information
- Application Number
- CN202510808677.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional ECG abnormality prediction methods have problems such as insufficient accuracy, low computational efficiency, and easy falling into local optimality when dealing with high temporal correlation, multi-dimensional morphological features and noise interference, making it difficult to meet the needs of real-time monitoring.
An ECG abnormality prediction method based on the distributed improved whale optimization algorithm was adopted. The Spark platform was used for data processing and preprocessing. The whale optimization algorithm was improved by combining Lévy flight and Brownian motion, and a distributed LBWOA-RF model was constructed. The parameters of the random forest model were optimized to avoid local optimality and improve the global search capability.
It significantly improves the accuracy and computational efficiency of ECG abnormality prediction, shortens training time, enhances the stability and real-time performance of the model, and can provide accurate prediction results in a timely manner, supporting early diagnosis and real-time monitoring of cardiovascular diseases.
Smart Images

Figure CN120708897A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical system technology and relates to an electrocardiogram (ECG) anomaly prediction method based on a distributed improved whale optimization algorithm. The method is suitable for scenarios such as arrhythmia warning and myocardial ischemia prediction in a medical big data environment, providing technical support for early clinical diagnosis and real-time monitoring. Background Art
[0002] In the field of smart healthcare, ECG anomaly prediction has become a core component for early diagnosis and real-time monitoring of heart disease. Accurate ECG anomaly prediction provides critical evidence for clinical decision-making, emergency resource allocation, and personalized treatment, improving the efficiency of early intervention for cardiovascular disease. However, practical ECG prediction scenarios present numerous challenges, such as the high temporal correlation, multi-dimensional morphological characteristics, and noise interference of ECG data. These issues hinder the accuracy and efficiency of ECG anomaly prediction.
[0003] Currently, traditional ECG anomaly prediction methods have exposed numerous limitations when addressing these complex changes. Traditional time series models, such as autoregressive models and autoregressive moving average models, are constructed based on linear assumptions and struggle to accurately capture the nonlinear and complex dynamic features of ECG data. Consequently, they have low recognition rates for anomalies such as ST-segment deviation and T-wave inversion. The Whale Optimization Algorithm (WOA), an emerging representative of swarm intelligence optimization algorithms, is often used to optimize the parameters of various prediction models, such as neural networks and support vector machines, to improve their predictive performance due to its simple principle, small number of parameters, and strong optimization capabilities. However, WOA still has significant drawbacks in ECG prediction applications. First, faced with massive amounts of ECG data, stand-alone algorithms like WOA, due to the concentration of iterative computation on a single node, result in long training times and are unsuitable for real-time monitoring scenarios. Second, WOA is prone to falling into local optima. When processing high-noise ECG data, WOA is prone to falling into local optima due to the concentrated distribution of initial solutions. WOA is particularly limited in its ability to search for anomalies in signals with low signal-to-noise ratios. Finally, WOA's convergence speed is slow. It can't meet the real-time requirements of ECG anomaly prediction. In real-world medical scenarios, failure to obtain accurate ECG prediction results in a timely manner can lead to delayed adjustments to treatment plans and reduced diagnostic efficiency for cardiovascular disease.
[0004] To address the above issues, the present invention utilizes a distributed computing framework, improves upon Lévy flights and Brownian motion, and proposes an ECG anomaly prediction method based on a distributed improved whale optimization algorithm. By leveraging the powerful parallel computing capabilities of the distributed computing framework, the processing of large-scale ECG data is distributed to multiple computing nodes for simultaneous execution, significantly increasing data processing speed and significantly reducing computation time. Furthermore, the Lévy flight method is utilized to optimize the ECG data initialization process, expanding the scope of solution exploration and enhancing global search capabilities. During the position update phase, the principle of Brownian motion is combined to randomly perturb the ECG's changing trend, avoiding local optima in complex ECG prediction scenarios. Summary of the Invention
[0005] In view of this, the technical problem that the present invention needs to solve is to propose an ECG abnormality prediction method based on the distributed improved whale optimization algorithm. The invention can solve the problems of insufficient accuracy, low computational efficiency and the algorithm easily falling into local optimality existing in traditional ECG abnormality prediction methods.
[0006] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0007] An ECG anomaly prediction method based on a distributed improved whale optimization algorithm, including:
[0008] Step 1) Build a distributed environment and data preprocessing module based on the Spark platform, use Spark's distributed computing capabilities to efficiently process ECG-related data, and complete data cleaning, feature engineering, and segmentation;
[0009] Step 2) Construct a Whale Optimization Algorithm combined with Levy flight and Brownian motion (LBWOA) module based on the improvement of Levy flight and Brownian motion. Combined with Levy flight to improve the initial population diversity, and with the help of Brownian motion to prevent the algorithm from falling into the local optimum, the random forest (RF) model parameters are optimized.
[0010] Step 3) A random forest prediction model module based on distributed LBWOA optimization (distributed LBWOA-RF for short) was constructed, the optimized random forest model was applied to ECG abnormality prediction, and the model was evaluated and verified.
[0011] Furthermore, the step 1) specifically includes the following steps:
[0012] Step 1.1) Data Acquisition and Storage: Collect multi-lead ECG signals (e.g., limb leads, chest leads), patient physiological parameters (age, heart rate variability), and diagnostic labels (e.g., normal / ventricular premature beats / atrial fibrillation), and store the data in the Hadoop Distributed File System (HDFS).
[0013] Step 1.2) Data cleaning and preprocessing: Using Spark's distributed computing capabilities, we used polynomial fitting to remove baseline drift, bandpass filtering to suppress myoelectric noise, and interpolation to fill in missing sampling points.
[0014] Step 1.3) Data feature engineering. Leveraging Spark's powerful data processing capabilities, feature engineering is performed to discretize continuous data, such as dividing the ECG signal's amplitude data into intervals. Categorical data is also encoded, such as encoding lead types as digital features. Furthermore, based on medical domain knowledge, new features are constructed, such as calculating the ECG signal's RR interval coefficient of variation and QRS axis deviation, as well as building comprehensive features based on patient age and gender, to enhance the data's ability to capture patterns of ECG abnormalities.
[0015] Step 1.4) Data Partitioning. Using Spark's distributed dataset operations, the preprocessed data is divided into a 70% training set and a 30% test set. The training set is used for model training and parameter optimization, while the test set is used to evaluate model performance and ensure its generalization ability on unseen data.
[0016] Furthermore, the step 2) specifically includes the following steps:
[0017] Step 2.1) Data initialization strategy based on Lévy flight. First, randomly generate the first whale individual, whose position vector represents a set of initial parameter values of the random forest model (including the maximum subtree depth max_depth, the maximum number of candidate features max_features, and the minimum number of samples required to construct a leaf node min_samples_leaf). For each subsequent whale individual, its position is determined based on the position of the previous individual through Lévy flight motion. The Mantegna method is used to simulate the Lévy flight process. The individual X i+1 The position is based on X i The position of , plus an offset determined by u and v. Among them, u and v obey the normal distribution σ with parameters and u and σ v The formula for determining these normal distribution parameters based on the characteristics of the Lévy flight is shown in the following example. In this method, β is set to an empirical value of 1.5. This increases the diversity of the initial population, providing a wider range of starting points for the algorithm to search for optimal solutions.
[0018] Step 2.2) Fitness Calculation and Optimal Solution Update. In each iteration, Spark's parallel computing capabilities are leveraged to perform a distributed calculation of the fitness of the random forest model corresponding to each individual whale on the training set. The fitness function is used to measure the performance of each individual using the random forest model's prediction accuracy on the training set. Based on the calculated fitness value, the optimal solution for the current iteration is updated. Specifically, the random forest model parameter combination representing the whale with the highest fitness value is selected as the current optimal solution.
[0019] Step 2.3) Parameter calculation and position update. Based on the current optimal solution, calculate the parameter A in the WOA and generate a random number p. Depending on the values of A and p, perform spiral update, shrinkage, or search and foraging to adjust the parameter combination. At the same time, combined with the principle of Brownian motion, add random perturbations to the optimal and suboptimal parameter combinations in the current iteration to generate new parameter combinations, replace the poorly performing individuals, and prevent the algorithm from falling into the local optimum. The specific perturbation calculation is based on the Brownian motion expression, and the set is sorted in descending order using the sort function. In this method, σ and t are set to 0.25 and 10, respectively.
[0020] Step 2.4) Iteration termination determination. Check whether the current iteration round has reached the preset maximum number of iterations. If not, continue with the next round of fitness value calculation, parameter calculation, and position update. If so, output the currently found optimal solution, i.e., the optimal random forest model parameter combination.
[0021] Furthermore, the step 3) specifically includes the following steps:
[0022] The optimal random forest model parameter combination obtained in step 2 was applied to the random forest model to construct a random forest ECG abnormality prediction model based on distributed LBWOA optimization. Leveraging Spark's distributed computing capabilities, the training set data was trained in parallel, generating multiple decision trees that were integrated to form the final prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0024] Figure 1 This is the flow chart of the method;
[0025] Figure 2 This is the overall model framework diagram of this method;
[0026] Figure 3 It is the basic framework of Spark in this method;
[0027] Figure 4This is the architecture diagram of the Hadoop distributed file system in this method;
[0028] Figure 5 The Levy flight initialization module constructed for this method;
[0029] Figure 6 The parameter update module of the spiral update method constructed for this method;
[0030] Figure 7 A parameter update module for the shrinking and surrounding method constructed for this method;
[0031] Figure 8 The parameter update module of the search and foraging method constructed for this method;
[0032] Figure 9 The Brownian motion perturbation module constructed for this method;
[0033] Figure 10 The following is a comparison chart of the root mean square error of distributed WOA-RF, traditional RF, and distributed LBWOA-RF.
[0034] Figure 11 Comparison of the mean absolute error between distributed WOA-RF, traditional RF, and distributed LBWOA-RF. DETAILED DESCRIPTION
[0035] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0036] The present invention provides an ECG abnormality prediction method based on a distributed improved whale optimization algorithm. The flow chart and overall model block diagram of the method are as follows: Figure 1 and Figure 2 As shown, the method includes the following steps:
[0037] Step 1) Hardware configuration and parameter initialization;
[0038] Step 2) Initialize the data using the Levi flight initialization module;
[0039] Step 3) Fitness calculation and optimal solution update;
[0040] Step 4) Use modules such as search and foraging, shrinkage and encirclement, and spiral update to update parameters;
[0041] Step 5) using the Brownian motion perturbation module to randomly perturb the parameter combination;
[0042] Step 6) Iterate control and output the optimal parameters.
[0043] Further, step 1) specifically includes the following steps:
[0044] Step 1.1) Use the Spark distributed computing framework to build a cluster consisting of one master node and multiple worker nodes. The master node is responsible for resource scheduling, while the worker nodes execute computing tasks in parallel, supporting distributed data storage and parallel processing. Configure HDFS to store large amounts of ECG data.
[0045] In step 1.2), set the parameter combination to 30 and the maximum number of iterations to 100.
[0046] Further, step 2) specifically includes the following steps:
[0047] Step 2.1) Randomly generate the first individual X0.
[0048] Step 2.2) Figure 5 As shown, the position of each individual is based on the position of the previous individual to perform Lévy flight motion, as shown in formula (1):
[0049]
[0050] Where: l is the step length control variable; Levy(λ) is the random search path, which satisfies formula (2):
[0051] Levy(λ)~u=i -λ 1<λ≤3 (2)
[0052] Since the Lévy flight described by formulas (1) and (2) is too complex and difficult to implement in real scenarios, the Mantegna method is often used to simulate the Lévy flight process. i+1 The position calculation process is shown in formula (3):
[0053]
[0054] Where u and v are respectively subject to the parameters σ u and σ v Normal distribution. u and σ v The definition of is shown in formula (4):
[0055]
[0056] In this method, β in formula (4) is set to an empirical value of 1.5.
[0057] Further, step 3) specifically includes the following steps:
[0058] Step 3.1) Calculate the fitness value of the model for each parameter combination: At the beginning of each iteration, this module evaluates the fitness value of each parameter combination. These values are calculated based on the fitness function and are used to measure the performance of an individual on a given problem. The root mean square error (RMSE) is a statistic used to measure the deviation between the predicted and actual results, which can provide a more accurate model evaluation result. The lower the value of this metric, the closer the predicted value is to the actual value, and the more accurate the model prediction result. The RMSE calculation formula is as follows:
[0059]
[0060] Among them, p i Represents the predicted value of the i-th test sample, y i Represents the actual true value of the i-th test sample.
[0061] Mean Absolute Error (MAE) is a statistic used to measure the absolute difference between the predicted value and the actual value. It can objectively grasp the mean square error of the model's prediction results and is not sensitive to the value. This indicator reflects the quality of the model's prediction ability. The smaller the value, the more accurate the model. The MAE calculation formula is as follows:
[0062]
[0063] Among them, p i Represents the predicted value of the i-th test sample, y i Represents the actual true value of the i-th test sample.
[0064] Step 3.2) In the Spark cluster, partition the training data to each Worker node, calculate the fitness value of each individual model in parallel, and obtain the global optimal fitness value through aggregation operation.
[0065] Further, step 4) specifically includes the following steps:
[0066] Step 4.1) Check if the current round i is less than n i , which is the key step of iterative control. If the current round i is less than n i , this module will continue to execute subsequent steps; otherwise, this module will stop iterating and output the optimal parameter combination.
[0067] Step 4.2) In each iteration, this module calculates the parameter A and generates a random number p based on the current optimal combination. A is a coefficient vector, and its calculation process is shown in formula (6):
[0068]
[0069] Among them, A is used to control the search direction and step size, and p is used to determine whether to adopt the spiral update or the shrinking enclosure strategy.
[0070] Step 4.3) As Figure 6 shown, when p < 0.5, the spiral update is adopted. The spiral update method mainly approaches the current optimal solution in a spiral manner. The optimal solution here represents the parameter combination that can minimize the prediction error of the random forest model for ECG abnormalities.
[0071]
[0072] Among them, the parameter b is the constant coefficient of the logarithmic spiral, l is a random number between [-1, 1], and D e represents the degree of difference between the current parameter combination of the random forest model and the optimal parameter combination.
[0073] Step 4.4) As [[ID=I9]] Figure 7 shown, when p ≥ 0.5 and |A| < 1, the shrinking enclosure is adopted. Under the shrinking enclosure method, as the iteration progresses, the algorithm can continuously improve the model parameters, reduce the prediction error of the model for ECG abnormalities, and improve the prediction accuracy. The specific position update is shown in formula (8):
[0074]
[0075] Among them D [[ID=I9]] m represents the distance between the current individual and the optimal individual after random coefficient adjustment.
[0076] Step 4.5) As Figure 8 shown, when p ≥ 0.5 and |A| ≥ 1, the search foraging is adopted. In the t-th iteration, the i-th model parameter combination is expressed as According to another randomly selected parameter combination it is updated. In the (t + 1)-th iteration of this stage, the position of the i-th best individual can be calculated by formula (9):
[0077]
[0078] Among them: is the position of a randomly selected parameter combination in the t-th iteration; A is the coefficient vector: D r is the random distance between the current parameter combination i and the randomly selected parameter combination r in the current round, and its calculation process is shown in formula (10):
[0079]
[0080] Here: r1 and r2 are random vectors with component values ranging from [0,1]; E represents the identity matrix; t max is the maximum number of iterations; Indicates a method of operation, such as
[0081] From formula (8), we can see that as the number of iterations t increases, the value of parameter a decreases linearly from 2 to 0. Correspondingly, the expected value of the absolute value of the coefficient vector A also decreases linearly from 2 to 0. In formula (9), the components of C take random numbers between [0, 2] to control the randomly selected parameter combination. For the current parameter combination The impact of location updates.
[0082] Further, step 5) specifically includes the following steps:
[0083] Step 5.1) Figure 9 As shown in the figure, LBWOA introduces Brownian motion to perturb the random forest parameter combination, avoiding the algorithm from falling into the local optimum and improving the global optimization ability. The specific method is to perform {X i}Sort in descending order by fitness, recorded as {Y k}=sort({X i}). Among them, the sorted and are the optimal and suboptimal combinations respectively; and It is the combination of the worst and the second worst.
[0084] Step 5.2) Brownian motion is achieved by the function f(x,σ,t)~N(0,σ 2 t) to simulate, where N(0,σ 2 t) means the mean is 0 and the variance is σ 2 Normal distribution of t. In this method, σ=0.25, t=10. The optimal parameter combination and suboptimal combinations Apply Brownian motion perturbation, the calculation formula is as follows:
[0085]
[0086]
[0087] Further, step 6) specifically includes the following steps:
[0088] The RMSE and MAE of distributed WOA-RF, traditional RF, and distributed LBWOA-RF are compared. The comparison of the RMSE and MAE values of the three models is shown in the figure. Figure 10 , Figure 11 shown.
[0089] analyze Figure 10 , Figure 11 As can be seen, compared with distributed WOA-RF and traditional RF, distributed LBWOA-RF performs better in terms of RMSE and MAE, with a smaller error dispersion range and smaller absolute error. The fluctuations reflected in the curves show that the fluctuations in both indicators of distributed LBWOA-RF are smaller, indicating greater model stability.
[0090] Finally, it should be noted that the aforementioned implementation examples are intended only to illustrate the technical solutions of the present invention and are not intended to impose any restrictions thereon. Although the present invention has been comprehensively and meticulously illustrated through the aforementioned implementation examples, those skilled in the art will appreciate that various modifications may be made to its form and specific details without departing from the scope of the present invention as defined by the claims.
Claims
1. An ECG abnormality prediction method based on a distributed improved whale optimization algorithm, characterized in that: The following steps are involved: Step 1: This method was implemented using Python 3.7.7 on a hardware platform equipped with a 2.2GHz quad-core Intel Core i7 processor and 16GB of 1600MHz DDR3 memory. A cluster was constructed using the Spark distributed computing framework, consisting of one master node and multiple worker nodes. The Hadoop Distributed File System (HDFS) was configured to store large-scale electrocardiogram (ECG) data. Parameters were first set to ensure the accuracy and stability of the results: the number of parameter combinations was set to 30, meaning that in each iteration, the method would simultaneously process 30 potential combinations. Furthermore, the maximum number of iterations was limited to 100 to ensure that the method converged to the optimal combination within a reasonable timeframe. To minimize the impact of random error on the final results, a replicated experiment strategy was employed: each set of experiments was independently run 10 times, and the results were statistically analyzed to obtain more reliable and stable conclusions. In step 2, during the initialization phase, this method randomly generates the first whale individual. Its position vector represents a set of initial parameter values for the Random Forest (RF) model (including the maximum subtree depth (max_depth), the maximum number of candidate features (max_features), and the minimum number of leaf nodes (min_samples_leaf). Subsequently, a Lévy flight initialization module is used to generate each subsequent individual. This module determines the position of subsequent individuals through Lévy flights, a heavy-tailed random walk process that occasionally produces very large step sizes, helping the method explore new areas in the search space. To simplify the implementation of Lévy flights, this method uses the Mantegna method to simulate the Lévy flight process, in which the position of an individual is calculated based on the position of the previous individual, plus an offset determined by the parameters of a specific normal distribution. In this method, one of the parameters of the normal distribution is set to an empirical value of 1.5 to increase the diversity of the initial population and provide a broad starting point for the algorithm to search for optimal solutions. Step 3, at the beginning of each iteration, the method uses the parallel computing ability of Spark to distributively calculate the fitness values of the random forest models corresponding to each whale individual on the ECG training set. These values are obtained through the fitness function, which is used to measure the quality of an individual in a given problem, and based on this, the optimal solution in the current iteration is updated. That is, the parameter combination of the random forest model represented by the whale individual with the highest fitness value is selected as the current optimal solution. Then, the method checks whether the current iteration round is less than the preset maximum number of iterations to decide whether to continue the iteration. If the iteration condition is met, parameters are calculated based on the current optimal solution and a random number is generated. These parameters will jointly determine the subsequent update strategy, aiming to find a better solution in the subsequent iterations. If the iteration condition is not met, the iteration is stopped and the currently found optimal solution is output as the final result. Step 4, during the execution of the method, different update strategies are adopted according to specific conditions to find the optimal solution. When p < 0.5, the spiral update method parameter combination update module proposed by this method is adopted, mainly for fine search near the optimal solution to optimize the model parameters. This module combines the constant coefficient of the logarithmic spiral, a random number in the [-1, 1] interval, and the difference degree between the current model parameter combination and the optimal parameter combination, realizing a fine search near the optimal solution. When p ≥ 0.5, |A| < 1, the contraction and enclosure method parameter combination update module proposed by this method is adopted. Under the contraction and enclosure method, as the iteration progresses, the algorithm can continuously improve the model parameters, reduce the prediction error of the model for ECG abnormalities, and improve the prediction accuracy. When p ≥ 0.5, |A| ≥ 1, the search and foraging method parameter combination update module proposed by this method is executed. In this module, each parameter combination is updated according to another randomly selected parameter combination. The update formula combines the position of a randomly selected parameter combination, the coefficient vector, and the random distance between parameter combinations. This random distance is calculated by search and foraging, involving parameters such as random vectors, identity matrices, and the maximum number of iterations. And as the number of iterations increases, the absolute value expectation of the coefficient vector linearly decreases to 0, which means that in the initial stage of iteration, the parameter combinations of the model have a larger movement range and can explore new parameter combinations in a wider area; while in the later stage of iteration, the movement range gradually shrinks, and the algorithm will focus more on the optimization of the local area. At the same time, by controlling the components of C (the components of C take random numbers between [0, 2]), the influence degree of the randomly selected parameter combination on the position update of the current parameter combination is adjusted, so as to actively search for new parameter combinations in the entire search space, maintain the diversity of the population, and increase the possibility of finding the global optimal solution. In step 5, the Brownian motion perturbation module, based on the improved Whale Optimization Algorithm combined with Levy flight and Brownian motion (LBWOA), employs a Brownian motion-based parameter update strategy. This module randomly perturbs the optimal and suboptimal model combinations in the current iteration. This module optimizes the search process by replacing the worst and second-worst parameter combinations. When executing this strategy, all parameter combinations are first sorted in descending order of fitness using the sort function to determine the optimal and suboptimal parameter combinations. To ensure the stability and effectiveness of the strategy, the relevant parameters are set to fixed values. In step 6, comparative experiments verified the performance. Compared with distributed WOA-RF and traditional RF, distributed LBWOA-RF outperformed the RMSE and MAE metrics, with a smaller error dispersion range and smaller absolute error. Specifically, the fluctuation amplitude of the two metric curves of distributed LBWOA-RF was smaller, indicating better model stability. Furthermore, LBWOA enhances its global optimal solution search capability and efficiency by introducing Brownian motion and Lévy flight methods.
2. The ECG abnormality prediction method based on the distributed improved whale optimization algorithm according to claim 1 is characterized in that: In the Lévy flight initialization module constructed in step 2, during the parameter iteration process of the optimization method, the position of each subsequent individual is determined based on the position of the previous individual by simulating the Lévy flight motion. A Lévy flight is a random walk process with unique properties. Its step size follows the Lévy distribution. This distribution is characterized by its heavy tail, which occasionally produces very large step sizes. This property helps the method effectively explore new potential areas in a broad search space. However, directly applying the Lévy distribution to generate step sizes is quite complex because the mathematical description of the Lévy distribution is quite cumbersome and difficult to directly sample. To simplify the implementation of the Lévy flight, this method uses the Mantegna method to simulate this process. In the Mantegna method, the position of the subsequent individual is calculated based on the position of the previous individual plus an offset determined by specific parameters. This offset is composed of two random numbers from a normal distribution whose parameters are determined based on the characteristics of the Lévy flight. This method effectively simulates the Lévy flight process, thereby leveraging this property in the optimization method to enhance global search capabilities.
3. The ECG abnormality prediction method based on the distributed improved whale optimization algorithm according to claim 1 is characterized in that: In steps 3 and 4, the fitness of the model parameter combination is evaluated to update the optimal solution. Parameters are calculated based on the current optimal solution to determine the next update module. These update modules include a spiral update parameter combination update module, a shrinking and encircling parameter combination update module, and a search and foraging parameter combination update module. In each iteration, Spark's parallel computing capabilities are leveraged to perform a distributed calculation of the fitness of each individual whale model on the training set. Iterative control is then performed to check whether the current iteration has reached the preset maximum number of iterations. If not, the iteration continues and the above process repeats. If so, the iteration stops and the currently found optimal solution is output. This solution, obtained through continuous evaluation and updating, represents the best parameter combination found within the current number of iterations.
4. The ECG abnormality prediction method based on the distributed improved whale optimization algorithm according to claim 1, characterized in that: The Brownian motion perturbation module constructed in step 5 is designed to enhance the method's global search capabilities and avoid trapping in local optima. Specifically, this method introduces Brownian motion to randomly perturb the positions of the optimal and suboptimal parameter combinations in the current iteration. This strategy is implemented by replacing the positions of the worst and second-worst individuals, thereby promoting diversity in parameter combinations. During implementation, the parameter combinations are first sorted in descending order of fitness using the sort function to determine the optimal and second-best individuals in the current iteration. The perturbation amount is then calculated based on the mathematical expression of Brownian motion. This process relies on fixed parameters to ensure the stability and repeatability of the method. By applying the calculated perturbation amount to the positions of the optimal and second-best individuals, new parameter combinations are generated as candidate solutions, replacing the worst and second-worst performing individuals, effectively improving the algorithm's global search capabilities. This update strategy not only preserves information about excellent individuals but also introduces new search directions through random perturbations, effectively enhancing the method's global search capabilities.