Construction and Application of an Esophageal Cancer Survival Prediction Model Based on an Improved Salp Swarm Algorithm

By improving the squid algorithm to optimize the BP neural network model, analyze the death factors of esophageal cancer patients, and build a survival prediction model, the error and limitations of esophageal cancer survival prediction in the existing technology are solved, and higher prediction accuracy and stability are achieved, helping to improve the survival rate of patients.

CN115662642BActive Publication Date: 2025-06-10ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211300762.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-23
Publication Date
2025-06-10
Estimated Expiration
2042-10-23

AI Technical Summary

Technical Problem

The prior art has errors and limitations in the prediction of survival of esophageal cancer patients, resulting in limited treatment methods and drug selection. The traditional detection methods are highly invasive and costly, and cannot effectively predict survival.

Method used

The BP neural network model based on the improved cassia algorithm is adopted to analyze the factors affecting death of esophageal cancer patients, optimize the initial weight and threshold, and build a survival prediction model. The algorithm exploration and mining capabilities are improved by using strategies such as division iteration strategies, leader fusion mutation strategies, optimal individual leading movement, dimensional Gaussian mutation and optimal domain perturbation.

Benefits of technology

It improves the accuracy and stability of the survival prediction model of esophageal cancer, reduces errors, and helps clinicians formulate personalized treatment plans and improves the survival rate of patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115662642B_ABST
    Figure CN115662642B_ABST
Patent Text Reader

Abstract

The present application discloses the construction and application of an esophageal cancer survival period prediction model based on an improved salp swarm algorithm. The present application uses a partitioned iteration strategy to divide the iteration of the algorithm into two periods, namely the early iteration period and the late iteration period; in the early iteration period, a leader fusion mutation strategy is used to update the leader position to increase the exploration ability of the algorithm; in the late iteration period, an optimal individual leading movement is adopted to update the leader position to optimize the exploitation ability of the algorithm; then a dimension-by-dimension Gaussian mutation strategy is used to update all salp swarm individuals to improve the diversity of the population; before the end of the algorithm iteration, the optimal salp position is subjected to an optimal neighborhood perturbation and the fitness values are compared, and a greedy strategy is used to select the optimal fitness value to improve its ability to jump out of the local optimal solution. The IPSSA-BP model of the present application has a good fitting effect, high accuracy, high prediction precision, and good prediction stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cancer survival prediction, and particularly relates to the construction and application of an esophageal cancer survival prediction model based on an improved salp swarm algorithm. Background Art

[0002] With the continuous progress of medical technology, the comprehensive treatment mode of esophageal cancer with surgical treatment as the core has brought good results to esophageal cancer patients. However, due to the complexity of esophageal cancer surgery and pathology, there are many postoperative complications of esophageal cancer, and then the 5-year survival rate of postoperative patients is only 10% to 30%. Then, the 5-year survival rate of early esophageal cancer patients after comprehensive treatment is higher than 70%. Also, because there are inevitably some errors and limitations in the prognosis inference of patients by artificial diagnosis and traditional statistical methods, the selection of treatment methods and drug types is restricted. Therefore, it is crucial to predict the survival prognosis of esophageal cancer patients in a timely and effective manner to improve the survival rate of patients. By analyzing the clinical data of the survival of esophageal cancer patients and constructing a prediction model to predict the survival period of esophageal cancer patients, it can help clinicians timely discover the survival prognosis of patients, and then be able to take better assistance in the treatment of patients, thereby improving the prognosis of esophageal cancer patients and further increasing the survival rate of esophageal cancer patients.

[0003] Although in the past few decades, with the introduction of new drugs and new technologies, the survival period of esophageal cancer patients has been significantly improved, but due to the complexity of pathology, there are certain errors in the artificial judgment of risk levels, resulting in esophageal cancer patients not being able to obtain appropriate treatment methods. To solve this problem, it is necessary to design a model that can predict the survival period of esophageal cancer patients.

[0004] The selection of traditional cancer treatment methods is based on the "gold standard" method, including three tests: clinical examination, radiological imaging, and pathological examination, and the choice of which treatment method to adopt is determined by the clinical experience and professional knowledge of doctors. The detection of traditional methods is invasive, which will bring physical discomfort and pain to the examined population. Secondly, the cost is high, and it is not suitable for large-scale promotion and use. At the same time, the examination results can only be used to prove the existence of cancer and cannot judge the risk level and survival period.

[0005] To solve the above problems, statistical analysis methods are used to predict the risk level of cancer patients. Statistical analysis methods only require some easily obtainable clinical characteristic indicators of patients and can be used to infer the relationships between variables. Meanwhile, they are highly efficient and low-cost. Commonly used statistical analysis methods include Kaplan-Meier (KM), Cox, survival tree, Lasso, etc. KM is a univariate analysis method for estimating survival probability from observed survival time. This method only describes the relationship between a single variable and survival while ignoring the influence of other variables and can intuitively show the survival rate or mortality rate of two or more groups. The survival tree method is applicable to large cohort data with many variables. It is difficult to meet the application conditions of classical survival analysis methods, and the results are presented in a tree structure diagram, which is more intuitive and easier to understand and interpret. Although statistical analysis methods can use the easily obtainable clinical data of patients and analyze the relationships between variables relatively quickly, statistical analysis has high requirements for the integrity and accuracy of historical statistical data, and its accuracy and reliability are poor when dealing with relatively complex data.

[0006] Compared with statistical models, machine learning has shown advantages in dealing with the complexity of large-scale data and discovering prognostic factors. Its learning process can generally be divided into four stages: data collection, data preprocessing, model training and prediction, and model evaluation. Machine learning technology has an absolute advantage over statistical analysis when facing datasets with a large quantity and high dimension. Through comprehensive collection of patient data, the data can be analyzed and utilized, and machine learning methods can be used to further explore the internal correlations between the data, thereby constructing a prognostic index and finally establishing a survival prediction model. The survival prediction model can help clinicians formulate targeted individualized treatment plans and better drug selections based on the survival prognosis of patients.

[0007] The information disclosed in this background art section is only used to deepen the understanding of the background art of the present disclosure and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention

[0008] In view of the problems of unbalanced exploration ability and exploitation ability of the salp swarm algorithm, slow convergence speed, low convergence accuracy, and easy to fall into local optimum, the inventor of this application proposes a salp swarm optimization algorithm based on a partitioning iteration strategy: using the partitioning iteration strategy to divide the iteration of the algorithm into two periods, the early iteration period and the late iteration period; in the early iteration period, use the leader fusion mutation strategy to update the leader position to increase the exploration ability of the algorithm; in the late iteration period, adopt the optimal individual leading movement to update the leader position to optimize the exploitation ability of the algorithm; then use the dimension-by-dimension Gaussian mutation strategy to update all salp individuals to improve the diversity of the population; before the end of the algorithm iteration, perform an optimal neighborhood perturbation on the optimal salp position and compare the fitness values, and use the greedy strategy to select the optimal fitness value to improve its ability to jump out of the local optimal solution. The algorithm is used to optimize the esophageal cancer survival prediction model and compared with other six models. The mean absolute error (MAE), mean square error (MSE), and mean absolute percentage error (MAPE) are used as evaluation criteria to analyze the model. The results show the effectiveness of the proposed model.

[0009] According to one aspect of the present disclosure, there is provided a method for constructing an esophageal cancer survival prediction model based on an improved salp swarm algorithm, including:

[0010] (1) Analyze and screen out the influencing factors of the death of esophageal cancer patients, use the screened factors as input variables, and the survival time as the output variable to establish a survival prediction model based on a BP neural network;

[0011] (2) Use the improved salp swarm algorithm to optimize the initial weights and thresholds of the BP neural network. The improved salp swarm algorithm includes the following steps:

[0012] S1: Set the relevant parameters of the salp swarm algorithm: population size N, maximum number of iterations L, dimension Dim of individuals, and initialize the population;

[0013] S2: Calculate the fitness value of each salp individual according to the objective function, and sort the fitness values; select the position with the best fitness as the position of the food source, and set the current iteration number l = 1;

[0014] S3: Judge whether the current iteration number l is less than the partitioning coefficient. If it is less, go to step S4; if it is greater, go to step S5;

[0015] S4: Use the leader fusion mutation strategy to update the leader position, and introduce an adaptive inertia weight when updating the follower position;

[0016] S5: Update the leader's position using the optimal individual to lead the motion strategy, and introduce an adaptive inertia weight when updating the followers' positions;

[0017] S6: Calculate the fitness value and update the positions of the salps. Update the positions of all salps using a dimension-by-dimension Gaussian mutation strategy;

[0018] S7: Update the position of the optimal salp using the optimal neighborhood perturbation strategy, compare it with the current fitness value, and select the optimal fitness value using a greedy strategy;

[0019] S8: If the current iteration number l is less than the maximum iteration number L, increment the current iteration l by 1 and go to step S2; otherwise, output the optimal solution and end the algorithm iteration.

[0020] In some embodiments of the present disclosure, in the step S2, the objective function expression is:

[0021]

[0022] where X is the number of training samples of the esophageal cancer survival model, y i is the output value of the network model of the i-th esophageal cancer patient, y' i is the actual output value of the i-th esophageal cancer patient.

[0023] In some embodiments of the present disclosure, in the step S4, the leader position update process is as follows:

[0024]

[0025] a + b + c = 1 ③;

[0026] In the formula, represents the position of the first salp individual (leader) in the j-th dimension; are two randomly selected salp individuals in the population; F j is the food source position in the j-th dimension; a, b, c are random numbers, taking values in the range [0, 1].

[0027] In some embodiments of the present disclosure, in the step S5, the leader position is updated according to the following formula:

[0028]

[0029] In the formula, F j is the food source position in the j-th dimension, R 2 R 3 is a random number in the range [0, 1]; R is a random number in the range [-1, 1], and the calculation formula of B is as follows:

[0030]

[0031] Where \(l\) is the current iteration number and \(L\) is the maximum iteration number.

[0032] In some embodiments of the present disclosure, in the step S6, all salp positions are updated according to the following formula:

[0033] \(q = 1-(l - L)\) 2 ⑥;

[0034] \(x(j)=q\times x(j)+randn\times x(j)\) ⑦;

[0035] Where \(q\) is the adaptive inertia weight; \(l\) is the current iteration number; \(L\) is the maximum iteration number; \(x(j)\) is the position of the \(j\)-th dimensional salp; \(randn\) is the Gaussian mutation operator.

[0036] In some embodiments of the present disclosure, in the step S7, the optimal salp position is updated using the following formula:

[0037]

[0038] Where \(x(best)\) is the optimal position during global update; \(x(newbest)\) is the newly generated position; where \(v\) and \(m\) are random numbers between \([0, 1]\).

[0039] In some embodiments of the present disclosure, in the step S7, for the generated neighborhood positions, the following formula is used to determine whether to retain them:

[0040]

[0041] Where \(f(x(best))\) is the optimal position fitness value during global update. If the generated position is better than the original position, it will replace the original position to make it the global optimum; otherwise, the optimal position remains unchanged.

[0042] In some embodiments of the present disclosure, in the step (1), five blood indexes related to the survival period of esophageal cancer patients are screened by single-factor COX regression analysis: white blood cell count, monocyte count, neutrophil count, prothrombin time, and international normalized ratio as inputs. The structure of the survival period model of the BP neural network constructed is 5-11-1, and the optimized dimension is:

[0043] \(Dim=(inputnum + 1)\times hiddennum+(hiddennum + 1)\times outputnum\) ⑩;

[0044] Among them, inputnum, hiddennum, and outputnum are the number of input layers, hidden layers, and output layers of the BP neural network, respectively.

[0045] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0046] 1. First, a partitioning iteration strategy balance algorithm is adopted to balance the exploration ability and exploitation ability of the algorithm. The iteration of the algorithm is divided into two periods using a partitioning coefficient. In the early stage of iteration, a leader fusion mutation strategy with stronger exploration ability is used to update the leader position; in the later stage of iteration, an optimal individual leading movement with stronger exploitation ability is used to update the leader position. Secondly, all salp individuals are subjected to one-dimensional Gaussian mutation to enrich the diversity of the population. Then, an optimal neighborhood perturbation strategy is used to update the optimal salp position, and a greedy strategy is used to obtain the optimal solution, avoiding the premature convergence of the algorithm and improving the ability of the algorithm to jump out of the local optimum.

[0047] 2. The IPSSA algorithm adopted has better convergence speed and convergence accuracy compared with different swarm intelligence algorithms and different improved salp algorithms while ensuring stability, and has a certain competitive advantage.

[0048] 3. The IPSSA-BP model of the present application has a good fitting effect, high accuracy, high prediction accuracy, and good prediction stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a flowchart of an improved salp algorithm in an embodiment of the present application.

[0050] Figure 2 It is a convergence curve of the improved salp algorithm and other swarm intelligence algorithms in an embodiment of the present application.

[0051] Figure 3 It is a convergence curve graph of the salp algorithm and other improved salp algorithms in an embodiment of the present application.

[0052] Figure 4 It is a comparison graph of the expected value and the actual value of the test set of the esophageal cancer survival period prediction model in an embodiment of the present application.

[0053] Figure 5 It is a comparison graph of the simulation evaluation criteria of 7 models in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to better understand the technical solutions of the present application, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0055] Example 1: Salp Swarm Optimization Algorithm Based on Partition Iteration Strategy

[0056] For swarm intelligence algorithms, balancing the exploration and exploitation capabilities of the algorithm during the search process can effectively improve the performance of the algorithm. According to the characteristics of the salp swarm algorithm, this example introduces a partition iteration strategy to balance the exploration and exploitation capabilities of the algorithm, and uses a dimension-by-dimension Gaussian mutation strategy to improve the diversity of the population. Finally, the ability of the algorithm to jump out of local optima is improved through optimal neighborhood perturbation and greedy strategy.

[0057] 1. Partition Iteration Strategy

[0058] The partition iteration strategy is a strategy for balancing algorithm exploration and exploitation. It divides the iteration of the algorithm into different periods, and different search equations are used for solution according to the characteristics of different periods. In the original salp swarm algorithm, more exploration capabilities are required in the early stage of the algorithm, and more exploitation capabilities are required in the later stage of the algorithm. In this example, the algorithm is partitioned and iterated into two periods, the early iteration period and the late iteration period, using a partition coefficient. In the early iteration period, an equation with strong exploration ability is used, and in the late iteration period, an algorithm equation with strong exploitation ability is used. In this way, the exploration and balancing capabilities of the algorithm can be balanced. In this example, the leader fusion mutation strategy is selected for update in the early iteration period, and an adaptive inertia weight is introduced into the follower position to improve the exploitation ability of the algorithm. In the late iteration period, the optimal individual is used to lead the movement to update the leader position, and an inertia weight is introduced into the follower position. Table 1 shows the pseudocode of the partition iteration strategy, and the setting of the partition coefficient is as follows:

[0059] L h = L / 2 (1);

[0060] In the formula, L is the maximum number of iterations. In this example, the iteration is divided into two periods with equal numbers of iterations.

[0061] 1.1 Leader Fusion Mutation Strategy

[0062] In the original salp swarm algorithm, the position of the leader is mainly guided by the food source, and the food source position is updated in each iteration, which is likely to fall into local extrema, resulting in a decline in the exploration ability of the algorithm. To improve this problem, this example introduces a leader fusion mutation strategy. Two salp individuals are randomly selected from all salp individuals, and the food source position and the positions of the two randomly selected salp individuals are fused and mutated to generate a new leader position, thereby accelerating the update of the leader position and affecting the search directions of the remaining individuals. The updated leader position is affected not only by the food source position but also by random individuals, thereby accelerating the convergence speed and exploration ability of the algorithm. The specific update process is as follows:

[0063]

[0064] a + b + c = 1 (3);

[0065] In the formula, represents the position of the first salp individual (leader) in the j-th dimension; are two salp individuals randomly selected from the population, F j is the food source position in the j-th dimension; a, b, and c are random numbers, and in this paper, their values are set to be between [0, 1].

[0066] 1.2 Optimal individual-guided movement

[0067] In the later stage of the algorithm iteration, the search of the algorithm requires more exploitation ability. In this paper, the optimal individual leadership-guided movement is introduced. The team needs to move towards a target area (the best area), so the leaders need to update their positions closer to the target position (the best area). While the followers move closer to the leaders, the leaders also need to continuously approach the optimal area. The leader position update can find the best position near the food source according to formula (4). Sometimes, the leaders must leave the current best position in order to find a better position, which improves the exploitation ability of the algorithm. The specific formula update is as follows:

[0068]

[0069] In the formula, F j is the food source position in the j-th dimension, R 2 R 3 is a random number in the interval [0, 1]; R is a random number in the interval [-1, 1]. The calculation formula of B is as follows:

[0070]

[0071] In the formula, l is the current iteration number, and L is the maximum iteration number. 2×R 3 generates more random movements, so the algorithm will not fall into local optimum, which means that the algorithm is also exploring during the exploitation stage. cos(2Rπ) searches for the best individuals with different radii to find better positions around the individuals.

[0072] 1.3 Follower position update

[0073] In the salp swarm algorithm, the followers move following the leaders. The followers update their positions according to Newton's laws of motion, and the formula is

[0074]

[0075] is the position of the i-th follower in the j-th dimension, i…2, a is the acceleration, v 0is the initial velocity, and t is the time; since the number of iterations in the algorithm is the time, the difference between each iteration is 1, the initial velocity is 0, and there is an inertia weight w(l) that changes with the number of iterations during each update, the movement formula update of the follower is as follows:

[0076]

[0077]

[0078] In the formula, l is the current number of iterations, and L is the maximum number of iterations. is the position of the i-th follower in the j-th dimension, is the position of the (i - 1)-th follower in the j-th dimension.

[0079] Table 1 Partition Iteration Strategy Pseudo-Code

[0080]

[0081] 2. Dimension-by-Dimension Gaussian Mutation

[0082] The Gaussian mutation strategy uses random numbers obeying the normal distribution to act on the original position vector to generate new positions. Most mutation operators are distributed around the original position, which is equivalent to performing a neighborhood search within a small range. This mutation not only improves the optimization accuracy of the optimization algorithm but also helps the algorithm jump out of the local optimal region. At the same time, a few operators are far from the current position, enhancing the diversity of the population and facilitating better search for potential regions, thereby improving the search speed and accelerating the convergence trend of the optimization algorithm. The dimension-by-dimension Gaussian mutation is performed on all salp individuals using the Gaussian mutation strategy to increase the diversity of the population. Table 2 shows the pseudo-code of the dimension-by-dimension Gaussian mutation strategy.

[0083] To better adjust the global exploration ability and local exploitation ability of the algorithm, an adaptive inertia weight q is introduced while performing Gaussian mutation. When the value of the adaptive inertia weight q is large, the exploration ability of the algorithm is strong, and when the value of the inertia weight is small, the exploitation ability of the algorithm is strong. The specific formula update is as follows:

[0084] q = 1 - (l - L) 2 (9);

[0085] x(j) = q × x(j) + randn × x(j) (10);

[0086] In the formula, l is the current number of iterations, L is the maximum number of iterations; x(j) is the position of the j-th salp; randn is the Gaussian mutation operator.

[0087] Table 2 Dimension-by-Dimension Gaussian Mutation Pseudo-Code

[0088]

[0089] 3. Optimal domain perturbation

[0090] When the salp swarm algorithm updates the position, it takes the current optimal position as the target of this iteration. In each iteration process, the optimal position is updated only when a position better than it appears, which results in not many updates of the optimal position before meeting the end condition, leading to low search efficiency of the algorithm. Therefore, in this example, an optimal domain perturbation strategy is introduced to perform random search near the optimal position to find a better optimal solution, which can not only improve the convergence speed of the algorithm, but also avoid the algorithm from premature convergence and improve the ability of the algorithm to jump out of the local optimum. Table 3 shows the pseudocode of the optimal domain perturbation.

[0091]

[0092] In the formula, x(best) is the optimal position during global update; x(newbest) is the generated new position; where v and m are random numbers between [0, 1] respectively. For the generated neighborhood positions, a greedy strategy is adopted

[27] to judge whether to retain them, and the formula is as follows:

[0093]

[0094] In the formula, f(x(best)) is the fitness value of the optimal position during global update. If the generated position is better than the original position, it will replace the original position with it to make it the global optimum. Otherwise, the optimal position remains unchanged.

[0095] Table 3 Pseudocode of the optimal domain perturbation

[0096]

[0097] 4. Algorithm flow

[0098] The flow chart of the salp swarm optimization algorithm based on the partition iteration strategy is as Figure 1 shown, and the detailed steps are as follows:

[0099] S1: Set the relevant parameters of the salp swarm algorithm: population size N, maximum iteration number L, dimension Dim of individuals, and initialize the population.

[0100] S2: Calculate the fitness value of each salp individual according to the objective function, and sort the fitness values. Select the position with the best fitness as the position of the food source, and set the current iteration number l = 1.

[0101] S3: Judge whether the current iteration number l is less than the partition coefficient. If it is less, go to step S4; if it is greater, go to step S5.

[0102] S4: Update the leader's position using the leader fusion mutation strategy, and introduce an adaptive inertia weight when updating the followers' positions.

[0103] S5: Update the leader's position using the optimal individual leading motion strategy, and introduce an adaptive inertia weight when updating the followers' positions.

[0104] S6: Calculate the fitness value and update the positions of the salps, and update the positions of all salps using the per-dimensional Gaussian mutation strategy.

[0105] S7: Update the position of the optimal salp using the optimal neighborhood perturbation strategy, compare it with the current fitness value, and select the optimal fitness value using the greedy strategy.

[0106] S8: If the current iteration number l is less than the maximum iteration number L, increment the current iteration l by 1 and go to step S2; otherwise, output the optimal solution and end the algorithm iteration.

[0107] 5. Function testing experiments

[0108] (1) Comparison with the original algorithm and different swarm intelligence algorithms

[0109] Compare the IPSSA algorithm in this example with the basic SSA algorithm and the selected ant lion algorithm (ALO), grey wolf algorithm (GWO), dragonfly algorithm (DA), and moth-flame optimization algorithm (MFO) with better convergence effects for optimization in the 30-dimensional and 100-dimensional of the F1 - F12 functions.

[0110] Figure 2 The convergence curves of IPSSA and other different swarm intelligence algorithms on 3 unimodal functions and 3 multimodal functions in the 30-dimensional and 100-dimensional are given. The convergence curves can show the number of times the algorithm falls into the local optimum and the convergence speed. The horizontal axis is the number of iterations, and the vertical axis is the optimal fitness value. Through Figure 3 It can be known that after 500 iterations, the IPSSA has certain advantages in terms of convergence speed and convergence accuracy compared with the traditional salp swarm algorithm (SSA), ant lion algorithm (ALO), grey wolf algorithm (GWO), dragonfly algorithm (DA), and moth-flame optimization algorithm (MFO). Specifically, IPSSA has found the theoretical optimal values for the functions F1, F3, F8, and F10, and has the fastest convergence speed. For the function F5 (Dim = 30), the convergence speed is relatively fast in the early stage but falls into the local optimum in the later stage, and the convergence accuracy is second only to GWO and SSA. For Figure 3 the other cases, IPSSA shows faster convergence speed and convergence accuracy.

[0111] (2) Optimization comparison with other improved salp swarm algorithms

[0112] To further verify the optimized performance of the IPSSA algorithm, IPSSA is compared with the three selected improved algorithms with better effects (i.e., ESSA, MSSA, SCSSA). Among them, ESSA was proposed by Mohammed H. Qais et al. in the paper "Enhanced salp swarm algorithm: Application to variable speed wind generators". In the paper, the coefficient c 1 is updated so that the leader moves towards the food source location according to the average exponential covariance variable c 1 , and the position formula of the followers is changed so that the followers can not only be responsible for exploring food but also help the leader make decisions, effectively improving the performance of the algorithm; MSSA was proposed by Chen Lianxing et al. in the paper "An improved salp swarm algorithm". In the paper, a weighted center of gravity is introduced for the leader to replace the optimal individual position to prevent premature aggregation near the optimal individual. For the followers, an adaptive inertia weight is introduced to balance the global search and local optimization capabilities of the algorithm. Finally, a per-dimensional random differential mutation is performed on the individuals to reduce the interference between dimensions and improve the diversity of the population. SCSSA was proposed by Chen Zhongyun et al. in the paper "Salp swarm algorithm with sine-cosine algorithm". In the paper, a Logistics chaotic sequence is introduced to generate the initial population to increase the diversity of the initial individuals; the sine-cosine algorithm is embedded as a local factor into the salp swarm algorithm to perform sine and cosine optimization on the salp individuals; a differential evolution mutation strategy is performed on the neighborhood space of the optimal salp to enhance the local search ability. The parameter settings of the comparison algorithms in this example are designed according to the above-mentioned literature.

[0113] Figure 3 It includes the comparison diagrams of the convergence curves of the IPSSA algorithm proposed in this example and other excellent improved algorithms in different dimensions of 3 unimodal functions and 3 multimodal functions. From Figure 4 it can be seen that for different types of test functions, IPSSA has better convergence speed and convergence accuracy. Specifically, it can be seen that IPSSA has found the theoretical optimal values of the functions in functions F1, F3, and F8, and the convergence speed is also relatively fast. In the 30-dimensional and 100-dimensional cases of function F5, the convergence speed is relatively fast in the early stage, but it falls into a local optimum in the later stage, and its convergence accuracy is second only to SCSSA. For other benchmark functions, IPSSA has better convergence accuracy and convergence speed.

[0114] Example 2: Esophageal cancer survival prediction model based on IPSSA-BP and its effect verification

[0115] The BP neural network continuously updates the initial weights and thresholds in a cyclic manner until the minimum of the initially set calculation error is reached or the initially set total number of training times is reached. Its advantage is that a higher accuracy can be obtained, and it can effectively solve complex problems. However, the initial weights and thresholds of the BP neural network are randomly generated, which has a great impact on the accuracy of the neural network and results in unreliable evaluation results.

[0116] In this example, the improved IPSSA algorithm is used to optimize the initial weights and thresholds of the BP neural network, and it is compared with the neural networks optimized by ESSA, SCSSA, ASSA, the traditional BP neural network, and the BP neural network optimized by the whale algorithm (WOA) to verify the effectiveness of the algorithm.

[0117] 1. Esophageal cancer survival prediction model based on BP neural network

[0118] Single-factor COX regression analysis was used to screen out the influencing factors of the death of esophageal cancer patients. The screened factors were used as input variables, and the survival time was used as the output variable to establish a survival prediction model of the BP neural network. A total of 500 esophageal cancer patients admitted to the Affiliated Hospital of Zhengzhou University from 2007 to 2018 were selected, including 316 males (63.2%) and 184 females (36.8%). The average age of the patients was 60.258 years. The blood indexes of these esophageal cancer patients were analyzed by single-factor COX regression analysis, and 5 blood indexes related to the survival period were obtained: white blood cell count (White Blood Cell Count, WBCC), monocyte count (Monocyte Count, Mono), neutrophil count (Neutrophil Count, Seg), prothrombin time (Prothrombin Time, PT), international normalized ratio (International Normalized Ratio, INR) as inputs. The structure of the survival model of the BP neural network constructed is 5-11-1, and the dimension to be optimized is:

[0119] Dim = (inputnum + 1) × hiddennum + (hiddennum + 1) × outputnum (12);

[0120] Among them, inputnum, hiddennum, and outputnum are the number of input layers, hidden layers, and output layers of the BP neural network, respectively. Therefore, the number of initial weights and thresholds to be optimized in this example is 5 × 11 + 11 × 1 = 66 weights and 12 thresholds that need to be optimized by the algorithm. Formula (16) is selected as the fitness function. The specific expression of the function is:

[0121]

[0122] Among them, X is the number of training samples of the esophageal cancer survival model, and y i is the output value of the network model of the i-th esophageal cancer patient, and y' i is the actual output value of the i-th esophageal cancer patient.

[0123] 2. Result Analysis

[0124] A model was established using the blood samples of 500 groups of esophageal cancer patients. Among them, 300 groups of data of esophageal cancer patients were randomly selected as the training data for the survival period prediction model of esophageal cancer patients for training, and the remaining 200 groups of data were used as the test data for the survival period prediction model of esophageal cancer patients to predict the prediction results of the survival level of esophageal cancer patients. The IPSSA algorithm was compared with 3 improved algorithms, as well as the traditional BP neural network and the WOA algorithm; the population size of all algorithms was uniformly set to 100, and other parameter settings were made according to the original literature. For the parameter configuration of the BP neural network, the maximum number of training times was set to 5000 times, the learning rate was set to 0.3, and the minimum error of the training target was set to le-9.

[0125] Figure 4 The comparison chart of the actual values and predicted values of the test set of the BP neural network optimized by each algorithm is given. In Figure 5 it can be roughly judged the prediction accuracy of the prediction model, but it cannot quantitatively reflect the prediction accuracy of the model. In order to better verify the prediction effect of the BP neural network optimized by the IPSSA algorithm, relevant error analysis was carried out on all models. The commonly used MAE, MSE, and MAPE were used as evaluation criteria to quantitatively analyze the models. MAE can accurately reflect the size of the actual prediction error. The smaller the value of MAE, the smaller the error of the prediction model; MSE can evaluate the degree of data change. The smaller the value of MSE, the better the accuracy of the prediction model in describing the experimental data; MAPS can measure the accuracy of the prediction. The smaller the value of MAPE, the better the fitting effect of the prediction model and the higher the accuracy. Figure 5 The corresponding comparison of the simulation evaluation criteria was plotted. It can also be intuitively seen from the figure that among the three evaluation indicators, the IPSSA-BP model in this example is at the minimum value, indicating that the IPSSA-BP model has better prediction accuracy, prediction stability and effect, indicating that the optimization performance of the IPSSA algorithm has certain competitive advantages.

[0126] Three groups of simulation experiments were carried out to analyze the convergence speed and convergence accuracy of the IPSSA algorithm proposed in the first embodiment: 1) comparison with different swarm intelligence algorithms; 2) comparison with other improved salp swarm algorithms; 3) optimization of the esophageal cancer survival prediction model. The first two groups of experiments were carried out under different dimensions of the F1-F12 functions. Through the convergence curve and Friedman test analysis, it can be seen that the IPSSA algorithm has better convergence speed and convergence accuracy compared with different swarm intelligence algorithms and different improved salp swarm algorithms, and has certain competitive advantages on the premise of ensuring stability. The third group of experiments was to optimize the esophageal cancer survival prediction model and compare it with other six models. MAE, MSE, and MAPE were used as evaluation criteria to analyze all models. The results show the effectiveness of the optimization performance of the IPSSA algorithm.

[0127] Although some preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0128] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present application and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A construction method of an esophageal cancer survival period prediction model based on an improved salp swarm algorithm, characterized in that, it includes the following steps: (1) Analyze and screen out the influencing factors of the survival period of esophageal cancer patients, use the screened factors as input variables, and the survival time as the output variable to establish a survival period prediction model based on a BP neural network; (2) Use the improved salp swarm algorithm to optimize the initial weights and thresholds of the BP neural network. The improved salp swarm algorithm includes the following steps: S1: Set the relevant parameters of the salp swarm algorithm: population size N, maximum number of iterations L, dimension Dim of individuals, and initialize the population; S2: Calculate the fitness value of each salp individual according to the objective function, and sort the fitness values; select the position with the best fitness as the position of the food source, and set the current iteration number l = 1; S3: Judge whether the current iteration number l is less than the division coefficient. If it is less, enter step S4; if it is greater, enter step S5; S4: Use the leader fusion mutation strategy to update the leader position, and introduce an adaptive inertia weight when updating the follower position; the leader position update process is as follows: a + b + c = 1 ③; wherein, represents the position of the 1st salp individual, i.e., the leader, in the j-th dimension; are two salp individuals randomly selected from the population; F j is the food source position in the j-th dimension; a, b, c are random numbers, taking values in [0, 1]; S5: Use the optimal individual leading movement strategy to update the leader position, and introduce an adaptive inertia weight when updating the follower position; the leader position update is carried out according to the following formula: where F j is the position of the j-th dimensional food source, and R 2 R 3 is a random number in the interval [0, 1]; R is a random number in the interval [-1, 1], and the calculation formula of B is as follows: where l is the current iteration number, and L is the maximum number of iterations; S6: Calculate the fitness value and update the positions of the salps, and use the dimension-by-dimension Gaussian mutation strategy to update the positions of all salps; update the positions of all salps according to the following formula: q = 1 - (l - L) 2 ⑥; x(j) = q × x(j) + randn × x(j) ⑦; where q is the adaptive inertia weight; l is the current iteration number; L is the maximum number of iterations; x(j) is the position of the j-th dimension salp; randn is the Gaussian mutation operator; S7: Use the optimal neighborhood perturbation strategy to update the position of the optimal salp, and compare it with the current fitness value, and use the greedy strategy to select the optimal fitness value; update the position of the optimal salp according to the following formula: where x(best) is the optimal position during global update; x(newbest) is the newly generated position; where v and m are random numbers between [0, 1]; S8: If the current iteration number l is less than the maximum number of iterations L, add 1 to the current iteration l and enter step S2; Otherwise, output the optimal solution and the algorithm iteration ends.

2. The construction method of the esophageal cancer survival period prediction model based on the improved salp swarm algorithm according to claim 1, characterized in that, in the step S2, the objective function expression is: Among them, X is the number of training samples of the esophageal cancer survival model, and y i is the output value of the network model of the i-th esophageal cancer patient, and y' i is the actual output value of the i-th esophageal cancer patient.

3. The construction method of the esophageal cancer survival period prediction model based on the improved salp swarm algorithm according to claim 1, characterized in that, in the step S7, for the generated neighborhood positions, use the following formula to judge whether to retain: Wherein, f(x(best)) is the optimal position fitness value during global update. If the generated position is better than the original position, it will replace the original position to become the global optimum; otherwise, the optimal position remains unchanged.

4. The method for constructing an esophageal cancer survival period prediction model based on an improved salp swarm algorithm according to claim 1, characterized in that, in the step (1), the blood indexes related to the survival period of esophageal cancer patients are screened by using univariate COX regression analysis: white blood cell count, monocyte count, neutrophil count, prothrombin time, international normalized ratio are used as inputs, and the structure of the survival period model of the BP neural network is 5-11-1, and the optimized dimension is: Dim = (inputnum + 1) × hiddennum + (hiddennum + 1) × outputnum ⑩; wherein, inputnum, hiddennum, and outputnum are the number of input layers, the number of hidden layers, and the number of output layers of the BP neural network, respectively.

5. A method for predicting the survival period of esophageal cancer patients, characterized in that, obtain the blood indexes related to the survival period of the esophageal cancer patient to be predicted, input them into the esophageal cancer survival period prediction model constructed in claim 1, and output the survival prediction period of the esophageal cancer patient.