Method for evaluating risk of ship navigation safety under influence of marine environment in high-dimensional small sample situation
By coupling BR high-dimensional associated risk sample expansion with the generalized regression neural network model, the nonlinear mapping problem of ship navigation safety risk assessment under high-dimensional small sample conditions is solved, efficient and accurate risk assessment and risk zoning map generation are achieved, and severe weather risk avoidance decision-making is supported.
Patent Information
- Application Number
- CN202411785155.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-06
AI Technical Summary
In the case of high-dimensional small samples, existing technologies find it difficult to accurately assess the impact of the marine environment on ship navigation safety. Traditional methods cannot characterize the nonlinear mapping relationship between risk assessment indicators and risks, and machine learning models require a large number of training samples and cannot be directly applied to ship navigation safety risk assessment.
The BR high-dimensional associated risk sample expansion model is coupled with the generalized regression neural network model. By expanding the high-dimensional indicator data and optimizing the model parameters, a BR-GRNN risk assessment model is established. The quantum genetic algorithm is used to solve the smoothing factor, realize nonlinear optimization, and perform risk assessment and clustering.
The accuracy and reliability of risk assessment of the impact of the marine environment on ship navigation safety under high-dimensional and small sample conditions are improved, which is significantly better than traditional methods, especially in the case of a small number of training samples, and a high assessment accuracy and precision can still be achieved.
Smart Images

Figure CN119740739B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of safety risk assessment, more specifically, relates to a marine environment influence ship navigation safety risk assessment method under high-dimensional small sample condition. BACKGROUND
[0002] Safe and efficient maritime transportation is an important guarantee for the sustained development of international trade and the security of energy supply. With the global warming, the marine environment becomes complex and severe, which seriously affects the safety of ship navigation. In 2023, there were 1867 accidents of ships above 3000 tons, which increased by 330 compared with 2013. The safety of ship navigation in severe sea conditions is the basic goal of meteorological navigation. Objective and accurate assessment of the influence of marine environment on ship navigation safety and scientific and reasonable judgment of the risk of marine environment influence are important contents of meteorological and hydrological guarantee.
[0003] Accurate and sufficient data information is the premise of correct risk assessment, but since the risk is often difficult to observe or record, the correlation data information and case samples between the marine environment influence risk and marine hydro-meteorological factors are extremely rare or seriously damaged. The available data information is only a small amount of test data samples or qualitative experience knowledge. On the other hand, there are many marine environmental risk factors affecting ship navigation safety, and these factors have extremely complex nonlinear correlations. In summary, high-dimensional small sample is a difficult problem affecting and restricting the risk assessment of ship navigation safety.
[0004] The mainstream method of ship navigation safety risk assessment (such as weighted comprehensive method, fuzzy comprehensive evaluation, grey correlation analysis, etc.) is still a linear weighted method in nature. The weighted comprehensive method is simple and rough, and fuzzy comprehensive evaluation or grey correlation analysis is often used in actual application. However, these two methods do not use the information of risk value when modeling, and the risk assessment model established cannot describe the nonlinear mapping relationship between risk evaluation index and risk. Moreover, these methods all need to solve the weight, and the analytic hierarchy process is a commonly used weighting method, but the comparison and judgment results of the analytic hierarchy process are relatively rough. Generally, the comparison and judgment matrix is established according to the prior knowledge of experts, and the human manipulation trace is too strong, which cannot avoid the subjective influence.
[0005] The atmosphere and ocean are highly complex nonlinear systems, making their behavior difficult to describe using simple linear relationships. With the development of artificial intelligence, machine learning models such as generalized regression neural networks are increasingly being applied to risk assessment. These models can automatically learn high-level abstract features from data, eliminating the need for manual feature engineering and better capturing the complex nonlinear relationships between meteorological and oceanographic data and navigation safety risks. However, machine learning models require a large number of training samples, and the high-dimensional small sample size problem prevents their direct application in the field of ship navigation safety risk assessment. Consequently, methods for assessing the impact of the marine environment on ship navigation safety risks are in urgent need of improvement, development, and innovation.
[0006] Therefore, there is an urgent need for a new risk assessment method for the impact of marine environment on ship navigation safety under high-dimensional and small sample conditions. Summary of the Invention
[0007] Aiming at the objective fact and scientific problem that there are many factors affecting the risk of marine environment affecting ship navigation safety and few case samples, this paper proposes a risk assessment method for the impact of marine environment on ship navigation safety under high-dimensional and small sample conditions.
[0008] In order to solve at least one of the above technical problems, according to one aspect of the present invention, a method for assessing the impact of the marine environment on ship navigation safety in a high-dimensional small sample situation is provided, comprising the following steps:
[0009] S1. Establish a BR high-dimensional correlation risk sample expansion model;
[0010] The associated risk sample refers to the high-dimensional risk assessment index value X and the corresponding risk value Y. The i-th high-dimensional associated risk sample is expressed as:
[0011] (X i ,Y i )=(x i,1 ,x i,2 ,…,x i,n ,y i ) (1)
[0012] Where n represents the number of high-dimensional risk assessment indicators;
[0013] S2. Couple the BR high-dimensional correlation risk sample expansion model with the generalized regression neural network model, optimize the model parameters, and obtain the BR-GRNN risk assessment model under high-dimensional small sample conditions;
[0014] S3. The BR-GRNN risk assessment model under high-dimensional small sample conditions is used to assess the risk of the marine environment affecting the navigation safety of ships, cluster the risk assessment values, and obtain a risk zoning map of the navigation area.
[0015] Furthermore, step S1 specifically includes the following steps:
[0016] S11. Establish a high-dimensional expansion index array;
[0017] S12. Establish a relationship model between high-dimensional index value X and risk value Y, and new Calculate the corresponding risk value Y new .
[0018] Furthermore, in S11, the steps for establishing the high-dimensional expansion index array are as follows:
[0019] S111, vertically cut the high-dimensional risk assessment index data and generate n labeled arrays (l1, l2, ···, l n ), let l j Represents the row number of the j-th column data, i.e. l j =(1,2,…,n) T ;
[0020] Assume that the matrix of high-dimensional small sample risk assessment indicators is X, the size of X is m×n, where m≤30, and X can be expressed as:
[0021]
[0022] S112, bind the two columns of data with strong intra-group correlation, that is, (l1, l2, ···, l n ) are combined with several indicator data with a Pearson correlation coefficient greater than or equal to 0.6 to obtain a new index array
[0023] S113, using the Bootstrap method Each column of is resampled, and m rows of new samples are obtained. Repeat the operation until the desired sample size is reached. The expanded new sample is recorded as X PP , X PP Simulate the strong correlation information contained in the original high-dimensional risk data, X PP The structure is shown below:
[0024]
[0025] S114, the original high-dimensional index data X and the expanded X PP The new data array X after merging and expansion new ; After expansion, X new The size of becomes (k+1)m×n, where k is the expansion multiple and k>2.
[0026] Furthermore, in S12, the specific steps of the relationship model between the high-dimensional indicator value X and the risk value Y are:
[0027] S121, let each high-dimensional correlation sample be (X i ,Y i )=(x i,1 ,x i,2 ,…,x i,n ,y i ), where i = 1, 2, ..., m, and these m associated samples are used as training data for the multivariate linear regression model, thereby obtaining the relationship f between the high-dimensional index value X and the risk value Y;
[0028] S122, the expanded new sample X PP Input into the relation f, thus obtaining PP The corresponding risk value vector Y PP ;
[0029] S123, put X PP and Y PP Merge to obtain a high-dimensional correlation sample matrix (X PP ,Y PP ), and (X PP ,Y PP ) is combined with the original high-dimensional correlated small sample matrix (X, Y) to obtain a new high-dimensional correlated sample matrix (X new ,Y new ).
[0030] Furthermore, the specific steps of S2 are:
[0031] S21, based on the expanded high-dimensional correlation sample matrix (X new ,Y new ) Training the generalized regression neural network model and establishing a nonlinear optimization model for the smoothing factor σ; thereby having sufficient training samples to optimize the parameters of the generalized regression neural network model;
[0032] S22. Design an efficient algorithm to solve the nonlinear optimization model about the smoothness factor σ and obtain the optimal smoothness factor.
[0033] Furthermore, the specific steps of S21 are as follows: In order to effectively avoid underfitting and overfitting in the σ optimization process, the objective function of the nonlinear optimization model is first established with the help of k-fold cross validation. The process is as follows:
[0034] For the sample data set Grouping, let the kth subset be Among them, m kis the sample size of the subset, each subset data is used as a validation set, and the remaining k-1 groups of subset data are used as training sets. Thus, k models are obtained, and the average mean square error of the k model validation sets is used as the nonlinear optimization objective function E of the cubic smoothing factor σ. The calculation formulas are shown in (7) and (8):
[0035]
[0036] Furthermore, according to the nonlinear optimization objective function of the smoothing factor σ, a nonlinear optimization model of the smoothing factor σ can be established:
[0037]
[0038] According to another aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the steps of the method for assessing the risk of marine environment affecting ship navigation safety under high-dimensional small sample conditions of the present invention are implemented.
[0039] According to another aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method for assessing the risk of marine environment impact on ship navigation safety under high-dimensional small sample conditions of the present invention are implemented.
[0040] Compared with the existing technology, the beneficial effect of the above method of the present invention is: compared with the existing technology model, the present invention greatly improves the accuracy of the risk assessment of the impact of the marine environment on ship navigation safety under high-dimensional small sample conditions, and has greater reliability and application potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present invention, but are not intended to limit the present invention.
[0042] Figure 1 A flow chart of a method according to a preferred embodiment of the present invention;
[0043] Figure 2 This is a flow chart of the BR expansion technology of a preferred embodiment of the present invention;
[0044] Figure 3 A flow chart of a quantum genetic algorithm for solving the optimal smoothing factor according to a preferred embodiment of the present invention;
[0045] Figure 4 This is a schematic diagram showing the comparison between the risk assessment values obtained by the four models and the actual values under 30 training samples in a preferred embodiment of the present invention;
[0046] Figure 5 This is a schematic diagram showing the comparison between the risk assessment values obtained by the four models and the actual values under 20 training samples in a preferred embodiment of the present invention;
[0047] Figure 6 This is a schematic diagram showing the comparison between the risk assessment values obtained by the four models and the actual values under 10 training samples in a preferred embodiment of the present invention;
[0048] Figure 7 Schematic diagram showing comparison between risk assessment values obtained by four models and actual values under five training samples in a preferred embodiment of the present invention;
[0049] Figure 8 This is a schematic diagram showing a comparison of the average relative errors of four models under different small sample scenarios according to a preferred embodiment of the present invention;
[0050] Figure 9 This is a schematic diagram comparing the mean square errors of four models under different small sample scenarios according to a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0051] To make the purpose, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0052] Unless otherwise defined, technical or scientific terms used herein shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0053] Example 1:
[0054] like Figure 1-9 As shown, the present invention proposes a method for assessing the risk of marine environment affecting ship navigation safety under high-dimensional small sample conditions, comprising the following steps:
[0055] S1. Establish a BR high-dimensional correlation risk sample expansion model.
[0056] like Figure 2 The figure shows the process flow of BR high-dimensional associated risk sample expansion technology.
[0057] The associated risk sample refers to the high-dimensional risk assessment index value X and the corresponding risk value Y. The i-th high-dimensional associated risk sample can be expressed as:
[0058] (X i ,Y i )=(x i,1 ,x i,2 ,…,x i,n ,yi ) (1)
[0059] Where n represents the number of high-dimensional risk assessment indicators.
[0060] Step S1 includes:
[0061] S11. Establish a high-dimensional expansion indicator array.
[0062] S111, vertically cut the high-dimensional risk assessment index data and generate n labeled arrays (l1, l2, ···, l n ), let l j Represents the row number of the j-th column data, i.e. l j =(1,2,…,n) T ;
[0063] Suppose the matrix of high-dimensional small sample risk assessment indicators is X, the size of X is m×n, where m≤30, X can be expressed as:
[0064]
[0065] S112, bind the two columns of data with strong intra-group correlation, that is, (l1, l2, ···, l n ) are combined with several indicator data with a Pearson correlation coefficient greater than or equal to 0.6 to obtain a new index array
[0066] S113, using the Bootstrap method Each column of is resampled, and m rows of new samples are obtained. Repeat the operation until the desired sample size is reached. The expanded new sample is recorded as X PP , we can see that X PP It can simulate the strong correlation information contained in the original high-dimensional risk data, X PP The structure is shown below:
[0067]
[0068] S114, the original high-dimensional index data X and the expanded X PP The new data array X after merging and expansion new After the above expansion process, X new The size of becomes (k+1)m×n, where k is the expansion multiple and k>2.
[0069] S12. Establish a relationship model between high-dimensional index value X and risk value Y, and new Calculate the corresponding risk value Y new .
[0070] S121, let each high-dimensional correlation sample be (X i ,Y i )=(x i,1 ,x i,2 ,…,x i,n ,y i ), where i = 1, 2, ..., m, and these m associated samples are used as training data for the multivariate linear regression model, thereby obtaining the relationship f between the high-dimensional index value X and the risk value Y.
[0071] S122, the expanded new sample X PP Input into the relation f, thus obtaining PP The corresponding risk value vector Y PP .
[0072] S123, put X PP and Y PP Merge to obtain a high-dimensional correlation sample matrix (X PP ,Y PP ), and (X PP ,Y PP ) is combined with the original high-dimensional correlated small sample matrix (X, Y) to obtain a new high-dimensional correlated sample matrix (X new ,Y new ), which solves the technical problem of high-dimensional small sample correlation data in risk assessment modeling.
[0073] S2. The BR high-dimensional associated risk sample expansion model is coupled with the generalized regression neural network (GRNN) model, and the model parameters are optimized to obtain a BR-GRNN risk assessment model under high-dimensional small sample conditions.
[0074] The basic principles and network topology of the GRNN model are explained:
[0075] Assume that the joint probability density function of random variables X and Y is f(x,y), and the observed value of X is When the input is When , the evaluation data of Y relative to X is:
[0076]
[0077] Based on the Parzen-window nonparametric estimation method, using the sample data set The probability density function can be Make an estimate:
[0078]
[0079] Where n is the sample size, P is the dimension of the random variable X, and σ is the smoothing factor. The prediction output formula is as follows:
[0080]
[0081] in, The numerator in formula (6) is y obtained from the training sample data set i The weighted sum of
[0082] The network topology of the generalized regression neural network includes the input layer, the pattern layer, the summation layer and the output layer. The input is the X observation value The output is The sample dataset is
[0083] Step S2 includes:
[0084] S21, based on the expanded high-dimensional correlation sample matrix (X new ,Y new )The generalized regression neural network model is trained and a nonlinear optimization model about the smoothing factor σ is established.
[0085] After using the above expansion model to expand the original high-dimensional small sample data, the high-dimensional small sample can be expanded into a large sample, meeting the requirements of the generalized regression neural network model for the amount of training sample data, thereby having sufficient training samples to optimize the parameters of the generalized regression neural network model.
[0086] From formula (5), we can see that the generalized regression neural network model has an unknown parameter σ (smooth factor). The value of σ has a great influence on the evaluation performance of the network. If the value of σ is too large, the evaluation output value will be If the value of σ is too small, the output value will be too low. The value of σ is close to the value of y, which makes the evaluation effect of the test set worse, and overfitting occurs. In order to effectively avoid underfitting and overfitting in the process of σ optimization, k-fold cross validation is used to first establish the objective function of the nonlinear optimization model. The specific establishment process is as follows:
[0087] For the sample data set Grouping, let the kth subset be (m k is the sample size of the subset), each subset data is used as a validation set, and the remaining k-1 groups of subset data are used as training sets, thereby obtaining k models. The average mean square error of the k validation sets in these models is used as the nonlinear optimization objective function E of the cubic smoothing factor σ. The calculation formulas are shown in (7) and (8).
[0088]
[0089] The above operation allows the samples in each round to be used to train the model, which can be as close to the distribution of the original samples as possible. The introduction of the validation set can also effectively avoid overfitting. Therefore, according to the nonlinear optimization objective function of the smoothing factor σ, a nonlinear optimization model of the smoothing factor σ can be established:
[0090]
[0091] S22. Design an efficient algorithm to solve the nonlinear optimization model of the smoothness factor σ and obtain the optimal smoothness factor. The process of using quantum genetic algorithm to solve the optimal smoothness factor is as follows: Figure 3 shown.
[0092] The present invention designs a quantum genetic algorithm (QGA) to solve the nonlinear optimization model about the smoothing factor σ. The quantum genetic algorithm is based on the quantum state vector and the genetic algorithm, integrates the quantum state vector expression with the genetic coding, encodes the chromosome based on the probability amplitude of the quantum bit, and realizes the evolution of the chromosome with the help of quantum logic gates, thereby realizing the solution of the nonlinear optimization model and effectively avoiding the local optimum.
[0093] First, the basic principles of quantum genetic algorithm are explained:
[0094] A quantum bit is a physical medium that acts as an information storage unit in a quantum computer and can be in a superposition of two quantum states at the same time. Where α and β are the probability amplitudes of the quantum state, a pair of complex constants satisfying |α| + |β| = 1; |0> and |1> represent the spin-down and spin-up states, respectively. Quantum gates are the execution structures of evolutionary operations in quantum genetic algorithms. They can transform the aforementioned qubit states to implement specific logical operations. Commonly used gates include quantum rotation gates and quantum Hadamard gates.
[0095] The QGA algorithm uses the qubit chromosome representation method. A qubit represents a gene and can express any superposition state at the same time. Based on the qubit encoding, the chromosome structure can be represented as:
[0096]
[0097] Where q j represents the chromosome of the jth individual, k represents the number of quantum bits encoded for each gene, m is the number of chromosome genes, α and β are the probability amplitudes of |0> and |1>, respectively, and satisfy the normalization condition |α| 2 +|β| 2 =1,|α| 2 represents the probability that the quantum measurement value is 0, |β| 2 represents the probability that the quantum measurement value is 1.
[0098] The update process based on the quantum revolving door is:
[0099] Among them, (α i ,β i ) and (α i ′,β i ′) represents the probability amplitude before and after the update of the chromosome i-th quantum bit rotating gate; θ i =Δθ i s(α i β i ) represents the rotation angle; Δθ i Indicates the angle of rotation; s(α i β i ) represents the direction of rotation, which is determined by the initial conditions and other parameter strategies and is primarily used to ensure algorithm convergence. The quantum update process iterates the chromosome code generated by the qubit code through a quantum rotation gate to search for the optimal code value.
[0100] Furthermore, the optimization steps of the smoothing factor σ are as follows:
[0101] Step 1: Set the basic parameters of the algorithm, including the population size N, the maximum number of iterations T, the binary length L of each variable, and the value range of the parameter σ to be optimized [σ min ,σ max ].
[0102] Step 2: Population initialization, set the quantum population to Where t is the current genetic generation; is a chromosome of an individual in the population, defined as is the parameter to be optimized, i.e. the quantum bit of the smoothing factor. Initialize Q(t) to get Q(t0), and set the probability amplitude of all chromosomes to be
[0103] Step 3: Measure each individual in the initial population Q(t0) once to generate a binary solution set P(t0).
[0104] Step 4: Perform fitness evaluation on each determined solution in P(t0), using the average mean square error of k-fold cross validation as the individual fitness value to reflect the prediction ability of GRNN. The fitness function formula is FD = E, where E is the average mean square error of k-fold cross validation.
[0105] Step 5: Determine whether the termination condition is met. If so, terminate; otherwise, proceed to the next step.
[0106] Step 6: Measure each individual in the population Q(t) once, generate a binary solution set P(t), and evaluate the fitness of each determined solution.
[0107] Step 7: Use the new quantum revolving gate U′(t) to adjust the individuals and obtain a new population Q(t+1). By measuring each individual, we obtain P(t+1). After evaluating the fitness of each determined solution, we record the optimal individual and its corresponding fitness value.
[0108] Step 8: Increase the number of iterations by 1 and return to step 5.
[0109] The method for assessing the risk of marine environment impact on ship navigation safety under high-dimensional small sample conditions of the present invention further comprises the following steps:
[0110] S3. Use the BR-GRNN model to assess the impact of the marine environment on the navigation safety of ships, cluster the risk assessment values, and obtain a risk zoning map of the navigation area.
[0111] To validate the effectiveness of our proposed model (BR-GRNN), we used the data from Table 7.12 on page 332 of the monograph "Diagnosis of Marine Environmental Characteristics and Risk Assessment of Maritime Military Activities" (Table 1) to verify the model's effectiveness. We also compared the risk assessment performance of the BR-GRNN with traditional risk assessment methods (GRNN, fuzzy comprehensive evaluation, and gray relational analysis) using different training sample sizes, verifying the BR-GRNN model's effectiveness and robustness under high-dimensional, small sample conditions. As Table 1 shows, these data include marine environmental risk assessment indicators (wind speed, wave height, visibility, thunderstorm probability, and low cloud cover) and their corresponding risk values. These risk assessment indicators and corresponding risk value data are referred to as high-dimensional relational samples. The specific approach is as follows: First, high-dimensional correlated small samples of different capacities are selected as training data, namely 30, 20, 10 and 5, and the remaining 6, 10, 20 and 31 samples are used as validation data; secondly, the BR-GRNN model is used to expand the training data of different capacities by 100 times, so that the original 30, 20, 10 and 5 samples become 3000, 2000, 1000 and 500 samples respectively; thirdly, based on the expanded samples, the smoothing factor of the BR-GRNN model is optimized using an efficient algorithm to obtain the optimal smoothing factor under different small sample conditions; finally, the optimized BR-GRNN model is used to evaluate the risk of the remaining 6, 10, 20 and 31 samples to obtain the risk assessment value, and the traditional methods (GRNN, fuzzy comprehensive evaluation, grey correlation analysis) are used to evaluate the risk of these samples. The comparison results of the risk assessment values under different small samples (30, 20, 10 and 5) with the actual values are shown as follows. Figures 4 to 7The average relative error and mean square error between the risk assessment values obtained by the four assessment models and the actual values under different small sample conditions are shown in Tables 2 to 5. The average relative error and mean square error between the risk assessment values obtained by the four assessment models and the actual values change with the sample size as shown in Tables 2 to 5. Figure 8 and 9 shown.
[0112] Table 1 Marine environmental risk assessment test samples
[0113]
[0114]
[0115] Depend on Figure 4 As can be seen, when using 30 training samples, the risk assessment values obtained by the BR-GRNN model are highly consistent with the experimental values, with very little deviation. However, the GRNN model exhibits significant deviations at the 33rd, 34th, 35th, and 36th sample points. This indicates that the BR-GRNN model significantly outperforms the GRNN model under small sample conditions. Fuzzy comprehensive evaluation and gray relational analysis, however, exhibit significant deviations at all sample points except the 32nd. Table 2 shows that when using 30 training samples, the BR-GRNN model achieves very low mean relative error and mean square error, ranging from 63.5% to 86.1% lower than the other three models. The BR-GRNN model achieves a mean relative error of only approximately 0.08, with an assessment accuracy of 92%, while the GRNN model achieves an accuracy of only 77%. The fuzzy comprehensive evaluation and gray relational analysis achieve accuracies of only approximately 39% and 45%, respectively. Overall, the fuzzy comprehensive evaluation and gray relational analysis have significant assessment errors. Compared with the GRNN model, the BR-GRNN model offers significant improvements, achieving high accuracy and precision.
[0116] Table 2 Error comparison between the BR-GRNN model and the other three models when using 30 training samples
[0117]
[0118] Depend on Figure 5As can be seen, when using 20 training samples, the risk assessment values obtained by the BR-GRNN model deviate very little from the actual values, except for the 26th point, where the deviation is large. The GRNN model exhibits significant deviations at the 26th, 30th, 34th, and 35th points. This indicates that the BR-GRNN model significantly outperforms the GRNN model under small sample conditions. Fuzzy comprehensive evaluation and gray relational analysis, however, exhibit significant deviations at all other points, except for the 24th, 25th, 29th, and 32nd points. Table 3 shows that when using 20 training samples, the BR-GRNN model achieves very low mean relative error and mean squared error, ranging from 42% to 65.9% lower than the other three models. The BR-GRNN model achieves a mean relative error of only 0.165, achieving an assessment accuracy of 83.5%, while the GRNN model achieves an accuracy of only 71.6%, and the fuzzy comprehensive evaluation and gray relational analysis achieve accuracies of approximately 52% and 53% respectively. Overall, the BR-GRNN model demonstrates significant improvement over the other three models, achieving very high assessment accuracy.
[0119] Table 3 Error comparison between the BR-GRNN model and the other three models when using 20 training samples
[0120]
[0121]
[0122] Depend on Figure 6 As can be seen, when using 10 training samples, the risk assessment values obtained by the BR-GRNN model generally maintain good consistency with the actual values, with the exception of some deviations at points 18, 21, and 26. However, the GRNN model exhibits significant deviations at most points starting from point 23. Considering that the BR-GRNN model uses only 10 training samples, it can be concluded that the model's assessment performance is still good. Therefore, under small sample conditions, the BR-GRNN model significantly outperforms the GRNN model; however, the fuzzy comprehensive evaluation and grey relational analysis exhibit significant assessment deviations at most sample points. As shown in Table 4, when using 10 training samples, the BR-GRNN model's mean relative error and mean squared error are very small, ranging from 22.7% to 54% lower than those of the other three models. The BR-GRNN model's mean relative error is only 0.215, and its assessment accuracy reaches 78.5%, while the GRNN model's assessment accuracy is only 66.8%, and the fuzzy comprehensive evaluation and grey relational analysis' assessment accuracy are only approximately 54% and 58% respectively. In general, compared with the other three models, the BR-GRNN model has obvious advantages and higher evaluation accuracy.
[0123] Table 4 Error comparison between the BR-GRNN model and the other three models when using 10 training samples
[0124]
[0125] Depend on Figure 7 As can be seen, when using only five training samples, the risk assessment values obtained by the BR-GRNN model deviate somewhat from the actual values, but the risk assessment values still show a good correlation with the actual values, with a Pearson correlation coefficient of 0.623. The risk assessment values obtained by the GRNN model deviate significantly from the actual values, with a correlation coefficient of only 0.418. The evaluation results of the fuzzy comprehensive evaluation and grey relational analysis remain very poor, but perform better than the GRNN model. Table 5 shows that when using only five training samples, the BR-GRNN model has the lowest mean relative error and mean square error, which are 19% to 51% lower than those of the other three models. The BR-GRNN model has a mean relative error of 0.347, and its evaluation accuracy still reaches 65.3%, while the GRNN model's evaluation accuracy is only 48.3%, and the fuzzy comprehensive evaluation and grey relational analysis' evaluation accuracy are only approximately 51% and 55% respectively. Overall, considering the limited number of training samples, the BR-GRNN model's evaluation performance is still acceptable and significantly outperforms the other three models.
[0126] Table 5 Error comparison between the BR-GRNN model and the other three models when using 5 training samples
[0127]
[0128] Figure 8 and 9The trend of the mean relative error and mean squared error of the four models under different training samples is shown. For the BR-GRNN and GRNN models, the mean relative error and mean squared error both increase with decreasing training samples, while those for fuzzy comprehensive evaluation and grey relational analysis tend to decrease. This is primarily due to the fact that fuzzy comprehensive evaluation and grey relational analysis do not utilize risk value information in their modeling. The modeling process primarily involves weighting, which is determined subjectively and does not require training samples. When using different numbers of training samples, the BR-GRNN model consistently achieves the best performance, followed by the GRNN model. The evaluation errors of fuzzy comprehensive evaluation and grey relational analysis are similar, both showing significant differences. When using only five training samples, the mean relative error and mean squared error of the BR-GRNN model are 0.347 and 0.409, respectively, which are smaller than the errors of fuzzy comprehensive evaluation and grey relational analysis with 30 training samples. This means that the BR-GRNN model still achieves an evaluation accuracy of 65.3% with five training samples. When using 10 training samples, the BR-GRNN model achieved an average relative error (ARE) and a mean squared error (MSE) of 0.215 and 0.311, respectively, significantly lower than the other three models. When using 20 and 30 training samples, respectively, the BR-GRNN model achieved an average relative error of only 0.165 and 0.085, and a mean squared error of only 0.197 and 0.117, with evaluation accuracies as high as 83.5% and 91.5%, respectively. This indicates that the BR-GRNN model achieves satisfactory evaluation results when the number of training samples is 20 or greater, and still achieves acceptable results when the number of training samples is only 5 or 10.
[0129] In summary, the experimental results show that the BR-GRNN model of the present invention has great reliability and application potential in high-dimensional small sample scenarios:
[0130] (1) Regardless of whether the number of small samples is 30, 20, 10, or 5, the BR-GRNN model is significantly better than the GRNN model, fuzzy comprehensive evaluation, and grey relational analysis.
[0131] (2) According to the calculation results of the average relative errors of different small sample situations, the BR-GRNN model has improved by about 32.8% to 63.5% compared with the GRNN model; compared with the fuzzy comprehensive evaluation, the BR-GRNN model has improved by about 30.0% to 86.1%; compared with the grey relational analysis, the BR-GRNN model has improved by about 21.6% to 84.8%.
[0132] (3) When the sample size is between 20 and 30, the average evaluation accuracy of the BR-GRNN model is 83.5% to 91.5%; when the sample size is 10, the average evaluation accuracy of the BR-GRNN model reaches 78.5%, which is still a good result; when the sample size is 5, the average evaluation accuracy of the BR-GRNN model is 65.3%.
[0133] The obtained BR-GRNN model can be used to evaluate the impact of the marine environment on the navigation safety of ocean-going ships, and a risk zoning map of the navigation area can be obtained, which can provide auxiliary decision-making support for ocean-going ships to avoid severe weather risks.
[0134] Example 2:
[0135] The computer-readable storage medium of this embodiment stores a computer program thereon, which, when executed by a processor, implements the steps of the method for assessing the risk of marine environment affecting ship navigation safety under high-dimensional small sample conditions in Example 1.
[0136] The computer-readable storage medium of this embodiment may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal; the computer-readable storage medium of this embodiment may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card, a secure digital card, a flash memory card, etc. equipped on the terminal; further, the computer-readable storage medium may also include both an internal storage unit of the terminal and an external storage device.
[0137] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0138] Example 3:
[0139] The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method for risk assessment of the impact of the marine environment on ship navigation safety under high-dimensional small sample conditions in Example 1 are implemented.
[0140] In this embodiment, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The memory can include read-only memory and random access memory, and provide instructions and data to the processor. A part of the memory can also include non-volatile random access memory. For example, the memory can also store information about the device type.
[0141] Those skilled in the art will appreciate that the disclosed contents of the embodiments may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0142] The present invention is described with reference to the flowcharts and / or block diagrams of the methods and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of the processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions; these computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0143] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0145] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0146] The examples described in the present invention are merely descriptions of the preferred embodiments of the present invention and are not intended to limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various modifications and improvements made to the technical solutions of the present invention by engineers and technicians in this field should fall within the scope of protection of the present invention.
Claims
1. A method for assessing the impact of marine environment on ship navigation safety risk in a high-dimensional small sample situation, characterized by: The steps include: S1. Establish a BR high-dimensional correlation risk sample expansion model; Associated risk samples refer to high-dimensional risk evaluation index values and the corresponding risk value , No. A high-dimensional associated risk sample is expressed as: (1) Where, Represents the number of high-dimensional risk assessment indicators; specifically, the risk assessment indicators include: wind speed, wave height, visibility, thunderstorm probability, and low cloud cover; Wherein, step S1 specifically includes the following steps: S11. Establish a high-dimensional expansion index array; S12. Establish high-dimensional index values and risk value The relationship model between Calculate the corresponding risk value ; Among them, in S11, the steps for establishing the high-dimensional expansion index array are as follows: S111, vertically cut the high-dimensional risk assessment index data and generate An array of labels ,make Representative The row number of the column data, that is ; Assume that the matrix of high-dimensional small sample risk assessment indicators is , Size ,in , It can be expressed as: (2) S112, bind two columns of data with strong correlation within the group, Combine several indicator data with Pearson correlation coefficient greater than or equal to 0.6 to obtain a new label array ; S113, using the Bootstrap method Each column of is resampled separately, thus obtaining Run new samples and repeat until the desired sample size is reached. The expanded new samples are recorded as , Simulate the strong correlation information contained in the original high-dimensional risk data, The structure is shown below: (3) S114, the original high-dimensional index data and the expanded The new data array after merging and expansion ; After expansion, The size becomes ,in is the multiple of the expansion and ; S2. Couple the BR high-dimensional correlation risk sample expansion model with the generalized regression neural network model, optimize the model parameters, and obtain the BR-GRNN risk assessment model under high-dimensional small sample conditions; Among them, the specific steps of S2 are: S21, based on the expanded high-dimensional correlation sample array Train the generalized regression neural network model and establish the smooth factor Nonlinear optimization model; thus there are enough training samples to optimize the parameters of the generalized regression neural network model; S22. Design an efficient algorithm to solve the smoothness factor The nonlinear optimization model is used to obtain the optimal smoothing factor; S3. The BR-GRNN risk assessment model under high-dimensional small sample conditions is used to assess the risk of the marine environment affecting the navigation safety of ships, cluster the risk assessment values, and obtain a risk zoning map of the navigation area.
2. The method according to claim 1, characterized in that In S12, high-dimensional index values and risk value The specific steps of the relationship model are: S121. Let each high-dimensional correlation sample be ,in , using this The associated samples are used as training data for the multivariate linear regression model, thereby obtaining high-dimensional index values. and risk value The relationship between ; S122, the expanded new sample Input to Relationship , thus obtaining The corresponding risk value vector ; S123, put and Merge to obtain a high-dimensional correlation sample matrix , and Small sample matrix associated with the original high-dimensional Merge to get a new high-dimensional correlation sample matrix .
3. The method according to claim 1, characterized in that The specific steps of S21 are: To solve the underfitting and overfitting in the optimization process, we first establish the objective function of the nonlinear optimization model with the help of k-fold cross validation. The process is as follows: For the sample data set Group by The subset is ,in, is the sample size of the subset, each subset data is used as a validation set, and the remaining The subset data is used as the training set, and k models are obtained. The average mean square error of the k model validation set is used as the smoothing factor. The nonlinear optimization objective function , the calculation formulas are shown in (7) and (8): (7) (8) Furthermore, according to the smoothness factor The nonlinear optimization objective function can be used to establish the smoothing factor Nonlinear optimization model of: (9)。 4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for risk assessment of the impact of the marine environment on ship navigation safety in a high-dimensional small sample situation as described in any one of claims 1 to 3 are implemented.
5. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for assessing the risk of marine environment affecting ship navigation safety in a high-dimensional small sample situation according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Bootstrap-Tobit ship accident economic loss prediction method considering ship accident data missing report problem
CN110443450A
Ship navigation safety performance evaluation method based on fuzzy neural network
CN118410344A