A method for classifying shallow sea seabed sediments
By combining the high-frequency and low-frequency echo data of the dual-frequency shallow formation profiler and optimizing the low-frequency signal sequence with a random forest classifier, the problem of low-frequency substratum classification efficiency in the existing technology is solved, and efficient and accurate substratum classification is achieved.
Patent Information
- Application Number
- CN202210509518.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-11
AI Technical Summary
The prior art is difficult to efficiently and accurately classify shallow seabed bottoms, especially when using shallow formation profile echo data, there is a problem of low signal processing and classification efficiency.
By combining the high-frequency and low-frequency echo data of the dual-frequency shallow formation profiler, the low-frequency signal sequence is optimized and intercepted and parameter optimization is optimized to build the optimal bottom classification model.
It realizes efficient and accurate classification of shallow seabed bottoms, improves classification efficiency and accuracy, can quickly process original data and obtain accurate bottom classification results.
Smart Images

Figure CN114861798B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of marine surveying and mapping technology. Specifically, it relates to a method for classifying the seabed sediment in shallow waters based on the echo data of a dual-frequency shallow stratigraphic profiler. Background Art
[0002] The type of seabed sediment is an important marine environmental parameter and the basis for seabed scientific research. Seabed sediment classification has important scientific and practical significance for marine engineering construction, seabed scientific research, and marine resource exploitation. A shallow stratigraphic profiler is an acoustic detection device used to obtain information on the shallow subsurface below the seabed. When it works, it vertically emits low-frequency acoustic wave signals to the seabed through a transmitting transducer. The acoustic waves penetrate the water layer and the seabed strata. During the downward process, they are filtered by each layer of medium, and reflection and transmission occur at the interfaces of adjacent layers with a certain difference in acoustic impedance. Part of the energy forms a reflected signal that reaches the transducer and is recorded, and the other part of the energy is transmitted and continues to travel downward until it cannot be detected. Its echo signals return in chronological order, are received by the transducer and converted into electrical signals and transmitted to the host computer. After signal calculation, the echo intensity sampling information of the shallow stratigraphic profile is obtained. The time series of the acoustic reflection signals returned by each acoustic pulse is called 1Ping data. Through transformation and according to a certain color mapping rule, it can be arranged in sequence according to the serial number of each Ping to form a shallow profile image. When a dual-frequency shallow stratigraphic profiler works, it can simultaneously emit signals of two different frequency bands, high-frequency and low-frequency, and the number of high-frequency and low-frequency sampling points in each Ping is the same. The high-frequency wave has weak penetration and can accurately measure the water depth, while the low-frequency signal is used for shallow stratigraphic profile measurement. Compared with other acoustic devices, the shallow profiler has the characteristics of low emission frequency and strong sediment penetration. The obtained echo signals can carry richer sediment characteristics. By using the different characteristics of the high-frequency part and the low-frequency part in the dual-frequency shallow profiler echo signals and combining a suitable classifier for sediment classification, the efficiency and accuracy of sediment classification can be effectively improved. Summary of the Invention
[0003] The purpose of this application is to provide a method that can comprehensively utilize the different characteristics of the high-frequency signal part and the low-frequency signal part in the shallow profiler echo data to efficiently and accurately classify the seabed sediment in shallow waters.
[0004] The embodiments of this application can be realized through the following technical solutions:
[0005] A method for classifying the seabed sediment in shallow waters classifies the seabed sediment based on the shallow profiler echo data of N pings. Among them, the shallow profiler echo data of each ping includes a high-frequency signal sequence of M sampling points and a corresponding low-frequency signal sequence of M sampling points. The method is characterized by including the following steps:
[0006] S100: Determine the alignment sequence number of each ping according to the high-frequency signal sequence of each ping, align and intercept the low-frequency signal sequence of each ping according to the alignment sequence number, and construct the initial data set with the low-frequency signal sequences of N pings after alignment and interception;
[0007] S200: Optimize the initial data set using a random forest classifier. The optimization is to optimize and intercept the low-frequency signal sequences of N pings after alignment and interception, and construct the optimized data set with the low-frequency signal sequences of N pings after optimized interception;
[0008] S300: Optimize the parameters of the random forest classifier using the optimized data set to obtain the optimal random forest classifier;
[0009] S400: Process the shallow profile echo data to be classified using the optimal random forest classifier to obtain the classification result of the seabed sediment.
[0010] Preferably, step S100 further includes the following steps:
[0011] S110: Convert according to the minimum water depth H of the test sea area min to obtain the corresponding effective echo sequence number h_min;
[0012] S120: For the high-frequency signal sequence P of the i-th ping i ={p i,1 , p i,2 ,... p i,j ,...,, p i,M}, i = 1.....N, where p i,j is the high-frequency signal of the j-th sampling point. Find the first high-frequency signal greater than the echo intensity threshold from its subsequence {p i,h_min , p i,h_min ,...,, p i,M} Take h i as the alignment sequence number of the i-th ping;
[0013] S130: For the low-frequency signal sequence Q of the i-th ping i ={q i,1 , q i,2 ,... q i,j ,...,, q i,M}, i = 1.....N, where q i,j is the low-frequency echo signal of the j-th sampling point. Intercept the low-frequency signal of its h i -th sampling point and the subsequent m sampling points to obtain the low-frequency signal sequence of the i-th ping after alignment and interception i=1...N, wherein m is determined based on the estimation of the seabed characteristics of the test sea area and m≤M;
[0014] S140: Construct an initial data set {Q′1, Q′2, Q′ i , ..., Q′ N}.
[0015] Preferably, the following steps are further included between step S120 and step S130:
[0016] When i>L, judge Is it greater than a preset error threshold? If the judgment result is true, the shallow profile echo data of the i-th ping is removed and new shallow profile echo data is added, where L is a positive integer greater than or equal to 2.
[0017] Preferably, the step S200 further includes the following steps:
[0018] S210: Setting the estimated range E of the starting sequence number start = {0, 1, 2, 3} and the estimated range E of the end sequence number end ={4, 5, ..., m};
[0019] S220: Nest the start sequence number and the end sequence number to traverse the E start and E end For each value of the starting sequence number and the ending sequence number, a random forest classifier is constructed, and each Q′ in the initial data set is i The low-frequency signal from the sampling point of the starting sequence number to the sampling point of the ending sequence number is intercepted, and then the training set and the test set are randomly divided according to the preset ratio. The training set is used to train the random forest classifier, and then the classification result accuracy is tested using the test set. After the traversal is completed, the optimal starting sequence number h_start and the optimal ending sequence number h_end are determined according to the classification result accuracy;
[0020] S230: Use the optimal start sequence number h_start and the optimal end sequence number h_end to compare {Q′1, Q′2, ...Q′ N} to optimize the interception and obtain the optimized data set {Q″1, Q″2, Q″ i , ..., Q″ N},in
[0021] Preferably, the step S300 further includes the following steps:
[0022] S310: Optimizing the number of decision trees of the random forest classifier using the optimized data set;
[0023] S320: Optimize the min_samples_split of the random forest classifier using the optimized dataset, where min_samples_split is the minimum number of samples required for further partitioning of the internal nodes of the decision trees in the random forest classifier;
[0024] S330: Optimize the min_samples_leaf of the random forest classifier using the optimized dataset, where min_samples_leaf is the minimum number of samples in the leaf nodes of the random forest classifier.
[0025] Further, the step S310 further includes the following steps:
[0026] S311: Determine the rough estimation range and rough estimation step size of the number of decision trees;
[0027] S312: Traverse the rough estimation range with the number of decision trees according to the rough estimation step size. For each value of the number of decision trees, construct a random forest classifier, divide the optimized dataset into a training set and a test set according to a preset ratio, train the random forest classifier using the training set and then perform a classification result accuracy test using the test set. After the traversal is completed, determine the rough estimated value of the number of decision trees according to the classification result accuracy;
[0028] S313: Determine the fine estimation range and fine estimation step size of the number of decision trees according to the rough estimated value of the number of decision trees;
[0029] S314: Traverse the fine estimation range with the number of decision trees according to the fine estimation step size. For each value of the number of decision trees, construct a random forest classifier, divide the optimized dataset into a training set and a test set according to a preset ratio, train the random forest classifier using the training set and then perform a classification result accuracy test using the test set. After the traversal is completed, determine the optimal value of the number of decision trees according to the classification result accuracy.
[0030] Further, the step S320 further includes the following steps:
[0031] S321: Determine the estimation range and estimation step size of min_samples_split;
[0032] S322: Traverse the estimated range of min_samples_split according to the estimated step size of min_samples_split. For each value of min_samples_split, construct a random forest classifier, where the number of decision trees in the random forest classifier is the optimal value of the number of decision trees determined in step S314. Divide the optimized dataset into a training set and a test set according to a preset ratio. After training the random forest classifier using the training set, use the test set to test the accuracy of the classification results. After the traversal is completed, determine the optimal value of min_samples_split according to the classification result accuracy.
[0033] Further, the step S330 further includes the following steps:
[0034] S331: Determine the estimated range and estimated step size of min_samples_leaf;
[0035] S332: Traverse the estimated range of min_samples_leaf according to the estimated step size of min_samples_leaf. For each value of min_samples_leaf, construct a random forest classifier, where the number of decision trees in the random forest classifier is the optimal value of the number of decision trees determined in step S314, and min_samples_split of the random forest classifier is the optimal value of min_samples_split determined in step S322. Divide the optimized dataset into a training set and a test set according to a preset ratio. After training the random forest classifier using the training set, use the test set to test the accuracy of the classification results. After the traversal of the optimized dataset is completed, determine the optimal value of min_samples_leaf according to the classification result accuracy;
[0036] S333: Determine the optimal random forest classifier. The number of decision trees, min_samples_split, and min_samples_leaf of the optimal random forest classifier are respectively the optimal value of the number of decision trees determined in step S314, the optimal value of min_samples_split determined in step S322, and the optimal value of min_samples_leaf determined in step S332.
[0037] Preferably, the step S400 further includes the following steps:
[0038] S410: Respectively extract the high-frequency signal sequence D to be classified = {d1, d2,... d M} and the low-frequency signal sequence to be classified C = {c1, c2, ..., c M};
[0039] S420: From the subsequence {d h_min , d h_min+1 , ..., d M} to find the first high-frequency signal d that is greater than the echo intensity threshold h , taking h as the alignment sequence number of the shallow section echo data to be classified;
[0040] S430: Using the h, the h_start and the h_end to optimize and intercept the C, and obtain the optimized and intercepted low-frequency signal sequence to be classified C′={c h+h_start , ..., c h+h_end};
[0041] S440: Input the C′ into the optimal random forest classifier to classify the seabed sediment.
[0042] The shallow seabed sediment classification method provided by the embodiment of the present application has at least the following beneficial effects:
[0043] (1) The method provided by the present application can effectively utilize shallow profile echo data of two different frequencies. According to the different characteristics of high-frequency and low-frequency echo data, the high-frequency echo signal can accurately obtain the information of the interface between seawater and bottom sediment, thereby eliminating the invalid information of the seawater part in the low-frequency echo signal and accurately separating the part of the low-frequency echo signal rich in bottom sediment characteristics;
[0044] (2) The method provided in this application constructs a bottom sediment classification model based on a random forest classifier, and according to the characteristics of the data, uses a random forest classifier to optimize the starting and ending numbers of the sampling points of the low-frequency echo signal used for classification, which can effectively reduce the amount of data used for training, testing and classification, and improve the training and classification speed;
[0045] (3) The method provided in this application optimizes the parameters such as the number of decision trees of the random forest classifier by combining rough estimation with fine estimation, which can greatly improve the optimization speed and optimization effect. This method can classify the collected original data by simply processing it, with fast classification speed and high efficiency, and has important practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a specific example of shallow profile echo data obtained by a dual-frequency shallow formation profiler;
[0047] Figure 2Flowchart of the shallow sea bottom sediment classification method according to an embodiment of the present application;
[0048] Figure 3 Shallow profile echo data for 1 ping according to an embodiment of the present application;
[0049] Figure 4 Optimized data set according to an embodiment of the present application. Detailed implementation manners
[0050] Hereinafter, the present application will be further described based on preferred implementation manners with reference to the accompanying drawings.
[0051] In addition, for the convenience of understanding, various components in the drawings are enlarged or reduced, but this is not intended to limit the protection scope of the present application.
[0052] Singular terms also include plural meanings, and vice versa.
[0053] In the description of the embodiments of the present application, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship when the products in the embodiments of the present application are usually placed. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation of the present application. In addition, in the description of the present application, in order to distinguish different units, terms such as first and second are used in this specification, but these are not limited by the manufacturing order and cannot be understood as indicating or implying relative importance. In the detailed description and claims of the present application, their names may be different.
[0054] The terms in this specification are used to describe the embodiments of the present application, but are not intended to limit the present application. It should also be noted that unless otherwise clearly defined and limited, if terms such as "set", "connected", "connected to" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present application can be specifically understood.
[0055] The embodiments of the present application provide a shallow sea bottom sediment classification method for classifying the shallow sea bottom sediment based on the shallow profile echo data of N pings, where the shallow profile echo data of each ping includes a high-frequency signal sequence of M sampling points and a corresponding low-frequency signal sequence of M sampling points.
[0056] In the embodiments of the present application, the values of N and M are determined according to the data volume of the actual shallow profile echo data and the number of sampling points of the shallow profile echo data for each ping.
[0057] Figure 1 An example of shallow profile echo data obtained by a dual-frequency shallow layer profiler is shown. The upper part is the high-frequency signal sequence part, and the lower part is the low-frequency signal sequence part. Each part contains 6500 ping signal sequences (i.e., N = 6500), and the echo signal of each ping signal sequence contains 291 sampling points (i.e., M = 291), with gray scale representing the signal intensity. From Figure 1 it can be seen that when the dual-frequency shallow layer profiler works, it can simultaneously transmit signals of two different frequency bands, high-frequency and low-frequency. And the number of high-frequency and low-frequency sampling points within a single Ping is the same. The high-frequency penetration ability is weak and it can accurately measure the water depth. The low-frequency signal carries rich sediment characteristics. By combining the high-frequency signal and the low-frequency signal, an accurate classification of the shallow sea seabed sediment can be obtained.
[0058] The implementation process of the shallow sea seabed sediment classification method in the embodiments of the present application is as Figure 2 shown and includes the following steps:
[0059] S100: Determine the alignment sequence number of each ping according to the high-frequency signal sequence of each ping, perform alignment and interception on the low-frequency signal sequence of each ping according to the alignment sequence number, and construct the initial data set from the N ping low-frequency signal sequences after alignment and interception;
[0060] S200: Optimize the initial data set using a random forest classifier. The optimization is to perform optimized interception on the N ping low-frequency signal sequences after alignment and interception, and construct the optimized data set from the N ping low-frequency signal sequences after optimized interception;
[0061] S300: Optimize the parameters of the random forest classifier using the optimized data set to obtain the optimal random forest classifier;
[0062] S400: Process the shallow profile echo data to be classified using the optimal random forest classifier to obtain the classification result of the seabed sediment.
[0063] The following will elaborate on steps S100 to S400 in detail.
[0064] Step S100 is a step of aligning the low-frequency signal sequences in the shallow profile echo data to obtain an initial data set. For the shallow profile echo signals obtained at different positions, since the water depths at the positions are different, the starting sampling point numbers of the effective seabed echo parts are also different. Before classifying using the low-frequency signal sequences, it is necessary to remove the signals before reflection from the sea water / seabed interface in the low-frequency signals of different pings, so that the low-frequency signal sequences for classification are aligned starting from the sampling points reaching the seabed, thereby reducing the influence of the water depth variable on the true sediment classification.
[0065] In some preferred embodiments of the present application, the alignment and interception of the low-frequency signals are performed through the following steps:
[0066] S110: According to the minimum water depth H of the test sea area min calculate the corresponding effective echo sequence number h_min;
[0067] S120: For the high-frequency signal sequence P of the i-th ping i ={p i,1 , p i,2 ,... p i,j ,... p i,M}, i = 1......N, where p i,j is the high-frequency signal at the j-th sampling point, find the first high-frequency signal greater than the echo intensity threshold from its subsequence {p i,h_min , p i,h_min+1 ,... p i,M} Take h i as the alignment sequence number of the i-th ping;
[0068] S130: For the low-frequency signal sequence Q of the i-th ping i ={q i,1 , q i,2 ,... q i,j ,... q i,M}, i = 1.....N, where q i,j is the low-frequency echo signal at the j-th sampling point, intercept the low-frequency signal of its h i -th sampling point and the subsequent m sampling points to obtain the low-frequency signal sequence of the i-th ping after alignment and interception i = 1...N, where m is estimated and determined based on the seabed characteristics of the test sea area and m ≤ M;
[0069] S140: Construct an initial data set {Q′1, Q′2, Q′ i ,... Q′ N};
[0070] Specifically, first, according to the minimum water depth H of the sea area where the test is conducted, min , combined with the propagation speed of sound waves in seawater and the time interval between adjacent sampling points, the effective echo sequence number h_min of the sea area can be obtained. Subsequent alignment and interception steps are all performed after the h_minth sampling point, thereby effectively eliminating the interference of seawater surface echoes;
[0071] Since the high-frequency signal sequence has an obvious intensity mutation at the seawater / seabed boundary, the starting position of the seabed echo can be determined more accurately. For the high-frequency signal sequence P of the i-th ping i , starting from its h_min+1th sampling point, find the first high-frequency signal p that is greater than the preset echo intensity threshold i,hi , the sampling point number h corresponding to the signal i As the alignment sequence number of the i-th ping, in some embodiments of the present application, the echo strength threshold may be predetermined by an empirical formula according to the strength of the transmitted signal, water depth and other data;
[0072] After determining the alignment sequence number h of the i-th ping i After that, use h i Intercept the low-frequency signal sequence Q of the i-th ping i h i The signal of the sampling point and the subsequent m sampling points is obtained by aligning and intercepting the low-frequency signal sequence of the i-th ping, Q′ i = Among them, the value of m is estimated by testing the seabed characteristics of the sea area, and the Q′ of each ping is i The number of sampling points is unified as m+1, which can eliminate the interference of multiple echoes and deep bottom echoes in the original low-frequency signal sequence, and effectively improve the accuracy of bottom classification;
[0073] Q′ for each ping i After alignment and interception, the initial data set {Q′1, Q′2, Q′ i , ..., Q′ N}.
[0074] In some preferred embodiments of the present application, the following steps are further included between step S120 and step S130:
[0075] When i>L, judge Is it greater than a preset error threshold? If the judgment result is true, the shallow profile echo data of the i-th ping is removed and new shallow profile echo data is added, where L is a positive integer greater than or equal to 2.
[0076] In the above steps, for the shallow profile echo data of the i-th ping, the alignment sequence number determined by the high-frequency signal sequence is compared with the mean value of the previous L alignment sequence numbers, and the data that significantly deviates from the mean value is removed, thereby reducing the influence of abnormal data on the accuracy of the classifier.
[0077] Step S200 is a step for optimizing parameters in the initial data set. For the low-frequency signal sequence after alignment and truncation, although the starting sequence number obtained according to the high-frequency signal characteristics can roughly align the seabed echo signal, there may still be slight deviations in the starting position of its low-frequency signal, and when the length of each ping is determined by the estimated value m, there may still be deep bottom sediment echoes in the second half, which affects the classification accuracy. Therefore, it is necessary to further optimize the starting position and ending position of the signal.
[0078] In some preferred embodiments of the present application, the initial data set is optimized through the following steps to obtain an optimized data set:
[0079] S210: Set the estimated range E start of the starting sequence number = {0, 1, 2, 3} and the estimated range E end of the ending sequence number = {4, 5,..., m};
[0080] S220: Nest and traverse the E start and E end , for each value of the starting sequence number and the ending sequence number, construct a random forest classifier, intercept the low-frequency signal of each Q' i in the initial data set from the sampling point of the starting sequence number to the sampling point of the ending sequence number, then randomly divide the training set and the test set according to a preset ratio, use the training set to train the random forest classifier, and then use the test set to test the accuracy of the classification result. After the traversal is completed, determine the optimal starting sequence number h_start and the optimal ending sequence number h_end according to the classification result accuracy;
[0081] S230: Use the optimal starting sequence number h_start and the optimal ending sequence number h_end to optimize and intercept {Q'1, Q'2,... Q' N} to obtain the optimized data set {Q''1, Q''2, Q'' i ,..., Q'' N}, where
[0082] Specifically, in some preferred embodiments of the embodiments of the present application, during the process of traversing the estimation range in step S220, for each estimated value of the starting serial number and the ending serial number, multiple trainings and tests can be performed, and the classification result accuracy corresponding to the starting serial number and the ending serial number is determined according to the mean value of the multiple test results.
[0083] Specifically, after the traversal process ends, compare the classification result accuracies corresponding to the estimated values of each starting serial number and ending serial number to determine the optimal starting serial number h_start and the optimal ending serial number h_end, and then optimize and intercept the initial data set to obtain an optimized data set. During the subsequent process of optimizing the parameters of the random forest classifier, the optimized data set is used for training and testing.
[0084] Step S300 is a step of optimizing the parameters of the random forest classifier using the optimized data set. In some preferred embodiments of the present application, the number of decision trees, min_samples_split (i.e., the minimum number of samples required for further division of the internal nodes of the decision tree), and min_samples_leaf (i.e., the minimum number of samples in the leaf nodes) of the random forest classifier are optimized in sequence through step S310, step S320, and step S330 to obtain the optimal values of the above respective parameters.
[0085] In some preferred embodiments of the present application, step S310 further includes the following steps:
[0086] S311: Determine the rough estimation range and the rough estimation step size of the number of decision trees;
[0087] S312: Traverse the rough estimation range with the number of decision trees according to the rough estimation step size. For each value of the number of decision trees, construct a random forest classifier, divide the optimized data set into a training set and a test set according to a preset ratio, use the training set to train the random forest classifier, and then use the test set to test the classification result accuracy. After the traversal is completed, determine the rough estimated value of the number of decision trees according to the classification result accuracy;
[0088] S313: Determine the fine estimation range and the fine estimation step size of the number of decision trees according to the rough estimated value of the number of decision trees;
[0089] S314: Traverse the fine estimation range with the number of decision trees according to the fine estimation step size. For each value of the number of decision trees, construct a random forest classifier, divide the optimized data set into a training set and a test set according to a preset ratio, use the training set to train the random forest classifier, and then use the test set to test the classification result accuracy. After the traversal is completed, determine the optimal value of the number of decision trees according to the classification result accuracy.
[0090] Specifically, since the estimation range of the number of decision trees is relatively large, first traverse according to the rough estimation range and step size. Based on the obtained rough estimation value of the number of decision trees, further divide the fine estimation range and step size and traverse them. Finally, obtain the optimal value of the number of decision trees. By first making a rough estimate and then a fine estimate, the estimation range can be effectively screened quickly, thereby improving the speed of parameter optimization.
[0091] Specifically, in some preferred embodiments of the present application, during the process of traversing the estimation range in steps S312 and S314, for each estimated value of the number of decision trees, multiple trainings and tests can be performed, and the accuracy rate of the classification result corresponding to the number of decision trees can be determined according to the mean value of the multiple test results.
[0092] In some preferred embodiments of the present application, step S320 further includes the following steps:
[0093] S321: Determine the estimation range and estimation step size of min_samples_split;
[0094] S322: Traverse the estimation range of min_samples_split according to the estimation step size of min_samples_split. For each value of min_samples_split, construct a random forest classifier, where the number of decision trees in the random forest classifier is the optimal value of the number of decision trees determined in step S314. Divide the optimized data set into a training set and a test set according to a preset ratio. After training the random forest classifier using the training set, use the test set to test the accuracy rate of the classification result. After the traversal is completed, determine the optimal value of min_samples_split according to the accuracy rate of the classification result.
[0095] Specifically, for the random forest classifier constructed in step S322, the value of the number of decision trees is the optimal value of the number of decision trees determined in step S314.
[0096] Specifically, in some preferred embodiments of the present application, during the process of traversing the estimation range in step S322, for each estimated value of min_samples_split, multiple trainings and tests can be performed, and the accuracy rate of the classification result corresponding to the min_samples_split can be determined according to the mean value of the multiple test results.
[0097] In some preferred embodiments of the present application, step S330 further includes the following steps:
[0098] S331: Determine the estimation range and estimation step size of min_samples_leaf;
[0099] S332: Traverse the estimated range of min_samples_leaf according to the estimated step size of min_samples_leaf. For each value of min_samples_leaf, construct a random forest classifier, where the number of decision trees of the random forest classifier is the optimal value of the number of decision trees determined in step S314, and the min_samples_split of the random forest classifier is the optimal value of min_samples_split determined in step S322. Divide the optimized dataset into a training set and a test set according to a preset ratio. After training the random forest classifier using the training set, use the test set to test the accuracy of the classification results. Determine the optimal value of min_samples_leaf according to the classification result accuracy after traversing the optimized dataset;
[0100] S333: Determine the optimal random forest classifier. The number of decision trees, min_samples_split, and min_samples_leaf of the optimal random forest classifier are respectively the optimal value of the number of decision trees determined in step S314, the optimal value of min_samples_split determined in step S322, and the optimal value of min_samples_leaf determined in step S332.
[0101] Specifically, for the random forest classifier constructed in step S332, the value of the number of decision trees is the optimal value of the number of decision trees determined in step S314, and the value of its min_samples_split is the optimal value of min_samples_split determined in step S322.
[0102] Specifically, in some preferred embodiments of the present application, during the process of traversing the estimated range in step S332, for each estimated value of min_samples_leaf, multiple trainings and tests can be performed, and the classification result accuracy corresponding to this min_samples_leaf is determined according to the mean value of the multiple test results.
[0103] After optimizing the number of decision trees, min_samples_split, and min_samples_leaf of the random forest classifier respectively, the optimal random forest classifier for classifying the shallow - sea seabed sediment is obtained.
[0104] Step S400 is the process of classifying the shallow - sea seabed sediment using the optimal random forest classifier.
[0105] In some preferred embodiments of the present application, step S400 further includes the following steps:
[0106] S410: respectively extract a high-frequency signal sequence D = {d1, d2,..., d M} and a low-frequency signal sequence C = {c1, c2,..., c M} from the shallow profile echo data to be classified;
[0107] S420: find the first high-frequency signal d h_min , d h_min+1 ,..., d M greater than the echo intensity threshold from the subsequence {d h} of D, and take h as the alignment sequence number of the shallow profile echo data to be classified;
[0108] S430: use the h, the h_start, and the h_end to optimize the interception of C, and obtain an optimized intercepted low-frequency signal sequence C' = {c h+h_start ,..., c h+h_end};
[0109] S440: input the C' into the optimal random forest classifier for classifying the seabed sediment.
[0110] Embodiment 1
[0111] In this embodiment, a dual-frequency shallow stratigraphic profiler is used to collect the echo acoustic signal data of the shallow subsurface layer in an experimental pool. The low-frequency transmitted acoustic wave frequency of the above dual-frequency shallow profiler is 20 kHz, and the high-frequency transmitted acoustic wave frequency is 100 kHz. The number of high-frequency and low-frequency sampling points in each ping is the same. The measurement range is selected as 1 - 20 m, and its signal pulse width is set to 20 us. Actually, each ping has 291 sampling points (i.e., M = 291). There are three types of bottom sediments in the experimental pool, namely cement, silt, and gravel. Different bottom sediment areas have been clearly distinguished. 20 collection points are randomly selected in each bottom sediment area. At each collection point, the dual-frequency shallow profiler is used to collect 1 minute of echo signal data and convert it into digital signals as the original data set. 300 pings of each type of bottom sediment are randomly selected from the original data set, and a total of 900 pings are used to form the shallow profile echo data with N = 900.
[0112] Figure 3 shows the shallow profile echo data converted into digital signals of one of the pings. From Figure 3 it can be seen that there are signal maxima at the sampling points near the water surface, which will interfere with the determination of the alignment sequence number. Therefore, it is necessary to first determine according to the minimum water depth H miDetermine the effective echo sequence number h_min. The water depth of the experimental water tank used in this experiment is 2 meters. According to the sampling point interval, the effective echo sequence number h_min = 29. In the subsequent process of searching for the aligned sequence number, start searching from the 29th sampling point of the high-frequency signal sequence of each ping, so as to filter out the echo interference on the water surface.
[0113] Next, find the first signal greater than the echo intensity threshold K from the signals of the 29th sampling point to the 291st sampling point of the high-frequency signal sequence of each ping. And use the corresponding sampling point number h i as the aligned sequence number of this ping. The size of the echo intensity threshold K is determined in advance according to factors such as signal intensity, water depth, and equipment range. In this embodiment, K = 4000.
[0114] In this embodiment, for the shallow profile echo data after the 5th ping, after determining the aligned sequence number of this ping, compare it with the average value of the aligned sequence numbers of the previous 5 pings (i.e., L = 5). If the difference between the two is greater than the preset error threshold, it is considered that there is a large error in this ping signal, and this ping signal is excluded and new shallow profile echo data is supplemented.
[0115] After determining the aligned sequence number h of each ping i further estimate the signal length m to be intercepted according to the seabed characteristics, and align and intercept the low-frequency signal sequence Q i according to h i and m of each ping. In this embodiment, m = 19, that is, for the low-frequency signal sequence of each ping, intercept the signals of its 20 sampling points starting from the aligned sequence number h i to form the aligned and intercepted low-frequency signal sequence i = 1...900, and finally form the initial data set {Q′1, Q′2, Q′ i ,..., Q′ 900}.
[0116] After constructing the initial data set, the start sequence number and end sequence number of the low-frequency signal are optimized according to steps S210 to S230. In this embodiment, for each set of start sequence number and end sequence number in the nested traversal process, 630 pings (70%) are randomly selected from the initial data set of 900 pings as the training set, and the remaining 270 pings (30%) are used as the test set for training the random forest classifier and testing the accuracy of the classification results. The parameters of the random forest classifier are set according to the default values. After repeating the training and testing 10 times, the average value is taken as the corresponding classification result accuracy. After the nested traversal ends, the classification result accuracy determines h_start = 1 and h_end = 18, and the initial data set is optimized and intercepted using them to obtain the optimized data set {Q″1, Q″2, Q″ i ,..., Q″ 900}, where Figure 4 shows the situation of the optimized data set obtained through the above steps in this embodiment.
[0117] After obtaining the optimized data set, the number of decision trees of the random forest classifier is optimized using the optimized data set according to steps S311 to S314, where the optimization steps are similar to steps S210 to S220 and will not be elaborated here. In this embodiment, the rough estimation range of the number of decision trees is 0 to 800, the rough estimation step size is 10, and the rough estimation value obtained through steps S311 to S312 is 230. The fine estimation range of the number of decision trees is further determined to be 220 to 240, and the fine estimation step size is 1. The optimal value of the number of decision trees is finally determined to be 237 through steps S313 to S314.
[0118] After determining the optimal value of the number of decision trees, the min_samples_split of the random forest classifier is optimized using the optimized data set according to steps S321 to S322, where the number of decision trees of the random forest classifier is set to the optimal value of the number of decision trees obtained through the previous steps. In this embodiment, the number of decision trees of the random forest classifier is set to 237, and the estimation range of min_samples_split is set to 2 to 10, and the estimation step size is 1. The optimal value of min_samples_split obtained through steps S321 to S322 is 2.
[0119] After determining the optimal value of the number of decision trees, the optimized dataset is used to optimize the min_samples_leaf of the random forest classifier according to steps S331 to S332, where the number of decision trees and min_samples_split of the random forest classifier are respectively set to the optimal values obtained through the foregoing steps. In this embodiment, the number of decision trees of the random forest classifier is set to 237, min_samples_split is set to 2, and the estimated range of min_samples_leaf is set to 1-10, with an estimated step size of 1. The optimal value of min_samples_leaf obtained through steps S331 to S332 is 2.
[0120] After obtaining the optimal values of the number of decision trees, min_samples_split, and min_samples_leaf of the above random forest classifier, the corresponding parameters of the random forest classifier are set to the above optimal values, that is, the optimal random forest classifier is obtained.
[0121] Step S400 is the step of classifying the shallow profile echo data to be classified using the above optimal forest classifier. Specifically, the high-frequency signal sequence and the low-frequency signal sequence are extracted from the shallow profile echo data to be classified. The steps of determining the alignment serial number through the high-frequency signal sequence, aligning and intercepting the low-frequency signal sequence, and optimizing the interception have been described in detail and will not be repeated here.
[0122] In this embodiment, after constructing the optimal random forest classifier through the above steps, randomly selected shallow profile echo data is used. After performing alignment interception and optimized interception operations respectively, it is input into the classifier for classification accuracy testing. The test results show that the classification accuracies for the three bottom sediments of gravel, silt, and cement are 95.56%, 98.89%, and 100% respectively.
[0123] The specific implementation manners of the present application have been described in detail above. For those skilled in the art of this technology, without departing from the principle of the present application, several improvements and modifications can still be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for classifying the seabed sediment in shallow waters, which classifies the seabed sediment based on the shallow profile echo data of N pings, where, The shallow profiling echo data of each ping includes a high-frequency signal sequence of M sampling points and a corresponding low-frequency signal sequence of M sampling points. It is characterized in that the method includes the following steps: S100: Determine the alignment sequence number of each ping according to the high-frequency signal sequence of each ping. Intercept the low-frequency signal sequence of each ping according to the alignment sequence number, and construct an initial data set from the low-frequency signal sequences of N pings after alignment interception; S200: Optimize the initial data set using a random forest classifier. The optimization is to perform optimized interception on the low-frequency signal sequences of N pings after alignment interception, and construct an optimized data set from the low-frequency signal sequences of N pings after optimized interception; S300: Optimize the parameters of the random forest classifier using the optimized data set to obtain an optimal random forest classifier; S400: Process the shallow profiling echo data to be classified using the optimal random forest classifier to obtain the classification result of the seabed sediment; The step S100 further includes the following steps: S110: Calculate the corresponding effective echo sequence number h_min according to the minimum water depth H in the test sea area min and convert it S120: For the high-frequency signal sequence P of the i-th ping i ={p i,1 , p i,2 ,… p i,j ,…, p i,M}, i = 1……N, where p i,j is the high-frequency signal at the j-th sampling point. Find the first high-frequency signal greater than the echo intensity threshold from its subsequence {p i,h_min , p i,h_min+1 ,…, p i,M} Take h i as the alignment sequence number of the i-th ping; S130: For the low-frequency signal sequence Q of the i-th ping i ={q i,1 , q i,2 , … q i,j , …, q i,M}, i = 1 …… N, where q i,j is the low-frequency echo signal of the j-th sampling point. Intercept the low-frequency signals of its h i -th sampling point and the subsequent m sampling points to obtain the low-frequency signal sequence of the i-th ping after alignment and interception The m is estimated and determined based on the seabed characteristics of the test sea area and m ≤ M; S140: Construct an initial data set {Q′1, Q′2, Q′ i , …, Q′ N}.
2. The shallow sea seabed sediment classification method according to claim 1, characterized in that: The following steps are further included between the step S120 and the step S130: When i>L, judge Is it greater than a preset error threshold? If the judgment result is true, the shallow profile echo data of the i-th ping is removed and new shallow profile echo data is added, where L is a positive integer greater than or equal to 2.
3. The method for classifying shallow sea bottom sediments according to claim 1, characterized in that, The step S200 further includes the following steps: S210: Set the estimated range E of the starting sequence number start ={0, 1, 2, 3} and the estimated range E of the ending sequence number end ={4, 5, …, m}; S220: Nest and traverse the starting serial number and the ending serial number for the E start and E end , for each value of the starting serial number and the ending serial number, construct a random forest classifier, and for each Q i ' in the initial dataset, intercept its low-frequency signal from the sampling point of the starting serial number to the sampling point of the ending serial number, then randomly divide the training set and the test set according to a preset ratio, use the training set to train the random forest classifier and then use the test set to test the accuracy of the classification result. After the traversal is completed, determine the optimal starting serial number h_start and the optimal ending serial number h_end according to the classification result accuracy; S230: Optimally intercept {Q1′, Q′2, … Q′ N} using the optimal starting sequence number h_start and the optimal ending sequence number h_end to obtain the optimized dataset {Q″1, Q″2, Q″ i , …, Q″ N}, where 4. The method for classifying the shallow sea seabed sediment according to claim 3, characterized in that, The step S300 further includes the following steps: S310: Optimize the number of decision trees of the random forest classifier using the optimized data set; S320: Optimize the min_samples_split of the random forest classifier using the optimized data set. The min_samples_split is the minimum number of samples required for further splitting of the internal nodes of the decision tree of the random forest classifier; S330: Optimize the min_samples_leaf of the random forest classifier using the optimized data set. The min_samples_leaf is the minimum number of samples of the leaf nodes of the random forest classifier.
5. The method for classifying shallow sea bottom sediments according to claim 4, wherein The step S310 further includes the following steps: S311: Determine the rough estimation range and rough estimation step size of the number of decision trees; S312: Traverse the rough estimation range with the number of decision trees according to the rough estimation step size. For each value of the number of decision trees, construct a random forest classifier, divide the optimized data set into a training set and a test set according to a preset ratio, use the training set to train the random forest classifier and then use the test set to test the accuracy of the classification result. After the traversal, determine the rough estimation value of the number of decision trees according to the classification result accuracy; S313: Determine the fine estimation range and fine estimation step size of the number of decision trees according to the rough estimation value of the number of decision trees; S314: Traverse the fine estimation range according to the fine estimation step size for the number of decision trees. For each value of the number of decision trees, construct a random forest classifier, divide the optimized dataset into a training set and a test set according to a preset ratio, use the training set to train the random forest classifier, and then use the test set to test the accuracy of the classification results. After the traversal is completed, determine the optimal value of the number of decision trees according to the accuracy of the classification results.
6. The method for classifying shallow sea seabed sediments according to claim 5, characterized in that, The step S320 further includes the following steps: S321: Determine the estimation range and estimation step size of min_samples_split; S322: Traverse the estimation range of min_samples_split according to the estimation step size of min_samples_split. For each value of min_samples_split, construct a random forest classifier, where the number of decision trees in the random forest classifier is the optimal value of the number of decision trees determined in step S314. Divide the optimized dataset into a training set and a test set according to a preset ratio, use the training set to train the random forest classifier, and then use the test set to test the accuracy of the classification results. After the traversal is completed, determine the optimal value of min_samples_split according to the accuracy of the classification results.
7. The method for classifying shallow sea seabed sediment according to claim 6, wherein The step S330 further includes the following steps: S331: Determine the estimation range and estimation step size of min_samples_leaf; S332: Traverse the estimation range of min_samples_leaf according to the estimation step size of min_samples_leaf. For each value of min_samples_leaf, construct a random forest classifier, where the number of decision trees in the random forest classifier is the optimal value of the number of decision trees determined in step S314, and the min_samples_split of the random forest classifier is the optimal value of min_samples_split determined in step S322. Divide the optimized dataset into a training set and a test set according to a preset ratio, use the training set to train the random forest classifier, and then use the test set to test the accuracy of the classification results. After the traversal of the optimized dataset is completed, determine the optimal value of min_samples_leaf according to the accuracy of the classification results; S333: Determine the optimal random forest classifier. The number of decision trees, min_samples_split, and min_samples_leaf of the optimal random forest classifier are the optimal value of the number of decision trees, the optimal value of min_samples_split, and the optimal value of min_samples_leaf determined in step S314, step S322, and step S332, respectively.
8. The shallow sea seabed sediment classification method according to any one of claims 4 to 7, characterized in that The step S400 further includes the following steps: S410: Extract the high-frequency signal sequence D to be classified = {d1, d2, …, d M} and the low-frequency signal sequence C to be classified = {c1, c2, …, c M} from the shallow profile echo data to be classified; S420: Find the first high-frequency signal \(d\) greater than the echo intensity threshold from the subsequence \(\{d\) h_min , d h_min+1 , …, d M \} of \(D\), and use \(h\) as the alignment serial number of the shallow profile echo data to be classified; h S430: Optimally intercept the C using the h, the h_start, and the h_end to obtain an optimally intercepted low-frequency signal sequence C' = {c h+h_start , …, c h+h_end} for classification; S440: Input the C′ into the optimal random forest classifier for classifying the seabed sediment.
Citation Information
Patent Citations
A human body fall detection system based on ground vibration signals
BE1026660B1
Towed deep sea seabed shallow structure acoustic detection system and method
CN111308474A