Marine data extraction method of k-medoids based on neighborhood variance optimization

Through the k-medoids method of domain variance optimization, the neighborhood variance of marine environmental data points was selected as the initial clustering center point, which solved the problem of difficulty in integrating multi-source ocean data, and achieved efficient and accurate data extraction and management.

CN120429672APending Publication Date: 2025-08-05NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510471321.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing marine environmental data processing technologies have problems of inefficiency and insufficient accuracy in multi-source data integration and database management, especially due to the difficulty of data integration caused by diverse data sources and inconsistent formats and complexity of database index tables.

Method used

The k-medoids method based on the optimization of the domain variance is adopted. By calculating the neighborhood variance of each data point, the data points with smaller variance are selected as the initial clustering center point, clustering and building a search database to improve the data extraction efficiency and accuracy.

Benefits of technology

It improves the efficiency and accuracy of marine environmental data extraction, reduces the impact of noise and outliers, adapts to the dynamic changes of marine environmental data, and quickly locates target data points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429672A_ABST
    Figure CN120429672A_ABST
Patent Text Reader

Abstract

The invention relates to an ocean data extraction method of k-medoids based on neighborhood variance optimization. The method comprises the following steps: calculating a plurality of first neighborhood data points of a first data point in a target ocean environment data set; determining a variance between the first data point and a plurality of first neighborhood data points; selecting K first data points with variances smaller than a second threshold as K initial clustering center points; clustering the target marine environment data set according to the K initial clustering center points; constructing a retrieval database according to a clustering result; and extracting target data points according to the retrieval database. According to the method, multiple first neighborhood data points corresponding to each data point are found for each data point, the variance of each data point in the neighborhood of the data point is calculated based on the multiple first neighborhood data points, the compactness of data distribution in the neighborhood is quantized through the variance, and part of data points with relatively small variance are selected as initial clustering center points; the method can select an initial clustering center point with a scientific basis, and improves the extraction efficiency and accuracy of the marine environment data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of ocean data processing, and in particular to a method for extracting ocean data based on k-medoids with neighborhood variance optimization.

[0002] Background technology With the continuous deepening of marine scientific research and the development of various marine-related industries, the demand for marine environmental data is increasing. Marine environmental data processing technology mainly involves the collection, organization, analysis, and storage of marine environmental data from various channels, in order to provide accurate and reliable data support for marine scientific research, marine resource development, marine environmental protection, marine disaster warning, and other aspects.

[0003] Several technologies and methods exist for efficiently extracting and managing multi-source marine environmental data. However, current marine environmental data processing technologies still face challenges. This is due to diverse data sources and inconsistent data formats and standards, which make data integration difficult, database index tables overly complex, and data extraction inefficient. Furthermore, due to the large volume and rapid update rate of marine environmental data, traditional data storage and management methods struggle to meet the requirements for efficient extraction. Summary of the Invention

[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0005] The main purpose of the embodiments of the present disclosure is to propose a k-medoids ocean data extraction method based on neighborhood variance optimization, which can improve the extraction efficiency and accuracy of marine environmental data.

[0006] A first aspect of an embodiment of the present application proposes a method for extracting ocean data based on k-medoids with neighborhood variance optimization, the method comprising: Acquire a target marine environment data set; the target marine environment data set includes a plurality of data points of target marine environment elements; Calculating a plurality of first neighboring data points of a first data point in the target marine environment dataset; the first data point is any data point in the target marine environment dataset; the first neighboring data point is a data point in the target marine environment dataset whose distance from the first data point is less than a first threshold; determining a variance between the first data point and the plurality of first neighborhood data points; Selecting K first data points whose variance is less than a second threshold as K initial cluster center points; K is a positive integer greater than 1; Clustering the target marine environment data set according to the K initial cluster center points to obtain a clustering result; constructing a retrieval database according to the clustering results; Target data points are extracted according to the search database.

[0007] The present disclosure provides a method for extracting ocean data based on k-medoids with neighborhood variance optimization, which has at least the following beneficial effects: This method finds multiple first-neighborhood data points corresponding to each data point, calculates the variance of each data point in its neighborhood based on the multiple first-neighborhood data points, quantifies the compactness of the data distribution in the neighborhood through the variance, and then selects some data points with relatively small variance as initial clustering centers. This method can select initial clustering centers with a scientific basis; then, the selected initial clustering centers are used to cluster the target marine environment data set and construct a retrieval database, thereby improving the efficiency and accuracy of marine environment data extraction.

[0008] In some embodiments, the variance is a local variance; the formula for calculating the local variance is: ; in, The first data point The local variance of to is the first to the last first neighborhood data point among the plurality of first neighborhood data points, The first data point and the first neighboring data point The distance between is the total number of the plurality of first neighborhood data points; .

[0009] In some embodiments, obtaining to The process includes: Calculate the ;in The calculation formula is: ; Based on the In order from small to large, select first neighborhood data points; wherein the first threshold and the largest one Related, for The transpose of .

[0010] In some embodiments, after determining the variance between the first data point and the plurality of first neighborhood data points, the method further includes: Calculating a standard deviation based on the local variance; the standard deviation is the square root of the local variance; Determining a plurality of second neighboring data points of the first data point from the plurality of first neighboring data points according to the standard deviation; wherein the distance between any of the second neighboring data points and the first data point is less than the standard deviation; The step of selecting K first data points whose variances are smaller than a second threshold as K initial cluster center points comprises the following steps: sorting the first data points in the target ocean environment dataset in ascending order according to the size of the variance to obtain an intermediate ocean environment dataset; Performing a cluster center point selection operation, the cluster center point selection operation comprising: selecting a first data point from the intermediate marine environment dataset, using the first data point as an initial cluster center point, and updating the target marine environment dataset; wherein, updating the target marine environment dataset comprises: removing the plurality of second neighborhood data points corresponding to the first data point from the target marine environment dataset; Performing the cluster center point selection operation according to the updated intermediate ocean environment data set; And so on, until the selection of K initial cluster centers is completed.

[0011] In some embodiments, obtaining a target marine environment dataset includes: Acquire multiple initial ocean environment data points; Based on the target marine environment elements, nonlinear least squares fitting is performed on the multiple initial marine environment data points to obtain the multiple data points; the data points are data vectors; The target ocean environment dataset is constructed based on the multiple data points.

[0012] In some embodiments, clustering the target marine environment dataset according to the K initial cluster center points to obtain a clustering result includes: Performing a first operation, the first operation comprising: assigning a second data point in the target marine environment dataset to a cluster containing the initial cluster center point with the shortest Euclidean distance, to obtain K clusters; the second data point is any one of the remaining data points in the target marine environment dataset except the K initial cluster center points; Performing a second operation, the second operation comprising: updating the center points of the K clusters so that the sum of the Euclidean distances between the center points of the K clusters and the remaining second data points in the clusters is minimized; Continue to perform the first operation and the second operation, and so on, until the clustering result does not change or reaches a preset number of iterations, and obtain a clustering result.

[0013] In some embodiments, constructing a search database based on the clustering results includes: Constructing K database index tables based on the K clusters in the clustering result; wherein the database index tables contain environmental information and pointers corresponding to the data points; A retrieval database is constructed according to the K database index tables.

[0014] A second aspect of the embodiments of the present application provides a k-medoids ocean data extraction system based on neighborhood variance optimization, the system comprising: A data acquisition unit is used to acquire a target marine environment data set; the target marine environment data set has multiple data points of target marine environment elements; a neighborhood point extraction unit, configured to calculate a plurality of first neighborhood data points of a first data point in the target marine environment dataset; the first data point is any data point in the target marine environment dataset; the first neighborhood data point is a data point in the target marine environment dataset whose distance from the first data point is less than a first threshold; a variance calculation unit, configured to determine a variance between the first data point and the plurality of first neighborhood data points; A cluster point selection unit, configured to select K first data points whose variance is less than a second threshold as K initial cluster center points; K is a positive integer greater than 1; a data clustering unit, configured to cluster the target marine environment data set according to the K initial cluster center points to obtain a clustering result; A database construction unit, configured to construct a retrieval database according to the clustering results; The ocean data retrieval unit is used to extract target data points according to the retrieval database.

[0015] A third aspect of an embodiment of the present application proposes an electronic device, at least one controller and a memory for communicating with the at least one controller; the memory stores instructions that can be executed by the at least one controller, and the instructions are executed by the controller to enable the controller to perform the ocean data extraction method based on k-medoids with neighborhood variance optimization as described in the first aspect.

[0016] A fourth aspect of an embodiment of the present application proposes a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned k-medoids ocean data extraction method based on neighborhood variance optimization.

[0017] It can be understood that the beneficial effects of the above-mentioned second to fourth aspects compared with the relevant technologies are the same as the beneficial effects of the above-mentioned first aspect compared with the relevant technologies. Please refer to the relevant description in the above-mentioned first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 This is a flowchart of a method for extracting ocean data based on k-medoids with neighborhood variance optimization provided by one embodiment of the present application; Figure 2 is a flowchart of a method for extracting ocean data based on k-medoids with neighborhood variance optimization provided by another embodiment of the present application; Figure 3 This is a schematic diagram of the query results and query time of the bottom sediment raw data provided by this application; Figure 4 This is a schematic diagram of the bottom sediment data query results and query time using this method provided by this application; Figure 5 Schematic diagram of the ocean data extraction system based on k-medoids with neighborhood variance optimization provided by this application; Figure 6 is a schematic diagram of an electronic device provided in this application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0021] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application. One embodiment of the present application provides a method for extracting ocean data based on k-medoids with neighborhood variance optimization, the method comprising the following steps S110 to S170: Step S110 , obtaining a target marine environment data set; the target marine environment data set includes a plurality of data points of target marine environment elements.

[0023] Step S120 , calculating a plurality of first neighborhood data points of a first data point in the target ocean environment dataset.

[0024] Step S130: determining the variance between the first data point and a plurality of first neighborhood data points.

[0025] Step S140 : Select K first data points whose variances are smaller than a second threshold as K initial cluster center points.

[0026] Step S150 , clustering the target marine environment dataset according to the K initial cluster center points to obtain a clustering result.

[0027] Step S160: constructing a retrieval database based on the clustering results.

[0028] Step S170: extract target data points according to the search database.

[0029] In step S110, the target marine environment dataset contains multiple marine environment data points (hereinafter referred to as data points), which can be in vector form. Target marine environment factors include, but are not limited to, seawater temperature, salinity, density, sound velocity, water depth, ocean currents, and geology. For example, the target marine environment dataset includes data points collected to study geological factors. The purpose of returning the target marine environment dataset in step S110 is to provide a data source for subsequent clustering operations.

[0030] In step S120, the first data point is any data point in the target marine environment data set. In this embodiment, first, for each data point (i.e., the first data point), corresponding first neighborhood data points are selected. The specific operations include: First, determine the distance (such as Euclidean distance, etc.) between the first data point and each data point in the target marine environment dataset; Then, data points whose distance from the first data point is less than a first threshold are selected, and such data points are used as a plurality of first neighborhood data points.

[0031] The first threshold value can be set based on experience, and similarly, the number of selected first neighborhood data points can be set based on experience. For details, please refer to the introduction of the subsequent embodiments.

[0032] In step S130 , after obtaining the plurality of first neighborhood data points for each data point, this embodiment continues to calculate the variance between each data point and the corresponding plurality of first neighborhood data points.

[0033] In some embodiments of the present application, this variance is a local variance, which is calculated as follows: ; ; in, The first data point The local variance of to is the first first neighborhood data point to the last first neighborhood data point among multiple first neighborhood data points, The first data point and the first neighboring data point The distance between is the total number of multiple first neighborhood data points.

[0034] The variance is calculated in this application for the subsequent selection of the initial cluster center point.

[0035] In step S140 , a second threshold is preset, and K first data points whose variances are smaller than the second threshold are selected as K initial cluster center points.

[0036] The second threshold here can be set based on experience, and the value of K can also be set based on experience.

[0037] In step S150, the target marine environment data set is clustered according to K initial cluster center points to obtain a clustering result. The clustering result is K clusters, and each cluster is a collection of multiple data points with similar characteristics of target marine environment elements.

[0038] The cluster centers selected by existing clustering algorithms are random, so the results are unstable. This method finds multiple first-neighborhood data points corresponding to each data point, calculates the variance of each data point in its neighborhood based on the multiple first-neighborhood data points, quantifies the compactness of the data distribution in the neighborhood through the variance, and then selects some data points with relatively small variance as the initial cluster centers. This method can select initial cluster centers with a scientific basis. The selected initial cluster centers are then used to cluster the target marine environment data set and construct a retrieval database, thereby improving the efficiency and accuracy of marine environment data extraction.

[0039] In some embodiments of the present application, obtaining a target marine environment dataset includes the following steps S210 to S230: Step S210: Acquire a plurality of initial ocean environment data points.

[0040] Step S220 , based on the target ocean environment elements, a plurality of initial ocean environment data points are subjected to nonlinear least square fitting to generate a plurality of data points; the data points are data vectors.

[0041] Step S230: constructing a target ocean environment dataset based on the multiple data points.

[0042] In step S210, initial ocean environment data points can be acquired from ocean observation stations, satellite remote sensing, drones, seabed observations, and other methods. This data covers a variety of ocean environmental factors, including seawater temperature, salinity, density, sound speed, water depth, currents (speed and direction), and bottom sediments. These data are spatiotemporal, multi-source, multi-dimensional, and involve large amounts of data.

[0043] In step S220, a genetic algorithm is used to perform a nonlinear least squares fit on the multiple initial marine environmental data points based on the target marine environmental factor (e.g., geology) to convert them into multiple data points. Because marine environmental data generally fluctuates significantly, a nonlinear least squares fit is performed on a single marine environmental factor in advance. This fits the original scattered marine environmental data into a smooth curve, thereby obtaining a characteristic vector that reflects the characteristics of each marine environmental factor.

[0044] In step S230, a target ocean environment dataset is constructed based on the multiple data points.

[0045] The marine environment feature vector obtained by this method based on the least squares fitting of the genetic algorithm can fully reflect the data characteristics of the corresponding marine environmental elements. Combined with the method of allocating data points based on the marine environment data feature vector, it can improve the accuracy of data point allocation and provide a more accurate index for subsequent data extraction.

[0046] In some embodiments of the present application, obtaining to The process includes: Step S310, calculate ;in The calculation formula is: ; Step S320, based on Select from the target marine environment dataset in order from small to large First neighborhood data points; among them, the first threshold and the largest one Related, for The transpose of .

[0047] In some embodiments of the present application, after determining the variance between the first data point and the plurality of first neighborhood data points in step S130, the following steps S410 to S420 are further included: Step S410, calculating the standard deviation based on the local variance; the standard deviation is the square root of the local variance; ; The standard deviation is .

[0048] Step S420, determining a plurality of second neighboring data points of the first data point from the plurality of first neighboring data points according to the standard deviation; the distance between any second neighboring data point and the first data point being less than the standard deviation; ; The first data point n is the total number of data points of the target marine environment data.

[0049] Step S140 includes the following steps S430 to S460: Step S430, sorting the first data points in the target ocean environment dataset in ascending order according to the size of the variance to obtain an intermediate ocean environment dataset; Step S440: performing a cluster center point selection operation, the cluster center point selection operation including: selecting a first data point from the intermediate marine environment dataset, using the first data point as the initial cluster center point, and updating the target marine environment dataset; wherein updating the target marine environment dataset includes: removing a plurality of second neighborhood data points corresponding to the first data point in the target marine environment dataset; Step S450, performing a cluster center point selection operation based on the updated intermediate ocean environment data set; Step S460, and so on, until the selection of K initial cluster centers is completed.

[0050] In some embodiments of the present application, step S150 clusters the target marine environment dataset according to K initial cluster centers to obtain a clustering result, including the following steps S510 to S530: Step S510, performing a first operation, the first operation comprising: assigning a second data point in the target marine environment dataset to a cluster containing an initial cluster center point having the shortest Euclidean distance, to obtain K clusters; the second data point is any one of the remaining data points in the target marine environment dataset except the K initial cluster centers; Step S520: performing a second operation, the second operation including: updating the center points of the K clusters so as to minimize the sum of the Euclidean distances between the center points of the K clusters and the remaining second data points in the clusters; Step S530 : Continue to perform the first operation and the second operation, and so on, until the clustering result does not change or reaches a preset number of iterations, and obtain the clustering result.

[0051] The following is an introduction to K-medoids: The K-medoids clustering algorithm is a partitioning clustering method used to partition a dataset into k clusters, where k is the number of clusters specified by the user. Unlike the K-means algorithm, the K-medoids algorithm selects actual data points as cluster centers (called medoids) rather than calculating the mean of the data points within a cluster. This makes the K-medoids algorithm more robust to outliers because it is not affected by extreme values.

[0052] This method selects actual data points as cluster centers and updates them by calculating the data point with the smallest sum of distances. This reduces the impact of noise and outliers, improving the stability and accuracy of the clustering results. Furthermore, the iterative optimization process of repeatedly assigning data points and updating cluster centers gradually stabilizes the clustering results, adapting to the dynamic changes in marine environmental data.

[0053] In some embodiments of the present application, constructing a search database according to the clustering results in step S160 includes the following steps S610 to S620: Step S610: construct K database index tables based on the K clusters in the clustering result; wherein the database index tables contain the environment information and pointers corresponding to the data points; Step S620: construct a search database based on the K database index tables.

[0054] The database index table established by this method based on the clustering results contains key information and data pointers, which can quickly locate the corresponding clusters and improve data extraction efficiency.

[0055] For ease of understanding, the present application provides an embodiment, which provides a method for extracting ocean data based on k-medoids with neighborhood variance optimization. The present application includes the following steps S910 to S960: Step S910 , acquiring multi-source ocean environment grid data and preprocessing the multi-source ocean environment grid data.

[0056] Obtain multi-source ocean environment grid data from ocean observation stations, satellite remote sensing, ocean drones, and seabed observations; The multi-source marine environment grid data were cleaned and standardized to remove outliers for subsequent cluster analysis.

[0057] Step S920 , performing genetic-based nonlinear least squares fitting on the pre-processed multi-source ocean environment grid data.

[0058] Since marine environmental data generally have large fluctuations, a single target marine environmental factor (such as seawater temperature, salinity, density, sound speed, water depth, ocean current, and geology) is pre-fitted with nonlinear least squares fitting, so that the original scattered multi-source marine environmental grid data is fitted into a smooth curve, thereby obtaining a characteristic vector that reflects the characteristics of the target marine environmental factor, as shown in the following formula: ; Where, is the depth value of the ocean environment data point of the target ocean environment element to be fitted, is the corresponding true value, is the objective function, that is, the environmental curve obtained by fitting the target marine environmental elements; is the fitting coefficient, which is used in subsequent clustering. is the total number of marine environment data points of the target marine environment element to be fitted.

[0059] In order to solve the defects of non-unique fitting results and large variation of fitting parameters in the fitting process, a genetic algorithm is introduced into the fitting process to find the global optimal solution of the marine environment data points (eigenvectors). The adaptability function of the genetic algorithm is designed as follows, and the definitions of the parameters are the same as above: ; Here, according to the needs, the genetic algorithm fitting can be run multiple times to obtain a stable feature vector combination.

[0060] Finally, the target marine environment dataset is obtained.

[0061] Step S930: clustering the target marine environment dataset.

[0062] Step S9310, determine the cluster center.

[0063] Unlike traditional random selection of initial centers, this method uses neighborhood variance optimization to select initial cluster centers. By calculating the neighborhood variance of each data point in the dataset, the data point with the smallest neighborhood variance is selected as the initial cluster center, improving the stability and accuracy of clustering.

[0064] Take the target marine environmental element as an example, which is bottom sediment data.

[0065] Through its characteristic parameter sample of The neighborhood is used to define the local variance.

[0066] Assume that the target marine environment dataset to be clustered is: ; in, The target marine environment dataset to be clustered , is the number of samples. Target marine environment dataset Two samples in and The distance can be described as: ; calculate and The distances of all other samples in , and the distances are arranged in ascending order according to the calculated results. The arrangement results are recorded as: ; in, To exclude A permutation of the remaining samples outside the samples.

[0067] exist Take the top of the dataset samples (different marine environmental factors in different sea areas adapted to The values are different, and the one with the highest accuracy needs to be verified by experiment. Value, usually After 20, the classification accuracy tends to be stable) of Neighborhood set, denoted as: ; At this time, the sample The local variance of can be expressed as: ; Where, Defined as: ; Define feature vector samples Standard deviation Neighborhood with radius: ; Then The standard deviation of a neighborhood with a radius of is defined as: ; During the algorithm implementation process, we first need to determine the number of clusters (Select according to user needs, for example, if it is determined that there are 10 types of substrate types, then Take 10, the temperature data needs to be classified into 6 different sea areas, then Take 6) and the neighborhood parameter .

[0068] Then calculate the sample data set Each data object in The local variance , sort the samples in the target ocean dataset in ascending order according to the local variance calculation results.

[0069] The data set after arrangement is recorded as , select the first eigenvector sample from For the initial cluster center point of a cluster, calculate its standard deviation , and use it as the radius to calculate Neighborhood , from the dataset After removing this part of the neighborhood sample data, continue to select the sample that ranks first at this time.

[0070] Repeat this process until you have selected Initial cluster centers. Step S9320, assigning data points: Based on the K-medoids algorithm principle and the Euclidean distance principle, the original data set is assigned to The samples in are assigned to each cluster center point, and the data set is obtained The initial cluster of .

[0071] Step S9330, updating the cluster center: updating the center point of each initial cluster so that the sum of the Euclidean distances between it and the remaining samples in the cluster is minimized.

[0072] Step S9340, iterative optimization: repeat the steps of allocating data points and updating cluster centers until the clustering result no longer changes or the preset number of iterations is reached.

[0073] During the iteration process, the distribution of cluster centers and data points is continuously adjusted to make the clustering results gradually stable.

[0074] Step S940: Create a database index table.

[0075] Based on the clustering results, a database index table is created for each cluster. The index table contains key information such as data type, resolution, data time, and pointers to all data in the cluster.

[0076] In this way, in the subsequent data extraction process, the corresponding cluster can be quickly located according to user needs, thereby improving the efficiency of data extraction.

[0077] Step S950: data extraction.

[0078] When a user makes a data extraction request, a search is performed in the newly created database index table based on the data type, resolution, data time and other conditions in the request.

[0079] After finding the corresponding cluster, extract the data that meets the conditions from the cluster. If you need to extract data for multiple conditions, you can do so by merging the results of multiple clusters.

[0080] Step S960: performance evaluation and optimization.

[0081] The performance of this method was evaluated, including the efficiency and accuracy of data extraction, the stability of clustering results, etc. Based on the evaluation results, the parameters of the clustering algorithm were adjusted and the structure of the database index table was optimized to continuously improve the performance of the method.

[0082] Based on the existing technology platform and taking bottom sediment data as an example, the original data query time and the data query time after introducing the method of steps S910 to S960 above were tested for the same sea area. The data query results and query time were compared and verified. The data query results using this method were basically consistent with the original data query results, and the data query time could be improved by 22.43%.

[0083] This method has at least the following beneficial effects: The neighborhood variance optimization adopted in this method to select the initial cluster center point can effectively improve the stability and accuracy of clustering.

[0084] The marine environment feature vector obtained by this method through least squares fitting based on genetic algorithm can fully reflect the data characteristics of the corresponding marine environmental elements. Combined with the method of allocating data points based on the marine environment data feature vector, it can improve the accuracy of data point allocation and provide a more accurate index for subsequent data extraction.

[0085] This method selects actual data points as cluster centers and updates them by calculating the data points with the smallest sum of distances, which can reduce the impact of noise and outliers and improve the stability and accuracy of clustering results.

[0086] This method repeatedly allocates data points and updates the iterative optimization process of cluster centers, which can make the clustering results gradually stable and adapt to the dynamic changes of marine environmental data.

[0087] The database index table established by this method based on the clustering results contains key information and data pointers, which can quickly locate the corresponding clusters and improve data extraction efficiency.

[0088] like Figure 5 One embodiment of the present application provides a k-medoids ocean data extraction system based on neighborhood variance optimization, the system comprising: The data acquisition unit 1100 is used to acquire a target ocean environment data set; the target ocean environment data set includes multiple data points of target ocean environment elements.

[0089] The neighborhood point extraction unit 1200 is used to calculate multiple first neighborhood data points of a first data point in the target marine environment data set; the first data point is any data point in the target marine environment data set; the first neighborhood data point is a data point in the target marine environment data set whose distance to the first data point is less than a first threshold.

[0090] The variance calculation unit 1300 is used to determine the variance between the first data point and a plurality of first neighborhood data points.

[0091] The cluster point selection unit 1400 is used to select K first data points whose variance is less than a second threshold as K initial cluster center points; K is a positive integer greater than 1.

[0092] The data clustering unit 1500 is used to cluster the target marine environment data set according to K initial cluster center points to obtain a clustering result.

[0093] The database construction unit 1600 is used to construct a retrieval database according to the clustering results.

[0094] The ocean data retrieval unit 1700 is used to extract target data points according to the retrieval database.

[0095] It should be noted that the ocean data extraction system based on k-medoids with neighborhood variance optimization provided in this embodiment and the ocean data extraction method based on k-medoids with neighborhood variance optimization described above are based on the same inventive concept. Therefore, the relevant content of the ocean data extraction method based on k-medoids with neighborhood variance optimization described above is also applicable to the ocean data extraction system based on k-medoids with neighborhood variance optimization. Therefore, they will not be repeated here.

[0096] like Figure 6 , an embodiment of the present application further provides an electronic device, the electronic device comprising: at least one memory; at least one processor; at least one program; The programs are stored in the memory, and the processor executes at least one program to implement the above-mentioned k-medoids ocean data extraction method based on neighborhood variance optimization in the present disclosure.

[0097] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.

[0098] The electronic device according to the embodiment of the present application is described in detail below.

[0099] The processor 1600 may be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention. Memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). Memory 1700 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in memory 1700 and is called by processor 1600 to execute the ocean data extraction method based on k-medoids with neighborhood variance optimization according to the embodiments of the present invention.

[0100] Input / output interface 1800, used for information input and output; Communication interface 1900, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.); Bus 2000 , which transmits information between various components of the device (e.g., processor 1600 , memory 1700 , input / output interface 1800 , and communication interface 1900 ); The processor 1600 , the memory 1700 , the input / output interface 1800 , and the communication interface 1900 are connected to each other in communication within the device via the bus 2000 .

[0101] An embodiment of the present invention also provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, which are used to enable a computer to execute the above-mentioned k-medoids ocean data extraction method based on neighborhood variance optimization.

[0102] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs. Furthermore, memory can include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device.

[0103] In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and the remote memory may be connected to the processor via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0104] The embodiments described in the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0105] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0106] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0107] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0108] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0109] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0110] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0111] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0112] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0114] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above-mentioned implementation methods. Technical personnel familiar with the art can also make various equivalent modifications or substitutions without violating the spirit of the embodiments of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the embodiments of the present application.

Claims

1. A k-medoids ocean data extraction method based on neighborhood variance optimization, characterized in that: The method comprises: Acquire a target marine environment data set; the target marine environment data set includes a plurality of data points of target marine environment elements; Calculating a plurality of first neighboring data points of a first data point in the target marine environment dataset; the first data point is any data point in the target marine environment dataset; the first neighboring data point is a data point in the target marine environment dataset whose distance from the first data point is less than a first threshold; determining a variance between the first data point and the plurality of first neighborhood data points; Selecting K first data points whose variance is less than a second threshold as K initial cluster center points; K is a positive integer greater than 1; Clustering the target marine environment data set according to the K initial cluster center points to obtain a clustering result; constructing a retrieval database according to the clustering results; Target data points are extracted according to the search database.

2. The ocean data extraction method based on k-medoids with neighborhood variance optimization according to claim 1, characterized in that: The variance is the local variance; the formula for calculating the local variance is: ; in, The first data point The local variance of to is the first to the last first neighborhood data point among the plurality of first neighborhood data points, The first data point and the first neighboring data point The distance between is the total number of the plurality of first neighborhood data points; 。 3. The ocean data extraction method based on k-medoids with neighborhood variance optimization according to claim 2, characterized in that: Get to The process includes: Calculate the ;in The calculation formula is: ; Based on the In order from small to large, select first neighborhood data points; wherein the first threshold and the largest one Related, for The transpose of .

4. The ocean data extraction method based on k-medoids with neighborhood variance optimization according to claim 2, characterized in that: After determining the variance between the first data point and the plurality of first neighborhood data points, the method further includes: Calculating a standard deviation based on the local variance; the standard deviation is the square root of the local variance; Determining a plurality of second neighboring data points of the first data point from the plurality of first neighboring data points according to the standard deviation; wherein the distance between any of the second neighboring data points and the first data point is less than the standard deviation; The step of selecting K first data points whose variances are smaller than a second threshold as K initial cluster center points comprises the following steps: sorting the first data points in the target ocean environment dataset in ascending order according to the size of the variance to obtain an intermediate ocean environment dataset; Performing a cluster center point selection operation, the cluster center point selection operation comprising: selecting a first data point from the intermediate marine environment dataset, using the first data point as an initial cluster center point, and updating the target marine environment dataset; wherein, updating the target marine environment dataset comprises: removing the plurality of second neighborhood data points corresponding to the first data point from the target marine environment dataset; Performing the cluster center point selection operation according to the updated intermediate ocean environment data set; And so on, until the selection of K initial cluster centers is completed.

5. The ocean data extraction method based on k-medoids with neighborhood variance optimization according to claim 1, characterized in that: The acquiring of the target marine environment data set comprises: Acquire multiple initial ocean environment data points; Based on the target marine environment elements, nonlinear least squares fitting is performed on the multiple initial marine environment data points to obtain the multiple data points; the data points are data vectors; The target ocean environment dataset is constructed based on the multiple data points.

6. The ocean data extraction method based on k-medoids with neighborhood variance optimization according to claim 1, characterized in that: Clustering the target marine environment data set according to the K initial cluster center points to obtain a clustering result includes: Performing a first operation, the first operation comprising: assigning a second data point in the target marine environment dataset to a cluster containing the initial cluster center point with the shortest Euclidean distance, to obtain K clusters; the second data point is any one of the remaining data points in the target marine environment dataset except the K initial cluster center points; Performing a second operation, the second operation comprising: updating the center points of the K clusters so that the sum of the Euclidean distances between the center points of the K clusters and the remaining second data points in the clusters is minimized; Continue to perform the first operation and the second operation, and so on, until the clustering result does not change or reaches a preset number of iterations, and obtain a clustering result.

7. The ocean data extraction method based on k-medoids with neighborhood variance optimization according to claim 1, characterized in that: The step of constructing a retrieval database according to the clustering results includes: Constructing K database index tables based on the K clusters in the clustering result; wherein the database index tables contain environmental information and pointers corresponding to the data points; A retrieval database is constructed according to the K database index tables.

8. A k-medoids ocean data extraction system based on neighborhood variance optimization, characterized in that: The system comprises: A data acquisition unit is used to acquire a target marine environment data set; the target marine environment data set has multiple data points of target marine environment elements; a neighborhood point extraction unit, configured to calculate a plurality of first neighborhood data points of a first data point in the target marine environment dataset; the first data point is any data point in the target marine environment dataset; the first neighborhood data point is a data point in the target marine environment dataset whose distance from the first data point is less than a first threshold; a variance calculation unit, configured to determine a variance between the first data point and the plurality of first neighborhood data points; A cluster point selection unit, configured to select K first data points whose variance is less than a second threshold as K initial cluster center points; K is a positive integer greater than 1; A data clustering unit, configured to cluster the target marine environment data set according to the K initial cluster center points to obtain a clustering result; A database construction unit, configured to construct a retrieval database according to the clustering results; The ocean data retrieval unit is used to extract target data points according to the retrieval database.

9. An electronic device, characterized in that: include: at least one controller and a memory for communicatively coupling with the at least one controller; The memory stores instructions that can be executed by the at least one controller, and the instructions are executed by the controller to enable the controller to perform the k-medoids ocean data extraction method based on neighborhood variance optimization according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the k-medoids ocean data extraction method based on neighborhood variance optimization according to any one of claims 1 to 7.