A method and system for fusion processing of multi-source data of marine environment
Through wireless communication and genetic algorithms, data allocation is optimized, and data collection and processing in traditional marine environmental monitoring is solved, efficient and accurate data processing and environmental monitoring is achieved, and scientific decision-making and risk warning are supported.
Patent Information
- Application Number
- CN202411794739.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-12-09
AI Technical Summary
The lack of efficient timing control mechanisms in traditional marine environmental monitoring leads to conflicts and interference in the data acquisition process, reducing the accuracy and efficiency of data acquisition. In addition, traditional data processing methods are inefficient in the face of large-scale multi-source data, and cannot achieve balanced distribution of data and efficient utilization of computing resources.
Wireless communication protocol is used to connect multiple types of sensor nodes, optimize data distribution through data sharding and dynamic allocation, and use genetic algorithms to allocate sharded data sets, and optimize data processing flow with performance evaluation indicators and data fusion technology.
It improves the real-time and accuracy of data processing, enhances the reliability and flexibility of the system, can more comprehensively reflect the changing laws of the marine environment, supports scientific decision-making in environmental protection and management, and achieves a timely warning of potential environmental risks.
Smart Images

Figure CN119271519B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for fusion processing of multi-source data of an ocean environment. Background Art
[0002] In the field of marine environmental monitoring, some traditional data acquisition methods lack an efficient timing control mechanism, which may lead to conflicts and interference during the data acquisition process, reducing the accuracy and efficiency of data acquisition.
[0003] This issue resulted in inconsistent data quality, impacting the reliability of subsequent data analysis. In the past, marine environmental monitoring often relied on a single type of sensor, which failed to fully reflect the multidimensional characteristics of the ocean environment. This limitation resulted in incomplete monitoring results, making it difficult to capture the complex changes in the marine environment.
[0004] Furthermore, when processing data after monitoring, traditional data processing methods sometimes suffer from inefficiencies when dealing with large-scale, multi-source data. Consequently, they are unable to achieve balanced data distribution and efficient utilization of computing resources. This leads to bottlenecks in data processing capabilities and makes it impossible to meet the growing demand for data processing. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and system for fusion processing of multi-source data of marine environment, which optimizes data distribution through data sharding and dynamic allocation, thereby improving processing efficiency.
[0006] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0007] In a first aspect, a method for fusion processing of multi-source data of an ocean environment is provided, the method comprising:
[0008] Step 11, using wireless communication protocols to connect various types of sensor nodes;
[0009] Step 12: Isolate the sensor during the acquisition process to obtain ocean environment data and data from the edge nodes of the device;
[0010] Step 13: Fusing the ocean environment data with the edge node data to obtain a fused data set;
[0011] Step 14: divide the fused data set into sharded data sets, and preliminarily distribute the sharded data sets to each heterogeneous node to obtain a distribution result;
[0012] Step 15: Set performance evaluation indicators, perform performance tests on each heterogeneous node according to the performance evaluation indicators, and record the performance indicators of each node when processing different sharded data sets to obtain test results;
[0013] Step 16: Analyze the test results and identify abnormal nodes and abnormal shard data sets;
[0014] Step 17: Optimize based on the abnormal nodes and the abnormal sharded data set to obtain an optimization result.
[0015] Furthermore, the sensors are relatively isolated during the collection process to obtain marine environmental data and data at the edge nodes of the equipment, including:
[0016] Set a corresponding sampling period for each sensor and calculate the lowest common multiple of the sampling periods of all sensors to find a time point at which all sensors complete their respective sampling periods.
[0017] Determine the sampling time of the first sensor as the starting point of the entire acquisition process;
[0018] Calculate the delay waiting time for the remaining sensors based on the lowest common multiple of all sensor sampling periods, so that each sensor will collect data isolated from the remaining sensors within the corresponding sampling period;
[0019] After calculating their respective delay waiting times, each sensor starts collecting ocean environment data and data at the sensor edge nodes;
[0020] After the ocean environment data and the data at the edge nodes of the sensor are collected, preprocessing is performed to obtain the final ocean environment data and the data at the edge nodes of the device.
[0021] Furthermore, the marine environment data and the edge node data are fused to obtain a fused data set, including:
[0022] Map the final ocean environment data and the edge node data of the device to the same time grid to obtain time-aligned data;
[0023] The time-aligned data is passed through Perform data fusion to obtain a fused data set, where is the fusion result, and They are marine environment data and edge node data, is the weight of historical data, is the error variance of the marine environmental data, is the error variance of edge node data; is the smoothing factor; is the index, is the number of historical moments; Represents time.
[0024] Furthermore, the fused dataset is divided into sharded datasets, and the sharded datasets are preliminarily distributed to various heterogeneous nodes to obtain distribution results, including:
[0025] Analyze the fused dataset to identify key characteristics between data columns, including correlation, data size, and update frequency;
[0026] Based on key features, a vertical sharding strategy is adopted to assign related data columns to the same node;
[0027] Evaluate the current load of each heterogeneous node and, based on the current load, Calculate the total dynamic load of each node;
[0028] in Representation node Total dynamic load; 、 and Over time The weight coefficient of the change; 、 and Respectively represent the assignment to nodes Shard CPU usage, storage capacity consumption, and data transfer volume; Is the allocation decision variable, if the shard Assigned to the node ,but ;otherwise, ; 、 and Node Real-time CPU usage, storage capacity utilization, and network bandwidth utilization; Indicates the total number of shards; The index representing the shard; Represents the index of the node;
[0029] According to the dynamic total load of each node, a genetic algorithm is used to preliminarily distribute the sharded data set to each heterogeneous node to obtain the distribution result.
[0030] Furthermore, based on the dynamic total load of each node, a genetic algorithm is used to preliminarily distribute the sharded data set to each heterogeneous node to obtain the distribution results, including:
[0031] An initial population is randomly generated, each individual represents a distribution scheme for the sharded data set, and the individual encoding is binary;
[0032] Define a fitness function to evaluate the quality of each individual. Based on the value of the fitness function, select the corresponding individual to enter the next generation through roulette. Randomly select two individuals and perform a crossover operation through single-point crossover to generate a new individual.
[0033] Perform mutation operations on the newly generated individuals, repeating the selection, crossover, and mutation operations until the preset number of iterations is reached. After the iteration is completed, the corresponding individual is output as the final solution;
[0034] The final decoder is converted into a distribution plan for the sharded dataset, and the sharded dataset is distributed according to the distribution plan to obtain the distribution result.
[0035] Furthermore, the calculation formula of the fitness function is:
[0036] ;
[0037] in, Represents the fitness value of an individual; It represents the total number of heterogeneous nodes, representing the number of different nodes participating in data distribution; Represents the node index, used to traverse all heterogeneous nodes; Indicates the assignment to the node The number of sharded datasets; and Represents the shard index; Indicates the assignment to The node The load of the sharded dataset; represents the weight coefficient; Indicates the number of all sharded data sets to be allocated; and Represents the sharded dataset index, used to traverse all sharded dataset pairs; Indicates the sharded datasets and The frequency of data interaction between sharded datasets; represents the weight coefficient; Represents the associated shard allocation indicator function. If sharded datasets and The sharded data sets are distributed to different nodes. 1; if they are assigned to the same node, then 0.
[0038] Furthermore, performance evaluation indicators include data processing speed, CPU usage, network transmission delay and throughput.
[0039] In a second aspect, a marine environment multi-source data fusion processing system includes:
[0040] an acquisition module for connecting various types of sensor nodes using wireless communication protocols;
[0041] The fusion module is used to relatively isolate the sensors during the acquisition process to obtain marine environmental data and data from the edge nodes of the equipment; the marine environmental data and the edge node data are fused to obtain a fused data set;
[0042] The processing module is used to set performance evaluation indicators, perform performance tests on each heterogeneous node based on the performance evaluation indicators, record the performance indicators of each node when processing different sharded data sets to obtain test results; analyze the test results, identify abnormal nodes and abnormal sharded data sets; and optimize based on the abnormal nodes and abnormal sharded data sets to obtain optimization results.
[0043] According to a third aspect, a computing device includes:
[0044] one or more processors;
[0045] The storage device is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method.
[0046] In a fourth aspect, a computer-readable storage medium stores a program, which implements the method when executed by a processor.
[0047] The above solution of the present invention includes at least the following beneficial effects:
[0048] Wireless communication enables real-time data acquisition from each sensor node, reducing data transmission latency. Combined with software-based delay waiting techniques, this allows for efficient data collection and processing. Fusion of data from diverse sensor types (such as temperature, salinity, dissolved oxygen, and turbidity) creates a rich and diverse dataset. Sharding the dataset and initially distributing it across heterogeneous nodes yields a distribution result. Dynamically allocating the sharded dataset based on node resource weights optimizes data distribution, fully utilizing the computing power of each node and improving overall data processing efficiency. Optimizing data distribution to ensure adjacent or related data is distributed across as many nodes as possible enhances data reliability and fault tolerance. Even if a node fails, the entire dataset is not lost or corrupted. Analyzing the characteristics of the sharded dataset based on this optimized data distribution can help researchers gain a deeper understanding of the changing patterns and trends of the marine environment. Fitting the data based on its characteristics yields more accurate data results, enabling predictions and early warnings of future marine environmental changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 The figure is a flow chart of a method for fusion processing of multi-source data of an ocean environment provided by an embodiment of the present invention.
[0050] Figure 2 This is a schematic diagram of a marine environment multi-source data fusion processing system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0051] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0052] like Figure 1 As shown, an embodiment of the present invention provides a method for fusion processing of multi-source data of marine environment, the method comprising the following steps:
[0053] Step 11, using wireless communication protocols to connect various types of sensor nodes;
[0054] Step 12: Isolate the sensor during the acquisition process to obtain ocean environment data and data from the edge nodes of the device;
[0055] Step 13: Fusing the ocean environment data with the edge node data to obtain a fused data set;
[0056] Step 14: divide the fused data set into sharded data sets, and preliminarily distribute the sharded data sets to each heterogeneous node to obtain a distribution result;
[0057] Step 15: Set performance evaluation indicators, perform performance tests on each heterogeneous node according to the performance evaluation indicators, and record the performance indicators of each node when processing different sharded data sets to obtain test results;
[0058] Step 16: Analyze the test results and identify abnormal nodes and abnormal shard data sets;
[0059] Step 17: Optimize based on the abnormal nodes and the abnormal sharded data set to obtain an optimization result.
[0060] In an embodiment of the present invention, the wireless communication protocol makes the deployment of sensor nodes more flexible and not restricted by wired connections. At the same time, it supports multiple types of sensor node connections, thereby enhancing the scalability of the system. Through relatively isolated acquisition methods, interference between data can be reduced, and the accuracy of marine environmental data and device edge node data can be improved. Data fusion can integrate information from different sources to provide a more comprehensive and complete marine environmental data set. By sharding and distributing the data set to different heterogeneous nodes, data can be processed in parallel, thereby improving data processing efficiency. Through performance testing, the processing power of each node can be understood, and by identifying and processing abnormal nodes and sharded data sets, the stability and reliability of the entire system can be improved. The optimized system can utilize computing resources more efficiently, reducing unnecessary energy consumption and time waste.
[0061] In the present invention, step 11, using wireless communication protocols to connect various types of sensor nodes, includes:
[0062] Use wireless communication protocols to connect various types of sensor nodes and various types of sensors, including sensors for key indicators such as temperature, salinity, dissolved oxygen, and turbidity;
[0063] In an embodiment of the present invention, real-time data collection and transmission are achieved through wireless communication protocols, which greatly reduces the delay in data transmission and improves the real-time nature and response speed of monitoring. The deployment of multiple types of sensors makes it possible to monitor multiple key environmental indicators at the same time, forming a comprehensive, multi-dimensional environmental monitoring system that more accurately reflects the actual conditions of the ecological environment. The use of wireless communication protocols makes the deployment of sensor nodes more flexible, and nodes can be added or reduced at any time as needed to expand the monitoring range or improve monitoring accuracy. Through data fusion technology, data from different sensors are integrated and analyzed, which can mine more valuable information and provide scientific decision-making support for environmental protection and management. The combination of real-time monitoring and early warning mechanisms enables timely warnings to be issued when environmental parameters are abnormal, providing a valuable time window for emergency response and reducing potential environmental risks.
[0064] In one embodiment of the present invention, step 11 uses a wireless communication protocol to connect various types of sensor nodes. The specific process includes: selecting various types of sensors for ecological environment monitoring, such as temperature sensors, salinity sensors, dissolved oxygen sensors, turbidity sensors, etc. These sensors can cover key indicators in the ecological environment, deploying sensor nodes in the monitored ocean area, ensuring that each node can effectively monitor the corresponding environmental parameters and form a certain monitoring network coverage, using an applicable wireless communication protocol to establish a wireless communication connection between the sensor node and the data center, configuring the communication parameters of each sensor node, ensuring that the data can be transmitted to the data center stably and efficiently, the sensor nodes collect environmental data in real time, such as the values of key indicators such as temperature, salinity, dissolved oxygen, turbidity, etc., and transmit the collected data to the data center in real time through the wireless communication protocol to reduce the delay in data transmission, and the data center receives data from each sensor node.
[0065] In the present invention, step 12, isolating the sensor relatively during the acquisition process to obtain ocean environment data and data at the edge node of the device, may include:
[0066] Step 121 , setting a corresponding sampling period for each sensor, and calculating the least common multiple of the sampling periods of all sensors to find a time point at which all sensors complete their respective sampling periods;
[0067] Step 122, determining the sampling time of the first sensor as the starting point of the entire acquisition process;
[0068] Step 123 , calculating respective delay waiting times for the remaining sensors based on the least common multiple of the sampling periods of all sensors, so that each sensor will perform isolated data collection relative to the remaining sensors within the corresponding sampling period;
[0069] Step 124: After calculating the respective delay waiting times, each sensor starts collecting ocean environment data and sensor edge node data;
[0070] Step 125 , after the ocean environment data and the data at the edge nodes of the sensors are collected, pre-processing is performed to obtain the final ocean environment data and the data at the edge nodes of the devices.
[0071] In an embodiment of the present invention, by setting the sampling period and calculating the least common multiple, it is possible to ensure that all sensors complete sampling synchronously at a certain point in time. The application of the least common multiple enables sensors with different sampling periods to coordinate work within a unified time frame, thereby improving the coordination of the entire data acquisition system. Setting a clear starting point for the entire acquisition process can simplify the operational flow of the data acquisition process and improve operational efficiency. By calculating the delay waiting time, it is possible to ensure that each sensor performs isolated data acquisition relative to other sensors within the corresponding sampling period, thereby avoiding data conflicts and mutual interference. Reasonable delay waiting time settings can enable sensors to use time and resources more efficiently during the data acquisition process. Each sensor performs data acquisition according to the set sampling period and delay waiting time, which can ensure the comprehensiveness and accuracy of marine environmental data and data at the sensor edge nodes. Sensors perform real-time data acquisition according to a predetermined plan, which helps to obtain and update marine environmental information in a timely manner. Through the preprocessing step, the raw data can be cleaned and sorted, noise and outliers can be removed, and the quality and reliability of the final data can be improved.
[0072] In one embodiment of the present invention, step 121 sets a corresponding sampling period for each sensor and calculates the least common multiple of all sensor sampling periods to find a time point at which all sensors complete their respective sampling periods. The specific process includes the following: In a multi-sensor system, each sensor has a sampling period, such as the sampling period of temperature, salinity, dissolved oxygen, and other sensors. To achieve synchronization and isolation of multi-sensor sampling, it is necessary to calculate the least common multiple of all sensor sampling periods. This value is the minimum time interval at which all sensor sampling periods occur simultaneously within a complete cycle. For each sensor, the least common multiple of its sampling period is calculated. The least common multiple of multiple numbers can be calculated using an algorithm. For example, the least common multiple L of two numbers a and b can be obtained using the formula L = |a×b|. The least common multiple of multiple numbers can be gradually calculated.
[0073] Step 122 , determining the sampling time of the first sensor as the starting point of the entire acquisition process. The specific process includes: the first sampling moment of each sensor, starting at time 0 or other set time.
[0074] Step 123, based on the least common multiple of all sensor sampling periods, calculate the delay waiting time for the remaining sensors, so that each sensor will collect data in isolation from the remaining sensors within the corresponding sampling period. The specific process includes: the delay waiting time is the key to ensure that each sensor samples at a relatively isolated moment; first, obtain the sampling time of the first sensor, which is usually set to 0; according to the sampling period and the first sampling time of each sensor, Calculate the delay waiting time of all sensors, where It is a sensor The sampling period, is the lowest common multiple of all sensor sampling periods, is the first sampling moment of the sensor, is the delay waiting time of other sensors, yes Divide by Each sensor adjusts its sampling time based on its delay time. Ensure that there is an appropriate time interval between different sensors at the same time to achieve the purpose of "isolation" and avoid mutual interference.
[0075] From step 124 to step 125, after calculating their respective delay waiting times, each sensor begins collecting ocean environment data and data at the sensor edge nodes. After the ocean environment data and sensor edge node data are collected, they are preprocessed to obtain the final ocean environment data and device edge node data. The specific process includes: after calculating their respective delay waiting times, each sensor adjusts the sampling time according to its delay time, and each sensor will collect data at its adjusted sampling time. The collected data may include ocean environment data (such as temperature, salinity, dissolved oxygen, etc.) and data at the device edge nodes. The collected data is subjected to denoising, outlier detection and elimination, error compensation, interpolation and missing value filling, and data normalization and standardization. Redundant data fusion processing is performed on the edge node data to eliminate random errors in the edge node data to obtain ocean environment data and device edge node data.
[0076] In the present invention, step 13, fusing the ocean environment data and the edge node data to obtain a fused data set, may include:
[0077] Step 131 maps the final ocean environment data and the data at the edge node of the device to the same time grid to obtain time-aligned data, specifically including: mapping the ocean environment data and the edge node data to the same time grid, assuming that the two data sets have different sampling frequencies or sampling time points, selecting the lowest common multiple of the sampling periods of the two data sets as the interval of the common time grid, and this time grid should cover the time range of all data; if the sampling time points of the two data sets do not overlap, interpolating the original data according to the time step of the lowest common multiple to ensure that the time series of each data set has corresponding values at the same time point, and aligning the data to the common time grid.
[0078] Step 132: Time-aligned data is passed through Perform data fusion to obtain a fused data set, where is the fusion result, and They are marine environment data and edge node data, is the weight of historical data, is the error variance of the marine environmental data, is the error variance of edge node data; is the smoothing factor; is the index, is the number of historical moments; Represents time. The specific process includes: after time alignment, data fusion is performed through the formula. During fusion, and The error variance of edge node data directly affects the fusion weight. Error variance is estimated for each dataset through regression analysis of historical data. The historical data weighting term in the fusion formula smooths the data using a weighted average. The weight of the historical data decays more rapidly as the historical time point moves further away from the current time point. Based on the given formula and parameters, the fusion result is iteratively calculated at each moment until the fusion results for all time points are complete. After data fusion, error metrics (such as mean squared error and bias) are calculated to check whether the fusion result meets the expected stability and accuracy.
[0079] In an embodiment of the present invention, time alignment is achieved by mapping data from different sources onto the same time grid, thereby eliminating data inconsistencies caused by time differences. Data fusion technology can integrate data from different sensors and nodes to fill in data gaps or anomalies that may exist in a single data source. The fused data set is more complete and more fully reflects the actual conditions of the marine environment. During the fusion process, by considering the weight of historical data and the error variance of different data sources, the fusion result can be optimized and noise and errors in the data can be reduced. The smoothing factor is adjusted according to the severity of data changes. It can maintain the stability of the fusion result when the data changes slowly, while increasing the focus on current data when the data changes drastically, thereby improving the dynamic adaptability and stability of the fused data. The fused data set has a unified time scale and higher quality. Through data fusion, it can reduce dependence on a single data source, fully utilize information from multiple sensors and nodes, achieve optimal allocation and efficient utilization of resources, help reduce monitoring costs, and improve overall monitoring efficiency.
[0080] In the present invention, step 14, dividing the fused data set into sharded data sets, and preliminarily distributing the sharded data sets to the heterogeneous nodes to obtain a distribution result, includes:
[0081] Step 141 analyzes the fused dataset to identify key characteristics between data columns, including correlation, data size, and update frequency. Specifically, for each pair of data columns in the dataset, the Pearson correlation coefficient is calculated. The calculated correlation coefficients are organized into a matrix, where each element represents the strength of the correlation between the corresponding data column pair. A diagonal element of 1 in the matrix indicates a perfect correlation between the data column and itself. A correlation threshold is set based on business needs and experience. For example, data column pairs with a correlation coefficient greater than 0.7 can be considered highly correlated. Based on the set threshold, highly correlated data column pairs are filtered from the correlation matrix. The data type of each column in the dataset is determined, such as integer (int), floating-point number (float), string (string), etc. Based on the data type and the number of elements in the column, the storage size of each column is estimated. For example, an integer typically occupies 4 bytes, while the storage size of a string depends on the length of its content and the encoding method. The storage sizes of all data columns are aggregated for subsequent analysis to check whether the dataset contains timestamps or change records, which can reflect the update status of the data column. For each column of data, calculate the update frequency based on its timestamp or change history. Update frequency can be expressed in various ways, such as the number of updates per day, week, or month, or the average interval between updates. Data columns can be categorized based on their update frequency. For example, columns that are frequently updated can be grouped into one category, while columns that are rarely or never updated can be grouped into another.
[0082] Step 142 uses a vertical sharding strategy based on key characteristics, assigning related data columns to the same node. Specifically, this includes determining sharding principles based on key characteristics of the data columns identified in step 141, such as correlation, data size, and update frequency. Formulating sharding principles, for example, places highly correlated data columns in the same shard, while also considering the impact of data size and update frequency on shard size. Review the correlation matrix constructed in step 141 to identify highly correlated data column pairs. Based on the correlation matrix, group the data columns into several groups, ensuring that the data columns within each group are highly correlated.
[0083] Convert each highly correlated group of data columns into a shard, ensuring that the data within each shard is logically closely related. Based on the data size and update frequency of the data columns, preliminarily estimate the storage resources, computing resources, and network bandwidth required for each shard. Evaluate the processing power and storage capacity of each heterogeneous node, including CPU performance, memory capacity, storage space, and network performance. Assign related shards to the same node to optimize local data access and processing and reduce cross-node communication overhead. If a shard requires higher processing performance due to large data volume or high update frequency, it should be assigned to a higher-performance node to ensure overall system performance and load balancing.
[0084] Step 143, evaluate the current load of each heterogeneous node, and according to the current load, Calculate the total dynamic load of each node;
[0085] in Representation node Total dynamic load; 、 and Over time The weight coefficient of the change; 、 and Respectively represent the assignment to nodes Shard CPU usage, storage capacity consumption, and data transfer volume; Is the allocation decision variable, if the shard Assigned to the node ,but ;otherwise, ; 、 and Node Real-time CPU usage, storage capacity utilization, and network bandwidth utilization; Indicates the total number of shards; The index representing the shard; Represents the index of the node;
[0086] Step 144 : Based on the dynamic total load of each node, a genetic algorithm is used to preliminarily distribute the sharded data set to each heterogeneous node to obtain a distribution result.
[0087] In an embodiment of the present invention, by analyzing key characteristics between data columns, such as correlation, it is possible to ensure that in subsequent data distribution, related data columns are allocated to the same node, which helps to reduce the complexity of cross-node data query and processing, thereby improving data processing efficiency. By evaluating the current load of each heterogeneous node and calculating the dynamic total load of each node, the real-time resource usage of the node can be more accurately reflected, which helps to avoid resource overload or idleness during data distribution and achieve balanced resource utilization. By comprehensively considering multiple factors such as CPU utilization, storage capacity consumption and the amount of data transmitted, it is possible to take into account the cost of storage and network resources while ensuring data processing performance, which helps to reduce overall operating costs while improving system performance. The use of genetic algorithms for preliminary data shard allocation can adapt to heterogeneous node environments of different sizes and configurations. When the system needs to be expanded or adjusted, it can quickly adapt to the new node configuration to maintain the flexibility and scalability of the system.
[0088] In a preferred embodiment of the present invention, the above step 144, based on the dynamic total load of each node, preliminarily distributes the sharded data set to each heterogeneous node using a genetic algorithm to obtain a distribution result, may include:
[0089] Step 1441 randomly generates an initial population, where each individual represents a sharded dataset allocation scheme. Individual encoding uses binary code. Specifically, the initial population size is set, i.e., the number of individuals in the population; each individual represents a sharded dataset allocation scheme. Using binary encoding, each binary bit represents whether a shard is assigned to a node. For example, if there are N shards and M nodes, the encoding length of each individual can be N×M, where each M bit represents the allocation of a shard across all nodes. Using a random number generator, the initial population is randomly generated according to the set population size and individual encoding length. Ensure that the binary code of each individual is unique to represent a different allocation scheme.
[0090] In step 1442, a fitness function is defined to evaluate the performance of each individual. Based on the fitness function value, the corresponding individual is selected for the next generation using a roulette wheel selection method. Two individuals are randomly selected for a single-point crossover operation to generate new individuals. Specifically, the fitness function is used to evaluate the performance of each individual (i.e., the allocation solution). This function should take into account the node's dynamic total load, including factors such as computing resources, storage resources, and network bandwidth. A higher fitness function value indicates a better allocation solution. For each individual in the population, its fitness value is calculated. This involves decoding the individual's binary code into the actual allocation solution and calculating the fitness value based on the node load. Based on each individual's fitness value, a roulette wheel selection method is used to select the individual for the next generation. Specifically, the fitness values of all individuals are summed up, and the probability of selection for each individual is equal to its fitness value divided by the sum. This increases the probability of selection for individuals with higher fitness. Two of the selected individuals are randomly selected for a single-point crossover operation. A point is randomly selected in the binary code of an individual, and then all genes (binary bits) after these two points are swapped to produce two new individuals.
[0091] Step 1443 mutates the newly generated individuals. The selection, crossover, and mutation operations are repeated until the preset number of iterations is reached. At the end of each iteration, the corresponding individual is output as the final solution. Specifically, the mutation operation is performed on the newly generated individuals to increase population diversity. One or more genes (binary bits) in the individuals are randomly selected and flipped (0 to 1, 1 to 0). The selection, crossover, and mutation operations are repeated until the preset number of iterations is reached. Each iteration generates a new population containing the optimized allocation plan. At the end of each iteration, the individual with the highest fitness is selected as the final solution. This individual represents the optimal allocation plan for the sharded dataset, optimized by the genetic algorithm.
[0092] Step 1444 decodes the final decoded data set into a sharded data set allocation plan. The sharded data sets are allocated according to the allocation plan to obtain the allocation result. This specifically includes: decoding the final binary code into the actual sharded data set allocation plan. Based on the encoding rules, the node to which each shard is assigned is determined; and based on the decoded allocation plan, the sharded data sets are actually allocated to the heterogeneous nodes. This ensures that the allocation process complies with the plan requirements, and monitors node load to ensure system stability and performance. Through the detailed steps above, a genetic algorithm-based sharded data set allocation process to heterogeneous nodes is implemented, resulting in an optimized allocation result.
[0093] In the embodiments of the present invention, a genetic algorithm is an optimization algorithm based on the principles of biological evolution and possesses powerful global search capabilities. Through a randomly generated initial population and continuous selection, crossover, and mutation operations, the algorithm can extensively explore the solution space, effectively avoiding local optima and ultimately finding a sharded dataset allocation solution that is closer to the global optimum. The genetic algorithm uses a fitness function to evaluate the performance of each individual and makes selections based on the evaluation results. This means the algorithm can automatically adapt to the characteristics of different sharded datasets and heterogeneous node environments without requiring excessive manual intervention, thereby improving the adaptability and flexibility of the allocation solution. By defining a suitable fitness function, the dynamic total load of each heterogeneous node, including computing resources, storage resources, and network bandwidth, can be fully considered. This allows the genetic algorithm to automatically balance the load during the allocation process, preventing some nodes from being overloaded while others are idle, thereby improving overall system resource utilization and performance. Although the genetic algorithm requires multiple selection, crossover, and mutation operations during the iteration process, through appropriate parameter settings and iteration control, the algorithm can find a satisfactory solution within an acceptable timeframe. Furthermore, due to the parallel nature of the genetic algorithm, parallel computing can be used to accelerate the solution process and improve the efficiency of generating allocation solutions.
[0094] In a preferred embodiment of the present invention, the calculation formula of the fitness function is:
[0095] ;
[0096] in, Represents the fitness value of an individual; It represents the total number of heterogeneous nodes, representing the number of different nodes participating in data distribution; Represents the node index, used to traverse all heterogeneous nodes; Indicates the assignment to the node The number of sharded datasets; and Represents the shard index; Indicates the assignment to The node The load of the sharded dataset; represents the weight coefficient; Indicates the number of all sharded data sets to be allocated; and Represents the sharded dataset index, used to traverse all sharded dataset pairs; Indicates the sharded datasets and The frequency of data interaction between sharded datasets; represents the weight coefficient; Represents the associated shard allocation indicator function. If sharded datasets and The sharded data sets are distributed to different nodes. 1; if they are assigned to the same node, then 0.
[0097] In this embodiment of the present invention, the fitness function evaluates load balancing by calculating the load difference between sharded datasets on each node. This helps the algorithm favor solutions that achieve more even load distribution during the search process, thereby avoiding overloading certain nodes and improving overall system performance and stability. The function also takes into account the frequency of data interaction between sharded datasets. When two shards with frequent interactions are assigned to the same node, data interaction between them is more efficient, reducing network transmission overhead and latency. This helps improve data processing speed and response time. By introducing a correlation shard allocation indicator function, the fitness function identifies and rewards solutions that assign highly correlated sharded datasets to the same node. This aggregation helps reduce the additional computational and communication overhead caused by inter-shard correlations, improving data processing locality and efficiency. The multiple weight coefficients in the fitness function (such as G and H) allow the relative importance of different factors to be adjusted based on actual conditions. This makes the algorithm more flexible and adaptable to different application scenarios and optimization objectives. Taking these factors into account, the fitness function provides a comprehensive optimization guide for the genetic algorithm, enabling it to comprehensively consider the influence of multiple factors during the search process, thereby increasing the likelihood of finding a globally optimal or near-globally optimal allocation solution.
[0098] In a preferred embodiment of the present invention, in step 15, performance evaluation indicators are set, and performance tests are performed on each heterogeneous node according to the performance evaluation indicators. The performance indicators of each node when processing different sharded data sets are recorded to obtain test results. The performance evaluation indicators include data processing speed, CPU usage, network transmission delay and throughput, and specifically include:
[0099] Determine evaluation indicators, including data processing speed (for example, the amount of data processed per second), CPU usage (CPU usage of the node when processing data), network transmission delay (delay time during data transmission), and throughput (the amount of data successfully transmitted per unit time).
[0100] Design a test plan: For each heterogeneous node, design a detailed performance test plan. This includes determining the type and scale of test data, the test environment configuration (such as network conditions), and the test duration. Deploy the necessary performance testing tools on the heterogeneous nodes. These tools can monitor and record the aforementioned performance metrics. Perform performance testing on each heterogeneous node according to the test plan. During the test, record the node's performance metrics when processing different sharded data sets. After the test, collect and record the performance metrics of each node to form a complete set of test results.
[0101] In a preferred embodiment of the present invention, the above step 16, analyzing the test results and identifying abnormal nodes and abnormal shard data sets, specifically includes:
[0102] Remove duplicate or invalid data records, verify the test data for each node to ensure it matches the test plan, verify data integrity, and ensure that all predefined performance metrics have been collected. Convert raw data into a unified format and standardize the data, for example, converting data in different units to the same metric. Store the organized data in a reliable database or data warehouse for easy access and analysis. Calculate the average data processing speed for all nodes to understand the overall performance level and the standard deviation to assess fluctuations in data processing speed. Calculate the average CPU usage for each node to identify high-load nodes, analyze the time points when CPU usage peaks, and correlate them with data processing tasks. Calculate the average and standard deviation of network transmission latency to identify transmission paths with high latency, analyze the distribution of latency data, and determine whether there are stable latency patterns. Calculate the average throughput for each node to understand network transmission capacity, analyze throughput trends over time, and determine whether there are bottlenecks. Compare the performance metrics of each node to identify outliers that are significantly below or above the average. Combine historical data with predefined performance thresholds to determine whether performance is abnormal. For nodes with abnormal data processing speeds, check whether their hardware configuration (such as CPU, memory, and storage) meets processing requirements. For nodes with excessively high CPU utilization, analyze running processes and services to identify the cause of excessive resource consumption. For paths with excessive network transmission delays, check network bandwidth, router configuration, and possible network congestion. For nodes with abnormal throughput, analyze data transmission protocols, concurrent connections, and possible network bottlenecks. Combine the data analysis results of multiple performance indicators to comprehensively evaluate the performance of each node.
[0103] In a preferred embodiment of the present invention, the above step 17 is to optimize based on the abnormal nodes and abnormal shard data sets to obtain an optimization result, specifically including:
[0104] Based on the analysis in step 16, determine the nodes that need to be upgraded and their hardware configurations. Select appropriate hardware devices, such as adding memory modules, replacing more efficient CPUs, or increasing storage capacity. Develop a detailed upgrade plan and schedule to ensure that the upgrade process does not affect normal system operation. Analyze abnormal network transmission latency and throughput data to locate network bottlenecks. Adjust network device configurations, such as increasing bandwidth, optimizing routing, or improving load balancing strategies. Consider adopting more efficient data transmission protocols to reduce network transmission overhead.
[0105] Based on the abnormal sharded dataset, reevaluate the current data sharding strategy and adjust the shard size and distribution to balance the load across nodes. Perform a gradual hardware upgrade according to the established plan and schedule. During the upgrade, closely monitor the system's operational status to ensure a smooth upgrade. After the upgrade is complete, conduct necessary system testing and verification to ensure compatibility between the new hardware and the system. Based on the optimization plan, gradually adjust the network device configuration and monitor changes in network performance to ensure that the adjusted configuration effectively improves network transmission efficiency. Redistribute and adjust the sharded dataset according to the new sharding strategy to ensure that the new sharding strategy effectively balances the load across nodes and improves data processing speed. After the optimization is complete, perform performance testing again, including metrics such as data processing speed, CPU utilization, network transmission latency, and throughput. Compare the performance metrics before and after the optimization to evaluate the actual results of the optimization. Analyze the performance test results to ensure that the performance metrics of the abnormal nodes and sharded datasets have significantly improved. If the performance metrics do not meet the expected optimization goals, analyze the reasons and adjust the optimization plan. If the initial optimization results are unsatisfactory, provide feedback and analysis based on the performance test results. Return to step 16 to reanalyze the test results and adjust the optimization plan, performing iterative optimization until the desired performance improvement is achieved. Through the above implementation process, you can systematically optimize abnormal nodes and sharded datasets, and verify the effectiveness of the optimization through performance testing.
[0106] like Figure 2 As shown, an embodiment of the present invention further provides a marine environment multi-source data fusion processing system 20, comprising:
[0107] An acquisition module 21 is configured to connect various types of sensor nodes using a wireless communication protocol;
[0108] The fusion module 22 is used to relatively isolate the sensors during the acquisition process to obtain the ocean environment data and the edge node data of the device; and fuse the ocean environment data and the edge node data to obtain a fused data set;
[0109] The processing module 23 is used to set performance evaluation indicators, perform performance tests on each heterogeneous node according to the performance evaluation indicators, record the performance indicators of each node when processing different sharded data sets to obtain test results; analyze the test results, identify abnormal nodes and abnormal sharded data sets; and optimize based on the abnormal nodes and abnormal sharded data sets to obtain optimization results.
[0110] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for fusion processing of multi-source data of marine environment, characterized in that: The method comprises: Step 11, using wireless communication protocols to connect various types of sensor nodes; Step 12: Isolate the sensor during the acquisition process to obtain ocean environment data and data from the edge nodes of the device; Step 13, fusing the ocean environment data and the edge node data to obtain a fused data set, including: mapping the final ocean environment data and the edge node data of the device to the same time grid to obtain time-aligned data; Perform data fusion to obtain a fused data set, where is the fusion result, and They are marine environment data and edge node data, is the weight of historical data, is the error variance of the marine environmental data, is the error variance of edge node data; is the smoothing factor; is the index, is the number of historical moments; Represents time; Step 14: Divide the fused dataset into sharded datasets, and preliminarily distribute the sharded datasets to the heterogeneous nodes to obtain a distribution result, including: Analyze the fused dataset to identify key characteristics between data columns, including correlation, data size, and update frequency; Based on key features, a vertical sharding strategy is adopted to assign related data columns to the same node; Evaluate the current load of each heterogeneous node and, based on the current load, , Calculate the total dynamic load of each node; in, Representation node Total dynamic load; 、 and Over time The weight coefficient of the change; 、 and Respectively represent the assignment to nodes Shard CPU usage, storage capacity consumption, and data transfer volume; Is the allocation decision variable, if the shard Assigned to the node ,but ;otherwise, ; 、 and Node Real-time CPU usage, storage capacity utilization, and network bandwidth utilization; Indicates the total number of shards; The index representing the shard; Represents the index of the node; According to the dynamic total load of each node, a genetic algorithm is used to preliminarily distribute the sharded data set to each heterogeneous node to obtain the distribution result; Step 15: Set performance evaluation indicators, perform performance tests on each heterogeneous node according to the performance evaluation indicators, and record the performance indicators of each node when processing different sharded data sets to obtain test results; Step 16: Analyze the test results and identify abnormal nodes and abnormal shard data sets; Step 17, performing optimization based on the abnormal nodes and abnormal shard data sets to obtain optimization results, including: An initial population is randomly generated, each individual represents a distribution scheme for the sharded data set, and the individual encoding is binary; Define a fitness function to evaluate the quality of each individual. Based on the value of the fitness function, select the corresponding individual to enter the next generation through roulette. Randomly select two individuals and perform a crossover operation through single-point crossover to generate a new individual. Perform mutation operations on the newly generated individuals, repeating the selection, crossover, and mutation operations until the preset number of iterations is reached. After the iteration is completed, the corresponding individual is output as the final solution; The final decoder is converted into a distribution plan for the sharded dataset, and the sharded dataset is distributed according to the distribution plan to obtain the distribution result. The calculation formula of the fitness function is: ; in, Represents the fitness value of an individual; It represents the total number of heterogeneous nodes, representing the number of different nodes participating in data distribution; Represents the node index, used to traverse all heterogeneous nodes; Indicates the assignment to the node The number of sharded datasets; and Represents the shard index; Indicates the assignment to The node The load of the sharded dataset; represents the weight coefficient; Indicates the number of all sharded data sets to be allocated; and Represents the sharded dataset index, used to traverse all sharded dataset pairs; Indicates the sharded datasets and The frequency of data interaction between sharded datasets; represents the weight coefficient; Represents the associated shard allocation indicator function. If sharded datasets and The sharded data sets are distributed to different nodes. 1; if they are assigned to the same node, then 0.
2. The method for fusion processing of multi-source data of marine environment according to claim 1, characterized in that: The sensors are relatively isolated during the acquisition process to obtain marine environmental data and data from the edge nodes of the equipment, including: Set a corresponding sampling period for each sensor and calculate the lowest common multiple of the sampling periods of all sensors to find a time point at which all sensors complete their respective sampling periods. Determine the sampling time of the first sensor as the starting point of the entire acquisition process; Calculate the delay waiting time for the remaining sensors based on the lowest common multiple of all sensor sampling periods, so that each sensor will collect data isolated from the remaining sensors within the corresponding sampling period; After calculating their respective delay waiting times, each sensor starts collecting ocean environment data and data at the sensor edge nodes; After the ocean environment data and the data at the edge nodes of the sensor are collected, preprocessing is performed to obtain the final ocean environment data and the data at the edge nodes of the device.
3. The method for fusion processing of multi-source data of marine environment according to claim 2, characterized in that: Performance evaluation indicators include data processing speed, CPU usage, network transmission delay and throughput.
4. A marine environment multi-source data fusion processing system, characterized by: The system implements the method according to any one of claims 1 to 3, comprising: an acquisition module for connecting various types of sensor nodes using wireless communication protocols; The fusion module is used to relatively isolate the sensors during the acquisition process to obtain marine environmental data and data from the edge nodes of the equipment; the marine environmental data and the edge node data are fused to obtain a fused data set; The processing module is used to set performance evaluation indicators, perform performance tests on each heterogeneous node based on the performance evaluation indicators, record the performance indicators of each node when processing different sharded data sets to obtain test results; analyze the test results, identify abnormal nodes and abnormal sharded data sets; and optimize based on the abnormal nodes and abnormal sharded data sets to obtain optimization results.
Citation Information
Patent Citations
Internet of Things data processing method, system and device based on edge computing and medium
CN119094581A