Marine observation data sharing platform based on privacy calculation

By designing a marine observation data sharing platform based on privacy computing, technical obstacles and data security issues in marine observation data sharing are solved, and the privacy protection and full utilization of data are achieved.

CN120012155APending Publication Date: 2025-05-16SOUTHERN MARINE SCI & ENG GUANGDONG LAB (ZHUHAI)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510094352.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The sharing of marine observation data faces technical obstacles and data security issues, and it is difficult for the existing technology to effectively share and utilize data while ensuring data privacy.

Method used

A marine observation data sharing platform based on privacy computing is designed, including the application layer, the privacy computing layer and the data storage layer. Through the algorithm module and the computing engine module in the privacy computing layer, the processing and analysis of marine observation data is realized without exposing the original data.

Benefits of technology

It realizes processing and analysis of ocean observation data without exposing the original data, ensuring data privacy, and fully leveraging the value of data, realizing the availability of data invisible and inseparable from domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012155A_ABST
    Figure CN120012155A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of marine data sharing application, in particular to a marine observation data sharing platform based on privacy computing, comprising a data sharing module used for issuing a data directory, retrieval data and application data; the data privacy calculation module is used for creating a privacy calculation task, and a plurality of algorithm models are stored in the algorithm module; and the calculation engine module adopts a secure multi-party calculation technical route, realizes secret state calculation based on plaintext and ciphertext calculation equipment of an open source privacy calculation framework, and improves the secret state calculation performance by utilizing distributed calculation task scheduling. The data layer stores original ocean observation data, and authorized access data is read by the privacy calculation module and then participates in privacy calculation. Through the algorithm module and the calculation engine module, the platform can process and analyze the data on the premise of not exposing the original data, so that the data can be available, invisible and non-out-of-domain sharing is realized, and the value of the ocean observation data is fully exerted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ocean data sharing applications, and in particular to an ocean observation data sharing platform based on privacy computing. Background Art

[0002] With the increasing development and utilization of marine resources, the open sharing of marine observation data has become particularly important. These data not only cover the basic information of the marine environment, but also involve information on the health of the marine ecosystem, marine disaster warnings, and the impact of human activities on the marine environment. Therefore, marine observation data are of great value to scientific research, policy making, environmental protection, and economic activities.

[0003] However, in practical applications, the sharing of ocean observation data faces many challenges. On the one hand, since the collection of ocean observation data often involves multiple institutions or projects, the data format, storage method and data quality vary, resulting in technical barriers to data sharing. On the other hand, ocean observation data often contain sensitive information, such as geographic location, distribution of marine resources, etc. Once leaked, this data may cause significant damage to national security, personal privacy and commercial interests. Therefore, how to share and utilize ocean observation data while ensuring data security has become an urgent problem to be solved.

[0004] Traditional data security measures, such as data encryption and access control, can protect data privacy to a certain extent, but they also limit the sharing and use of data. Summary of the invention

[0005] (I) Purpose of the invention

[0006] The purpose of the present invention is to provide an ocean observation data sharing platform based on privacy computing that can both protect the privacy of ocean observation data and give full play to the value of ocean observation data.

[0007] (II) Technical solution

[0008] To solve the above problems, the present invention provides an ocean observation data sharing platform based on privacy computing, including:

[0009] Application layer, privacy computing layer, and data storage layer;

[0010] The application layer includes a data sharing module and a data privacy calculation module;

[0011] The data sharing module is used to publish data catalogs, retrieve data and apply for data;

[0012] The data privacy computing module is used to create privacy computing tasks;

[0013] The privacy computing layer includes an algorithm module and a computing engine module;

[0014] The algorithm module stores a plurality of algorithm models, and the plurality of algorithm models are used to be called when creating a privacy computing task;

[0015] The computing engine module is used to perform confidential computing based on the created privacy computing task;

[0016] The data storage layer stores ocean observation data, and the data disk is mounted on the privacy computing node. The authorized ocean observation data can be read and then participate in confidential computing.

[0017] In another aspect of the present invention, preferably,

[0018] The ocean observation data is stored based on a preset data sharing standard;

[0019] The preset data sharing standards include: metadata standards, data record format standards and data quality control standards;

[0020] The metadata standard includes that the ocean observation data is published based on metadata fields;

[0021] The data recording format standard includes recording the ocean observation data according to a unified table structure;

[0022] The data quality control standard refers to the use of quality control symbols to identify the ocean observation data values ​​with qualified or abnormal evaluation results according to a preset data quality evaluation algorithm. In another aspect of the present invention, preferably,

[0023] The preset data quality assessment algorithm includes range test, gradient test and peak test;

[0024] The range test is performed using the following formula:

[0025] X min ≤x i ≤X max

[0026] Among them, x i represents the i-th observation value of the x element; X min Indicates the minimum value of the x element, X max Indicates the maximum value of the x element;

[0027] The gradient test is performed using the following formula:

[0028]

[0029] Among them, xi-1 represents the i-1th observation value of the x element, H x represents the gradient test parameter of the x element. The peak test is performed using the following formula:

[0030] |x i -(x i-1 +x i+1 ) / 2|-|x i+1 -x i-1 | / 2≤H' x

[0031] Among them, x i represents the i-th observation value of the x element, x i-1 represents the i-1th observation value of the x element, x i+1 represents the i+1th observation value of the x element, H' x Indicates the peak detection parameters of the x element.

[0032] In another aspect of the present invention, preferably,

[0033] The privacy computing module is used to create a privacy computing task process, including:

[0034] Select the authorization data to be used;

[0035] Utilize several algorithm models stored in the algorithm module, or a custom algorithm model, to configure a calculation model;

[0036] Generate a calculation code according to the calculation model, and review the calculation code;

[0037] After the calculation code passes the review, the selected authorization data is used to perform the secret calculation to obtain the calculation result;

[0038] Review the calculation results and provide them to the data user after confirmation;

[0039] Generate a corresponding privacy computing process traceability report for the calculation results.

[0040] In another aspect of the present invention, preferably,

[0041] The computing engine module is used to perform confidential computing based on the created privacy computing task, including:

[0042] Use secure multi-party computing technology based on distributed multi-node computing to achieve secret computing;

[0043] The secure multi-party computing is based on the open source privacy computing framework cryptography for computing;

[0044] The cipher has multiple cryptographic computing virtual devices built in, and the cryptographic computing virtual devices include plaintext computing devices and ciphertext computing devices;

[0045] The privacy computing task generates a data analysis code through a plaintext computing device, and the data analysis code is converted into a computing code of a ciphertext computing device through a protocol, and the computing code is used to perform secret computing.

[0046] In another aspect of the present invention, preferably,

[0047] During the confidential computing process, the data user has task management authority, which includes starting a task, canceling a task, pausing a task, and restarting a task;

[0048] During the confidential computing process, the privacy computing node has the authority to suspend the computing tasks on its own node.

[0049] In another aspect of the present invention, preferably, the calculation code is generated according to a custom algorithm model, and reviewing the calculation code includes:

[0050] Automatically reviewing by computer whether the computational code has operations for acquiring ocean observation data;

[0051] If there are any, mark the suspicious code lines and provide them to the data owner for manual review;

[0052] If not, the review ends.

[0053] In another aspect of the present invention, preferably,

[0054] The algorithm module stores a number of algorithm models, which are used to call when creating privacy computing tasks, including: calculating seawater density using CTD temperature and salinity data;

[0055] The seawater density is calculated using the following formula:

[0056]

[0057] Among them, ρ(s, t, p) represents the density of seawater, s represents the salinity, t represents the temperature of seawater, p represents the pressure, and K(s, t, p) is the secant bulk modulus of the thermodynamic parameters of seawater.

[0058] In another aspect of the present invention, preferably, the algorithm module stores a plurality of algorithm models, and the plurality of algorithm models are used to be called when creating a privacy computing task, including: frequency reduction of tide data;

[0059] The tide data frequency reduction includes:

[0060] According to the preset resampling frequency, the original tide data are grouped by time unit to form multiple data groups;

[0061] For each data set, the average value of the original tide level data in the data set is calculated, and the average value is used as the representative value of the tide level of the corresponding time unit;

[0062] The representative value of tide level is used to form the down-converted tide level data.

[0063] In another aspect of the present invention, preferably, the algorithm module stores a plurality of algorithm models, and the plurality of algorithm models are used to be called when creating a privacy computing task, including: interpolation of water depth data observed during ship-borne navigation;

[0064] The water depth data interpolation of the shipborne cruise observation includes:

[0065] Create two-dimensional grid coordinates based on the scattered latitude and longitude coordinates of the original navigation;

[0066] Interpolation is performed using a linear interpolation method or an inverse distance interpolation method according to the two-dimensional grid point coordinates.

[0067] (III) Beneficial effects

[0068] The above technical solution of the present invention has the following beneficial technical effects:

[0069] Through the algorithm module and computing engine module in the privacy computing layer of the present invention, the platform can process and analyze data without exposing the original data. The algorithm module in the privacy computing layer stores a variety of algorithm models for users to call. The algorithm module has a high degree of scalability and can develop new algorithms for general privacy computing scenarios according to needs. In addition, the algorithm module supports user-defined programming to meet unique data processing and analysis needs and has a high degree of flexibility. The computing engine module supports secret computing and can perform complex data processing and analysis tasks while protecting data privacy to meet users' data usage needs. Through privacy computing, the ocean observation data can be made available but invisible, shared within the domain, and the value of the data can be fully utilized. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is a schematic diagram of the overall structure of an embodiment of the present invention;

[0071] Figure 2 is a graph showing changes in temperature, salinity and calculated density with depth observed by CTD according to an embodiment of the present invention;

[0072] Figure 3 is a frequency reduction graph of tide data according to an embodiment of the present invention;

[0073] Figure 4This is an interpolation diagram of water depth data of shipborne cruise observation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0074] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.

[0075] The accompanying drawings show schematic diagrams of structures according to embodiments of the present invention. These figures are not drawn to scale, and some details are magnified and some details may be omitted for the purpose of clarity. The shapes of various regions and layers shown in the figures and the relative sizes and positional relationships therebetween are only exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art may additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0076] Obviously, the described embodiments are only some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0077] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0078] The present invention will be described in more detail below with reference to the accompanying drawings. In each of the accompanying drawings, the same elements are represented by similar reference numerals. For the sake of clarity, the various parts in the accompanying drawings are not drawn to scale.

[0079] Embodiment 1

[0080] A privacy-preserving ocean observation data sharing platform. Figure 1 FIG. 1 shows a schematic diagram of the overall structure of an embodiment of the present invention. Figure 1 As shown, including:

[0081] Application layer, privacy computing layer, and data storage layer;

[0082] The application layer includes a data sharing module and a data privacy calculation module;

[0083] The data sharing module is used to publish data directories, retrieve data and apply for data; the data provider publishes the data directory, and the data user can then retrieve the data, initiate an application for the required data, and obtain data use authorization after approval by the data provider.

[0084] The data privacy computing module is used to create privacy computing tasks; the privacy computing module creates privacy computing tasks and configures the algorithm model required for the privacy computing task. After the data user obtains the data use authorization, he can select the data to create a privacy computing task and configure the data algorithm model. The algorithm model includes two modes: selecting pre-stored and uploading custom code. Then, the computing task is started and the computing task is automatically scheduled to improve the computing performance. After the task is completed, the decrypted calculation results are output to the data user, and the visualization of the calculation results is provided for the user to preview and check the data.

[0085] The privacy computing layer includes an algorithm module and a computing engine module;

[0086] The algorithm module stores several algorithm models, which are used to call when creating privacy computing tasks; several algorithm models include common ocean observation data processing algorithms such as density calculation using CTD temperature and salinity data, frequency reduction of tide data, and interpolation of water depth data from shipborne cruise observations. The code implementation of the algorithm is supported by Python's third-party library, including Numpy for array calculations, Seawater for analysis of seawater physical states and dynamic processes, Geopy for geographic coordinate system conversion, and Metpy for marine meteorological data analysis.

[0087] The computing engine module is used to perform secret computing based on the created privacy computing tasks; the computing engine module is deployed on the local node of the data provider; the computing engine module is developed based on the open source privacy computing framework "Hidden Language", which has built-in multiple secret computing virtual devices, as well as out-of-the-box privacy protection data analysis, machine learning and other functions, which can help developers quickly build privacy computing applications. The privacy computing engine includes a plaintext computing device PYU (Plaintext Device) and a ciphertext computing device SPU (Secretflow Processing Unit). The data analysis code written in Python is first implemented through PYU, and then compiled into SPU-compilable code after protocol conversion to achieve secret computing.

[0088] The data storage layer stores ocean observation data. The original ocean observation data is stored locally by the data provider, and the data security is high. The data disk is mounted to the privacy computing node to participate in the calculation. The authorized data can be read into the privacy computing node memory to participate in the privacy computing.

[0089] Furthermore, in this embodiment, due to the multi-source heterogeneity of ocean observation data and the uneven data quality, before sharing through privacy computing, the data observed by different instruments must first be converted into data record formats, and then quality control must be performed on the standard format data to ensure that the observation data involved in the calculation are all high-quality data, so as to ensure the accuracy and credibility of the calculation results. In this embodiment, the ocean observation data is stored based on preset data sharing standards; the preset data sharing standards include: metadata standards, data record format standards, and data quality control standards;

[0090] The metadata standard includes the release of ocean observation data based on metadata fields; it stipulates 10 types of metadata fields, including data title, data summary, data elements, data category, subject category, time range, spatial range, data format, data quality, and data sample, so that data users can quickly understand the basic information of the data. All data released and shared on the sharing platform must follow the metadata standard, fill in the metadata content, and generate a data resource directory for data retrieval and viewing.

[0091] The data recording format standard includes recording the ocean observation data according to a unified table structure; for common ocean observation data, including observations of ocean temperature, salinity, ocean currents, waves, tides, water depth and other elements, according to the different observation instruments of the elements (such as temperature, salinity and depth data observed by Conductive-Temperature-Depth (CTD), ocean current data observed by Acoustic Doppler Current Profiler (ADCP), wave height and wave direction data observed by buoys, tide level data observed by tide gauges, water depth data observed by multi-beams, etc.), the ocean observation data recording format standard is formulated respectively. In order to facilitate calculation, the data is uniformly stipulated as a table structure, and the first row of records is the table header: the element name and its unit are uniformly stipulated, and the data column of the element will be sorted according to the unified regulations. In order to facilitate data upload by different types of users such as data institutions and individuals, it supports data storage in relational databases or text files (CSV, TXT and other files) that conform to the data recording format.

[0092] Table 1 shows the CTD temperature-salinity-depth data recording format standard. As shown in Table 1, the first column of the data standardized according to the recording format requirements records the pressure pres, the second column records the temperature temp, and the third column records the salinity sal.

[0093] Table 1 Data recording format standard using CTD temperature, salinity and depth data as an example

[0094] Serial number Variable naming Data Types Data accuracy unit 1 pres float 3 dbar 2 temp float 4 ℃ 3 sal float 4 PSU

[0095] The data quality control standard refers to the use of quality control symbols to mark the ocean observation data values ​​with qualified or abnormal evaluation results according to the preset data quality evaluation algorithm. Due to the harsh ocean observation environment, encountering harsh marine meteorological environments such as typhoons, large waves, storm surges, or observation instruments being damaged by ships, fishing nets, etc., as well as the system errors of the observation instruments themselves, there are widespread data quality problems in ocean observation data, and it is necessary to correct the abnormal values ​​before the data can be further used for analysis and calculation in scientific research and business applications. By formulating data quality control standards, a unified quality evaluation method is provided for shared ocean observation data. After the standard record format data uploaded by different users is corrected for abnormal values ​​according to the standard quality control method, the shared data quality is high and unified. Ocean observation data usually uses range checks, peak checks, gradient checks and other inspection methods to distinguish abnormalities. Abnormal values ​​will be identified by quality control symbols (Quality Control Flag), and data users can refer to the quality control symbols to identify abnormal values, thereby avoiding the influence of abnormal values ​​on data calculation.

[0096] Further, in this embodiment, the preset data quality assessment algorithm includes range check, gradient check, and peak check;

[0097] The range test is performed using the following formula:

[0098] X min ≤x i ≤X max

[0099] Among them, x i represents the i-th observation value of the x element; X min Indicates the minimum value of the x element, X max Indicates the maximum value of the x factor; the minimum and maximum values ​​of the x factor can be obtained based on the corresponding historical data. If the observed value exceeds the corresponding range, the data is abnormal.

[0100] The ocean observation data has continuity within a certain time or space range. The change values ​​of observation elements that are close in time or adjacent in position should be within a certain range, otherwise the data is considered abnormal. The specific method is: Assume that the current observation value is x i , the previous correct value adjacent to it is x (i-1) , the following formula should be satisfied, otherwise the gradient will be too large and the data will be suspicious.

[0101] |x i -x i-1 |≤H x

[0102] x i -x i-1 ≥0

[0103] Among them, x i-1 H represents the i-1th observation value of the x element; x It represents the gradient test parameter of the x element, which is determined by factors such as element type, observation time interval or spatial interval, observation time and area.

[0104] The peak test is performed using the following formula: Assuming the current observation value is x i , and the first correct value adjacent to its left and right or top and bottom is x i-1 and x i+1 , then the current observation value should satisfy the following formula, otherwise x i It is a spike value, and the data is abnormal:

[0105] |x i -(x i-1 +x i+1 ) / 2|-|x i+1 -x i-1 | / 2≤H' x

[0106] Among them, x i represents the i-th observation value of the x element, x i-1 represents the i-1th observation value of the x element, x i+1 represents the i+1th observation value of the x element, H' x It represents the peak detection parameter of the x element, which is determined by factors such as element type, observation time interval or spatial interval, observation time and area.

[0107] Table 2 shows the data quality assessment algorithm using CTD temperature, salinity and depth data as an example, including the value range of range test, gradient test and peak test, to achieve quality control of CTD temperature, salinity and depth data in standard recording format.

[0108] Table 2 Data quality assessment algorithm using CTD temperature, salinity and depth data as an example

[0109]

[0110] Furthermore, in this embodiment, the privacy computing module is used to create a privacy computing task process, including:

[0111] Select the authorized data to be used; select the authorized standard record format data to participate in the calculation. Support users to explore data and view the statistical indicators of data, including extreme values, means, variances, etc., to assist users in judging data quality; and support users to view sample data to intuitively understand the data structure, paving the way for configuring calculation models or customizing code.

[0112] The calculation model is configured using several algorithm models stored in the algorithm module, or a custom algorithm model (custom programming); the system built-in operators are selected by dragging and dropping in the graphical interface, and then the operators are combined according to the algorithm steps on the canvas, the calculation process is drawn, and the calculation model is formed.

[0113] Generate calculation code based on the calculation model; the system will automatically generate calculation code based on the calculation model configured by the user, and the code will be open to users for review to improve the transparency and credibility of the calculation process.

[0114] Conduct a security review of the computing code to ensure that the code does not contain system attacks, snooping of raw data, etc. The review method is mainly based on the AI ​​large language model code review, supplemented by manual review. Use a computer to automatically review whether the computing code has operations to obtain ocean observation data; if so, mark the suspicious code lines and provide them to the data owner for manual review; if not, the review ends.

[0115] After the calculation code passes the review, the calculation task can be started, and the selected authorized data can be used to perform confidential calculations to obtain calculation results.

[0116] During the confidential computing process, the data user has task management authority, which includes starting tasks, canceling tasks, pausing tasks, and restarting tasks; during the confidential computing process, the privacy computing node has the authority to pause tasks calculated on its own node.

[0117] The calculation results support previewing some record lines in table form; support visualization: for example, drawing temperature-depth and salinity-depth profiles of CTD data, support users to customize the aspect ratio and color of the graph, and support exporting pictures at the specified resolution, which can be directly used as illustrations for scientific research papers. Finally, users are supported to export the calculation results as CSV, TXT and other text files with table structures.

[0118] Generate a corresponding privacy computing process traceability report for the calculation results. Issue a privacy computing process traceability report, including: the start and end time of the calculation task, the name of the original data involved in the privacy calculation, the name of the calculation model used, the methods and technologies used in data encryption transmission and the calculation process, etc., as a certificate of the source of the privacy calculation result data, to prove the credibility of the calculation result.

[0119] Further, in this embodiment,

[0120] The computing engine module is used to perform confidential computing based on the created privacy computing task, including:

[0121] Utilize the secure multi-party computing technology route based on distributed multi-node computing to realize secret computing; combine the actual needs of typical privacy computing scenarios of ocean observation data, and use the secure multi-party computing technology route to realize secret computing. Compared with the other two technical routes of trusted hardware execution environment and federated learning, secure multi-party computing has strong control over data and supports a wide range of privacy algorithms (including customized data analysis, machine learning, etc.), but the computing performance has declined. In order to improve computing performance, distributed multi-node computing will be used to increase data processing speed through intelligent big data task scheduling.

[0122] The secure multi-party computing is based on the open source privacy computing framework cryptography for computing;

[0123] The cipher has multiple secret computing virtual devices built in, including a plaintext computing device PYU (Plaintext Device) and a secretflow computing device SPU (Secretflow Processing Unit);

[0124] The privacy computing task generates data analysis code through a plaintext computing device, which is converted into a computing code of a ciphertext computing device through a protocol, and the computing code is used to perform secret computing. The written data analysis code is first implemented through PYU, and then converted through a protocol and compiled into a code that can be compiled by SPU, thereby realizing secret computing.

[0125] Furthermore, in this embodiment, three typical ocean observation data privacy computing scenarios are built, and their algorithm models are encapsulated in the system for users to call. Scenario 1-Calculate seawater density using CTD temperature and salinity data: Density is an important variable reflecting the state of seawater, and the density field can be used to diagnose and analyze the state and dynamic characteristics of seawater; Scenario 2-Tide data frequency reduction: High-frequency tide data can be desensitized after frequency reduction; Scenario 3-Interpolation of water depth data observed by shipboard navigation: Scattered water depth data observed along the survey line usually needs to be spatially interpolated to form raster data for use.

[0126] The algorithm model of the privacy computing scenario "calculating seawater density using CTD temperature and salinity data" is as follows:

[0127] The seawater density is calculated using the following formula:

[0128]

[0129] Among them, ρ(s, t, p) represents the density of seawater, s represents the salinity, t represents the temperature of seawater, p represents the pressure, and K(s, t, p) is the secant bulk modulus of the thermodynamic parameters of seawater, which can be approximated by a polynomial function of seawater temperature and salinity.

[0130] The above algorithm can be implemented through the seawater.eos80.dens(s,t,p) method of the seawater library in Python. You need to use the Cryptography IDE to develop the privacy computing program of the seawater library, and after deployment, a built-in operator is formed for users to call.

[0131] Figure 2 FIG. 4 shows a graph showing the temperature, salinity and density calculated by CTD according to an embodiment of the present invention as a function of depth. Figure 2 As shown, temperature and salinity participate in privacy calculation, the original observation data is not leaked, and the inferred density can be directly exported in plain text for users to use.

[0132] The algorithm model of the privacy computing scenario "tide data frequency reduction" is as follows:

[0133] According to the preset resampling frequency, the original tide data are grouped by time unit to form multiple data groups; the preset resampling frequency, such as 'D' (daily), 'M' (monthly), 'Q' (quarterly), 'Y' (yearly), etc., the original tide data are grouped by time unit. For example, if the frequency is 'D', the data will be grouped by day.

[0134] For each data set, the average value of the original tide data in the data set is calculated, and the average value is used as the representative value of the tide level for the corresponding time unit; when calculating the average value, missing values ​​(NaN) are ignored by default. If necessary, missing values ​​can be filled after resampling.

[0135] The tide level representative value is used to form the downsampled tide level data. A new time series is returned, whose index is the preset resampling frequency and the value is the average value in each frequency interval.

[0136] The above algorithm can be implemented by calling the resample method of the Python Pandas library supported by lingo.

[0137] For example, the high-frequency tide data with a sampling frequency of 1 minute is downsampled to a sampling interval of 1 hour. Figure 3 FIG. 4 shows a frequency reduction diagram of tide data according to an embodiment of the present invention. Figure 3 As shown in the figure, the original high-frequency high-dimensional data and the tide data curve that is down-converted by privacy computing. The down-converted tide time series can effectively extract the characteristic information of the original high-frequency tide time series, which can meet the needs of most scientific research and business applications.

[0138] The algorithm model of the privacy computing scenario "interpolation of water depth data observed during the ship's navigation" is as follows: two-dimensional grid coordinates are generated based on the scattered latitude and longitude coordinates of the original navigation;

[0139] Interpolation is performed using a linear interpolation method or an inverse distance interpolation method according to the two-dimensional grid point coordinates.

[0140] Linear interpolation algorithm, assuming that the coordinates and coordinate values ​​of the three observation points are (x1, y1, z1), (x2, y2, z2), (x3, y3, z3), the coordinate value of the interpolated point is (x, y, z), the interpolated point two-dimensional grid coordinates (x, y), the interpolated point is located between the triangles formed by the coordinates of the three observation points, and the z value of the interpolated point is calculated using the following formula:

[0141]

[0142] The inverse distance interpolation algorithm is a distance-based interpolation method. The closer the point is to the known data point, the greater the impact on the interpolation result, and the farther the point is, the smaller the impact. Assume that all the n observation points are (x i ,y i ,z i ), 1≤i≤n. For any interpolation grid coordinate (x, y), the function value z is calculated as follows:

[0143]

[0144] Among them, d i is the distance from the interpolation point to the i-th observation point.

[0145] The above algorithm can be implemented by calling the griddata method of the Python Scipy library supported by the cipher.

[0146] Figure 4 FIG. 4 shows an interpolation diagram of water depth data of shipborne cruise observation according to an embodiment of the present invention. Figure 4 As shown in Figure 1, the water depth data observed by the shipborne single beam is interpolated to the longitude and latitude grid coordinates with a horizontal resolution of 0.001° x 0.001°. Figure 4 The left side is the original data. Figure 4 The right side of is the interpolated data. The interpolated grid data can better reflect the spatial distribution of the observed water depth and can meet the needs of most scientific research and business applications.

[0147] Through the algorithm module and computing engine module in the privacy computing layer of the present invention, the platform can process and analyze the data without exposing the original data. The algorithm module in the privacy computing layer stores a variety of algorithm models for users to call, including algorithm models for three typical ocean observation data privacy computing scenarios: density calculation using CTD temperature and salinity data, tide data frequency reduction, and water depth data interpolation of shipborne navigation observation. The algorithm module has a high degree of scalability, and algorithms for general privacy computing scenarios can be newly developed according to demand. In addition, the algorithm module supports user-defined programming to meet unique data processing and analysis needs, and has a high degree of flexibility. The computing engine module supports secret computing, and adopts a secure multi-party computing technology route to realize the conversion of plaintext computing code. It can perform complex data processing and analysis tasks while protecting data privacy, meeting the user's data usage needs. Through privacy computing, the ocean observation data can be "available but invisible, and shared within the domain", giving full play to the value of the data.

[0148] It should be understood that the above specific embodiments of the present invention are only used to illustrate or explain the principles of the present invention, and do not constitute a limitation of the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention should be included in the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modifications that fall within the scope and boundaries of the appended claims, or the equivalent forms of such scope and boundaries.

[0149] In the above description, the technical details of patterning and etching of each layer are not described in detail. However, those skilled in the art should understand that various means in the prior art can be used to form layers, regions, etc. of desired shapes. In addition, in order to form the same structure, those skilled in the art can also design methods that are not completely the same as the methods described above.

[0150] The present invention has been described above with reference to the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, a person skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

[0151] Although the embodiments of the present invention have been described in detail, it should be understood that the various changes, substitutions, and alterations could be made hereto without departing from the spirit and scope of the invention.

[0152] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.

Claims

1. An ocean observation data sharing platform based on privacy computing, characterized in that: include: Application layer, privacy computing layer, and data storage layer; The application layer includes a data sharing module and a data privacy calculation module; The data sharing module is used to publish data catalogs, retrieve data and apply for data; The data privacy computing module is used to create privacy computing tasks; The privacy computing layer includes an algorithm module and a computing engine module; The algorithm module stores a plurality of algorithm models, and the plurality of algorithm models are used to be called when creating a privacy computing task; The computing engine module is used to perform confidential computing based on the created privacy computing task; The data storage layer stores ocean observation data, and the data disk is mounted on the privacy computing node. The authorized ocean observation data can be read and then participate in confidential computing.

2. The ocean observation data sharing platform based on privacy computing according to claim 1 is characterized in that: The ocean observation data is stored based on a preset data sharing standard; The preset data sharing standards include: metadata standards, data record format standards and data quality control standards; The metadata standard includes that the ocean observation data is published based on metadata fields; The data recording format standard includes recording the ocean observation data according to a unified table structure; The data quality control standard includes using quality control symbols to mark the ocean observation data values ​​with qualified or abnormal evaluation results according to a preset data quality evaluation algorithm.

3. The ocean observation data sharing platform based on privacy computing according to claim 2 is characterized in that: The preset data quality assessment algorithm includes range test, gradient test and peak test; The range test is performed using the following formula: X min ≤x i ≤X max Among them, x i represents the i-th observation value of the x element; X min Indicates the minimum value of the x element, X max Indicates the maximum value of the x element; The gradient test is performed using the following formula: Among them, x i-1 represents the i-1th observation value of the x element, H x represents the gradient test parameter of the x factor; The peak test is performed using the following formula: |x i -(x i-1 +x i+1 ) / 2|-|x i+1 -x i-1 | / 2≤H' x Among them, x i represents the i-th observation value of the x element, x i-1 represents the i-1th observation value of the x element, x i+1 represents the i+1th observation value of the x element, H' x Indicates the peak detection parameters of the x element.

4. The ocean observation data sharing platform based on privacy computing according to claim 1 is characterized in that: The privacy computing module is used to create a privacy computing task process, including: Select the authorization data to be used; Utilize several algorithm models stored in the algorithm module, or a custom algorithm model, to configure a calculation model; Generate a calculation code according to the calculation model, and review the calculation code; After the calculation code passes the review, the selected authorization data is used to perform the secret calculation to obtain the calculation result; Review the calculation results and provide them to the data user after confirmation; Generate a corresponding privacy computing process traceability report for the calculation results.

5. The ocean observation data sharing platform based on privacy computing according to claim 4 is characterized in that: The computing engine module is used to perform confidential computing based on the created privacy computing task, including: Use secure multi-party computing technology based on distributed multi-node computing to achieve secret computing; The secure multi-party computing is based on the open source privacy computing framework cryptography for computing; The cipher has multiple cryptographic computing virtual devices built in, and the cryptographic computing virtual devices include plaintext computing devices and ciphertext computing devices; The privacy computing task generates a data analysis code through a plaintext computing device, and the data analysis code is converted into a computing code of a ciphertext computing device through a protocol, and the computing code is used to perform secret computing.

6. The ocean observation data sharing platform based on privacy computing according to claim 4 is characterized in that: During the confidential computing process, the data user has task management authority, which includes starting a task, canceling a task, pausing a task, and restarting a task; During the confidential computing process, the privacy computing node has the authority to suspend the computing tasks on its own node.

7. The ocean observation data sharing platform based on privacy computing according to claim 4 is characterized in that: The calculation code is generated according to a custom algorithm model, and the review of the calculation code includes: Automatically reviewing by computer whether the computational code has operations for acquiring ocean observation data; If there are any, mark the suspicious code lines and provide them to the data owner for manual review; If not, the review ends.

8. The ocean observation data sharing platform based on privacy computing according to claim 1 is characterized in that: The algorithm module stores a number of algorithm models, which are used to call when creating privacy computing tasks, including: calculating seawater density using CTD temperature and salinity data; The seawater density is calculated using the following formula: Among them, ρ(s, t, p) represents the density of seawater, s represents the salinity, t represents the temperature of seawater, p represents the pressure, and K(s, t, p) is the secant bulk modulus of the thermodynamic parameters of seawater.

9. The ocean observation data sharing platform based on privacy computing according to claim 1 is characterized in that: The algorithm module stores a number of algorithm models, which are used to call when creating privacy computing tasks, including: tidal data frequency reduction; The tide data frequency reduction includes: According to the preset resampling frequency, the original tide data are grouped by time unit to form multiple data groups; For each data set, the average value of the original tide level data in the data set is calculated, and the average value is used as the representative value of the tide level of the corresponding time unit; The representative value of tide level is used to form the down-converted tide level data.

10. The ocean observation data sharing platform based on privacy computing according to claim 1, characterized in that: The algorithm module stores a number of algorithm models, which are used to call when creating privacy computing tasks, including: interpolation of water depth data observed by shipboard navigation; The water depth data interpolation of the shipborne cruise observation includes: Create two-dimensional grid coordinates based on the scattered latitude and longitude coordinates of the original navigation; Interpolation is performed using a linear interpolation method or an inverse distance interpolation method according to the two-dimensional grid point coordinates.