Marine biological diversity big data collaborative analysis method based on federal knowledge graph

By identifying and screening robust map nodes in the federal knowledge graph, the problems of difficulty in collecting marine biodiversity data and insufficient data are solved, the accuracy and stability of marine risk prediction are achieved, and the risk of false predictions and deviations is reduced.

CN120336768AActive Publication Date: 2025-07-18GUANGDONG OCEAN UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510803345.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In the prior art, the difficulty in collecting marine biodiversity data and the high deviation problems caused by insufficient data are prone to false short-term prediction and inference bias amplification, and the deep learning model's generalization ability deteriorates in the sample sparse interval.

Method used

By presetting the marine risk type, the graph nodes are identified from the federal knowledge graph, the inference confidence is obtained, and reliable nodes are selected through validity judgment, and robust graph nodes are screened out using Fourier transform and Pearson correlation coefficients to perform distributed modeling and voting warning.

Benefits of technology

It effectively reduces the reasoning bias in the non-aligned data scenario, improves the accuracy and stability of marine risk prediction, and ensures the robustness and accuracy of federal knowledge graph collaborative analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336768A_ABST
    Figure CN120336768A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and provides a marine biological diversity big data collaborative analysis method based on a federated knowledge graph, and the method specifically comprises the steps: firstly, presetting a marine risk type, recognizing graph nodes from the federated knowledge graph, obtaining the reasoning confidence from each graph node, and carrying out the calculation of the reasoning confidence; then, effectiveness judgment is carried out according to the reasoning confidence degree obtained by each map node in a unit time period; and finally, effective map nodes are selected according to the effectiveness judgment result to obtain the ocean risk confidence degree. Confidence sequence comparison is effectively utilized, discontinuous features and unstable features of a prediction sequence structure brought by a non-aligned data model are highlighted, and meanwhile prediction differences of different map nodes are further utilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a collaborative analysis method for marine biodiversity big data based on a federated knowledge graph. Background Art

[0002] Marine biodiversity itself has a high sensitivity and early responsiveness to ecological disturbances, and has a mechanism explanation function from the perspective of the ecological chain for various marine risks such as red tides, hypoxia, and species invasions. However, the collection of biodiversity data in the same region has problems of difficult collection and limited coverage of biological types. Although physical data monitoring can reduce the operation bottleneck caused by the difficulty of biodiversity monitoring to a certain extent, the models constructed by the data collection in a single region usually still have the problem of high deviation caused by insufficient original data. The big data collaborative analysis based on the federated knowledge graph can effectively integrate the biodiversity data of different geographical regions and the inference models refined from this database.

[0003] However, the measured data and sampling data considered in the marine biodiversity data. The measured data is the monitoring of objective physical quantities, and usually the real-time measurement operation can be realized more frequently. The sampling data refers to the monitoring data of scientific research surveys. Due to the high dependence on manual labor or high-precision equipment, the monitoring interval is large and the data acquisition frequency is low. When such measured data and sampling data act on the big data collaborative analysis of the federated knowledge graph at the same time, the data scenario of this non-aligned data is prone to false short-term prediction schemes compared with the inference models constructed by ordinary data scenarios; in addition, the density of the measured data samples is much higher than that of the sampling data, which causes the generalization ability of deep learning-based inference models to degenerate in the sparse sample interval. This makes the data scenario of non-aligned data have the risk of amplified deviation in the big data collaborative analysis of the federated knowledge graph. Therefore, it is necessary to judge the effectiveness of the collaboration between different graph nodes in the federated knowledge graph to prevent the risk of amplified inference deviation. Summary of the Invention

[0004] The purpose of the present invention is to propose a collaborative analysis method for marine biodiversity big data based on a federated knowledge graph to solve one or more technical problems existing in the prior art, and at least provide a beneficial choice or create conditions.

[0005] To achieve the above object, according to one aspect of the present invention, there is provided a collaborative analysis method for marine biodiversity big data based on a federated knowledge graph, and the method includes the following steps: S100, preset marine risk types, and identify graph nodes from the federated knowledge graph; S200, obtain inference confidence levels from each graph node; S300, determine the validity based on the inference confidence levels obtained from each graph node within a unit time period; S400, select valid graph nodes according to the validity judgment result to obtain the marine risk confidence level.

[0006] Furthermore, in step S100, the preset marine risk types and the method of identifying graph nodes from the federated knowledge graph are as follows: The federated knowledge graph consists of several graph nodes, and different graph nodes represent biodiversity monitoring databases in different geographical locations; preset marine risk types and filter out the graph nodes with the labels of these risk types.

[0007] The marine risk types include red tides, alien invasive species, sea temperature anomalies, ocean acidification, coral bleaching, or aquaculture leakage, etc.; The core of the federated knowledge graph here is the collaborative inference of multiple databases to build the knowledge graph. Because there are defects in constructing the marine risk model with a single data source, including: first, for unmonitored risk types, data annotation needs to wait until the time point when the risk actually occurs, and there are obvious shortcomings in the annotation opportunity; second, the lack of annotation opportunities leads to insufficient robustness of the constructed inference model and low model accuracy.

[0008] In the step of presetting marine risk types, only one marine risk type is selected to make the inference objectives of the graph nodes consistent, and prevent the inference objectives of the subsequent constructed inference model from being generalized and blurred, reducing the accuracy. Each data record in the biodiversity monitoring database has a timestamp, making each feature data have a time attribute; for any graph node, when it is determined that a certain marine risk occurs in the corresponding sea area during a certain time period, then this graph node has the label of the marine risk during this time period. After presetting or selecting the only marine risk type, select the graph nodes from the federated knowledge graph that meet the condition of having the preset marine risk label and exclude the graph nodes without the preset marine risk label.

[0009] Furthermore, in step S200, the method of obtaining the inference confidence level from each graph node includes: obtaining the scheduling feature list of the preset marine risk in the current graph node. The scheduling feature list is a set of features used to construct the prediction model. Send the scheduling feature list and the marine risk type to each graph node for distributed modeling and obtain the prediction model. Input the prediction model into the local historical data to obtain the inference confidence level at any time point.

[0010] The feature data includes measured data and sampled data. The measured data is real-time monitoring type data of objective physical quantities through sensors, such as temperature, pH, hyperspectral data, remote sensing data, or image data, etc.; the sampled data refers to the monitoring data relying on scientific research surveys, such as species distribution, individual quantity, phytoplankton species and quantity, zooplankton species and quantity, etc.

[0011] In an embodiment, an instance of a distributed modeling request is as follows: Modeling request {risk label: {red tide}, scheduling feature set: {phytoplankton density, water temperature, pH, salinity, NO3, chlorophyll concentration}}.

[0012] The process of sending the scheduling feature list and the marine risk type to each graph node and constructing a prediction model respectively is as follows: send a unified request containing the scheduling feature list and the marine risk type to each graph node through the FL-Coordinator federated controller or the SPARQL-Over-Federation federated coordination protocol. After receiving the request, any graph node instantiates a local model according to the risk type: extracts the data of the features corresponding to the scheduling feature list from the local historical data as training data, constructs a local model with the marine risk type as the prediction feature. The local model adopts any one of the decision tree model, the random forest model or the deep learning model, and returns the finally trained model as the prediction model to the current graph node that initiates the scheduling; Here, the decision tree model or the random forest model is preferably used because they have more tolerant modeling characteristics for tasks with a small number of features or missing features. Therefore, they are very suitable for this distributed modeling task where the held feature data type does not completely conform to the scheduling feature list. The more missing features there are, the more suitable it is. If the feature data type completely meets the scheduling feature list, deep learning is preferred because the neural network has the characteristics of strong fitting ability and supports dynamic feature combination, which can integrate associated data to improve the prediction accuracy.

[0013] The inference confidence refers to the confidence that the prediction feature is true, representing the possible degree of the current risk occurrence.

[0014] Furthermore, in step S300, the method for judging the validity of the inference confidence obtained by each graph node within a unit time period is as follows: Set a time period as the acquisition period Copt, Copt ∈ [0.5, 2.5] years. During the acquisition period, obtain the inference confidence every 5 to 10 natural days, and record the time scale for obtaining the inference confidence as the measurement point; During the acquisition period, the inference confidences of each measurement point form an inference sequence; Calculate the standard deviation of the inference sequence of each node and record it as a floating value. If the floating value of any node is greater than the upper quartile of the floating values of all nodes, mark this node as a high-noise node, otherwise mark it as a node to be estimated. The node validity of the high-noise node is false, and this node is unreliable; The inference confidences of each node to be estimated form a reference sequence Rf.Ls; Denote the maximum value in the Rf.Ls sequence as the reference inflection point, and record each measurement point between any reference inflection point and the first reference inflection point in its reverse time direction as a confidence analysis domain; Among them, for the measurement points at the beginning and end of the Rf.Ls sequence that are not covered by any confidence analysis domain, they are respectively assigned to the confidence analysis domain with the closest time to them.

[0015] In the Rf.Ls sequence, calculate the robust reliability value of the current node according to the inter-domain change rate and inter-domain consistency of each confidence analysis domain; Among them, calculate the inter-domain change rate Rt.Cb of each confidence analysis domain, Rt.Cb k =max{((Cb k,i -Mea.Cb k-1 ) / Mea.Cb k-1 )}; where k is the serial number of each confidence analysis domain in the Rf.Ls sequence, i is the serial number of the elements in each confidence analysis domain in the Rf.Ls sequence, Rt.Cb k is the inter-domain change rate of the k-th confidence analysis domain, Cb k,i is the i-th element in the k-th confidence analysis domain, Mea.Cb k-1 is the average value of the (k - 1)-th confidence analysis domain, and max{} is the function to obtain the maximum value; Calculate the inter-domain consistency Ovr.ap of each confidence analysis domain, Ovr.ap k =|CI k ∩CI k-1 | / min(|CI k |,|CI k-1 |); where Ovr.ap k is the inter-domain consistency of the k-th confidence analysis domain, CI k is the interval of the inference confidence of the k-th confidence analysis domain, and min{} is the function to obtain the minimum value; Among them, the previous confidence analysis domain of the first confidence analysis domain is the last confidence analysis domain in the Rf.Ls sequence; Calculate the robust reliability value of the current node according to the inter-domain change rate and inter-domain consistency of each confidence analysis domain, and the specific implementation formula is mean{(1 - |Rt.Cb k |)+ Ovr.ap k )}; where mean{} is the average value function; The inter-domain change rate can reflect the degree of change of a node in each confidence parsing domain and measure the deviation of the current confidence parsing domain from its historical trend. If the inter-domain change rates are generally large, it indicates that the inference confidence of the node fluctuates greatly and the node has a risk of being unreliable; the inter-domain consistency can measure the consistency of the inference confidence time series in different time periods. If the overlapping of the inference confidence intervals of two inference confidence intervals is large, it indicates that the change of the inference confidence of the node is stable. The robust and reliable value can reflect the stability of the node. The smaller the inter-domain change rate, the more stable it is. The larger the inter-domain consistency, the more stable it is. The higher the overall robust and reliable value, the more reliable the inference confidence of the node can be indicated.

[0016] Obtain the difference between the average value and the standard deviation of the robust and reliable values of all nodes to be estimated, and record it as the reliable boundary value. If the robust and reliable value of any node to be estimated is greater than the reliable boundary value, the node validity of the node to be estimated is true and the node is reliable; otherwise, the node validity is false and the node is unreliable.

[0017] Since the validity judgment of each node needs to be obtained by processing the confidence parsing domain, it can effectively quantify the fluctuation of the inference confidence of the marine biodiversity node, identify potential abnormal nodes in advance, and avoid the risk that the inference deviation is amplified in the collaborative analysis process, thereby reducing the interference of abnormal nodes on the overall inference result of the federated knowledge graph. However, the acquisition of the robust and reliable value overly relies on the local fluctuation characteristics of the change of the inference confidence, which easily leads to the sensitivity of the node stability evaluation to the sample density, and further causes a significant deviation between the obtained node validity and the actual stability level, resulting in decision-making errors, especially in the abnormal dense period, making the problem more prominent. However, the existing technologies cannot effectively compensate for the phenomenon of stability misjudgment caused by sample sparsity and misaligned data. To eliminate this influence, the present invention proposes a more preferable solution as follows: Further, in step S300, the method for judging the validity by the inference confidence obtained from each graph node within a unit time period is: Set a time period as the acquisition period Copt, Copt ∈ [2, 3] years. Within the acquisition period, obtain the inference confidence every 5 to 10 natural days, and record the time scale for obtaining the inference confidence as the measurement point; Within the acquisition period, the inference confidences of each measurement point form an inference sequence ER.Ls; Use the Fourier transform algorithm for the inference sequence to obtain the energy spectral density after Fourier transform and form a fluctuation energy sequence; among them, in the result of the Fourier transform algorithm, the frequencies are arranged from low to high, and the corresponding energy spectral densities are also arranged in ascending order of frequency.

[0018] The elements in the second half of the fluctuation energy sequence are summed and recorded as the high-frequency energy value, and the ratio of the high-frequency energy value to the sum of all elements in the fluctuation energy sequence is recorded as the fluctuation proportion Fpo; Since measured data in marine biodiversity monitoring are frequent and sampled data are sparse, when fused and modeled in the federated knowledge graph, too much inference fluctuation will be generated in the short term. Such fluctuations do not reflect real ecological disturbances, but are "false signals" caused by data holes or insufficient model fitting. Therefore, the fluctuation ratio is introduced as a metric to judge the "short-term instability" of nodes. In essence, it is to characterize the proportion of "short-term abnormal fluctuations caused by non-ecological disturbance factors" in the node reasoning sequence from the frequency domain perspective; By performing Fourier transform on the reasoning sequence of each node and converting it to the frequency domain, we can get a complex sequence in the frequency domain. The square of the modulus of the complex number represents the energy of each frequency component. To judge the fluctuation of reasoning confidence, we mainly observe the distribution of energy in frequency. If the energy is concentrated in low frequency, it means that the signal is stable as a whole. If the energy is distributed in high frequency, it means that the signal has significant fluctuations and the reliability of the node will be affected by its fluctuation. Then, we use the fluctuation ratio of each node to screen out nodes with violent fluctuations. A high fluctuation ratio means that the reasoning result of the node has violent, unstable, and short-term impact-like jumps in timing, and often does not have the stable response characteristics at the ecosystem level, and its credibility is low. Special processing of the screened nodes can ensure the reasoning stability and robustness of the collaborative modeling of the federated knowledge graph in a "cross-regional, heterogeneous, and non-aligned data environment."

[0019] If the fluctuation ratio of any node is greater than the upper quartile of the fluctuation ratio of all nodes, the node is marked as a post-node, otherwise it is marked as a pre-node; According to the inference confidence, obtain the stability level of each node; The specific implementation formula of the stationary level is mean(DVF(ER.Ls)) / max(DVF(ER.Ls)); the DVF(ER.Ls) function is a set of absolute values of the difference between each element in the ER.Ls sequence and the previous element, mean() and max() are functions for calculating the average value and obtaining the maximum value, respectively; The Pearson correlation coefficient between any post-node and all pre-nodes is calculated using the inference sequence. If there is a pre-node with a Pearson correlation coefficient greater than 0.7, the stationary level of the pre-node is updated to the minimum value between its original stationary level and the stationary level of the post-node. Among them, if there are multiple preceding nodes with Pearson correlation coefficients greater than 0.7, update the stability level of the preceding node corresponding to the maximum value of the Pearson correlation coefficient; a Pearson correlation coefficient greater than 0.7 can indicate a strong positive correlation between two nodes.

[0020] Obtain the preceding node corresponding to the average value of the stability levels of all preceding nodes that is numerically closest, and denote it as the credible preceding node; The node validity of the credible preceding node is true, and this node is reliable; Calculate the fluctuation risk Faus of other nodes based on the credible preceding nodes: ; Among them, j1 is the cumulative variable, Hqq is the number of measurement points obtained during the acquisition period, ER.Ls(j1) and MER.Ls(j1) are the inference confidence levels of the j1-th measurement point in the inference sequences of the current node and the credible preceding node respectively, MER and P.MER are the stability levels of the current node and the credible preceding node respectively, and exp() is the exponential function with the natural constant e as the base; If the fluctuation risk of any node is less than the 80th percentile of the fluctuation risks of all nodes, the node validity of this node is true and this node is reliable; otherwise, the node validity is false and this node is unreliable.

[0021] Beneficial effects: Since the process of validity judgment is fed back by each graph node, the inference confidence level of predicting local data is not only the prediction result at one time point, but the overall decision-making of multiple inference confidence levels in a continuous time period. Therefore, it can effectively utilize the confidence level sequence comparison and highlight the discontinuous and unstable characteristics of the prediction sequence structure brought by the non-aligned data model. At the same time, it also utilizes the prediction differences of different graph nodes.

[0022] Furthermore, in step S400, the method for obtaining the ocean risk confidence level by selecting valid graph nodes according to the validity judgment result is: Obtain the graph nodes with true node validity for each node as valid nodes, vote on the occurrence of ocean risk at the current position for each valid node, obtain the result with a rate higher than 50% and send it to the client. If it is higher than 50%, the risk occurrence is warned; otherwise, the risk does not occur.

[0023] Preferably, among them, for all variables not defined in the present invention, if there is no clear definition, they can all be artificially set thresholds.

[0024] The present invention also provides a collaborative analysis system for marine biodiversity big data based on a federated knowledge graph. The collaborative analysis system for marine biodiversity big data based on a federated knowledge graph includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the collaborative analysis method for marine biodiversity big data based on a federated knowledge graph are implemented. The collaborative analysis system for marine biodiversity big data based on a federated knowledge graph can run on computing devices such as desktop computers, laptop computers, palm computers, and cloud data centers. The operable system may include, but is not limited to, a processor, a memory, and a server cluster. The processor executes the computer program and runs in the units of the system.

[0025] The beneficial effects of the present invention are as follows: Since the process of validity judgment is fed back by each graph node, the inference confidence of local data is predicted not only by the prediction result at one time point, but by taking the multiple inference confidences in a continuous time period for overall decision-making. Therefore, it can effectively utilize the comparison of confidence sequences and highlight the discontinuous and unstable characteristics of the prediction sequence structure brought by the non-aligned data model. At the same time, it also utilizes the prediction differences of different graph nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] By describing the embodiments shown in the accompanying drawings in detail, the above and other features of the present invention will become more obvious. The same reference numerals in the drawings of the present invention represent the same or similar elements. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 It shows a flowchart of a collaborative analysis method for marine biodiversity big data based on a federated knowledge graph; Figure 2 It shows a structure diagram of a collaborative analysis system for marine biodiversity big data based on a federated knowledge graph. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The following will clearly and completely describe the concept, specific structure, and technical effects generated by the present invention in combination with the embodiments and the drawings, so as to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0028] As Figure 1 It shows a flowchart of a collaborative analysis method for marine biodiversity big data based on a federated knowledge graph. The following will be combined with Figure 1To elaborate on the collaborative analysis method of marine biodiversity big data based on a federated knowledge graph according to an embodiment of the present invention, the method includes the following steps: S100, preset marine risk types, and identify graph nodes from the federated knowledge graph; S200, obtain the inference confidence from each graph node; S300, perform validity judgment on the inference confidence obtained from each graph node within a unit time period; S400, select valid graph nodes according to the validity judgment result to obtain the marine risk confidence.

[0029] Further, in step S100, the method of presetting marine risk types and identifying graph nodes from the federated knowledge graph is as follows: The federated knowledge graph consists of several graph nodes, and different graph nodes represent biodiversity monitoring databases in different geographical locations; preset marine risk types and filter out the graph nodes with the marks of these risk types.

[0030] The marine risk types include red tides, alien invasive species, sea temperature anomalies, ocean acidification, coral bleaching, or aquaculture leakage, etc.; The core of the federated knowledge graph here is the collaborative inference of multiple databases to build a knowledge graph. Because there are defects in building a marine risk model with a single data source, including: first, for unmonitored risk types, data annotation needs to wait until the actual occurrence time point of the risk, and there are obvious shortcomings in the annotation opportunity; second, the insufficient annotation opportunity leads to insufficient robustness of the constructed inference model and low model accuracy.

[0031] In the step of presetting marine risk types, only one marine risk type is selected to make the inference target of the graph nodes consistent, and prevent the inference target of the subsequent constructed inference model from being generalized and blurred, reducing the accuracy. Each data in the biodiversity monitoring database has a timestamp, so that each feature data has a time attribute; for any graph node, when it is determined that a certain marine risk occurs in the corresponding sea area during a certain time period, then the graph node has the mark of the marine risk during that time period. After presetting or selecting the only marine risk type, select the graph nodes from the federated knowledge graph that have the condition of the preset marine risk mark and exclude the graph nodes without the preset marine risk mark.

[0032] Further, in step S200, the method of obtaining the inference confidence from each graph node includes: obtaining the scheduling feature list of the preset marine risk in the current graph node, where the scheduling feature list is a set of features used to build a prediction model, sending the scheduling feature list and the marine risk type to each graph node for distributed modeling to obtain a prediction model, and inputting the prediction model into the local historical data to obtain the inference confidence at any time point.

[0033] Characteristic data include measured data and sampled data. Measured data refers to real-time monitoring data of objective physical quantities through sensors, such as temperature, pH, hyperspectral data, remote sensing data or image data; sampled data refers to monitoring data that relies on scientific research investigations, such as species distribution, number of individuals, types and numbers of phytoplankton, types and numbers of zooplankton, etc.

[0034] In an embodiment, an example of a distributed modeling request is as follows: modeling request {risk label: {red tide}, scheduling feature set: {phytoplankton density, water temperature, pH, salinity, NO3, chlorophyll concentration}}.

[0035] The process of sending the dispatch feature list and ocean risk type to each graph node and building a prediction model respectively is to send a unified request containing the dispatch feature list and ocean risk type to each graph node through the FL-Coordinator federal controller or the SPARQL-Over-Federation federal coordination protocol. After receiving the request, any graph node instantiates the local model according to the risk type: extract the data corresponding to the features of the dispatch feature list in the local historical data as training data, and build a local model with the ocean risk type as the prediction feature. The local model adopts any one of the decision tree model, random forest model or deep learning model, and returns the final model obtained by training as the prediction model to the current graph node that initiates the dispatch. Decision tree models or random forest models are preferred here because they have more tolerant modeling characteristics for tasks with a small number of features or missing features. Therefore, they are very suitable for distributed modeling tasks that have incomplete feature data types and scheduling feature lists. The more missing features, the more suitable it is. If the feature data type completely satisfies the scheduling feature list, deep learning is preferred because neural networks have strong fitting capabilities and support dynamic feature combinations, integrating related data to improve prediction accuracy.

[0036] Inference confidence refers to the confidence that the predicted feature is true, which represents the possibility of the current risk occurring.

[0037] Furthermore, in step S300, the method for judging the validity of the inference confidence obtained by each graph node in a unit time period is: Assume a time period as the collection period Copt, Copt∈[0.5,2.5] years. During the collection period, the inference confidence is obtained every 5 to 10 natural days, and the time scale of obtaining the inference confidence is recorded as the measurement point; During the acquisition period, the inference confidence of each measurement point constitutes an inference sequence; Calculate the standard deviation of the inference sequences of each node and record it as a floating value. If the floating value of any node is greater than the upper quartile of the floating values of all nodes, mark the node as a high-noise node; otherwise, mark it as an estimated node. The node validity of a high-noise node is false, indicating that the node is unreliable. The inference confidence levels of each estimated node form a reference sequence Rf.Ls. Record the maximum value in the Rf.Ls sequence as the reference inflection point, and mark each measurement point between any reference inflection point and the first reference inflection point in its reverse time direction as a confidence analysis domain. Among them, for the measurement points at the beginning and end of the Rf.Ls sequence that are not covered by any confidence analysis domain, they are respectively assigned to the confidence analysis domain with the closest time to them.

[0038] In the Rf.Ls sequence, calculate the robust reliability value of the current node according to the inter-domain change rate and inter-domain consistency of each confidence analysis domain. Among them, calculate the inter-domain change rate Rt.Cb of each confidence analysis domain. Rt.Cb k = max{((Cb k,i - Mea.Cb k-1 ) / Mea.Cb k-1 )}; where k is the serial number of each confidence analysis domain in the Rf.Ls sequence, i is the serial number of the elements in each confidence analysis domain in the Rf.Ls sequence, Rt.Cb k is the inter-domain change rate of the k-th confidence analysis domain, Cb k,i is the i-th element in the k-th confidence analysis domain, Mea.Cb k-1 is the average value of the (k - 1)-th confidence analysis domain, and max{} is the function to obtain the maximum value. Calculate the inter-domain consistency Ovr.ap of each confidence analysis domain. Ovr.ap k = |CI k ∩CI k-1 | / min(|CI k |, |CI k-1 |); where Ovr.ap k is the inter-domain consistency of the k-th confidence analysis domain, CI k is the interval of the inference confidence level of the k-th confidence analysis domain, and min{} is the function to obtain the minimum value. Among them, the previous confidence analysis domain of the first confidence analysis domain is the last confidence analysis domain in the Rf.Ls sequence. Calculate the robust reliability value of the current node according to the inter-domain change rate and inter-domain consistency of each confidence analysis domain. The specific implementation formula is mean{(1 - |Rt.Cb k |) + Ovr.ap k)}; where mean{} is the mean function; The inter-domain change rate can reflect the degree of change of a node in each confidence parsing domain and measure the degree of deviation of the current confidence parsing domain from its historical trend. If the inter-domain change rates are generally large, it indicates that the inference confidence of the node fluctuates greatly and the node has the risk of being unreliable; the inter-domain consistency can measure the consistency of the inference confidence time series in different time periods. If the overlapping of the inference confidence intervals of two inference confidence intervals is large, it indicates that the change of the inference confidence of the node is stable. The robust reliability value can reflect the stability of the node. The smaller the inter-domain change rate, the more stable it is. The larger the inter-domain consistency, the more stable it is. The higher the overall robust reliability value, the more reliable the inference confidence of the node can be indicated.

[0039] Obtain the difference between the mean and the standard deviation of the robust reliability values of all nodes to be estimated, and denote it as the reliable boundary value. If the robust reliability value of any node to be estimated is greater than the reliable boundary value, the node validity of the node to be estimated is true and the node is reliable; otherwise, the node validity is false and the node is unreliable.

[0040] Since the validity judgment of each node needs to be obtained by processing the confidence parsing domain, it can effectively quantify the inference confidence fluctuation of the marine biodiversity node, identify potential abnormal nodes in advance, and avoid the risk that the inference deviation is amplified in the collaborative analysis process, thereby reducing the interference of abnormal nodes on the overall inference result of the federated knowledge graph. However, the acquisition of the robust reliability value overly relies on the local fluctuation characteristics of the inference confidence change, which easily leads to the sensitivity of the node stability evaluation to the sample density, and further causes a significant deviation between the obtained node validity and the actual stability level, resulting in decision-making errors, especially in the abnormally dense period, making the problem more prominent. However, the existing technologies cannot effectively compensate for the phenomenon of stability misjudgment caused by sample sparsity and misaligned data. To eliminate this influence, the present invention proposes a more preferable solution as follows: Further, in step S300, the method for judging the validity by the inference confidence obtained from each graph node within a unit time period is: Set a time period as the acquisition period Copt, Copt ∈ [2, 3] years. During the acquisition period, obtain the inference confidence once every 5 to 10 natural days, and record the time scale of obtaining the inference confidence as the measurement point; During the acquisition period, the inference confidences of each measurement point form an inference sequence ER.Ls; Use the Fourier transform algorithm for the inference sequence to obtain the energy spectral density after Fourier transform and form a fluctuation energy sequence; among them, in the result of the Fourier transform algorithm, the frequencies are arranged from low to high, and the corresponding energy spectral densities are also arranged from small to large according to the frequencies.

[0041] The elements in the second half of the fluctuation energy sequence are summed and recorded as the high-frequency energy value, and the ratio of the high-frequency energy value to the sum of all elements in the fluctuation energy sequence is recorded as the fluctuation proportion Fpo; Since measured data in marine biodiversity monitoring are frequent and sampled data are sparse, when fused and modeled in the federated knowledge graph, too much inference fluctuation will be generated in the short term. Such fluctuations do not reflect real ecological disturbances, but are "false signals" caused by data holes or insufficient model fitting. Therefore, the fluctuation ratio is introduced as a metric to judge the "short-term instability" of nodes. In essence, it is to characterize the proportion of "short-term abnormal fluctuations caused by non-ecological disturbance factors" in the node reasoning sequence from the frequency domain perspective; By performing Fourier transform on the reasoning sequence of each node and converting it to the frequency domain, we can get a complex sequence in the frequency domain. The square of the modulus of the complex number represents the energy of each frequency component. To judge the fluctuation of reasoning confidence, we mainly observe the distribution of energy in frequency. If the energy is concentrated in low frequency, it means that the signal is stable as a whole. If the energy is distributed in high frequency, it means that the signal has significant fluctuations and the reliability of the node will be affected by its fluctuation. Then, we use the fluctuation ratio of each node to screen out nodes with violent fluctuations. A high fluctuation ratio means that the reasoning result of the node has violent, unstable, and short-term impact-like jumps in timing, and often does not have the stable response characteristics at the ecosystem level, and its credibility is low. Special processing of the screened nodes can ensure the reasoning stability and robustness of the collaborative modeling of the federated knowledge graph in a "cross-regional, heterogeneous, and non-aligned data environment."

[0042] If the fluctuation ratio of any node is greater than the upper quartile of the fluctuation ratio of all nodes, the node is marked as a post-node, otherwise it is marked as a pre-node; According to the inference confidence, obtain the stability level of each node; The specific implementation formula of the stationary level is mean(DVF(ER.Ls)) / max(DVF(ER.Ls)); the DVF(ER.Ls) function is a set of absolute values of the difference between each element in the ER.Ls sequence and the previous element, mean() and max() are functions for calculating the average value and obtaining the maximum value, respectively; The Pearson correlation coefficient between any post-node and all pre-nodes is calculated using the inference sequence. If there is a pre-node with a Pearson correlation coefficient greater than 0.7, the stationary level of the pre-node is updated to the minimum value between its original stationary level and the stationary level of the post-node. Among them, if there are multiple precursor nodes with Pearson correlation coefficients greater than 0.7, update the steady level of the precursor node corresponding to the maximum value of the Pearson correlation coefficient; a Pearson correlation coefficient greater than 0.7 can indicate a strong positive correlation between two nodes.

[0043] Obtain the precursor node corresponding to the average value of the steady levels of all precursor nodes that is numerically closest, and denote it as the credible precursor node; The node validity of the credible precursor node is true, and the node is reliable; Calculate the fluctuation risk Faus of other nodes based on the credible precursor nodes: ; Among them, j1 is the cumulative variable, Hqq is the number of measurement points obtained during the acquisition period, ER.Ls(j1) and MER.Ls(j1) are the inference confidence levels of the j1-th measurement point in the inference sequences of the current node and the credible precursor node respectively, MER and P.MER are the steady levels of the current node and the credible precursor node respectively, and exp() is the exponential function with the natural constant e as the base; If the fluctuation risk of any node is less than the 80th percentile of the fluctuation risks of all nodes, the node validity of this node is true and the node is reliable; otherwise, the node validity is false and the node is unreliable.

[0044] Beneficial effects: Since the process of validity judgment is fed back by each graph node, the inference confidence level of predicting local data is not only the prediction result at one time point, but the overall decision-making of multiple inference confidence levels in a continuous time period. Therefore, it can effectively use the confidence level sequence comparison and highlight the discontinuous and unstable characteristics of the prediction sequence structure brought by the non-aligned data model. At the same time, it also utilizes the prediction differences of different graph nodes.

[0045] Furthermore, in step S400, the method for obtaining the marine risk confidence level by selecting valid graph nodes according to the validity judgment result is: Obtain the graph nodes with true node validity for each node as valid nodes, vote on the occurrence of marine risks at the current position by each valid node, obtain the result with a rate higher than 50% and send it to the client. If it is higher than 50%, it warns of the occurrence of risks, otherwise the risks do not occur.

[0046] The marine biodiversity big data collaborative analysis system based on the federated knowledge graph provided by the embodiments of the present invention, such as Figure 2The following is a structural diagram of the marine biodiversity big data collaborative analysis system based on the federated knowledge graph of the present invention. The marine biodiversity big data collaborative analysis system based on the federated knowledge graph in this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the embodiment of the above-mentioned marine biodiversity big data collaborative analysis method based on the federated knowledge graph are implemented.

[0047] The system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it runs in the units of the system.

[0048] The marine biodiversity big data collaborative analysis system based on the federated knowledge graph can run on computing devices such as desktop computers, laptop computers, palmtop computers, and cloud servers. The marine biodiversity big data collaborative analysis system based on the federated knowledge graph, the system that can run may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above examples are only examples of the marine biodiversity big data collaborative analysis system based on the federated knowledge graph, and do not constitute a limitation on the marine biodiversity big data collaborative analysis system based on the federated knowledge graph. It may include more or fewer components than the examples, or combine some components, or different components. For example, the marine biodiversity big data collaborative analysis system based on the federated knowledge graph may also include input / output devices, network access devices, buses, etc.

[0049] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the operating system of the marine biodiversity big data collaborative analysis system based on the federated knowledge graph, and uses various interfaces and lines to connect all parts of the operating system of the marine biodiversity big data collaborative analysis system based on the federated knowledge graph.

[0050] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory, the processor can implement various functions of the collaborative analysis system for marine biodiversity big data based on the federated knowledge graph. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0051] Although the description of the present invention has been quite detailed and several of the described embodiments have been described in particular, it is not intended to be limited to any of these details or embodiments or any particular embodiment, so as to effectively cover the intended scope of the present invention. In addition, the present invention has been described above in terms of embodiments foreseeable by the inventor in order to provide a useful description, and non-substantive modifications to the present invention that are not currently foreseeable may still represent equivalent modifications of the present invention.

Claims

1. A collaborative analysis method for marine biodiversity big data based on a federated knowledge graph, characterized in that, The method includes the following steps: S100, preset marine risk types, and identify graph nodes from the federal knowledge graph; S200, obtain inference confidence levels from each graph node; S300, perform validity judgment based on the inference confidence levels obtained from each graph node within a unit time period; S400, select valid graph nodes according to the validity judgment result to obtain marine risk confidence levels.

2. The collaborative analysis method for marine biodiversity big data based on the federated knowledge graph according to claim 1, wherein In step S100, the method of presetting marine risk types and identifying graph nodes from the federal knowledge graph is as follows: The federal knowledge graph consists of several graph nodes, and different graph nodes represent biodiversity monitoring databases in different geographical locations; Preset marine risk types and filter out graph nodes marked with this risk type.

3. The collaborative analysis method for marine biodiversity big data based on the federated knowledge graph according to claim 1, wherein, In step S200, the method of obtaining inference confidence levels from each graph node includes: obtaining a preset scheduling feature list of the marine risk at the current graph node. The scheduling feature list is a set of features used to construct a prediction model, sending the scheduling feature list and the marine risk type to each graph node for distributed modeling to obtain a prediction model, and inputting the prediction model into local historical data to obtain the inference confidence level at any time point.

4. The collaborative analysis method for marine biodiversity big data based on a federated knowledge graph according to claim 1, wherein, In step S300, the method of performing validity judgment based on the inference confidence levels obtained from each graph node within a unit time period is: Set a time period as the acquisition period Copt, Copt ∈ [0.5, 2.5] years. During the acquisition period, obtain the inference confidence level every 5 to 10 natural days, and record the time scale of obtaining the inference confidence level as the measurement point; During the acquisition period, the inference confidence levels of each measurement point form an inference sequence; Calculate the standard deviation of the inference sequence of each node and record it as a floating value. If the floating value of any node is greater than the upper quartile of the floating values of all nodes, mark this node as a high-noise node, otherwise mark it as a node to be estimated. The node validity of the high-noise node is false, and this node is unreliable; The inference confidence levels of each node to be estimated form a reference sequence Rf.Ls; Record the maximum value in the Rf.Ls sequence as the reference inflection point, and record each measurement point between any reference inflection point and the first reference inflection point in its reverse time direction as a confidence analysis domain; Obtain the difference between the average value and the standard deviation of the robust and reliable values of all nodes to be estimated, and record it as the reliable boundary value. If the robust and reliable value of any node to be estimated is greater than the reliable boundary value, the node validity of this node to be estimated is true, and this node is reliable; otherwise, the node validity is false, and this node is unreliable.

5. The collaborative analysis method for marine biodiversity big data based on the federated knowledge graph according to claim 1, wherein In step S300, the method of performing validity judgment based on the inference confidence levels obtained from each graph node within a unit time period is: Set a time period as the acquisition period Copt, Copt ∈ [2, 3] years. During the acquisition period, obtain the inference confidence level every 5 to 10 natural days, and record the time scale of obtaining the inference confidence level as the measurement point; During the acquisition period, the inference confidence levels of each measurement point form an inference sequence ER.Ls; Use the Fourier transform algorithm on the inference sequence to obtain the energy spectral density after Fourier transform and form a fluctuation energy sequence; Sum the elements in the second half of the fluctuation energy sequence and denote it as the high-frequency energy value. Denote the ratio of the high-frequency energy value to the sum of all elements in the fluctuation energy sequence as the fluctuation proportion Fpo. If the fluctuation proportion of any node is greater than the upper quartile of the fluctuation proportions of all nodes, mark this node as a post node; otherwise, mark it as a pre node. Obtain the stability level of each node according to the inference confidence. Use the inference sequence to calculate the Pearson correlation coefficient between any post node and all pre nodes. If there is a pre node with a Pearson correlation coefficient greater than 0.7, update the stability level of this pre node to the minimum value between its original stability level and the stability level of this post node. Obtain the pre node corresponding to the average value of the stability levels of all pre nodes that is numerically closest, and denote it as the credible pre node. The node validity of the credible pre node is true, and this node is reliable. Calculate the fluctuation risk Faus of other nodes according to the credible pre node: ; where j1 is the cumulative variable, Hqq is the number of measurement points obtained during the acquisition period, ER.Ls(j1) and MER.Ls(j1) are the inference confidences of the j1-th measurement point in the inference sequences of the current node and the credible pre node respectively, MER and P.MER are the stability levels of the current node and the credible pre node respectively, and exp() is the exponential function with the natural constant e as the base. If the fluctuation risk of any node is less than the 80% quantile of the fluctuation risks of all nodes, the node validity of this node is true, and this node is reliable; otherwise, the node validity is false, and this node is unreliable.

6. The collaborative analysis method for marine biodiversity big data based on a federated knowledge graph according to claim 1, wherein, In step S400, the method for obtaining the marine risk confidence by selecting valid graph nodes according to the validity judgment result is: Obtain the graph nodes with true node validity for each node as valid nodes. Vote on the occurrence of the marine risk at the current position by each valid node, and send the result with a rate higher than 50% to the client. If it is higher than 50%, the risk is warned to occur; otherwise, the risk does not occur.

7. The collaborative analysis system for marine biodiversity big data based on the federated knowledge graph is characterized in that The marine biodiversity big data collaborative analysis system based on the federated knowledge graph includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the method for marine biodiversity big data collaborative analysis based on the federated knowledge graph according to any one of claims 1-6. The marine biodiversity big data collaborative analysis system based on the federated knowledge graph runs on computing devices such as desktop computers, laptop computers, palm computers, and cloud data centers.

Citation Information

Patent Citations

  • Inference method, system and device based on knowledge federation and graph network and medium

    CN112200321A

  • Method for constructing ecological conservation geographic knowledge graph

    CN113505234A

  • Knowledge federal undirected graph-based federal ring detection method, system and device, and medium

    CN113656802A

  • Power grid fault intelligent analysis and disposal method and system based on knowledge graph

    CN117992743A

  • Connectome Ensemble Transfer Learning

    US20240161017A1