Fishing boat behavior knowledge construction method and system based on multi-source data
By constructing a knowledge graph and using graph neural network for model training, the problem of data aggregation of multi-source fishing vessels and missing entity correlation information in the existing technology is solved, and high-accurate fishing vessel behavior prediction and online optimization are achieved.
Patent Information
- Application Number
- CN202510109203.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology fails to effectively collect multi-source fishing vessel data and does not introduce related information between entities, resulting in low accuracy in model prediction and inability to continuously optimize online.
By collecting fishing boat behavior data, building a knowledge graph, and using graph neural networks to train, predict and optimize the model to achieve accurate prediction and analysis of fishing boat behavior.
It improves the accuracy of fishing boat behavior prediction, supports automatic continuous online optimization, and can quickly adapt to new nodes to improve the prediction accuracy of unknown scenarios.
Smart Images

Figure CN120069026A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method for constructing fishing vessel behavior knowledge based on multi-source data.
Background Art
[0002] Currently, for the disaster prevention and mitigation requirements such as safety assessment, hidden danger investigation, risk early warning, and emergency rescue of fishing vessels and fishermen, traditional data collection and analysis methods are still mainly relied on; and these methods usually use a single data source, lacking the integrated application of multi-source data, resulting in the inability to fully utilize the potential of data in the decision-making process.
[0003] Certainly, there are also some knowledge model construction schemes based on multi-source data in the prior art. Their knowledge models are mostly obtained through deep learning training, and entities are represented as vectors, which is also a mainstream approach to constructing knowledge models. However, the existing schemes do not effectively converge and process multi-source fishing vessel data, and do not introduce the association information between entities. Only single entities are considered as samples for model training, resulting in the trained models being prone to false alarms and missed alarms, with low model prediction accuracy and the inability to continuously optimize online. In view of the above existing problems, the inventor of this case conducted in-depth research on this problem, and thus this case was born.
Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for constructing fishing vessel behavior knowledge based on multi-source data, which solves the problems in the prior art that multi-source fishing vessel data is not effectively converged and processed, the association information between entities is not introduced, only single entities are considered as samples for model training, resulting in the trained models being prone to false alarms and missed alarms, with low model prediction accuracy and the inability to continuously optimize online.
[0005] The present invention is implemented as follows:
[0006] In a first aspect, a method for constructing fishing vessel behavior knowledge based on multi-source data, the method includes the following steps:
[0007] Convergence of multi-source data of fishing vessel behavior: Using various sensors and standardized API interfaces to collect fishing vessel behavior data, and performing convergence processing on the collected fishing vessel behavior data;
[0008] Construction of fishing vessel behavior knowledge: Taking entities as nodes and the relationships between entities as edges, constructing a knowledge graph through entity recognition and relationship recognition, and assigning weights to the edges in the knowledge graph according to the importance of the relationships;
[0009] Construction of fishing vessel behavior model: Using the knowledge graph to train, predict, and optimize the graph neural network model, so as to obtain the final fishing vessel behavior knowledge model.
[0010] Further, when aggregating and processing the collected fishing vessel behavior data, it also includes:
[0011] Deploy a producer-consumer model based on Kafka, and push the fishing vessel behavior data received from various sensors and standardized API interfaces to the Kafka topic in real time, and use multiple downstream processing threads to consume the fishing vessel behavior data in the Kafka topic;
[0012] At the same time, during the data processing, the partition mechanism of Kafka is also used to perform load balancing on the data, the transaction management mechanism of Kafka is enabled for transaction management, and Kafka Streams or Apache Flink is introduced into the data stream.
[0013] Further, the aggregation and processing of the collected fishing vessel behavior data includes using a feature selection mechanism to screen the collected fishing vessel behavior data, specifically including:
[0014] Input the fishing vessel behavior data into the XGBoost model for fitting, set the target variable as the classification result of the fishing vessel behavior, and set the model parameters at the same time;
[0015] Set the screening threshold, use the feature_importances_ attribute of the XGBoost model to extract the importance scores of each feature, and eliminate the features with low importance according to the importance scores of each feature and the set screening threshold;
[0016] Retrain the XGBoost model with the remaining features and evaluate the performance metrics of the model to ensure that the screened feature subset can improve or maintain the model performance.
[0017] Further, the aggregation and processing of the collected fishing vessel behavior data also includes establishing a unified data format standard and processing the fishing vessel behavior data according to the unified data format standard, specifically including:
[0018] Unify the longitude and latitude format: adopt the WGS-84 coordinate system and accurate the longitude and latitude to six decimal places;
[0019] Convert the device signals into standardized fields by defining field mapping relationships: determine the field ranges of all device signals, list the field names and corresponding types of all device signals; design a field mapping table, define the standardized field names and corresponding conversion rules; use ETL tools to perform batch automatic mapping and conversion on the device signals; verify the conversion results through random sampling and rule checking.
[0020] Further, the aggregation and processing of the collected fishing vessel behavior data also includes:
[0021] Data association is performed based on unique identifiers to merge the behavior data of the same fishing vessel from different sources, and invalid data is filtered out using a time series anomaly detection algorithm;
[0022] For continuous variables, appropriate processing methods are selected according to the data distribution characteristics, including but not limited to using box plots or Z-scores to detect and remove outliers, using bucketing techniques to discretize data into multiple intervals, and using normalization or standardization to adjust data to a uniform range;
[0023] Category distribution balancing is used for discrete features, including but not limited to adjusting the data volume by upsampling or downsampling, and mapping categories to numerical values by using frequency encoding or target encoding.
[0024] Furthermore, the aggregating and processing the collected fishing boat behavior data also includes:
[0025] Use the SHA-256 hash function to encrypt the selected field data;
[0026] When sharing or displaying data externally, introduce differential privacy for geographic location information;
[0027] Use permission control and hierarchical storage strategies to restrict access to sensitive data in different levels.
[0028] Furthermore, the fishing vessel behavior knowledge construction specifically includes:
[0029] Entity recognition: For unstructured data, the pre-trained language model is used to perform semantic understanding and part-of-speech tagging of the text, identify key entities, and use context information to optimize the determination of entity boundaries; the labeled data is used to supervise the NER model, and the supervised learning NER model is used to extract entities from the unstructured data as nodes of the knowledge graph;
[0030] For structured data, field matching and rule mapping are used to directly parse the structured data into nodes of the knowledge graph;
[0031] Relationship identification: A combination of rules and machine learning is used to mine various relationships between entities and use them as edges in the knowledge graph.
[0032] Knowledge representation: Based on the identified entities and the relationships between them, the fishing vessel behavior knowledge is represented as a knowledge graph consisting of nodes and edges. The nodes and edges in the knowledge graph carry attributes, and the edges in the knowledge graph are weighted according to the importance of the relationship.
[0033] Knowledge storage: The Neo4j graph database is used to store the knowledge graph. When importing data, a batch import script is used to import node and edge data from the standardized CSV or JSON format into Neo4j, and Cypher statements are used to set indexes and constraints to optimize query performance. At the same time, for dynamic scenarios, the Apache Flink stream processing framework is integrated to achieve real-time updates or additions of nodes and edges.
[0034] Further, the construction of the fishing boat behavior model specifically includes:
[0035] Model training: The feature representations of nodes and edges in the knowledge graph are input into the graph neural network. The graph neural network uses a multi-layer graph convolutional network (GCN) or graph attention network (GAT). The feature representations of nodes are updated by aggregating the information of neighboring nodes. At the same time, during the training process, the model parameters are optimized by defining a multi-task loss function, and mini-batch stochastic gradient descent combined with a learning rate scheduling strategy is used for training and learning.
[0036] Model prediction: Predict the fishing boat behavior based on the final representations of nodes in the graph neural network.
[0037] Model optimization: Adjust the key parameters through hyperparameter search and combine regularization methods to prevent overfitting. At the same time, when new nodes are added to the knowledge graph, the graph neural network updates the weights of nodes through an incremental learning mechanism, so as to achieve rapid adaptation to new nodes.
[0038] Further, the graph neural network uses a multi-layer graph convolutional network (GCN) or graph attention network (GAT). The process of updating the feature representations of nodes by aggregating the information of neighboring nodes specifically includes:
[0039] Collection of neighbor node features: Identify the neighbors of the target node through the adjacency matrix of the knowledge graph and collect the features of the neighbor nodes corresponding to the target node.
[0040] Feature aggregation: Weighted aggregation of the features of neighbor nodes. Specifically, for the graph convolutional network (GCN), the standardized adjacency matrix is multiplied by the feature matrix to achieve aggregation; for the graph attention network (GAT), the attention weights between neighbor nodes and the target node are calculated, and the features are weighted averaged.
[0041] Nonlinear transformation: Apply an activation function to the aggregated features.
[0042] Feature update: Combine the aggregated features with the original features of the target node to update the feature representation of the target node.
[0043] Multi-layer processing: By stacking multiple layers of graph convolutional networks (GCNs) or graph attention networks (GATs), gradually integrate the information of neighbors at farther distances.
[0044] In a second aspect, a fishing vessel behavior knowledge construction system based on multi-source data, the system includes a data aggregation module, a knowledge graph construction module, and a model construction module;
[0045] The data aggregation module is used for aggregating multi-source data of fishing vessel behavior: collecting fishing vessel behavior data by using various sensors and standardized API interfaces, and performing aggregation processing on the collected fishing vessel behavior data;
[0046] The knowledge graph construction module is used for constructing fishing vessel behavior knowledge: taking entities as nodes and the relationships between entities as edges, constructing a knowledge graph through entity recognition and relationship recognition, and weighting the edges in the knowledge graph according to the importance of the relationships;
[0047] The model construction module is used for constructing a fishing vessel behavior model: using the knowledge graph to train, predict, and optimize the graph neural network model, so as to obtain the final fishing vessel behavior knowledge model.
[0048] The present invention uses technical means such as formatting and normalization to perform aggregation processing on multi-source fishing vessel behavior data; taking entities as nodes and the relationships between entities as edges, constructing a knowledge graph through entity recognition and relationship recognition on the aggregated fishing vessel behavior data, and weighting the edges in the knowledge graph according to the importance of the relationships; the graph neural network adopts a multi-layer graph convolutional network GCN or a graph attention network GAT, and updates the feature representation of the nodes by aggregating the information of neighboring nodes. By adopting the technical solution of the present invention, not only can more accurate prediction and analysis of fishing vessel behavior be realized, the accuracy of model prediction be improved, so as to provide efficient support for fishery management and reliable data support for risk monitoring and resource allocation in marine economic activities; but also it can well support automatic continuous online optimization, and at the same time can quickly adapt to newly added nodes, improving the prediction accuracy for unknown scenarios.
BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The following further describes the present invention with reference to the accompanying drawings in conjunction with embodiments.
[0050] Figure 1 is a principle block diagram of a method for constructing fishing vessel behavior knowledge based on multi-source data according to the present invention;
[0051] Figure 2 is a structural block diagram of a fishing vessel behavior knowledge construction system based on multi-source data according to the present invention.
DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] In order to better understand the technical solution of the present invention, the technical solution of the present invention will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0053] Example 1
[0054] Please refer to Figure 1 As shown in the figure, a method for constructing fishing vessel behavior knowledge based on multi-source data according to the present invention includes the following steps:
[0055] Convergence of multi-source data of fishing vessel behavior: Using various sensors and standardized API interfaces to collect fishing vessel behavior data, and performing convergence processing on the collected fishing vessel behavior data; In order to ensure the comprehensiveness and accuracy of the data as much as possible to correctly describe the fishing vessel state, in the specific implementation of the present invention, fishing vessel behavior data from multiple data sources such as fishing vessel positioning, fishing vessel basic information, fishing vessel certificate information, and fishing vessel terminal information are collected using various sensors and standardized API interfaces, and these fishing vessel behavior data cover various fields such as longitude, latitude, vessel name, vessel length, vessel width, vessel owner, vessel owner's contact phone number, fishing license number, nationality ownership number, equipment signal, equipment type, and equipment code number;
[0056] Construction of fishing vessel behavior knowledge: Using entities as nodes and the relationships between entities as edges, constructing a knowledge graph through entity recognition and relationship recognition, that is, representing each entity as a node in the knowledge graph, and representing the feature vector of the node according to its attributes (such as fishing vessel number, fishing vessel positioning, etc.), representing the relationship between entities as an edge in the knowledge graph, and assigning weights to the edges in the knowledge graph according to the importance of the relationship;
[0057] Construction of fishing vessel behavior model: Using the knowledge graph to train, predict, and optimize the graph neural network model to obtain the final fishing vessel behavior knowledge model.
[0058] In some embodiments of the present invention, when performing convergence processing on the collected fishing vessel behavior data, it further includes:
[0059] Deploying a producer-consumer model based on Kafka, pushing the fishing vessel behavior data received from various sensors and standardized API interfaces to the Kafka topic in real time, and using multiple downstream processing threads to consume the fishing vessel behavior data in the Kafka topic to ensure the fast transfer and processing of the data stream;
[0060] Meanwhile, during the data processing, the partitioning mechanism of Kafka is also used to perform load balancing on the data, the transaction management mechanism of Kafka is enabled for transaction management, and Kafka Streams or Apache Flink is introduced into the data stream. Among them, enabling the transaction management mechanism of Kafka for transaction management specifically means: by setting the enable.idempotence parameter to true, the transaction management mechanism of Kafka is enabled to ensure that the producer can achieve exactly-once delivery during message passing. At the same time, transaction logs are used at the consumer side to record unfinished processing tasks to support fault recovery and re-consumption. The partitioning mechanism of Kafka allows data to be scattered and stored on different nodes of the Kafka cluster, realizing horizontal expansion and load balancing of data to improve the throughput and scalability of Kafka; Kafka Streams is a lightweight stream processing library that allows developers to implement complex stream processing logics, such as real-time aggregation, event-driven processing, etc., by writing Java code; Apache Flink is a distributed stream processing engine that provides high performance, fault tolerance, and exactly-once processing guarantees. Flink supports event-time-based processing, can handle delayed and out-of-order data, and ensures accurate processing results.
[0061] By deploying the Kafka-based producer-consumer model in the aggregation processing of data, the present invention can well achieve real-time data access under high concurrency, thus adapting to high-frequency data acquisition scenarios and ensuring the fast transfer and processing of the data stream; at the same time, the partitioning mechanism of Kafka is used to perform load balancing on the data, and combined with the transaction function of Kafka, it can effectively avoid data loss caused by node failures, thus ensuring data consistency; in addition, Kafka Streams or Apache Flink is introduced into the data stream, which can realize real-time filtering, aggregation, and transformation of high-frequency data to better meet the subsequent analysis requirements and is beneficial to further optimizing the real-time performance.
[0062] In some embodiments of the present invention, the aggregation processing of the collected fishing boat behavior data includes using a feature selection mechanism to screen the collected fishing boat behavior data, specifically including:
[0063] A1. Input the fishing boat behavior data into the XGBoost model for fitting, specifically input the cleaned fishing boat behavior data into the XGBoost model for fitting, set the target variable as the classification result of the fishing boat behavior, and set the model parameters at the same time; wherein, the set model parameters include but are not limited to the maximum depth of the tree (max_depth), the learning rate (learning_rate) and the number of weak classifiers (n_estimators). As a specific implementation of the present invention, the initial value of the maximum depth of the tree can be set to 6, the learning rate can be set to 0.1, the number of weak classifiers can be set to 100, and the model can be optimized by grid search or random search; the XGBoost model is an optimized gradient boosting decision tree algorithm, which has made many improvements on the basis of GBDT (Gradient Boosting Decision Tree), including the second-order Taylor expansion of the loss function, the addition of regularization terms, parallel computing and missing value processing, thereby significantly improving the training speed and model generalization ability;
[0064] A2. Setting a screening threshold. In the specific implementation of the present invention, the screening threshold can be determined according to the feature importance distribution (for example, retaining features with a cumulative importance of 95%); the feature_importances_ attribute of the XGBoost model is used to extract the importance score of each feature, and the features with low importance are eliminated according to the importance score of each feature and the set screening threshold; wherein feature_importances_ is an attribute in the XGBoost model, which represents the importance score of each feature in the model. This attribute is calculated by statistical characteristics such as the splitting benefit of the decision tree node, and is specifically used to measure the contribution of the feature in the model;
[0065] A3. Retrain the XGBoost model using the retained features and evaluate the model's performance indicators (such as accuracy, F1 score, etc.) to ensure that the filtered feature subset can improve or maintain the model performance.
[0066] The present invention calculates the importance scores of features by adopting the XGBoost model, and screens out fields that have a significant impact on the prediction of fishing vessel behavior according to the importance scores, effectively reducing the interference of data redundancy and noise on subsequent model training, thereby helping to improve the accuracy of model prediction.
[0067] In some embodiments of the present invention, the aggregating and processing the collected fishing vessel behavior data further includes establishing a unified data format standard, and processing the fishing vessel behavior data according to the unified data format standard, specifically including:
[0068] Unify the longitude and latitude format: Adopt the WGS-84 coordinate system and accurate the longitude and latitude to six decimal places; among them, the WGS-84 coordinate system is a three-dimensional, spherical coordinate system used to determine the position of any point on the earth's surface, including longitude, latitude and altitude information;
[0069] Convert device signals into standardized fields by defining field mapping relationships: Determine the field ranges of all device signals, list the field names of all device signals and their corresponding types; design a field mapping table to define the standardized field names and corresponding conversion rules (such as type conversion, value range mapping); use an ETL tool (such as Apache Nifi or a custom script) to batch and automatically map and convert device signals; verify the conversion results through random sampling and rule checking to ensure the accuracy of the standardized fields.
[0070] By establishing a unified data format standard and processing fishing vessel behavior data according to the unified data format standard, the present invention can ensure the consistency of multi-source data.
[0071] In some embodiments of the present invention, in order to ensure the data quality and stability of the input model, the aggregation processing of the collected fishing vessel behavior data further includes:
[0072] Perform data association based on unique identifiers, merge the behavior data of the same fishing vessel from different sources, and at the same time use a time series anomaly detection algorithm to filter out invalid data; in the specific implementation of the present invention, the behavior data of the same fishing vessel from different sources can be associated and merged based on unique identifiers such as fishing vessel numbers and device code numbers; at the same time, for noise and abnormal data, methods such as ARIMA or LSTM-based can be used to filter out invalid data.
[0073] For continuous variables (such as the length and width of fishing vessels), select appropriate processing methods according to the data distribution characteristics for processing, specifically including but not limited to using box plot method or Z-score method for outlier detection and elimination, using binning technology to discretize data into multiple intervals, and using normalization or standardization to adjust data to a unified range. That is, in the specific implementation of the present invention, if there are obvious outliers in the data, the box plot method or Z-score method can be used for outlier detection and elimination; for the case where the numerical distribution is relatively concentrated but has a long-tailed distribution, binning technology can be used to discretize data into multiple intervals, for example, dividing the ship length into "small", "medium" and "large" categories; for data presenting a normal distribution or close to a normal distribution, normalization (Min-Max Scaling) or standardization (Z-Score) can be used to adjust the data to a unified range to ensure the comparability of features during model training.
[0074] For discrete features, class distribution balancing processing is adopted, specifically including but not limited to adjusting the data volume by upsampling or downsampling, and mapping classes to numerical values using frequency encoding or target encoding. That is, when the present invention is specifically implemented, when statistically analyzing the class distribution, if the number of samples in certain classes is too small, upsampling (such as SMOTE) or downsampling methods can be used to adjust the data volume to reduce the interference of class imbalance on model learning; if the number of classes is too large and there is redundant information, frequency encoding (Frequency Encoding) or target encoding (Target Encoding) can be used to map classes to numerical values.
[0075] In some embodiments of the present invention, the aggregation processing of the collected fishing vessel behavior data further includes:
[0076] Using the SHA-256 hash function to encrypt the selected field data; for example, for fields of unique identification information such as the ID number and contact phone number of the shipowner, the SHA-256 hash function can be used to encrypt them to ensure that the data cannot be reversely restored;
[0077] When sharing or displaying data externally, differential privacy (Differential Privacy) is introduced for geographical location information (such as longitude and latitude). Differential privacy technology means adding a certain degree of noise to the data so that attackers cannot obtain the privacy information of specific individuals through data analysis; by introducing differential privacy for geographical location information when sharing or displaying data externally, the present invention can protect privacy while retaining the effectiveness of the data;
[0078] Adopting a permission control and hierarchical storage strategy to restrict hierarchical access to sensitive data to ensure data security and compliance, specifically including:
[0079] Permission control: Adopting a role-based access control (RBAC) model to define user roles and their corresponding access permissions; for example, the administrator role has full access rights, while ordinary users can only view non-sensitive information; the access scope of each type of data is clearly defined through a permission matrix;
[0080] Data classification: Classify and label the data (such as public, restricted, confidential, etc.), and set hierarchical storage rules, such as storing restricted data in an isolated database instance and enabling advanced encryption for confidential data;
[0081] Audit mechanism: Introduce a logging function to audit each data access to ensure the traceability of data access;
[0082] Dynamic permission management: Dynamically adjust permissions in combination with user behavior analysis (UBA), for example, triggering a permission tightening mechanism through abnormal login or access behaviors.
[0083] By adopting multiple protection measures to process sensitive data, the present invention can well protect privacy.
[0084] In some embodiments of the present invention, the construction of fishing vessel behavior knowledge specifically includes:
[0085] Entity recognition: For unstructured data (such as fishing logs, transaction texts, etc.), through a pre-trained language model (such as BERT or GPT) to perform semantic understanding and part-of-speech tagging on the text, key entities are identified, and the determination of entity boundaries is optimized using context information. The specific method of optimizing the determination of entity boundaries using context information is: by introducing a bidirectional attention mechanism (such as BiDAF) to capture the context correlation of the text before and after the entity, thereby improving the accuracy of boundary determination. At the same time, a dynamic conditional random field (CRF) is combined to globally decode the output sequence to ensure the consistency and accuracy of entity boundaries; a labeled data is used to perform supervised learning on the NER model, and a CRF layer is often used during the supervised learning process to improve the sequence labeling accuracy to ensure the accuracy of entity extraction; the NER model after supervised learning is used to extract entities (such as crew members, ports, etc.) from unstructured data as nodes of the knowledge graph.
[0086] For structured data, field matching and rule mapping are used to directly parse the structured data into nodes of the knowledge graph; for example, the "fishing vessel number" field is directly mapped to a unique identifier node, and the "port name" field is directly generated into a port entity node; after parsing, the consistency of the nodes is ensured through a unified coding standard to avoid duplication and ambiguity; at the same time, in combination with the data cleaning step, redundant or incorrect field values are removed to ensure the accuracy and reliability of node attributes; the process of vectorizing the features of the nodes includes converting categorical features into embedding vectors, for example, mapping the port name to a 128-dimensional feature representation through an Embedding layer, and numerical features (such as ship length, fishing value, etc.) are directly embedded as attributes.
[0087] Relationship identification: A method combining rules and machine learning is used to mine various relationships between entities for use as edges in the knowledge graph. In the specific implementation of the present invention, a method combining rules and machine learning is used to mine various relationships between entities, including various relationships such as fishing boat-trading, fishing boat-terminal, fishing boat-peer, fishing boat-port, etc. The rule method is specifically implemented through template matching. For example, if fishing boats A and B are moored at the same port and the time interval is less than one day, a "peer" relationship is generated; the machine learning method is to establish predictive relationships between entities through the GNN model. These relationships are represented as edges in the graph. The weights of the edges represent the interaction frequency and semantic similarity between different entities. The graph neural network learns the weights of these edges through convolution operations. These weights can be regarded as the importance of a certain relationship. The difference from the existing method is that the existing method manually assigns weights to different relationships. The introduction of such subjective information may be detrimental to the objectivity and accuracy of the system. The following lists some definition rules for relationships for illustration:
[0088] 1. "Fishing boat-buyer" relationship: By analyzing the transaction record data, if the number of transactions between the same buyer or seller and a fishing boat exceeds a certain threshold, a buyer-seller relationship is established; specifically, natural language processing technology can be used to match key entities in the unstructured text in the transaction record to ensure the correct identification of the transaction participants;
[0089] 2. "Fishing boat-traveling" relationship: Based on the fishing boat positioning data, if the longitude and latitude tracks of two fishing boats remain close within a certain time window (such as the distance is less than 500 meters) and the overlap time exceeds 1 hour, a traveling relationship is generated; this process is achieved through a density clustering algorithm based on DBSCAN, which automatically identifies close interactions between tracks;
[0090] 3. "Fishing boat-port" relationship: When the longitude and latitude of a fishing boat matches the location of a known port (within 1 km of the allowable error) and the berthing time exceeds 30 minutes, the association between the fishing boat and the port is established; this rule combines the geographic location matching algorithm and berthing time analysis.
[0091] Knowledge Representation: According to the identified entities and the relationships between them, the knowledge of fishing vessel behavior is represented as a knowledge graph composed of nodes and edges. The nodes and edges in the knowledge graph carry attributes, and weights are assigned to the edges in the knowledge graph according to the importance of the relationships. By adopting this knowledge representation method of the present invention, non-linear complex relationships can be presented in an intuitive and easy-to-process manner. The following further introduces the knowledge graph of the present invention through a specific example: For example, in the knowledge graph, there are three nodes, namely "Fishing Vessel A", "Port X", and "Crew Member B". The node attributes of "Fishing Vessel A" can include basic information such as fishing vessel number, positioning information (latitude and longitude), and captain. The node attributes of "Port X" can include port name, geographical location, berth capacity, etc. The node attributes of "Crew Member B" can include crew member name, position, contact phone number, etc. The relationships established between these nodes are such as "Fishing Vessel A - Docked at - Port X", "Crew Member B - Belongs to - Fishing Vessel A", etc. The weights of the relationships can represent, for example, docking frequency, importance of the position, etc.
[0092] Knowledge Storage: The Neo4j graph database is used to store the knowledge graph. Since the Neo4j graph database has efficient query capabilities and a flexible attribute management mechanism, it is particularly suitable for storing and managing large-scale and complex knowledge graphs. And when importing data, batch import scripts are used to import the data of nodes and edges from the standardized CSV or JSON format into Neo4j, and Cypher statements are used to set indexes and constraints to optimize query performance. At the same time, for dynamic scenarios, the Apache Flink stream processing framework is integrated to achieve real-time update or addition of nodes and edges to keep the knowledge graph in the latest state.
[0093] In some embodiments of the present invention, the construction of the fishing vessel behavior model specifically includes:
[0094] Model Training: The feature representations of nodes and edges in the knowledge graph are input into a graph neural network. The graph neural network adopts a multi-layer graph convolutional network (GCN) or graph attention network (GAT). By aggregating the information of neighboring nodes, the feature representation of nodes is updated. In this way, the final representation of nodes integrates their own features and the information of neighboring nodes, and can more comprehensively describe the behavior characteristics of fishing boats. At the same time, through graph convolution operations, the model can gradually update the feature representation of nodes, thereby learning complex behavior patterns. For example, the behavior of a single fishing boat entering and leaving the port can be predicted through the change pattern of the fishing boat's position, while the behavior of multiple fishing boats traveling together depends on the relationship strength between fishing boats. That is, in the present invention, the constructed knowledge graph is used by a graph neural network (GNN) to capture the complex behavior patterns of fishing boats. The input of the graph neural network (GNN) is the feature representation of nodes and edges, and the processing performed by the graph neural network (GNN) is to aggregate the information of neighboring nodes to update the feature representation of nodes. The output of the graph neural network (GNN) can be the prediction result of the current state of the fishing boat (such as entering and leaving the port, fishing operation, multiple boats traveling together, etc.).
[0095] Meanwhile, during the training process, the model parameters are optimized by defining a multi-task loss function (such as behavior classification loss and edge weight prediction loss), and mini-batch stochastic gradient descent combined with a learning rate scheduling strategy is used for training and learning to improve the convergence speed of the model. As a specific implementation manner of the present invention, the mini-batch size is set to 32 to balance memory consumption and the stability of gradient updates; the initial learning rate is set to 0.01 and gradually decreased through a cosine annealing scheduling strategy. By adopting mini-batch stochastic gradient descent combined with a learning rate scheduling strategy, the present invention enables the model to quickly learn important patterns in the initial stage, and at the same time gradually reduces the learning rate, so that it can focus on refining the model parameters in the later stage to improve the prediction accuracy. In addition, to further enhance the stability of model training, the momentum technique can be combined, usually set to 0.9, to ensure the continuity of the gradient update direction and avoid drastic fluctuations.
[0096] Model Prediction: The behavior of the fishing boat (such as entering and leaving the port, fishing operation, etc.) is predicted based on the final representation of nodes in the graph neural network. In particular, in the present invention, when a new entity is added, according to the information of the entity, an edge can be established with the existing nodes in the knowledge graph. Based on the edge and the information of neighboring nodes, the model can make predictions for the newly added nodes, such as judging behaviors such as a single fishing boat entering and leaving the port, fishing operation, multiple boats traveling together, formation separation, etc. Therefore, compared with the existing methods, the present invention can make more accurate predictions for fishing boat behaviors not included in the training data. At the same time, by combining the attention mechanism, the model can explain the prediction results and identify the neighboring nodes or relationships that contribute the most to the current prediction, thereby providing a reference for decision-making.
[0097] Model Optimization: Adjust key parameters through hyperparameter search (such as Optuna, etc.), and combine regularization methods to prevent overfitting. At the same time, when new nodes (such as new fishing boats, crew members, etc.) are added to the knowledge graph, the graph neural network updates the weights of the nodes through an incremental learning mechanism, so as to achieve rapid adaptation to new nodes. In addition, the model can also optimize the prediction effect based on feedback, and gradually improve the prediction accuracy for unknown scenarios. For example, for ships that do not appear in the training data, the model can still infer through their features and connection relationships and predict their behavior patterns. As can be seen from the above, the model of the present invention can well support real-time prediction and dynamic optimization after going online.
[0098] More specifically, the graph neural network adopts a multi-layer graph convolutional network GCN or graph attention network GAT, and updates the feature representation of the nodes by aggregating the information of neighbor nodes, which specifically includes:
[0099] Collection of neighbor node features: Identify the neighbors of the target node through the adjacency matrix of the knowledge graph, and collect the features of the neighbor nodes corresponding to the target node;
[0100] Feature aggregation: Weighted aggregation of the features of neighbor nodes. Specifically, for the graph convolutional network GCN, use the normalized adjacency matrix to multiply the feature matrix to achieve aggregation; for the graph attention network GAT, calculate the attention weights between neighbor nodes and the target node, and perform weighted averaging on the features;
[0101] Nonlinear transformation: Apply an activation function to the aggregated features to introduce nonlinearity;
[0102] Feature update: Combine the aggregated features with the original features of the target node (such as by addition or concatenation) to update the feature representation of the target node;
[0103] Multi-layer processing: By stacking multi-layer graph convolutional networks GCN or graph attention networks GAT, gradually integrate the information of farther neighbors, so as to enrich the global representation ability of the nodes. At the same time, in the specific implementation of the present invention, the graph convolution operation of each layer can be combined with the weights of the edges to achieve weighted propagation of information.
[0104] In summary, the core inventive concept of the present invention includes: using technical means such as formatting and normalization to converge and process multi-source fishing vessel behavior data; taking entities as nodes and the relationships between entities as edges, constructing a knowledge graph by performing entity recognition and relationship recognition on the converged fishing vessel behavior data, and assigning weights to the edges in the knowledge graph according to the importance of the relationships; using a multi-layer graph convolutional network (GCN) or graph attention network (GAT) for the graph neural network to update the feature representation of the nodes by aggregating the information of neighboring nodes. By adopting the technical solution of the present invention, not only can more accurate prediction and analysis of fishing vessel behavior be realized, the accuracy of model prediction be improved, so as to provide efficient support for fishery management and reliable data support for risk monitoring, resource allocation, etc. in marine economic activities; but also automatic continuous online optimization can be well supported, and at the same time, new added nodes can be quickly adapted, improving the prediction accuracy for unknown scenarios.
[0105] Embodiment 2
[0106] Please refer to Figure 2 As shown, a fishing vessel behavior knowledge construction system based on multi-source data of the present invention, the system includes a data convergence module, a knowledge graph construction module, and a model construction module;
[0107] The data convergence module is used for converging multi-source data of fishing vessel behavior: collecting fishing vessel behavior data by using various sensors and standardized API interfaces, and performing convergence processing on the collected fishing vessel behavior data; in order to ensure the comprehensiveness and accuracy of the data as much as possible to correctly describe the state of the fishing vessel, in the specific implementation of the present invention, fishing vessel behavior data from multiple data sources such as fishing vessel positioning, fishing vessel basic information, fishing vessel certificate information, and fishing vessel terminal information are collected by using various sensors and standardized API interfaces, and these fishing vessel behavior data cover various fields such as longitude, latitude, vessel name, vessel length, vessel width, vessel owner, vessel owner contact phone number, fishing license number, nationality ownership number, equipment signal, equipment type, and equipment code number;
[0108] The knowledge graph construction module is used for constructing fishing vessel behavior knowledge: taking entities as nodes and the relationships between entities as edges, constructing a knowledge graph through entity recognition and relationship recognition, that is, representing each entity as a node in the knowledge graph, and representing the feature vector of the node according to its attributes (such as fishing vessel number, fishing vessel positioning, etc.), representing the relationship between entities as an edge in the knowledge graph, and assigning weights to the edges in the knowledge graph according to the importance of the relationships;
[0109] The model construction module is used for constructing a fishing vessel behavior model: using the knowledge graph to perform model training, model prediction, and model optimization on the graph neural network, so as to obtain the final fishing vessel behavior knowledge model.
[0110] In some embodiments of the present invention, when the data aggregation module performs aggregation processing on the collected fishing vessel behavior data, it is further configured to:
[0111] Deploy a Kafka-based producer-consumer model, and push the fishing vessel behavior data received from various sensors and standardized API interfaces to the Kafka topic in real time, and use multiple downstream processing threads to consume the fishing vessel behavior data in the Kafka topic to ensure the rapid transfer and processing of the data stream;
[0112] Meanwhile, during the data processing, the partition mechanism of Kafka is also used to perform load balancing on the data, the transaction management mechanism of Kafka is enabled for transaction management, and Kafka Streams or Apache Flink is introduced into the data stream. Among them, enabling the transaction management mechanism of Kafka for transaction management specifically means: by setting the enable.idempotence parameter to true, the transaction management mechanism of Kafka is enabled to ensure that the producer can achieve exactly-once delivery during message passing. At the same time, transaction logs are used at the consumer side to record the unfinished processing tasks to support fault recovery and re-consumption. The partition mechanism of Kafka allows data to be scattered and stored on different nodes of the Kafka cluster, realizing horizontal expansion and load balancing of the data to improve the throughput and scalability of Kafka; Kafka Streams is a lightweight stream processing library that allows developers to implement complex stream processing logics, such as real-time aggregation, event-driven processing, etc., by writing Java code; Apache Flink is a distributed stream processing engine that provides high performance, fault tolerance, and exactly-once processing guarantees. Flink supports event-time-based processing, can process delayed and out-of-order data, and ensures accurate processing results.
[0113] By deploying a Kafka-based producer-consumer model in the aggregation processing of data, the present invention can well achieve real-time data access under high concurrency, thus adapting to high-frequency data collection scenarios and ensuring the rapid transfer and processing of the data stream; at the same time, the partition mechanism of Kafka is used to perform load balancing on the data, and combined with the transaction function of Kafka, it can effectively avoid data loss caused by node failures, thus ensuring data consistency; in addition, Kafka Streams or Apache Flink is introduced into the data stream, which can realize real-time filtering, aggregation, and transformation of high-frequency data to better meet the subsequent analysis requirements and is conducive to further optimizing the real-time performance.
[0114] In some embodiments of the present invention, in the data aggregation module, the aggregation processing of the collected fishing vessel behavior data includes screening the collected fishing vessel behavior data using a feature selection mechanism, specifically including:
[0115] A1. Input the fishing boat behavior data into the XGBoost model for fitting, specifically input the cleaned fishing boat behavior data into the XGBoost model for fitting, set the target variable as the classification result of the fishing boat behavior, and set the model parameters at the same time; wherein, the set model parameters include but are not limited to the maximum depth of the tree (max_depth), the learning rate (learning_rate) and the number of weak classifiers (n_estimators). As a specific implementation of the present invention, the initial value of the maximum depth of the tree can be set to 6, the learning rate can be set to 0.1, the number of weak classifiers can be set to 100, and the model can be optimized by grid search or random search; the XGBoost model is an optimized gradient boosting decision tree algorithm, which has made many improvements on the basis of GBDT (Gradient Boosting Decision Tree), including the second-order Taylor expansion of the loss function, the addition of regularization terms, parallel computing and missing value processing, thereby significantly improving the training speed and model generalization ability;
[0116] A2. Setting a screening threshold. In the specific implementation of the present invention, the screening threshold can be determined according to the feature importance distribution (for example, retaining features with a cumulative importance of 95%); the feature_importances_ attribute of the XGBoost model is used to extract the importance score of each feature, and the features with low importance are eliminated according to the importance score of each feature and the set screening threshold; wherein feature_importances_ is an attribute in the XGBoost model, which represents the importance score of each feature in the model. This attribute is calculated by statistical characteristics such as the splitting benefit of the decision tree node, and is specifically used to measure the contribution of the feature in the model;
[0117] A3. Retrain the XGBoost model using the retained features and evaluate the model's performance indicators (such as accuracy, F1 score, etc.) to ensure that the filtered feature subset can improve or maintain the model performance.
[0118] The present invention calculates the importance scores of features by adopting the XGBoost model, and screens out fields that have a significant impact on the prediction of fishing vessel behavior according to the importance scores, effectively reducing the interference of data redundancy and noise on subsequent model training, thereby helping to improve the accuracy of model prediction.
[0119] In some embodiments of the present invention, in the data aggregation module, the aggregation process of the collected fishing vessel behavior data further includes establishing a unified data format standard, and processing the fishing vessel behavior data according to the unified data format standard, specifically including:
[0120] Unify the longitude and latitude format: Adopt the WGS-84 coordinate system, and accurate the longitude and latitude to six decimal places; among them, the WGS-84 coordinate system is a three-dimensional, spherical coordinate system used to determine the position of any point on the earth's surface, including longitude, latitude and altitude information;
[0121] Convert the device signal into a standardized field by defining the field mapping relationship: Determine the field range of all device signals, list the field names and corresponding types of all device signals; Design a field mapping table, define the standardized field names and corresponding conversion rules (such as type conversion, value range mapping); Use an ETL tool (such as Apache Nifi or a custom script) to perform batch automatic mapping and conversion on the device signals; Verify the conversion results through random sampling and rule checking to ensure the accuracy of the standardized fields.
[0122] By establishing a unified data format standard and processing the fishing vessel behavior data according to the unified data format standard, the present invention can ensure the consistency of multi-source data.
[0123] In some embodiments of the present invention, in order to ensure the data quality and stability of the input model, in the data aggregation module, the aggregation process of the collected fishing vessel behavior data further includes:
[0124] Perform data association based on the unique identifier, merge the same fishing vessel behavior data from different sources, and at the same time use the time series anomaly detection algorithm to filter out invalid data; In the specific implementation of the present invention, the same fishing vessel behavior data from different sources can be associated and merged based on unique identifiers such as fishing vessel numbers and device code numbers; At the same time, for noise and abnormal data, ARIMA or LSTM-based methods can be used to filter out invalid data.
[0125] For continuous variables (such as the length and width of fishing boats), appropriate processing methods are selected according to the data distribution characteristics, including but not limited to using box plot method or Z-score method for outlier detection and removal, using binning technology to discretize data into multiple intervals, and using normalization or standardization to adjust data to a unified range. That is, in the specific implementation of the present invention, if there are obvious outliers in the data, the box plot method or Z-score method can be used for outlier detection and removal; for the case where the numerical distribution is relatively concentrated but has a long-tailed distribution, the binning technology can be used to discretize data into multiple intervals, for example, dividing the ship length into "small", "medium" and "large" categories; for data presenting a normal distribution or close to a normal distribution, normalization (Min-Max Scaling) or standardization (Z-Score) can be used to adjust data to a unified range to ensure the comparability of features during model training.
[0126] For discrete features, class distribution balancing processing is adopted, including but not limited to using oversampling or undersampling to adjust the data volume, and using frequency encoding or target encoding to map categories to numerical values. That is, in the specific implementation of the present invention, when counting the class distribution, if the number of samples in some classes is too small, oversampling (such as SMOTE) or undersampling methods can be used to adjust the data volume to reduce the interference of class imbalance on model learning; if the number of classes is too large and there is redundant information, frequency encoding or target encoding can be used to map categories to numerical values.
[0127] In some embodiments of the present invention, in the data aggregation module, the aggregation processing of the collected fishing boat behavior data further includes:
[0128] Using the SHA-256 hash function to encrypt the selected field data; for example, for fields of unique identification information such as the ID number and contact phone number of the ship owner, the SHA-256 hash function can be used to encrypt them to ensure that the data cannot be reversely restored;
[0129] When sharing or displaying data externally, differential privacy is introduced for geographical location information (such as longitude and latitude). Differential privacy technology means adding a certain degree of noise to the data so that attackers cannot obtain the privacy information of specific individuals through data analysis; by introducing differential privacy for geographical location information when sharing or displaying data externally, the present invention can protect privacy while retaining the effectiveness of the data;
[0130] Adopt permission control and hierarchical storage strategies to restrict hierarchical access to sensitive data to ensure data security and compliance. Specifically, it includes:
[0131] Permission control: Adopt a role-based access control (RBAC) model to define user roles and their corresponding access permissions. For example, the administrator role has full access rights, while ordinary users can only view non-sensitive information. Clearly define the access scope of each type of data through a permission matrix.
[0132] Data classification: Classify and label data (such as public, restricted, confidential, etc.), and set hierarchical storage rules. For example, store restricted data in an isolated database instance and enable advanced encryption for confidential data.
[0133] Audit mechanism: Introduce a logging function to audit each data access to ensure the traceability of data access.
[0134] Dynamic permission management: Dynamically adjust permissions in combination with user behavior analysis (UBA). For example, trigger a permission tightening mechanism through abnormal login or access behavior.
[0135] By adopting multi-layer protection measures to process sensitive data, the present invention can well protect privacy.
[0136] In some embodiments of the present invention, the knowledge graph construction module is specifically used for:
[0137] Entity recognition: For unstructured data (such as fishery logs, transaction texts, etc.), through a pre-trained language model (such as BERT or GPT), perform semantic understanding and part-of-speech tagging on the text to identify key entities. Use context information to optimize the determination of entity boundaries. Specifically, introduce a bidirectional attention mechanism (such as BiDAF) to capture the context association of the text before and after the entity, thereby improving the accuracy of boundary determination. At the same time, combine a dynamic conditional random field (CRF) to globally decode the output sequence to ensure the consistency and accuracy of entity boundaries. Use labeled data for supervised learning of the NER model. During the supervised learning process, often use a CRF layer to improve the sequence labeling accuracy to ensure the accuracy of entity extraction. Use the supervised NER model to extract entities (such as crew members, ports, etc.) from unstructured data as nodes of the knowledge graph.
[0138] For structured data, field matching and rule mapping are used to directly parse the structured data into nodes of the knowledge graph; for example, the "fishing boat number" field is directly mapped to a unique identifier node, and the "port name" field is directly generated into a port entity node; after parsing, the consistency of the nodes is ensured through a unified coding standard to avoid duplication and ambiguity; at the same time, combined with the data cleaning step, redundant or incorrect field values are removed to ensure the accuracy and reliability of the node attributes; the process of vectorizing the features of the nodes includes converting categorical features into embedding vectors, for example, mapping the port name to a 128-dimensional feature representation through an Embedding layer, and numerical features (such as ship length, fishing value, etc.) are directly embedded as attributes.
[0139] Relationship identification: A method combining rules and machine learning is adopted to mine various relationships between entities for use as the edges of the knowledge graph; in the specific implementation of the present invention, a method combining rules and machine learning is used to mine various relationships between entities, including relationships such as fishing boat - buying and selling, fishing boat - terminal, fishing boat - peer, fishing boat - port, etc. The rule method is specifically implemented through template matching. For example, if fishing boats A and B are berthed at the same port and the time interval is less than one day, then a "peer" relationship is generated; the machine learning method is to establish a predictive relationship between entities through a GNN model. These relationships are represented as edges in the graph, and the weights of the edges represent the interaction frequency and semantic similarity between different entities. The graph neural network learns the weights of these edges through convolutional operations, and these weights can be regarded as the importance of a certain relationship. The difference from the existing method is that the existing method assigns weights to different relationships artificially, and the introduction of this subjective information may be unfavorable to the objectivity and accuracy of the system; some more relationship definition rules are listed below for illustration:
[0140] 1. "Fishing boat - buying and selling" relationship: By analyzing the transaction record data, if the number of transactions of the same buyer or seller with a certain fishing boat exceeds a certain threshold, then a buying and selling relationship is established; specifically, natural language processing technology can be used to perform key entity matching on the unstructured text in the transaction record to ensure the correct identification of the transaction participants.
[0141] 2. "Fishing boat - peer" relationship: Based on the fishing boat positioning data, if the latitude and longitude trajectories of two fishing boats are close within a certain time window (such as the distance is less than 500 meters) and the overlapping time exceeds 1 hour, then a peer relationship is generated; this process is implemented through a density clustering algorithm based on DBSCAN to automatically identify the close interactions between the trajectories.
[0142] 3. "Fishing boat - port" relationship: When the latitude and longitude where a certain fishing boat is berthed matches the location of a known port (allowing an error within 1 kilometer) and the berthing duration exceeds 30 minutes, then an association relationship between the fishing boat and the port is established; here, the rule combines a geographical location matching algorithm and berthing duration analysis.
[0143] Knowledge Representation: According to the identified entities and the relationships between them, the fishing vessel behavior knowledge is represented as a knowledge graph composed of nodes and edges. The nodes and edges in the knowledge graph carry attributes, and weights are assigned to the edges in the knowledge graph according to the importance of the relationships. By adopting this knowledge representation method of the present invention, non-linear complex relationships can be presented in an intuitive and easy-to-process manner. The following further introduces the knowledge graph of the present invention through a specific example: For example, in the knowledge graph, there are three nodes, namely "Fishing Vessel A", "Port X", and "Crew Member B". The node attributes of "Fishing Vessel A" may include basic information such as fishing vessel number, positioning information (latitude and longitude), and captain. The node attributes of "Port X" may include port name, geographical location, berth capacity, etc. The node attributes of "Crew Member B" may include crew member name, position, contact phone number, etc. The relationships established between these nodes are such as "Fishing Vessel A - Docked at - Port X", "Crew Member B - Belongs to - Fishing Vessel A", etc. The weights of the relationships can represent, for example, docking frequency, importance of the position, etc.
[0144] Knowledge Storage: The Neo4j graph database is used to store the knowledge graph. Since the Neo4j graph database has efficient query capabilities and a flexible attribute management mechanism, it is particularly suitable for storing and managing large-scale and complex knowledge graphs. And when importing data, a batch import script is used to import the data of nodes and edges from the standardized CSV or JSON format into Neo4j, and Cypher statements are used to set indexes and constraints to optimize query performance. At the same time, for dynamic scenarios, by integrating the Apache Flink stream processing framework, real-time updates or new nodes and edges are realized to keep the knowledge graph in the latest state.
[0145] In some embodiments of the present invention, the model construction module is specifically used for:
[0146] Model Training: The feature representations of nodes and edges in the knowledge graph are input into a graph neural network. The graph neural network adopts multi-layer graph convolutional networks (GCNs) or graph attention networks (GATs). By aggregating the information of neighboring nodes, it updates the feature representations of nodes. In this way, the final representations of nodes integrate their own features and the information of neighboring nodes, enabling a more comprehensive description of the behavior characteristics of fishing vessels. At the same time, through graph convolution operations, the model can gradually update the feature representations of nodes, thereby learning complex behavior patterns. For example, the behavior of a single fishing vessel entering or leaving the port can be predicted based on the change pattern of the fishing vessel's position, while the behavior of multiple fishing vessels traveling together depends on the strength of the relationship between fishing vessels. That is, in the present invention, the constructed knowledge graph is used to capture the complex behavior patterns of fishing vessels through a graph neural network (GNN). The input of the graph neural network (GNN) is the feature representations of nodes and edges, and the processing performed by the graph neural network (GNN) is to aggregate the information of neighboring nodes to update the feature representations of nodes. The output of the graph neural network (GNN) can be the prediction results of the current state of the fishing vessel (such as entering or leaving the port, fishing operations, multiple fishing vessels traveling together, etc.).
[0147] Meanwhile, during the training process, the model parameters are optimized by defining a multi-task loss function (such as behavior classification loss and edge weight prediction loss), and mini-batch stochastic gradient descent combined with a learning rate scheduling strategy is used for training and learning to improve the convergence speed of the model. As a specific implementation manner of the present invention, the mini-batch size is set to 32 to balance memory consumption and the stability of gradient updates. The initial learning rate is set to 0.01 and gradually decreased through a cosine annealing scheduling strategy. By adopting mini-batch stochastic gradient descent combined with a learning rate scheduling strategy, the present invention enables the model to quickly learn important patterns in the initial stage, while gradually reducing the learning rate, and can focus on refining the model parameters in the later stage to improve the prediction accuracy. In addition, to further enhance the stability of model training, the momentum technique can be combined, usually set to 0.9, to ensure the continuity of the gradient update direction and avoid drastic fluctuations.
[0148] Model Prediction: The behavior of the fishing vessel (such as entering or leaving the port, fishing operations, etc.) is predicted based on the final representations of nodes in the graph neural network. In particular, in the present invention, when a new entity is added, according to the information of the entity, edges can be established with the existing nodes in the knowledge graph. Based on the edges and the information of neighboring nodes, the model can make predictions for the newly added nodes, such as judging behaviors like a single fishing vessel entering or leaving the port, fishing operations, multiple fishing vessels traveling together, formation separation, etc. Therefore, compared with existing methods, the present invention can make more accurate predictions for the behaviors of fishing vessels that are not included in the training data. At the same time, by combining the attention mechanism, the model can explain the prediction results and identify the neighboring nodes or relationships that contribute the most to the current prediction, thereby providing a reference for decision-making.
[0149] Model Optimization: Adjust key parameters through hyperparameter search (such as Optuna, etc.) and combine regularization methods to prevent overfitting. At the same time, when new nodes (such as new fishing boats, crew members, etc.) are added to the knowledge graph, the graph neural network updates the weights of the nodes through an incremental learning mechanism, so as to achieve rapid adaptation to new nodes. In addition, the model can also optimize the prediction effect based on feedback and gradually improve the prediction accuracy for unknown scenarios. For example, for boats that do not appear in the training data, the model can still infer their behavior patterns through their features and connection relationships. As can be seen from the above, the model of the present invention can well support real-time prediction and dynamic optimization after going online.
[0150] More specifically, the graph neural network adopts a multi-layer graph convolutional network GCN or graph attention network GAT, and updates the feature representation of the nodes by aggregating the information of neighbor nodes, which specifically includes:
[0151] Collection of neighbor node features: Identify the neighbors of the target node through the adjacency matrix of the knowledge graph and collect the features of the neighbor nodes corresponding to the target node;
[0152] Feature aggregation: Weighted aggregation of the features of neighbor nodes. Specifically, for the graph convolutional network GCN, use the normalized adjacency matrix to multiply the feature matrix to achieve aggregation; for the graph attention network GAT, calculate the attention weights between neighbor nodes and the target node, and perform weighted averaging on the features;
[0153] Nonlinear transformation: Apply an activation function to the aggregated features to introduce nonlinearity;
[0154] Feature update: Merge the aggregated features with the original features of the target node (such as by addition or concatenation) to update the feature representation of the target node;
[0155] Multi-layer processing: By stacking multiple layers of graph convolutional network GCN or graph attention network GAT, gradually integrate the information of farther neighbors to enrich the global representation ability of the nodes. At the same time, in the specific implementation of the present invention, the graph convolution operation of each layer can be combined with the weights of the edges to achieve weighted propagation of information.
[0156] Although the specific implementation manners of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A method for constructing fishing vessel behavior knowledge based on multi-source data, characterized in that: The method comprises the following steps: Fishing vessel behavior multi-source data aggregation: Use various sensors and standardized API interfaces to collect fishing vessel behavior data, and aggregate and process the collected fishing vessel behavior data; Fishing vessel behavior knowledge construction: Entities are used as nodes, and the relationships between entities are used as edges. A knowledge graph is constructed through entity recognition and relationship recognition. The edges in the knowledge graph are weighted according to the importance of the relationships. Construction of fishing vessel behavior model: Use knowledge graph to perform model training, model prediction and model optimization on graph neural network to obtain the final fishing vessel behavior knowledge model.
2. A method for constructing fishing vessel behavior knowledge based on multi-source data as claimed in claim 1, characterized in that: When the collected fishing vessel behavior data is aggregated and processed, it also includes: Deploy a producer-consumer model based on Kafka to push the fishing boat behavior data received from various sensors and standardized API interfaces to Kafka topics in real time, and use multiple downstream processing threads to consume the fishing boat behavior data in the Kafka topic; At the same time, during the data processing process, Kafka's partitioning mechanism is used to load balance the data, Kafka's transaction management mechanism is enabled for transaction management, and Kafka Streams or Apache Flink is introduced into the data stream.
3. The method for constructing fishing vessel behavior knowledge based on multi-source data according to claim 1, characterized in that: The aggregation processing of the collected fishing vessel behavior data includes screening the collected fishing vessel behavior data using a feature selection mechanism, specifically including: The fishing boat behavior data is input into the XGBoost model for fitting, the target variable is set as the classification result of the fishing boat behavior, and the model parameters are set at the same time; Set a filtering threshold, use the feature_importances_ attribute of the XGBoost model to extract the importance score of each feature, and remove low-importance features based on the importance score of each feature and the set filtering threshold; Retrain the XGBoost model using the retained features and evaluate the model's performance metrics to ensure that the filtered feature subset can improve or maintain model performance.
4. The method for constructing fishing vessel behavior knowledge based on multi-source data according to claim 1, characterized in that: The aggregating and processing the collected fishing vessel behavior data also includes establishing a unified data format standard, and processing the fishing vessel behavior data according to the unified data format standard, specifically including: Unified longitude and latitude format: Use the WGS-84 coordinate system to accurately calculate longitude and latitude to six decimal places; Convert device signals into standardized fields by defining field mapping relationships: determine the field range of all device signals, list the field names and corresponding types of all device signals; design a field mapping table, define standardized field names and corresponding conversion rules; use ETL tools to automatically map and convert device signals in batches; verify the conversion results through random sampling and rule verification.
5. The method for constructing fishing vessel behavior knowledge based on multi-source data according to claim 1, characterized in that: The aggregating and processing the collected fishing boat behavior data also includes: Data association is performed based on unique identifiers to merge the behavior data of the same fishing vessel from different sources, and invalid data is filtered out using a time series anomaly detection algorithm; For continuous variables, appropriate processing methods are selected according to the data distribution characteristics, including but not limited to using box plots or Z-scores to detect and remove outliers, using bucketing techniques to discretize data into multiple intervals, and using normalization or standardization to adjust data to a uniform range; Category distribution balancing is used for discrete features, including but not limited to adjusting the data volume by upsampling or downsampling, and mapping categories to numerical values by using frequency encoding or target encoding.
6. The method for constructing fishing vessel behavior knowledge based on multi-source data according to claim 1, characterized in that: The aggregating and processing the collected fishing boat behavior data also includes: Use the SHA-256 hash function to encrypt the selected field data; When sharing or displaying data externally, introduce differential privacy for geographic location information; Use permission control and hierarchical storage strategies to restrict access to sensitive data in different levels.
7. The method for constructing fishing vessel behavior knowledge based on multi-source data according to claim 1, characterized in that: The fishing vessel behavior knowledge construction specifically includes: Entity recognition: For unstructured data, the pre-trained language model is used to perform semantic understanding and part-of-speech tagging of the text, identify key entities, and use context information to optimize the determination of entity boundaries; the labeled data is used to supervise the NER model, and the supervised learning NER model is used to extract entities from the unstructured data as nodes of the knowledge graph; For structured data, field matching and rule mapping are used to directly parse the structured data into nodes of the knowledge graph; Relationship identification: A combination of rules and machine learning is used to mine various relationships between entities and use them as edges in the knowledge graph. Knowledge representation: Based on the identified entities and the relationships between them, the fishing vessel behavior knowledge is represented as a knowledge graph consisting of nodes and edges. The nodes and edges in the knowledge graph carry attributes, and the edges in the knowledge graph are weighted according to the importance of the relationship. Knowledge storage: Neo4j graph database is used to store knowledge graphs. When importing data, a batch import script is used to import node and edge data from standardized CSV or JSON format into Neo4j. Indexes and constraints are set through Cypher statements to optimize query performance. At the same time, for dynamic scenarios, the Apache Flink stream processing framework is integrated to achieve real-time updates or new additions of nodes and edges.
8. The method for constructing fishing vessel behavior knowledge based on multi-source data according to claim 1, characterized in that: The construction of the fishing vessel behavior model specifically includes: Model training: The feature representations of nodes and edges in the knowledge graph are input into the graph neural network. The graph neural network uses a multi-layer graph convolutional network GCN or a graph attention network GAT to update the feature representation of nodes by aggregating the information of neighboring nodes. At the same time, during the training process, the model parameters are optimized by defining a multi-task loss function, and small batch stochastic gradient descent combined with a learning rate scheduling strategy is used for training and learning. Model prediction: predicting fishing boat behavior based on the final representation of the nodes in the graph neural network; Model optimization: adjust key parameters through hyperparameter search and combine regularization methods to prevent overfitting; at the same time, when new nodes are added to the knowledge graph, the graph neural network updates the node weights through an incremental learning mechanism, thereby achieving rapid adaptation to the new nodes.
9. A method for constructing fishing vessel behavior knowledge based on multi-source data as claimed in claim 8, characterized in that: The graph neural network uses a multi-layer graph convolutional network GCN or a graph attention network GAT to update the feature representation of the node by aggregating the information of neighboring nodes. Specifically, it includes: Neighbor node feature collection: The neighbors of the target node are identified through the adjacency matrix of the knowledge graph, and the features of the neighbor nodes corresponding to the target node are collected; Feature aggregation: perform weighted aggregation on the features of neighbor nodes. Specifically, for the graph convolutional network (GCN), aggregation is achieved by multiplying the standardized adjacency matrix by the feature matrix. For the graph attention network (GAT), the attention weights between the neighbor nodes and the target node are calculated, and the features are weighted averaged. Non-linear transformation: Apply activation function to the aggregated features; Feature update: merge the aggregated features with the original features of the target node to update the feature representation of the target node; Multi-layer processing: By stacking multiple layers of graph convolutional networks (GCN) or graph attention networks (GAT), the information of neighbors at a longer distance is gradually integrated.
10. A fishing vessel behavior knowledge construction system based on multi-source data, characterized in that: The system includes a data aggregation module, a knowledge graph construction module and a model construction module; The data aggregation module is used for aggregating multi-source data of fishing vessel behavior: using various sensors and standardized API interfaces to collect fishing vessel behavior data, and aggregating and processing the collected fishing vessel behavior data; The knowledge graph construction module is used for fishing boat behavior knowledge construction: entities are used as nodes, and the relationships between entities are used as edges. A knowledge graph is constructed through entity recognition and relationship recognition, and the edges in the knowledge graph are weighted according to the importance of the relationships. The model building module is used for fishing vessel behavior model construction: using the knowledge graph to perform model training, model prediction and model optimization on the graph neural network, so as to obtain the final fishing vessel behavior knowledge model.
Citation Information
Cited By
Flow dynamic processing method and system based on intelligent logistics scene multi-module cooperation
CN120494470A
Dynamic process processing method and system based on multi-module collaboration in smart logistics scenarios
CN120494470B