Ship shoreline illegal invasion and occupation identification method and system based on AIS big data
By combining AIS data and geographic information data, using multi-level clustering algorithms to identify ship mooring points, it solves the problem of difficult to identify ships illegally occupying shorelines in the existing technology, and achieves efficient and accurate supervision and resource allocation.
Patent Information
- Application Number
- CN202510591785.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
It is difficult for the existing technology to effectively identify and monitor whether ships illegally encroach on shorelines. Traditional regulatory methods have limitations in coverage, efficiency and cost.
By obtaining the ship's AIS data and geographic information data, preprocessing and geometric characterization construction are carried out, and the ship's mooring points are clustered in coarse and fine areas are identified by combining the multi-level clustering algorithm to identify the illegally occupied area of the ship's shoreline.
It has achieved accurate identification of illegally occupied areas of ship shorelines, improved the efficiency of port supervision, reduced manual intervention, optimized the rational allocation of shoreline resources, and ensured shipping safety.
Smart Images

Figure CN120104714A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of marine transportation technology, and in particular to a method and system for identifying illegal shoreline occupation by ships based on AIS big data. Background Art
[0002] In recent years, the management of global port resources has faced increasingly severe challenges, and the phenomenon of illegal occupation of shorelines has gradually increased in some areas. This kind of occupation not only leads to the disorderly use of port shoreline resources, but may also affect port planning, shipping safety and water environment management. Traditional supervision methods mainly rely on manual inspections and fixed monitoring equipment, but this method has certain limitations in coverage, efficiency and cost. Manual inspections are difficult to monitor the entire port area in real time and comprehensively, while fixed monitoring equipment has high installation and maintenance costs and can usually only cover limited key areas. Therefore, how to use big data analysis and intelligent algorithms to improve the efficiency of port supervision has become an important research direction in the current industry.
[0003] AIS (Automatic Identification System) is an important data source for real-time monitoring of global ships, providing rich information for ship behavior analysis. AIS data records dynamic data such as ship location information, speed, heading, berthing time, and static information such as captain and maritime mobile identification code MMSI, providing a large amount of information for ship data analysis. The document titled "A New Classification Method for Ship TrajectoriesBased on AIS Data" proposes a technical solution for ship trajectory classification based on AIS data. It proposes a new integrated classifier based on common machine learning classification algorithms, which is better than a single classifier. The trajectory classification results include five types, namely normal navigation trajectory, anchored or moored trajectory, trajectory with offset, trajectory with AIS signal loss, and irregular trajectory. Although the document has a good classification effect on navigation trajectory, it does not classify whether the ship's trajectory illegally occupies the coastline, and there is little identification of illegal occupation of the coastline in the existing technology. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for identifying illegal shoreline occupation by ships based on AIS big data, which can accurately identify illegally occupied shoreline areas.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] A method for identifying illegal occupation of a ship's shoreline based on AIS big data comprises the following steps: Obtain the ship's AIS data and perform preprocessing; Acquire geographic information data and construct geometric representation of geographic information; The pre-processed AIS data and the geometric representation of the geographic information are superimposed and matched to obtain the area where the ship behavior occurs, and combined with the set threshold, the abnormal ship behavior recognition result is obtained; Based on the abnormal behavior identification results of the ships, a multi-level clustering algorithm is used to roughly cluster the ship berthing points, and then each rough clustering group is finely clustered, and statistical analysis is performed on the pre-processed AIS data to obtain the identification results of illegal occupation of the ship shoreline; The AIS data of the ship includes static information data and dynamic information data of the ship, the static information data includes the Maritime Mobile Identity (MMSI), captain information, ship length, ship width and ship draft information, and the dynamic information data includes longitude, latitude and speed.
[0007] Furthermore, the pre-processing step includes: Decoding the AIS data of the ship, wherein the AIS data of the ship is sent and received in the form of ASCII code, including an identifier, a total number of sentences, a sentence sequence number, an identification code, a channel, an encapsulated message, a number of padding bits, and a check mark; Based on the decoded AIS data, outlier removal and missing fill processing are performed in sequence to complete the preprocessing process.
[0008] Furthermore, the steps of removing outliers and filling in missing values include: The decoded AIS data is traversed and checked, and the data points whose speed, longitude and latitude are not within the set range are marked as abnormal and removed to obtain the AIS data sequence to be filled, wherein the calculation expression of the speed is: ; ; In the formula, is the speed, For point and The distance between the latitude λ and longitude The distance between coordinates, is the radius of the Earth; Determine the time interval between adjacent data points under normal circumstances, traverse the AIS data sequence to be filled, and obtain the position of the missing value; According to the position of the missing value, the missing value of the AIS data sequence to be filled is filled by using the cubic spline interpolation method, and the missing value filling is performed once every set time until the traversal of the AIS data sequence to be filled is completed.
[0009] Furthermore, the geographic information data includes geographic information of VTS reporting lines, anchorages, waterways, port areas, and shorelines, and the step of constructing a geometric representation of the geographic information includes: 1) For anchorages, waterways and port areas: Based on the geographical information of waterways, ports and anchorages, the shape feature points are used as connecting control points to construct polygons, and the geometric features of waterways, ports and anchorages are obtained and defined as: Anchorage: Anchorage { Anchorage_name , Point_List}; channel: Fairway { Fairway_name , Point_List}; Port Area: Harbour { Harbour_name , Point_List}; Among them, for the navigation area waters in the port area that are not divided into functional areas, the shape feature points of the navigation area waters are connected to the control points to construct a polygon, and the geometric features of the navigation area outside the functional area are obtained and defined as: Navigation area outside the functional area: Navigable_Area { Navigable_Area_name , Point_List}; In the formula, Anchorage_name is the name of the anchorage. Fairway_name is the name of the waterway, Harbour_name is the name of the port area. Point_List is the polygon latitude and longitude control point matrix, Navigable_Area_name The name of the waters in the navigation area outside the functional area; 2) For the shoreline: Based on the geographic information of the coastline, the shape feature points of the coastline are used as connection control points, and the geometric line segment description method is used to express the irregular curve characteristics. The geometric characteristics of the coastline are obtained and defined as: Shoreline: Costline { Costline_name,Point_List}; In the formula, Costline_name is the name of the shoreline; 3) For VTS reporting lines: Based on the VTS reporting line geographic information, the geometric features of the circular or rectangular VTS reporting line are formed and defined as: Circular VTS reporting line: VTS {r Circle ,( x cen ,y cen )}; Rectangular VTS reporting lines: VTS { Point_List}; In the formula, r Circle The radius of the circular VTS reporting line, is the center control point; The closed area formed by the geometric features of the VTS reporting line and the geometric features of the coastline determines the scope of the port waters.
[0010] Furthermore, the step of obtaining the area where the ship behavior occurs includes: Based on the pre-processed AIS data and the geometric representation of the geographic information, a vector cross product method is used to perform superposition matching to determine the area where the ship behavior occurs, specifically including: 1) Geometric representation of geographic information for each polygon shape: Set the latitude and longitude of the ship's track point to , and the polygon latitude and longitude control point matrix is expressed as ; calculate ,judge Is it established? If yes, it means that the ship behavior has occurred. If no, it means that the ship behavior has not occurred, where: ; In the formula, d i For the ship's i Time to i Displacement length at moment +1; 2) Geometric representation of circular geographic information: Set the latitude and longitude of the ship's track point to , the center of the circle is ; Calculate the latitude and longitude of the trajectory point With the center The distance and the radius of the circular area r Circle Compare and judge Is it established? If so, it means that the ship behavior has occurred. If not, it means that the ship behavior has not occurred.
[0011] Furthermore, the step of obtaining the abnormal behavior identification result of the ship includes: Set speed thresholds, including speed thresholds for ships at anchor , Port area ship speed threshold , Channel ship speed threshold , Ship speed threshold in navigation area ; Set speed duration thresholds, including time thresholds for anchoring at anchorage , Time threshold for berthing in the port area , Time threshold for channel anchorage 2. Time threshold for ships to stay in the navigation area ; According to the ship behavior occurrence area and the ship navigation state, combined with the speed threshold and the speed duration threshold, the ship behavior is obtained, wherein the ship behavior includes the following: ① Anchorage navigation: The area where the ship's behavior occurs is an anchorage and meets the following conditions: , is the ship's speed; ② Anchorage: The area where the vessel behavior occurs is the anchorage and meets the following conditions: , T is the current stationary time of the ship; ③ Port navigation: The vessel behavior occurs in the port area and meets the following conditions: ; ④ Port berthing: The vessel behavior occurs in the port area and meets the following conditions: ; ⑤ Navigation in a waterway: The area where the ship's behavior occurs is a waterway and meets the following conditions: ; ⑥ Channel anchorage: The vessel behavior occurs in the channel and meets the following conditions: ; ⑦ Navigation in Navigation Area: The area where the ship behavior occurs is the navigation area and meets the following conditions: ; ⑧ Navigation area anchorage: The area where the vessel behavior occurs is a navigation area and meets the following conditions: ; According to the ship behavior, the abnormal behavior of the ship is determined to obtain a ship abnormal behavior identification result, wherein the abnormal behavior of the ship includes berthing behavior in non-navigation areas, non-port areas, and non-anchorage areas.
[0012] Furthermore, the K-Means++ algorithm is used to perform rough clustering of the parking points. The specific steps include: Cluster center initialization: Randomly select a data point from the data set as the first cluster center , wherein the data set is constructed based on the abnormal behavior recognition result of the ship, and a cluster represents a group; Minimum distance calculation: For the data set X Each data point in x, calculate the minimum distance to all current cluster centers, where the calculation expression of the minimum distance is: ; In the formula, is the minimum distance, is the number of cluster centers currently selected, For the i Cluster centers; Cluster center selection: As a probability distribution, randomly select the next cluster center from the remaining data points , where the expression is selected as: ; In the formula, Data Points x The probability of being selected as the next cluster center, For other points in the data set, it contains all point information; Repeat selection: Repeat the minimum distance calculation and cluster center selection steps until the k Cluster centers; Data point assignment: Calculate each data point to k The Euclidean distance of the cluster centers is used to classify each data point into the nearest cluster center, which is expressed as: ; In the formula, for x The final clustering of points, Cluster j the center of Cluster center update: By calculating the mean of the data points in the cluster, the new center of each cluster is obtained, where the calculation expression is: ; In the formula, Belong to the cluster j The set of all data points of ; Iterative optimization: Repeat the data point allocation and cluster center update until the iteration end condition is reached to obtain a coarse clustering group of multiple anchor points.
[0013] Furthermore, the DBSCAN clustering algorithm is used to perform detailed clustering on each coarse clustering group, and the DBSCAN parameters of each coarse clustering group are selected using the existing legal library capacity. The specific steps include: Parameter grid setting: Set the parameters of the DBSCAN clustering algorithm and further set the parameter grid, where the parameters include the neighborhood radius and the minimum number of points to form a cluster minPts The minimum neighborhood radius is The maximum value is , the step length is , the minimum value of the minimum number of points is The maximum value is , the step length is , the parameter grid is: ; In the formula, is the parameter grid, j For the j The smallest point, and Respectively represent the upper limit of the number of search steps in the two dimensions of neighborhood radius and minimum number of points, satisfying: ; Parameter selection: Traverse the parameter grid and select the neighborhood radius of each coarse cluster group and minimum points , for detailed clustering; Core point identification: For each stop point in the data set, determine its neighborhood radius Are all the points in the , if yes, then the anchor point is marked as a core point, if no, then it indicates that the anchor point is a boundary point or a noise point, wherein the data set is constructed based on each coarse clustering group; Cluster expansion: For each core point, take it as the starting point and expand its neighborhood radius All points within are grouped into the same cluster, and the neighborhood radius is recursively expanded. New core points within the area until no further expansion is possible; Boundary point and noise point processing: If a point is not within the neighborhood radius of any core point If a point is in the neighborhood of a core point but is not a core point, it is marked as a boundary point. Clustering is completed: all points in the data set are traversed, and one or more parking area clusters are finally formed, and noise points are distinguished; Cluster evaluation: based on each existing legal storage capacity center point and the cluster centers of each detailed cluster , calculate the evaluation index value and obtain the cluster evaluation result, wherein the calculation expression of the evaluation index value is: ; ; In the formula, For the l The evaluation results of the coarse clustering groups are n Indicates lThe number of known legal storage capacity centers within the coarse clustering range, For the l The first j Legal storage capacity score, For distance The straight-line distance to the nearest fine-scale cluster center, A.B is the set threshold.
[0014] Furthermore, the step of obtaining the identification result of illegal occupation of the ship's shoreline includes: For each coarse clustering group containing detailed clustering results, calculate the distance between detailed clustering groups in different coarse clustering groups, and determine whether the distance is less than the preset value. If so, merge the detailed clustering groups, otherwise, do not merge them. The distance calculation expression between coarse clustering groups is: ; In the formula, Rough clustering group i and j The distance between is the distance between the centers of the coarse cluster groups, and It represents the distance between the farthest points in the coarse clustering group; According to the merging results, the cluster center of the multi-level clustering algorithm is obtained, and the cluster center is matched with the existing legal storage capacity. If the straight-line distance between the cluster center and the existing legal storage capacity is within a certain threshold, it is determined to be a legal point, otherwise it is determined to be a legal point, and the illegal occupation area is obtained. The berthing volume of ships in the illegal occupation area at different time periods is calculated to complete the identification process, where the calculation expression of the berthing volume is: ; In the formula, is the number of ships moored in a certain period of time, for a certain moment i The number of ships at anchor, n Indicates time.
[0015] The present invention also provides a ship shoreline illegal occupation identification system based on AIS big data, comprising: Data acquisition and preprocessing module: used to acquire the ship's AIS data and perform preprocessing; Geometric representation construction module: used to obtain geographic information data and construct geometric representation of geographic information; Abnormal behavior identification module: used to overlay and match the pre-processed AIS data with the geometric representation of the geographic information, obtain the area where the ship behavior occurs, and combine the set threshold to obtain the abnormal behavior identification result of the ship; Illegal occupation identification module: based on the abnormal behavior identification result of the ship, a multi-level clustering algorithm is used to perform rough clustering on the ship berthing points, and then fine clustering is performed on each rough clustering group, and statistical analysis is performed in combination with the pre-processed AIS data to obtain the identification result of illegal occupation of the ship shoreline; The AIS data of the ship includes static information data and dynamic information data of the ship, the static information data includes the Maritime Mobile Identity (MMSI), captain information, ship length, ship width and ship draft information, and the dynamic information data includes longitude, latitude and speed.
[0016] Compared with the prior art, the present invention has the following beneficial effects:
[0017] (1) The present invention combines AIS data and geographic information data for in-depth analysis to identify the berthing behavior of ships, and analyzes the berthing points in combination with a multi-level clustering algorithm. The multi-level clustering algorithm can well identify high-density berthing points by combining coarse clustering and fine clustering, so as to accurately identify illegal occupation of shoreline areas.
[0018] (2) The present invention is based on the traditional DBSCAN clustering algorithm because of its good adaptability and can find clusters of any shape. In particular, for dense data sets such as ship trajectories, it can obtain good clustering effects. On this basis, improvements are made and a multi-level clustering algorithm combining the K-Means++ algorithm and the DBSCAN algorithm is proposed. This algorithm can not only greatly reduce the time complexity of the algorithm, but also has a good clustering effect. In addition, by double traversing coarse clusters with close distances and checking whether there are fine clusters that should be merged between the two coarse clusters, it can avoid points that should be in the same fine cluster being separated by the coarse clusters, thereby further improving the accuracy of identifying the location of ship shoreline encroachment.
[0019] Since the number of berthing areas is initially unknown, the clustering method cannot use the number of clusters as its initial parameter; however, some berthing points are temporary and cannot constitute a general berthing area like other berthing points. The multi-level clustering algorithm used in the present invention can obtain cluster points without the initial number of clusters and can also identify abnormal points.
[0020] The present invention combines big data processing, geographic information system (GIS) analysis and machine learning algorithms to provide port management departments with efficient and automated supervision tools, enabling regulatory departments to quickly lock suspicious areas based on data analysis, improve inspection efficiency, and reduce the workload of manual intervention. In addition, the present invention can also improve the rational allocation of shoreline resources and the ability to identify illegal docks, provide decision support for port planning, help optimize the rational allocation of shoreline resources, and avoid illegal docks from interfering with the operation order of regular ports, thereby optimizing port planning, ensuring water order, and improving shipping safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of the method flow of the present invention; Figure 2 AIS data analysis schematic diagram of the present invention; Figure 3 This is the ship trajectory clustering processing technical route of the present invention; Figure 4 Schematic diagram of the ship trajectory clustering algorithm of the present invention. DETAILED DESCRIPTION
[0022] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0023] Example 1
[0024] This example provides a method for identifying illegal shoreline occupation by ships based on AIS big data. Figure 1 As shown, the following steps are included: Step 1: Parse the given AIS file and obtain the ship track related data.
[0025] Step 1 is to obtain the ship track related data, specifically: The data collected by the AIS system is AIS data, which contains a wealth of static information (maritime mobile identification code MMSI, ship length, ship width, ship draft, etc.) and dynamic information (such as longitude, latitude, speed, heading, etc.). Under normal circumstances, ships send dynamic data every 30 seconds and static data every 5 minutes, so AIS data has a strong timeliness. In order to improve the transmission speed and efficiency of data, the original AIS information is transmitted and received in ASCII code, and the receiving end needs to decode the AIS data before it can read the specific information. A complete AIS message usually includes an identifier, a total number of sentences, a sentence sequence number, an identification code, a channel, an encapsulated telegram, a number of padding bits, and a checksum.
[0026] Decode AIS data. Figure 2 As shown in the figure, the decoding of AIS data mainly refers to the decoding of the encapsulated message. The encapsulated message is generally located in the sixth field of the original AIS data. Due to the different contents sent, the types of encapsulated messages are also different. The types of encapsulated messages are different, and their formats are also different. Therefore, in different types of AIS data, the fields corresponding to the same information are different.
[0027] Step 2: For the parsed AIS data, outliers are eliminated and missing values are filled.
[0028] The outlier removal and missing value filling preprocessing described in step 2 are specifically as follows: During the transmission of AIS data, data loss and data anomalies may occur due to human input errors, equipment failures, etc. In order to reduce the impact of abnormal AIS data on the experimental results, the present invention firstly performs a traversal check on the decoded AIS data, marks the data points with a speed greater than 50kn, a longitude not between [-180°, 180°], and a latitude not between [-90°, 90°] as abnormal, and determines the position of the missing value. The speed calculation formula is: ; in Indicate point and Use the Haversine formula to calculate the latitude ( λ ) and longitude ( ) distance between coordinates: ; in, is the radius of the Earth.
[0029] Then, cubic spline interpolation is used for interpolation replacement. The specific interpolation process is as follows: 1: Assume that the AIS data sequence after removing outliers is , for and The data sequence is traversed, and according to the time interval between two adjacent data points, two data points with a time interval exceeding 30s are found, and the position of the missing value is determined according to the position of the two points.
[0030] 2: Select the cubic spline interpolation algorithm to fill the trajectory data at the positions that need to be filled, and interpolate every 15 seconds until all data traversal is completed.
[0031] Step 3: Based on the geographic information system, establish the geometric representation of geographic information.
[0032] Step 3 describes establishing a geometric representation of geographic information, specifically: Geographic information includes the Vessel Traffic Management Center (VTS) reporting line, anchorage, waterway, port area, coastline, etc. The mathematical and geometric representation of the geographical information of the port area, anchorage and waterway is achieved by connecting control points (shape feature points) to construct polygons. The coastline expresses its irregular curve characteristics by connecting control points (shape feature points) and adopting the geometric line segment description method. The closed area formed by the coastline and the VTS reporting line can determine the scope of the port waters. In addition, for the navigable waters in the port area that are not divided into functional areas, the mathematical and geometric representation of the geographical information is achieved by connecting control points (shape feature points) to construct polygons.
[0033] By defining a class containing name, point set, radius or boundary point, etc., the names and ranges of different anchorages, channels, ports, etc. are defined, as follows: Anchorage: Anchorage { Anchorage_name , Point_List}; channel: Fairway { Fairway_name , Point_List}; Port Area: Harbour { Harbour_name , Point_List}; Shoreline: Costline { Costline_name,Point_List}; VTS reporting line (circle): VTS { r Circle ,( x cen ,y cen )}; VTS reporting line (rectangle): VTS { Point_List}; Navigation area outside the functional area: Navigable_Area { Navigable_Area_name , Point_List}; Where: Anchorage_name is the name of the anchorage; Fairway_name is the name of the waterway; Harbour_name is the name of the port area; Navigable_Area_name It is the name of the navigable waters outside the functional area; r Circle Reports line radius for circular VTS;( x cen ,y cen ) is the center control point; Point_List It is the longitude and latitude matrix of shape feature control points.
[0034] Step 4: Establish a ship behavior classification model.
[0035] The present invention adopts the method of superimposing and matching the longitude and latitude information of the ship with the geographical information of the shoreline to identify the position of the ship, judges the navigation status of the ship according to the speed of the ship, and finally defines the ship behavior by combining the navigation status of the ship and the position of the ship at that time. The specific process is as follows: Figure 3 The specific identification method is as follows: Step 4.1: According to the matching results of the ship's longitude and latitude and the shoreline geographic information, determine the area where the ship's behavior occurs. For polygonal geographic areas, the vector cross product method is used to determine whether the ship's behavior occurs in this area. The calculation formula is shown below. Assume that the longitude and latitude of the track point are , the polygon latitude and longitude control point matrix is ,like , then the ship behavior occurs in this area.
[0036] ; In the formula, d i For the ship's i Time to i The displacement length at the +1 moment.
[0037] For a circular area, calculate the trajectory points With the center The distance is calculated and compared with the radius of the circular area. If the following formula is satisfied, the ship behavior occurs in this area.
[0038] ; Step 4.2: Set the speed threshold. Speed threshold for ships at anchor 0.5, the threshold of ship speed in the port area 0.2kn, the speed threshold of the channel ship 0kn, the speed threshold of ships in the navigation area It is 0.5kn.
[0039] Step 4.3: Set the speed duration threshold. Since the ship's stop behavior depends on the continuous low speed, it is necessary to determine the duration of the ship's low speed. According to navigation experience, experimental statistics and literature research, the time threshold for anchoring is set. 0.5h; the time threshold for berthing in the port area is 1h; since ships are strictly prohibited from berthing in the waterway, the time threshold 0h; the time threshold for ships to stay in the navigation area 0.5h.
[0040] Step 4.4: Based on the ship's navigation status and the area where the ship's behavior occurs, combined with the speed threshold and speed duration threshold, define the ship's behavior as shown in the following table ( is the ship speed, T is the current stationary time of the ship).
[0041] Table 1 Ship behavior
[0042] Step 5: Classification of illegal shoreline occupation.
[0043] In this embodiment, the main research is on the problem of ships encroaching on the shoreline in the reservoir area, that is, illegally building berthing areas in the waters at the edge of the reservoir area. Therefore, this embodiment defines the berthing behavior of ships in non-navigation areas, non-ports, and non-anchorage areas as abnormal behavior, and focuses on the identification of such abnormal behavior.
[0044] Firstly, the historical AIS data of ships in the navigation area are screened, the historical tracks are drawn, the track timestamps are recorded, and the track points with anchoring characteristics such as long-term low speed and no obvious changes in track points for a long time are captured to realize the identification of anchoring behavior in the navigation area.
[0045] Step 6: Identify the location information of the reservoir area encroachment.
[0046] According to the results of ship abnormal behavior identification, data on ships anchoring in abnormal areas can be obtained. However, a small number of ships anchoring in a certain range may be caused by other reasons. Only when a large number of ships anchor here can this range be defined as an illegally built anchorage area. Therefore, a clustering algorithm is used to identify the location of ship shoreline encroachment based on the results of ship abnormal behavior identification.
[0047] The traditional DBSCAN clustering algorithm has good adaptability and can find clusters of any shape, especially for dense data sets such as ship trajectories, it can achieve good clustering results. However, for scenes such as large waters with large data volumes, directly using the DBSCAN clustering algorithm will result in a long clustering time. Therefore, this paper improves the clustering algorithm and proposes a multi-level clustering algorithm: first, the K-Means++ clustering algorithm is used to perform coarse clustering of the mooring points, and DBSCAN clustering is performed in each cluster, thereby greatly reducing the time complexity of the algorithm.
[0048] Step 6.1: First, use the K-Means++ clustering algorithm for rough classification. The K-Means++ clustering algorithm is an improved version of the K-Means clustering algorithm. It mainly optimizes the selection method of the initial cluster center to reduce the influence of the local optimal solution and speed up the convergence speed. A probability-based initialization strategy is adopted to make the initial cluster center as dispersed as possible, thereby improving the stability and effect of clustering. The algorithm flow is as follows: First, initialize the cluster center. Randomly select X Select a point as the first cluster center For each point in the data set , and calculate its minimum distance to the centers of all currently selected clusters: ; in Indicates the number of cluster centers currently selected. As a probability distribution, randomly select the next cluster center from the remaining data points ,Right now: ; In the formula, For data points x The probability of being selected as the next cluster center, for other points in the data set.
[0049] This process is repeated until the selected k Cluster center.
[0050] Next is the data point assignment. The Euclidean distance of each data point to all cluster centers is calculated and it is assigned to the nearest cluster center: ; In the formula, for x The final clustering of points, Cluster j center.
[0051] Cluster center update. Calculate the new center of each cluster, which is the mean of all data points belonging to the cluster: ; in, for x The final clustering of points, Belong to the cluster j The set of all data points.
[0052] Finally, iterative optimization is performed. The data point assignment and cluster center update steps are performed repeatedly until the cluster center no longer changes significantly or the maximum number of iterations is reached.
[0053] Step 6.2: Use the DBSCAN clustering algorithm to perform detailed clustering on the points in each coarse cluster, and use the existing legal storage capacity (terminal) to select the DBSCAN parameters for each coarse cluster group, with the goal that the clustering results can be well matched with the existing legal storage capacity points. Figure 4 As shown, the specific steps include: Parameter grid settings. DBSCAN parameters include the neighborhood radius ( Eps ) and the minimum number of points minPts ,in specifies how close points should be to each other to be considered part of a cluster. minPts Indicates the minimum number of points required to form a cluster to ensure that each candidate anchorage area has enough AIS anchorage points. minPts Fixed, calculate each point and its minPts The distance between nearest neighbors can automatically determine the optimal eps (Single-level density method). And set the minimum neighborhood radius to , the maximum value is , the step length is The minimum number of points is , the maximum value is , the step length is , then the parameter grid is: ; In the formula, is the parameter grid, j For the j The smallest point, and Respectively represent the upper limit of the number of search steps in the two dimensions of neighborhood radius and minimum number of points, satisfying: ; Parameter selection: traverse the parameter grid and select the neighborhood radius and minimum points , and based on these two parameters, the coarse clustering results are clustered in detail. The detailed clustering process includes the following core point identification, cluster diffusion, boundary point and noise point processing, and clustering completion.
[0054] Core point identification. For each anchor point in the data set, if The number of points inside is greater than or equal to , then the point is marked as a “core point”; otherwise, the point may be a “boundary point” or a “noise point”.
[0055] Cluster expansion. For each core point, start from this point and expand its neighborhood All points in the same cluster are grouped together, and the neighborhood is recursively expanded. New core points are added until further expansion is impossible.
[0056] Boundary point and noise point processing. If a point does not belong to the neighborhood of any core point , then it is marked as a “noise point”, indicating that it may be a random stop or an abnormal data point; if a point is in the neighborhood of a core point but is not a core point itself, it is marked as a “boundary point”.
[0057] Clustering is complete. The algorithm traverses all data points and eventually forms one or more parking area clusters and distinguishes noise points.
[0058] Cluster evaluation: After clustering is completed, the cluster is evaluated. This evaluation is based on the analysis of each existing legal storage capacity center point and each detailed cluster center. The evaluation indicators are designed as follows: ; ; In the formula, For the l The evaluation results of the coarse clustering groups are n Indicates l The number of known legal storage capacity centers within the coarse clustering range, For the l The first j Legal storage capacity score, For distance The straight-line distance to the nearest detailed cluster center, 50 、 100 is the set threshold.
[0059] Step 6.3: Merge clusters at the edge of coarse clustering. First, check the distance between coarse clustering groups that contain multiple fine clustering results. If the distance between groups is too far, it is impossible to merge. Excluding these groups from the calculation process can effectively ensure the reduction of computational complexity. The distance between groups is calculated by the following formula: ; In the formula, represents the distance between the group centers obtained by coarse clustering using K-Means++, and It represents the distance between the farthest points in the group. , it means that a merge is needed.
[0060] The cluster center of the multi-level clustering algorithm is obtained based on the merged results, and the obtained cluster center is matched with the legal storage capacity, and the remaining part is the illegally occupied area. Specifically, if the straight-line distance between the cluster center and the existing legal storage capacity is within a certain threshold, it is determined to be a legal point, otherwise it is an illegal point.
[0061] In summary, this embodiment combines AIS data and clustering results to extract the track point data of ships' illegal docking, count the time of ships' docking that occupy the shoreline, analyze the distribution of ships' stays, grasp the hot spots of ship docking, and identify hot spots of illegal areas. The identified hot spots of illegal areas should cover all the berthing points in this area. Some ships dock at temporary locations for specific reasons, so after filtering these out, it is necessary to find general berthing area points. Since the number of berthing areas is initially unknown, the clustering method cannot use the number of clusters as its initial parameter. However, some berthing points are temporary and cannot constitute a general berthing area like other berthing points. Therefore, a multi-level clustering algorithm is used to obtain cluster points without the initial number of clusters, but outliers can also be identified. Finally, each berthing area is uniquely identified, including the center point and radius of the area, that is, the distance from the farthest point to the center point.
[0062] The method to grasp the hot spots of ship berthing is to divide a day into different periods, count the number of ships berthing in different periods according to the berthing areas identified above, and finally perform visual analysis. The statistical model of berthing volume is: ; Where: Indicates the number of ships moored in a certain period of time. Indicates a moment i The number of ships at anchor, n Indicates time.
[0063] In summary, this embodiment can accurately identify the berthing behavior of ships and further analyze the berthing mode of ships through in-depth mining of AIS data. The berthing behavior analysis based on AIS big data can not only be used to monitor the use of legal docks, but also help discover abnormal berthing point clusters, and then identify possible illegal occupation of shoreline areas. In order to achieve this goal, this embodiment uses an improved multi-level clustering algorithm to analyze the berthing points of ships. By identifying high-density clustered berthing points, it is possible to effectively distinguish between normal operating docks and areas where violations may exist, and to find irregularly shaped clustered areas, which are suitable for data characteristics such as uneven spatial distribution of ship berthing points. The implementation of this embodiment will provide an efficient and automated technical means for port supervision, enabling regulatory authorities to quickly lock suspicious areas based on data analysis, improve patrol efficiency, and reduce the workload of manual intervention. In addition, the method can also provide decision support for port planning, help optimize the reasonable allocation of shoreline resources, and avoid interference of illegal docks with the operation order of regular ports.
[0064] Example 2
[0065] This embodiment provides a system for identifying illegal shoreline occupation by ships based on AIS big data, including: Data acquisition and preprocessing module: used to acquire the ship's AIS data and perform preprocessing; Geometric representation construction module: used to obtain geographic information data and construct geometric representation of geographic information; Abnormal behavior identification module: used to overlay and match the pre-processed AIS data with the geometric representation of the geographic information, obtain the area where the ship behavior occurs, and combine the set threshold to obtain the abnormal behavior identification result of the ship; Illegal occupation identification module: based on the abnormal behavior identification results of the ship, a multi-level clustering algorithm is used to perform coarse clustering of the ship berthing points, and then each coarse clustering group is finely clustered, and statistical analysis is performed on the pre-processed AIS data to obtain the identification results of illegal occupation of the ship shoreline.
[0066] The rest is the same as in Example 1.
[0067] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.
[0068] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The solutions in the embodiments of the present invention may be implemented in various computer languages, for example, object-oriented programming language Java and interpreted scripting language JavaScript, etc.
[0069] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0070] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0072] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0073] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for identifying illegal occupation of shoreline by ships based on AIS big data, characterized in that: The following steps are involved: Obtain the ship's AIS data and perform preprocessing; Acquire geographic information data and construct geometric representation of geographic information; The pre-processed AIS data and the geometric representation of the geographic information are superimposed and matched to obtain the area where the ship behavior occurs, and combined with the set threshold, the abnormal ship behavior recognition result is obtained; Based on the abnormal behavior identification results of the ships, a multi-level clustering algorithm is used to roughly cluster the ship berthing points, and then each rough clustering group is finely clustered, and statistical analysis is performed on the pre-processed AIS data to obtain the identification results of illegal occupation of the ship shoreline; The AIS data of the ship includes static information data and dynamic information data of the ship, the static information data includes the Maritime Mobile Identity (MMSI), captain information, ship length, ship width and ship draft information, and the dynamic information data includes longitude, latitude and speed.
2. According to claim 1, a method for identifying illegal occupation of shoreline by ships based on AIS big data is characterized in that: The pre-processing steps include: Decoding the AIS data of the ship, wherein the AIS data of the ship is sent and received in the form of ASCII code, including an identifier, a total number of sentences, a sentence sequence number, an identification code, a channel, an encapsulated message, a number of padding bits, and a check mark; Based on the decoded AIS data, outlier removal and missing fill processing are performed in sequence to complete the preprocessing process.
3. According to claim 2, a method for identifying illegal occupation of shoreline by ships based on AIS big data is characterized in that: The steps of removing outliers and filling missing values include: The decoded AIS data is traversed and checked, and the data points whose speed, longitude and latitude are not within the set range are marked as abnormal and removed to obtain the AIS data sequence to be filled, wherein the calculation expression of the speed is: ; ; In the formula, is the speed, For point and The distance between the latitude λ and longitude The distance between coordinates, is the radius of the Earth; Determine the time interval between adjacent data points under normal circumstances, traverse the AIS data sequence to be filled, and obtain the position of the missing value; According to the position of the missing value, the missing value of the AIS data sequence to be filled is filled by using the cubic spline interpolation method, and the missing value filling is performed once every set time until the traversal of the AIS data sequence to be filled is completed.
4. According to the method for identifying illegal occupation of shoreline by ships based on AIS big data in claim 1, it is characterized in that: The geographic information data includes VTS reporting lines, anchorages, waterways, port areas, and shoreline geographic information. The step of constructing a geometric representation of the geographic information includes: 1) For anchorages, waterways and port areas: Based on the geographical information of waterways, ports and anchorages, the shape feature points are used as connecting control points to construct polygons, and the geometric features of waterways, ports and anchorages are obtained and defined as: Anchorage: Anchorage { Anchorage_name , Point_List }; channel: Fairway { Fairway_name , Point_List }; Port Area: Harbour { Harbour_name , Point_List }; Among them, for the navigation area waters in the port area that are not divided into functional areas, the shape feature points of the navigation area waters are connected to the control points to construct a polygon, and the geometric features of the navigation area outside the functional area are obtained and defined as: Navigation area outside the functional area: Navigable_Area { Navigable_Area_name , Point_List }; In the formula, Anchorage_name is the name of the anchorage. Fairway_name is the name of the waterway, Harbour_name is the name of the port area. Point_List is the polygon latitude and longitude control point matrix, Navigable_Area_name The name of the waters in the navigation area outside the functional area; 2) For the shoreline: Based on the geographic information of the coastline, the shape feature points of the coastline are used as connection control points, and the geometric line segment description method is used to express the irregular curve characteristics. The geometric characteristics of the coastline are obtained and defined as: Shoreline: Costline { Costline_name,Point_List }; In the formula, Costline_name is the name of the shoreline; 3) For VTS reporting lines: Based on the VTS reporting line geographic information, the geometric features of the circular or rectangular VTS reporting line are formed and defined as: Circular VTS reporting line: VTS { r Circle ,( x cen ,y cen )}; Rectangular VTS reporting lines: VTS { Point_List }; In the formula, r Circle The radius of the circular VTS reporting line, is the center control point; The closed area formed by the geometric features of the VTS reporting line and the geometric features of the coastline determines the scope of the port waters.
5. According to claim 4, a method for identifying illegal occupation of shoreline by ships based on AIS big data is characterized in that: The steps to obtain the area where the ship behavior occurs include: Based on the pre-processed AIS data and the geometric representation of the geographic information, a vector cross product method is used to perform superposition matching to determine the area where the ship behavior occurs, specifically including: 1) Geometric representation of geographic information for each polygon shape: Set the latitude and longitude of the ship's track point to , and the polygon latitude and longitude control point matrix is expressed as ; calculate ,judge Is it established? If yes, it means that the ship behavior has occurred. If no, it means that the ship behavior has not occurred, where: ; In the formula, d i For the ship's i Time to i Displacement length at moment +1; 2) Geometric representation of circular geographic information: Set the latitude and longitude of the ship's track point to , the center of the circle is ; Calculate the latitude and longitude of the trajectory point With the center The distance and the radius of the circular area r Circle Compare and judge Is it established? If so, it means that the ship behavior has occurred. If not, it means that the ship behavior has not occurred.
6. According to the method for identifying illegal occupation of shoreline by ships based on AIS big data in claim 1, it is characterized in that: The step of obtaining the abnormal behavior identification result of the ship comprises: Set speed thresholds, including speed thresholds for ships at anchor , Port area ship speed threshold , Channel ship speed threshold , Ship speed threshold in navigation area ; Set speed duration thresholds, including time thresholds for anchoring at anchorage , Time threshold for berthing in the port area , Time threshold for channel anchorage 2. Time threshold for ships to stay in the navigation area ; According to the ship behavior occurrence area and the ship navigation state, combined with the speed threshold and the speed duration threshold, the ship behavior is obtained, wherein the ship behavior includes the following: ① Anchorage navigation: The area where the ship's behavior occurs is an anchorage and meets the following conditions: , is the ship's speed; ② Anchorage: The area where the vessel behavior occurs is the anchorage and meets the following conditions: , T is the current stationary time of the ship; ③ Port navigation: The vessel behavior occurs in the port area and meets the following conditions: ; ④ Port berthing: The vessel behavior occurs in the port area and meets the following conditions: ; ⑤ Navigation in a waterway: The area where the ship's behavior occurs is a waterway and meets the following conditions: ; ⑥ Channel anchorage: The vessel behavior occurs in the channel and meets the following conditions: ; ⑦ Navigation in Navigation Area: The area where the ship behavior occurs is the navigation area and meets the following conditions: ; ⑧ Mooring in a navigation area: The area where the vessel behavior occurs is a navigation area and meets the following conditions: ; According to the ship behavior, the abnormal behavior of the ship is determined to obtain a ship abnormal behavior identification result, wherein the abnormal behavior of the ship includes berthing behavior in non-navigation areas, non-port areas, and non-anchorage areas.
7. According to claim 1, a method for identifying illegal occupation of shoreline by ships based on AIS big data is characterized in that: The K-Means++ algorithm is used to perform rough clustering of the parking points. The specific steps include: Cluster center initialization: Randomly select a data point from the data set as the first cluster center , wherein the data set is constructed based on the abnormal behavior recognition result of the ship, and a cluster represents a group; Minimum distance calculation: For the data set X Each data point in x , calculate the minimum distance to all current cluster centers, where the calculation expression of the minimum distance is: ; In the formula, is the minimum distance, is the number of cluster centers currently selected, For the i Cluster centers; Cluster center selection: As a probability distribution, randomly select the next cluster center from the remaining data points , where the expression is selected as: ; In the formula, Data Points x The probability of being selected as the next cluster center, For other points in the data set, it contains all point information; Repeat selection: Repeat the minimum distance calculation and cluster center selection steps until the k Cluster centers; Data point assignment: Calculate each data point to k The Euclidean distance of the cluster centers is used to classify each data point into the nearest cluster center, which is expressed as: ; In the formula, for x The final clustering of points, Cluster j the center of; Cluster center update: By calculating the mean of the data points in the cluster, the new center of each cluster is obtained, where the calculation expression is: ; In the formula, Belong to the cluster j The set of all data points of ; Iterative optimization: Repeat the data point allocation and cluster center update until the iteration end condition is reached to obtain a coarse clustering group of multiple anchor points.
8. The method for identifying illegal occupation of shoreline by ships based on AIS big data according to claim 1 is characterized in that: The DBSCAN clustering algorithm is used to perform detailed clustering on each coarse clustering group, and the DBSCAN parameters of each coarse clustering group are selected using the existing legal library capacity. The specific steps include: Parameter grid setting: Set the parameters of the DBSCAN clustering algorithm and further set the parameter grid, where the parameters include the neighborhood radius and the minimum number of points to form a cluster minPts The minimum neighborhood radius is The maximum value is , the step length is , the minimum value of the minimum number of points is The maximum value is , the step length is , the parameter grid is: ; In the formula, is the parameter grid, j For the j The smallest point, and Respectively represent the upper limit of the number of search steps in the two dimensions of neighborhood radius and minimum number of points, satisfying: ; Parameter selection: Traverse the parameter grid and select the neighborhood radius of each coarse cluster group and minimum points , for detailed clustering; Core point identification: For each stop point in the data set, determine its neighborhood radius Are all the points in the , if yes, then the anchor point is marked as a core point, if no, then it indicates that the anchor point is a boundary point or a noise point, wherein the data set is constructed based on each coarse clustering group; Cluster expansion: For each core point, take it as the starting point and expand its neighborhood radius All points within are grouped into the same cluster, and the neighborhood radius is recursively expanded. New core points within the area until further expansion is impossible; Boundary point and noise point processing: If a point is not within the neighborhood radius of any core point If a point is in the neighborhood of a core point but is not a core point, it is marked as a boundary point. Clustering is completed: all points in the data set are traversed, and one or more parking area clusters are finally formed, and noise points are distinguished; Cluster evaluation: based on each existing legal storage capacity center point and the cluster centers of each detailed cluster , calculate the evaluation index value and obtain the cluster evaluation result, wherein the calculation expression of the evaluation index value is: ; ; In the formula, For the l The evaluation results of the coarse clustering groups are n Indicates l The number of known legal storage capacity centers within the coarse clustering range, For the l The first j Legal storage capacity score, For distance The straight-line distance to the nearest fine-scale cluster center, A.B is the set threshold.
9. The method for identifying illegal occupation of shoreline by ships based on AIS big data according to claim 1 is characterized in that: The steps of obtaining the identification result of illegal occupation of the ship's shoreline include: For each coarse clustering group containing detailed clustering results, calculate the distance between detailed clustering groups in different coarse clustering groups, and determine whether the distance is less than the preset value. If so, merge the detailed clustering groups, otherwise, do not merge them. The distance calculation expression between coarse clustering groups is: ; In the formula, Rough clustering group i and j The distance between is the distance between the centers of the coarse cluster groups, and It represents the distance between the farthest points in the coarse clustering group; According to the merging results, the cluster center of the multi-level clustering algorithm is obtained, and the cluster center is matched with the existing legal storage capacity. If the straight-line distance between the cluster center and the existing legal storage capacity is within a certain threshold, it is determined to be a legal point, otherwise it is determined to be a legal point, and the illegal occupation area is obtained. The berthing volume of ships in the illegal occupation area at different time periods is calculated to complete the identification process, where the calculation expression of the berthing volume is: ; In the formula, is the number of ships moored in a certain period of time, for a certain moment i The number of ships at anchor, n Indicates time.
10. A ship shoreline illegal occupation identification system based on AIS big data, characterized in that: include: Data acquisition and preprocessing module: used to acquire the ship's AIS data and perform preprocessing; Geometric representation construction module: used to obtain geographic information data and construct geometric representation of geographic information; Abnormal behavior identification module: used to overlay and match the pre-processed AIS data with the geometric representation of the geographic information, obtain the area where the ship behavior occurs, and combine the set threshold to obtain the abnormal behavior identification result of the ship; Illegal occupation identification module: based on the abnormal behavior identification result of the ship, a multi-level clustering algorithm is used to perform rough clustering on the ship berthing points, and then fine clustering is performed on each rough clustering group, and statistical analysis is performed in combination with the pre-processed AIS data to obtain the identification result of illegal occupation of the ship shoreline; The AIS data of the ship includes static information data and dynamic information data of the ship, the static information data includes the Maritime Mobile Identity (MMSI), captain information, ship length, ship width and ship draft information, and the dynamic information data includes longitude, latitude and speed.
Citation Information
Patent Citations
Channel depth information generating method based on electronic navigation chart and water level data
CN102982494A
Ship anomaly identification method and system and readable storage medium
CN115774804A
Intelligent navigation mark monitoring method and system based on big data
CN118503744A
Clustering and anomaly analysis method for ship trajectory data
CN119357727A
Ship anchor ground abnormal behavior detection method and system
CN119884801A
Cited By
Method for determining ship berthing water area based on AIS data
CN121120948A
A method for determining ship mooring areas based on AIS data
CN121120948B
Dynamic change detection result positioning and displaying method oriented to illegal invasion and occupation behaviors
CN121190692A
A dynamic change detection result positioning display method for illegal encroachment behavior
CN121190692B
Multi-source data fusion verification-based AIS pseudo signal discrimination system and method
CN121412744A