Method and System for Identifying Illegal Occupation of Ship Shorelines Based on AIS Big Data
By combining AIS data and geographical information, a multi-level clustering algorithm is used to identify ship mooring points, solving the efficiency and cost of illegal shoreline identification in traditional regulatory methods, and achieving efficient and automated port supervision and resource optimization.
Patent Information
- Application Number
- CN202510591785.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing technology is difficult to effectively identify whether ships illegally encroach on shorelines. The traditional supervision methods have limited coverage and high costs, and the efficiency of manual patrol and fixed monitoring equipment is low.
By acquiring ship AIS data and geographic information data, combining multi-level clustering algorithms, including K-Means++ and DBSCAN algorithms, we identify coarse clustering and detailed clustering of ship mooring points, and analyze ship behavior to identify illegally encroached areas.
It improves the accuracy and efficiency of illegal encroachment of shoreline identification, reduces manual intervention, optimizes port planning and resource allocation, and ensures shipping safety.
Smart Images

Figure CN120104714B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of maritime transportation, and in particular, to a method and system for identifying illegal occupation of shorelines by ships based on AIS big data. Background Art
[0002] In recent years, the management of global port resources has faced increasingly severe challenges, and the phenomenon of illegal occupation of shorelines has gradually increased in some regions. Such occupation behavior not only leads to the disorderly use of port shoreline resources, but may also affect port planning, shipping safety, and water environment management. Traditional supervision methods mainly rely on manual inspections and fixed monitoring equipment, but this method has certain limitations in terms of coverage, efficiency, and cost. Manual inspections are difficult to monitor the entire port area in real time and comprehensively, while the installation and maintenance costs of fixed monitoring equipment are relatively high, and usually only limited key areas can be covered. Therefore, how to use big data analysis and intelligent algorithms to improve the efficiency of port supervision has become an important research direction in the current industry.
[0003] AIS (Automatic Identification System) is an important data source for real-time monitoring of global ships and provides rich information for ship behavior analysis. AIS data records dynamic data such as the position information, speed, course, and berthing time of ships, as well as static information such as the ship's captain and Maritime Mobile Service Identity (MMSI), providing a large amount of information for ship data analysis. The literature titled "A New Classification Method for Ship Trajectories Based on AIS Data" proposed a technical solution for classifying ship trajectories based on AIS data. It proposed a new integrated classifier based on common machine learning classification algorithms, and the effect is better than that of a single classifier. The trajectory classification results include five types, namely normal navigation trajectories, anchoring or mooring trajectories, trajectories with offsets, trajectories with AIS signal loss, and irregular trajectories. Although this literature has a good classification effect on navigation trajectories, it does not classify whether the ship's travel trajectory illegally occupies the shoreline, and there is also little research on the identification of illegal occupation of shorelines in the existing technology. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for identifying illegal occupation of shoreline areas by ships based on AIS big data, which can accurately identify such areas.
[0005] The purpose of the present invention can be achieved through the following technical solutions:
[0006] A method for identifying illegal occupation of ship shorelines based on AIS big data, comprising the following steps:
[0007] Obtain the AIS data of the ship and perform preprocessing;
[0008] Obtain geographic information data and construct a geometric representation of the geographic information;
[0009] Overlay and match the preprocessed AIS data with the geometric representation of the geographic information to obtain the area where the ship's behavior occurs, and combine with a set threshold to obtain the identification result of the ship's abnormal behavior;
[0010] Based on the identification result of the ship's abnormal behavior, use a multi-level clustering algorithm to roughly cluster the ship's berthing points, then perform detailed clustering on each rough clustering group, and combine with the preprocessed AIS data for statistical analysis to obtain the identification result of illegal occupation of the ship shoreline;
[0011] The AIS data of the ship includes static information data and dynamic information data of the ship. The static information data includes the Maritime Mobile Service Identity (MMSI), ship length information, ship length, ship width, and ship draft information. The dynamic information data includes longitude, latitude, and speed.
[0012] Further, the steps of the preprocessing include:
[0013] Decode the AIS data of the ship, wherein the AIS data of the ship is sent and received in the form of ASCII code, including an identifier, total number of statements, statement serial number, identification code, channel, encapsulated telegram, number of padding bits, and checksum;
[0014] Based on the decoded AIS data, perform outlier removal and missing value filling processing in sequence to complete the preprocessing process.
[0015] Further, the steps of performing outlier removal and missing value filling processing include:
[0016] Traverse and check the decoded AIS data, mark the data points where the speed, longitude, and latitude are not within the set range as abnormal and remove them to obtain the AIS data sequence to be filled, wherein the calculation expression of the speed is:
[0017]
[0018] In the formula, v is the speed, ||p j+1 -p j || is the distance between point p j and p j+1 , representing the distance between the latitude λ and the longitude coordinates, R eis the radius of the Earth;
[0019] Determine the time interval between adjacent data points under normal circumstances, traverse the AIS data sequence to be filled, and obtain the positions of missing values;
[0020] According to the positions of the missing values, use cubic spline interpolation to fill the missing values in the AIS data sequence to be filled, and perform missing value filling every set time until the traversal of the AIS data sequence to be filled is completed.
[0021] Further, the geographic information data includes VTS reporting lines, anchorages, fairways, port areas, and shoreline geographic information. The steps for constructing the geometric representation of the geographic information include:
[0022] 1) For anchorages, fairways, and port areas:
[0023] Based on the geographic information of fairways, port areas, and anchorages, construct polygons by using the shape feature points as connection control points respectively, obtain the geometric features of fairways, port areas, and anchorages respectively, and define them as:
[0024] Anchorage: Anchorage{Anchorage_name,Point_List};
[0025] Fairway: Fairway{Fairway_name,Point_List};
[0026] Harbour: Harbour{Harbour_name,Point_List};
[0027] Among them, for the navigable waters in the port area that are not divided into functional areas, connect the shape feature points of the navigable waters to form a polygon to obtain the geometric features of the navigable area outside the functional area, and define it as:
[0028] Navigable_Area outside the functional area: Navigable_Area{Navigable_Area_name,Point_List};
[0029] In the formula, Anchorage_name is the name of the anchorage, Fairway_name is the name of the fairway, Harbour_name is the name of the port area, Point_List is the matrix of longitude and latitude control points of the polygon, and Navigable_Area_name is the name of the navigable waters outside the functional area;
[0030] 2) For the shoreline:
[0031] Based on the shoreline geographic information, the shape feature points of the shoreline are used as connection control points, and the geometric segment segmentation description method is adopted to express the irregular curve features, obtaining the geometric features of the shoreline, which are defined as:
[0032] Shoreline: Costline{Costline_name,Point_List};
[0033] In the formula, Costline_name is the shoreline name;
[0034] 3) For the VTS reporting line:
[0035] Based on the VTS reporting line geographic information, the geometric features of the circular or rectangular VTS reporting line are formed and defined as:
[0036] Circular VTS reporting line: VTS{r Circle ,(x cen ,y cen )};
[0037] Rectangular VTS reporting line: VTS{Point_List};
[0038] In the formula, r Circle is the radius of the circular VTS reporting line, (x cen ,y cen ) is the center control point;
[0039] Among them, the closed area formed by the geometric features of the VTS reporting line and the geometric features of the shoreline determines the port water area range.
[0040] Furthermore, the steps to obtain the ship behavior occurrence area include:
[0041] Based on the preprocessed AIS data and the geometric representation of the geographic information, the vector cross product method is adopted for superposition matching to determine the ship behavior occurrence area, specifically including:
[0042] 1) For the geometric representation of each polygon-shaped geographic information:
[0043] Set the longitude and latitude of the ship's trajectory point as (x, y), and represent the polygon longitude and latitude control point matrix as Point_List ={(x1, y1),...,(x n ,y n )};
[0044] Calculate d1×d2×...×d n Judge whether d1×d2×...×d n >0 holds. If so, it indicates that a ship behavior has occurred; if not, it indicates that no ship behavior has occurred, where:
[0045] d i = (y i - y) × (x i+1 - x) - (y i+1 - y) × (x i - x);
[0046] Wherein, d i is the displacement length of the ship from the i-th moment to the (i + 1)-th moment;
[0047] 2) Geometric representation of geographic information in a circular shape:
[0048] Set the longitude and latitude of the trajectory point of the ship as (x, y), and the center of the circle as (x cen , y cen );
[0049] Calculate the distance between the longitude and latitude (x, y) of the trajectory point and the center of the circle (x cen , y cen ), and compare it with the radius r of the circular area to determine whether (x - x Circle cen ) 2 o_limit + (y - y cen ) 2 < r 2 Circle holds. If so, it indicates that a ship behavior has occurred; if not, it indicates that no ship behavior has occurred.
[0050] Furthermore, the steps of obtaining the ship abnormal behavior recognition result include:
[0051] Set speed thresholds, including the speed threshold S a_limit for ships at anchor, the speed threshold S h_limit for ships in the port area, the speed threshold S f_limit for ships in the waterway, and the speed threshold S o_limit for ships in the navigation area water area;
[0052] Set speed duration thresholds, including the time threshold T a_limit for anchoring at anchor, the time threshold T h_limit for berthing in the port area, the time threshold T f_limit for berthing in the waterway, and the time threshold T o_limit for ships berthing in the navigation area water area;
[0053] According to the ship behavior occurrence area and the ship navigation state, combined with the speed threshold and the speed duration threshold, obtain ship behaviors, where the ship behaviors include the following:
[0054] ① Navigation at anchor: The ship behavior occurrence area is the anchor area, and it satisfies Sv >S a_limit ,S v is the ship speed;
[0055] ② Anchorage mooring: The area where the ship's behavior occurs is the anchorage, and it satisfies S v ≤S a_limit , T≥T a_limit , where T is the current stationary time of the ship;
[0056] ③ Port area navigation: The area where the ship's behavior occurs is the port area, and it satisfies S v >S h_limit ;
[0057] ④ Port area berthing: The area where the ship's behavior occurs is the port area, and it satisfies S v ≤S h_limit , T≥T h_limit ;
[0058] ⑤ Channel navigation: The area where the ship's behavior occurs is the channel, and it satisfies S v >S f_limit ;
[0059] ⑥ Channel berthing: The area where the ship's behavior occurs is the channel, and it satisfies S v ≤S f_limit , T≥T f_limit ;
[0060] ⑦ Navigation area navigation: The area where the ship's behavior occurs is the navigation area, and it satisfies S v >S o_limit ;
[0061] ⑧ Navigation area berthing: The area where the ship's behavior occurs is the navigation area, and it satisfies S v ≤S o_limit , T≥T o_limit ;
[0062] Based on the ship's behavior, the abnormal behavior of the ship is determined, and the ship abnormal behavior recognition result is obtained, where the abnormal behavior of the ship includes berthing behaviors in non-navigation areas, non-port areas, and non-anchorage areas.
[0063] Furthermore, the K-Means++ algorithm is used to perform rough clustering on the berthing points, and the specific steps include:
[0064] Initialization of cluster centers: Randomly select a data point from the data set as the first cluster center c1, where the data set is constructed based on the ship abnormal behavior recognition result, and one cluster represents a group;
[0065] Calculation of the minimum distance: For each data point x in the data set X, calculate the minimum distance from the current all cluster centers, where the calculation expression of the minimum distance is:
[0066]
[0067] Wherein, D(x) is the minimum distance, k' is the number of cluster centers already selected, and c i is the i-th cluster center;
[0068] Cluster center selection: Using D(x) 2 as the probability distribution, randomly select the next cluster center c' from the remaining data points k+1 , where the selection expression is:
[0069]
[0070] Wherein, P(x) is the probability that the data point x is selected as the next cluster center, and x' is other points in the dataset, including all point information;
[0071] Repeated selection: Repeat the above steps of minimum distance calculation and cluster center selection until k cluster centers are selected;
[0072] Data point assignment: Calculate the Euclidean distance from each data point to the k cluster centers, and classify each data point to the nearest cluster center, expressed as:
[0073]
[0074] Wherein, Cluster(x) is the final clustering of point x, and c j is the center of cluster j;
[0075] Cluster center update: Calculate the mean value of the data points in the cluster to obtain the new center of each cluster, where the calculation expression is:
[0076]
[0077] Wherein, S j is the set of all data points belonging to cluster j;
[0078] Iterative optimization: Repeat the above steps of data point assignment and cluster center update until the iteration end condition is reached, and obtain the rough clustering groups of multiple mooring points.
[0079] Furthermore, use the DBSCAN clustering algorithm to perform fine clustering on each rough clustering group, and use the existing legal storage capacity to select the DBSCAN parameters of each rough clustering group. The specific steps include:
[0080] Parameter grid setting: Set the parameters of the DBSCAN clustering algorithm, and further set the parameter grid, where the parameters include the neighborhood radius ε and the minimum number of points minPts that make up a cluster. The minimum value of the neighborhood radius is ε min, with a maximum value of ε max , with a step size of ε step , and the minimum value of the minimum number of points is minPts min , with a maximum value of minPts max , with a step size of minPts step , the parameter grid is as follows:
[0081]
[0082] where, is the parameter grid, j is the j-th minimum point, and n ε and n minPts respectively represent the upper limits of the search steps in two dimensions of the neighborhood radius and the minimum number of points, satisfying:
[0083]
[0084] Parameter selection: Traverse the parameter grid and select the neighborhood radius ε i and the minimum number of points minPts i of each coarse clustering group for fine clustering;
[0085] Core point identification: For each mooring point in the dataset, determine whether all points within its neighborhood radius ε i are greater than or equal to the minimum number of points minPts i . If so, mark the mooring point as a core point; if not, it means the mooring point is a boundary point or a noise point, where the dataset is constructed based on each coarse clustering group;
[0086] Clustering expansion: For each core point, use it as the starting point and assign all points within its neighborhood radius ε i to the same cluster, and recursively expand the new core points within the neighborhood radius ε i until no further expansion is possible;
[0087] Boundary point and noise point processing: If a point is not within the neighborhood radius ε i of any core point, it is marked as a noise point. If a point is within the neighborhood of a core point but is not a core point, it is marked as a boundary point;
[0088] Clustering completion: Traverse all points in the dataset to finally form one or more mooring area clusters and distinguish noise points;
[0089] Clustering evaluation: Based on each existing legal storage capacity center point l j and the clustering center c k of each fine clustering, calculate the evaluation index value to obtain the clustering evaluation result, where the calculation expression of the evaluation index value is:
[0090]
[0091] Wherein, score l is the evaluation result of the first rough clustering group, n represents the number of known legal storage capacity center points in the first rough clustering range, and s j is the score of the j-th legal storage capacity in the first rough clustering range, and d j is the straight-line distance to the closest fine clustering center to s j and A and B are set thresholds.
[0092] Furthermore, the steps of obtaining the identification result of illegal occupation of the ship shoreline include:
[0093] For each rough clustering group containing the fine clustering result, calculate the distance between the fine clustering groups within different rough clustering groups, and determine whether the distance is less than a preset value. If so, merge the fine clustering groups; if not, do not merge. The calculation expression for the distance between rough clustering groups is:
[0094]
[0095] Wherein, is the distance between rough clustering groups i and j, dis_KC ij is the distance between the centers of the rough clustering groups, dis_r i and dis_r j respectively represent the distance between the farthest points in the rough clustering groups;
[0096] According to the merging result, obtain the clustering center of the multi-level clustering algorithm, and match the clustering center with the existing legal storage capacity. If the straight-line distance between the clustering center and the existing legal storage capacity is within a certain threshold, it is determined as a legal point; otherwise, it is determined as an illegal point, obtain the illegal occupation area, and calculate the berthing volume of ships in the illegal occupation area at different time periods to complete the identification process. The calculation expression for the berthing volume is:
[0097]
[0098] Wherein, is the number of berthed ships within a certain time period, B i is the number of berthed ships at the i-th moment, and n represents time.
[0099] The present invention also provides a system for identifying illegal occupation of the ship shoreline based on AIS big data, including:
[0100] Data acquisition and preprocessing module: used to acquire the AIS data of ships and perform preprocessing;
[0101] Geometric representation construction module: used to obtain geographic information data and construct a geometric representation of geographic information;
[0102] Abnormal behavior recognition module: used to superimpose and match the preprocessed AIS data and the geometric representation of the geographic information to obtain the area where the ship behavior occurs, and combine with a set threshold to obtain the ship abnormal behavior recognition result;
[0103] Illegal occupation recognition module: used to perform rough clustering on ship berthing points based on the ship abnormal behavior recognition result by using a multi-level clustering algorithm, then perform detailed clustering on each rough clustering group, and combine with the preprocessed AIS data for statistical analysis to obtain the ship shoreline illegal occupation recognition result;
[0104] The AIS data of the ship includes static information data and dynamic information data of the ship. The static information data includes the Maritime Mobile Service Identity (MMSI), captain information, ship length, ship width, and ship draft information. The dynamic information data includes longitude, latitude, and speed.
[0105] Compared with the prior art, the present invention has the following beneficial effects:
[0106] (1) The present invention combines AIS data and geographic information data for in-depth analysis, identifies the berthing behavior of the ship, and analyzes the berthing points by combining a multi-level clustering algorithm. This multi-level clustering algorithm can well identify the berthing points with high-density aggregation through the combination of rough clustering and detailed clustering, so as to accurately identify the area of illegal occupation of the shoreline.
[0107] (2) Based on the traditional DBSCAN clustering algorithm, because of its good adaptability, it can discover clusters of any shape. Especially for a dense data set such as ship trajectories, it can obtain good clustering results. On this basis, an improved multi-level clustering algorithm combining the K-Means++ algorithm and the DBSCAN algorithm is proposed. It can not only greatly reduce the time complexity of the algorithm, but also has good clustering results. In addition, by double-traversing the rough clusters with close distances and checking whether there are detailed clusters that should be merged between these two rough clusters, it can avoid the points that should be in the same detailed cluster being separated by the rough clusters, thereby further improving the accuracy of identifying the location of ship shoreline occupation.
[0108] Since the number of berthing areas is initially unknown, the clustering method cannot use the number of clusters as its initial parameter; but some berthing points are temporary and cannot form a general berthing area like other berthing points. The multi-level clustering algorithm adopted by the present invention can obtain clustering points without the initial number of clusters and can also identify abnormal points.
[0109] The present invention combines big data processing, Geographic Information System (GIS) analysis, and machine learning algorithms to provide an efficient and automated supervision tool for port management departments. This enables the supervision departments to quickly lock down suspicious areas based on data analysis, improve the inspection efficiency, and reduce the workload of manual intervention. In addition, the present invention can also improve the rational allocation of shoreline resources and the identification ability of illegal wharves, provide decision-making support for port planning, help optimize the rational allocation of shoreline resources, avoid the interference of illegal wharves on the normal port operation order, thereby optimizing port planning, ensuring water area order, and enhancing shipping safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0110] Figure 1 is a schematic flow chart of the method of the present invention;
[0111] Figure 2 is a schematic diagram of the AIS data parsing of the present invention;
[0112] Figure 3 is the technical route of the ship trajectory clustering process of the present invention;
[0113] Figure 4 is a schematic diagram of the ship trajectory clustering algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0114] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0115] Embodiment 1
[0116] This example provides a method for identifying illegal occupation of the shoreline by ships based on AIS big data, as Figure 1 shown, including the following steps:
[0117] Step 1: Parse the given AIS file to obtain data related to the ship trajectory.
[0118] The obtaining of the data related to the ship trajectory in Step 1 is specifically:
[0119] The data collected through the AIS system is AIS data, which contains rich static information (such as Maritime Mobile Service Identity MMSI, ship length, ship width, draft, etc.) and dynamic information (such as longitude, latitude, speed, course, etc.). Generally, ships send dynamic data every 30 seconds and static data every 5 minutes. Therefore, AIS data has strong timeliness. To improve the data transmission speed and efficiency, the original AIS information is transmitted and received in ASCII code, and the receiving end needs to decode the AIS data to read the specific information. A complete AIS message usually includes an identifier, the total number of messages, the message sequence number, the identification code, the channel, the encapsulated message, the number of padding bits, and the checksum.
[0120] Decode the AIS data. As Figure 2 shown, the decoding of AIS data mainly refers to the decoding of the encapsulated message. The encapsulated message is generally located in the sixth field of the original AIS data. Due to different transmitted contents, the types of encapsulated messages are also different. Different types of encapsulated messages have different formats. Therefore, in different types of AIS data, the fields corresponding to the same information are different.
[0121] Step 2: For the parsed AIS data, perform outlier removal and missing value filling.
[0122] The outlier removal and missing value filling preprocessing described in Step 2 is specifically as follows:
[0123] During the transmission of AIS data, due to reasons such as human input errors and equipment failures, data loss and data anomalies may occur. To reduce the impact of abnormal AIS data on the experimental results, the present invention first traverses and checks the decoded AIS data, marks the data points with a speed greater than 50 kn, a longitude not within [-180°, 180°], and a latitude not within [-90°, 90°] as abnormal, and at the same time determines the positions of the missing values. Among them, the formula for speed is:
[0124]
[0125] where ||p j+1 -p j || represents the distance between point p j and p j+1 . The Haversine formula is used to calculate the distance between latitude (λ) and longitude coordinates:
[0126]
[0127] where, R e is the radius of the earth.
[0128] Then, cubic spline interpolation is used for interpolation replacement, and the specific interpolation process is as follows:
[0129] 1: Assume that the AIS data sequence after removing outliers is X = {x1, x2, x3,..., x n}, Δt j is the time interval between x i and x i-1 . Traverse the data sequence. According to the time interval between adjacent two data points, find two data points with a time interval exceeding 30s, and determine the position of the missing value according to the positions of the two points.
[0130] 2: Select the cubic spline interpolation algorithm to fill the trajectory data at the positions that need to be filled. Interpolate once every 15s until all data traversal is completed.
[0131] Step 3: Based on the geographic information system, establish the geometric representation of geographic information.
[0132] The establishment of the geometric representation of geographic information described in Step 3 is specifically as follows:
[0133] Geographic information includes vessel traffic service (VTS) reporting lines, anchorages, fairways, port areas, shorelines, etc. The port area, anchorage, and fairway are mathematically and geometrically represented by constructing polygons through connecting control points (shape feature points). The shoreline expresses its irregular curve characteristics by connecting control points (shape feature points) and adopting the geometric line segment segmentation description method. The closed area formed by it and the VTS reporting line can determine the port water area range. In addition, for the navigable waters within the port area that are not divided into functional areas, their mathematical and geometric representations of geographic information are realized by constructing polygons through connecting control points (shape feature points).
[0134] By defining a class that includes name, point set, radius, or boundary points, etc., define the names and scopes of different anchorages, fairways, port areas, etc. Specifically, it is as follows:
[0135] Anchorage: Anchorage{Anchorage_name,Point_List};
[0136] Fairway: Fairway{Fairway_name,Point_List};
[0137] Harbour: Harbour{Harbour_name,Point_List};
[0138] Costline: Costline{Costline_name,Point_List};
[0139] VTS reporting line (circular): VTS{rCircle , (x cen , y cen )};
[0140] VTS reporting line (rectangle): VTS{Point_List};
[0141] Navigation area outside the functional area: Navigable_Area{Navigable_Area_name, Point_List};
[0142] Where: Anchorage_name is the name of the anchorage; Fairway_name is the name of the fairway; Harbour_name is the name of the port area; Navigable_Area_name is the name of the navigable waters outside the functional area; r Circle is the radius of the circular VTS reporting line; (x cen , y cen ) is the center control point; Point_List is the longitude and latitude matrix of the shape feature control points.
[0143] Step 4: Establish a ship behavior classification model.
[0144] The present invention identifies the ship's position by overlaying and matching the ship's longitude and latitude information with the shoreline geographical information, determines the ship's navigation status based on the ship's speed, and finally defines the ship's behavior in combination with the ship's navigation status and the ship's position at that time. The specific process is as Figure 3 shown. The specific identification method is as follows:
[0145] Step 4.1: Determine the area where the ship's behavior occurs according to the matching result of the ship's longitude and latitude and the shoreline geographical information. For a polygonal geographical area, the vector cross product method is used to determine whether the ship's behavior occurs in this area. The calculation formula is as follows. Assume that the longitude and latitude of the trajectory point are (x, y), and the longitude and latitude control point matrix of the polygon is Point_List = {(x1, y1),...,(x n , y n )}, if d1 × d2 × … × d n > 0, then the ship's behavior occurs in this area.
[0146] d1 × d2 × … × d n
[0147] d1 × d2 × … × d n > 0;
[0148] Where, d i is the displacement length of the ship from the i-th moment to the i + 1-th moment.
[0149] For a circular area, calculate the distance between the trajectory point (x, y) and the center (xcen , y cen ), and compare it with the radius of the circular area. If the following formula is satisfied, the ship's behavior occurs in this area.
[0150] (x - x cen ) 2 +(y - y cen ) 2 <r 2 Circle ;
[0151] Step 4.2: Set the speed threshold. The speed threshold S a_limit for ships at anchor is 0.5, the speed threshold S h_limit for ships in the port area is 0.2 kn, the speed threshold S f_limit for ships in the waterway is 0 kn, and the speed threshold S o_limit for ships in the navigation area water is 0.5 kn.
[0152] Step 4.3: Set the speed duration threshold. Since the stay behavior of ships is judged by continuous low speed, it is necessary to determine the duration of the ship's low speed. According to navigation experience, experimental statistics and literature research, set the time threshold T a_limit for anchoring at anchor as 0.5 h; the time threshold T h_limit for berthing in the port area as 1 h; since ship berthing behavior is strictly prohibited in the waterway, the time threshold T f_limit is 0 h; the time threshold T o_limit for ships berthing in the navigation area water is 0.5 h.
[0153] Step 4.4: According to the ship's navigation status and the area where the ship's behavior occurs, combined with the speed threshold and the speed duration threshold, define the ship's behavior as shown in the following table (S v is the ship's speed, and T is the current stationary time of the ship).
[0154] Table 1 Ship Behavior
[0155]
[0156] Step 5: Classification of illegal shoreline occupation behavior.
[0157] In this embodiment, the main research is on the problem of ships occupying the shoreline in the reservoir area, that is, illegally building a berthing area at the edge of the reservoir area. Therefore, in this embodiment, the berthing behavior of ships in non-navigation areas, non-ports, and non-anchorage areas is defined as abnormal behavior, and the identification of this abnormal behavior is mainly concerned.
[0158] First, screen the historical AIS data of ships in the navigation area, draw the historical trajectory, record the trajectory timestamp, and capture the trajectory points with berthing characteristics such as long-term low speed and no obvious change in trajectory points for a long time to realize the identification of berthing behavior in the navigation area.
[0159] Step 6: Identification of the occupied position information in the reservoir area.
[0160] Based on the identification results of abnormal ship behaviors, the data of ships berthing in the abnormal area can be obtained. However, the berthing of a small number of ships within a certain range may be caused by other reasons, and only when a large number of ships berth here can this range be defined as an illegally constructed berthing area. Therefore, a clustering algorithm is used for the identification results of abnormal ship behaviors to identify the positions where the ship encroaches on the shoreline.
[0161] The traditional DBSCAN clustering algorithm has good adaptability and can discover clusters of any shape. Especially for dense datasets such as ship trajectories, it can obtain good clustering results. However, for scenarios with a large amount of data such as large water areas, directly using the DBSCAN clustering algorithm will lead to too long clustering time. Therefore, this paper improves the clustering algorithm and proposes a multi-level clustering algorithm: First, use the K-Means++ clustering algorithm to perform rough clustering on the berthing points, and perform DBSCAN clustering in each cluster, thereby greatly reducing the time complexity of the algorithm.
[0162] Step 6.1: First, use the K-Means++ clustering algorithm for rough classification. The K-Means++ clustering algorithm is an improved version of the K-Means clustering algorithm, which mainly optimizes the selection method of the initial cluster centers to reduce the influence of local optimal solutions and speed up the convergence rate. It adopts a probability-based initialization strategy to make the initial cluster centers as scattered as possible, thereby improving the stability and effect of clustering. Its algorithm process is as follows:
[0163] First, initialize the cluster centers. Randomly select a point from the dataset X as the first cluster center c1. For each point x in the dataset, calculate its minimum distance from all currently selected cluster centers:
[0164]
[0165] where k′ represents the number of currently selected cluster centers. Using D(x) 2 as the probability distribution, randomly select the next cluster center c′ k+1 from the remaining data points, that is:
[0166]
[0167] In the formula, P(x) is the probability that the data point x is selected as the next cluster center, and x′ is other points in the dataset.
[0168] This process is repeated until k cluster centers are selected.
[0169] Next is the data point allocation. Calculate the Euclidean distance from each data point to all cluster centers and classify it into the nearest cluster center:
[0170]
[0171] where Cluster(x) is the final clustering of point x, and c j is the center of cluster j.
[0172] Cluster center update. Calculate the new center of each cluster, which is the mean of all data points belonging to that cluster:
[0173]
[0174] where Cluster(x) is the final clustering of point x, and S j is the set of all data points belonging to cluster j.
[0175] Finally, perform iterative optimization. Repeatedly execute the data point allocation and cluster center update steps until the cluster centers no longer change significantly or the maximum number of iterations is reached.
[0176] Step 6.2: Use the DBSCAN clustering algorithm to perform a detailed clustering on the points in each rough clustering cluster, and select the DBSCAN parameters for each rough clustering group using the existing legal storage capacity (wharf), with the goal that the clustering results can fit well with the existing legal storage capacity points. Combining Figure 4 as shown, the specific steps include:
[0177] Parameter grid setting. The parameters of DBSCAN include the neighborhood radius ε (Eps) and the minimum number of points minPts. Among them, ε specifies how close the distance between points should be to be considered part of a cluster, and minPts represents the minimum number of points required to form a cluster to ensure that there are enough AIS mooring points in each candidate mooring area. Once minPts is fixed, calculating the distance between each point and its minPts nearest neighbors can automatically determine the optimal eps (single-level density method). Let the minimum value of the neighborhood radius be ε min and the maximum value be ε max with a step size of ε step and the minimum value of the minimum number of points be minPts min and the maximum value be minPts max with a step size of minPts step Then the parameter grid is:
[0178]
[0179] where is the parameter grid, j is the jth minimum point, n ε and nminPts They respectively represent the upper limits of the search steps in two dimensions of the neighborhood radius and the minimum number of points, satisfying:
[0180]
[0181] Parameter selection: Traverse the parameter grid and select the neighborhood radius ε i and the minimum number of points minPts i and perform fine-grained clustering on the rough clustering results based on these two parameters. The fine-grained clustering process includes the following core point identification, clustering expansion, boundary point and noise point processing, and clustering completion.
[0182] Core point identification. For each mooring point in the dataset, if the number of points within its neighborhood ε i is greater than or equal to minPts i , then this point is marked as a "core point"; otherwise, this point may be a "boundary point" or a "noise point".
[0183] Clustering expansion. For each core point, starting from this point, all points within its neighborhood ε i are grouped into the same cluster, and the new core points within the neighborhood ε i are recursively expanded until no further expansion is possible.
[0184] Boundary point and noise point processing. If a point does not belong to the neighborhood ε i of any core point, then it is marked as a "noise point", indicating that it may be a random mooring or abnormal data point; if a point is within the neighborhood of a certain core point but is not a core point itself, it is marked as a "boundary point".
[0185] Clustering completion. The algorithm traverses all data points and finally forms one or more mooring area clusters and distinguishes noise points.
[0186] Clustering evaluation: Evaluate the clustering after completion. This evaluation is based on the analysis of each existing legal storage capacity center point and the clustering center of each fine-grained clustering. The evaluation index is designed as:
[0187]
[0188]
[0189] In the formula, score l is the evaluation result of the first rough clustering group, n represents the number of known legal storage capacity center points in the first rough clustering range, s j is the score of the jth legal storage capacity in the first rough clustering range, d j is the distance from s jThe straight-line distance to the nearest fine clustering center, where 50 and 100 are the set thresholds.
[0190] Step 6.3: Merge the clusters at the coarse clustering edges. First, check the distances between the coarse clustering groups that contain multiple fine clustering results. If the distances between the groups are too far, merging is not possible. Excluding these groups from the calculation process can effectively ensure a reduction in the computational complexity. The inter-group distance is calculated by the following formula:
[0191]
[0192] In the formula, dis_KC ij represents the distance between the group centers obtained by coarse clustering using K-Means++, and dis_r i and dis_r j represent the distances between the farthest points in the groups. If it indicates that merging is required.
[0193] Based on the merging results, obtain the clustering centers of the multi-level clustering algorithm, match the obtained clustering centers with the legal storage capacity, and the remaining part is the illegally occupied area. Specifically, if the straight-line distance between the clustering center and the existing legal storage capacity is within a certain threshold, it is determined as a legal point; otherwise, it is an illegal point.
[0194] In summary, in this embodiment, by combining AIS data and clustering results, the data of the trajectory points of illegal ship berthing is extracted, the berthing time of the ships occupying the shoreline is counted, the distribution of ship stays is analyzed, the hot spots of ship berthing are grasped, and the hot spot illegal areas are identified. The identified hot spot illegal areas should cover all berthing points in this area. Since some ships berth at temporary locations for specific reasons, after filtering these out, it is necessary to find the general berthing area points. Since the number of berthing areas is initially unknown, the clustering method cannot use the number of clusters as its initial parameter. However, some berthing points are temporary and cannot form a general berthing area like other berthing points. Therefore, a multi-level clustering algorithm is adopted, which can obtain clustering points without the initial number of clusters, but can also identify abnormal points. Finally, each berthing area is uniquely identified, including the center point and radius of the area, that is, the distance from the farthest point to the center point.
[0195] The method for grasping the hot spots of ship berthing time is to divide a day into different time periods, count the number of berthing ships in different time periods according to the berthing areas identified above, and finally conduct visual analysis. The statistical model of the berthing volume is:
[0196]
[0197] In the formula: represents the number of berthing ships in a certain time period, and B irepresents the number of berthed ships at a certain moment i, and n represents time.
[0198] In summary, through in-depth mining of AIS data, this embodiment can accurately identify the berthing behavior of ships and further analyze the berthing patterns of ships. The berthing behavior analysis based on AIS big data can not only be used to monitor the use of legal terminals, but also help to discover abnormal berthing point clusters, and then identify possible illegal encroachment on the shoreline area. To achieve this goal, this embodiment uses an improved multi-level clustering algorithm to analyze the berthing points of ships. By identifying the berthing points with high-density aggregation, it can effectively distinguish normal operating terminals and areas where violations may exist, and can discover irregularly shaped aggregation areas, which is suitable for the data characteristics of uneven spatial distribution of ship berthing points. The implementation of this embodiment will provide an efficient and automated technical means for port supervision, enabling the supervision department to quickly lock suspicious areas based on data analysis, improve the inspection efficiency, and reduce the workload of manual intervention. In addition, this method can also provide decision-making support for port planning, help optimize the rational allocation of shoreline resources, and avoid the interference of illegal terminals on the normal operation order of regular ports.
[0199] Embodiment 2
[0200] This embodiment provides a system for identifying illegal encroachment of ship shorelines based on AIS big data, including:
[0201] Data acquisition and preprocessing module: used to acquire AIS data of ships and perform preprocessing;
[0202] Geometric representation construction module: used to acquire geographic information data and construct a geometric representation of geographic information;
[0203] Abnormal behavior identification module: used to overlay and match the preprocessed AIS data with the geometric representation of the geographic information to obtain the area where ship behavior occurs, and combine with a set threshold to obtain the ship abnormal behavior identification result;
[0204] Illegal encroachment identification module: used to perform rough clustering on ship berthing points based on the ship abnormal behavior identification result using a multi-level clustering algorithm, then perform detailed clustering on each rough clustering group, and combine with the preprocessed AIS data for statistical analysis to obtain the ship shoreline illegal encroachment identification result.
[0205] The rest is the same as in Embodiment 1.
[0206] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0207] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The solutions in the embodiments of the present invention can be implemented using various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0208] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0209] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0210] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 in one block or multiple blocks Figure 1 of the steps.
[0211] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0212] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for identifying illegal occupation of ship shorelines based on AIS big data, characterized in that, Including the following steps: Obtain the AIS data of the ship and perform preprocessing; Obtain the geographic information data and construct the geometric representation of the geographic information; Overlay and match the preprocessed AIS data and the geometric representation of the geographic information to obtain the area where the ship behavior occurs, and combine with the set threshold to obtain the ship abnormal behavior recognition result; Based on the ship abnormal behavior recognition result, use the multi-level clustering algorithm to roughly cluster the ship berthing points, then perform detailed clustering on each rough clustering group, and combine with the preprocessed AIS data for statistical analysis to obtain the ship shoreline illegal occupation recognition result; The AIS data of the ship includes the static information data and dynamic information data of the ship. The static information data includes the Maritime Mobile Service Identity (MMSI), captain information, ship length, ship width, and ship draft information. The dynamic information data includes longitude, latitude, and speed; The step of obtaining the ship shoreline illegal occupation recognition result includes: For each rough clustering group containing the detailed clustering result, calculate the distance between the detailed clustering groups within different rough clustering groups, and judge whether the distance is less than the preset value. If so, merge the detailed clustering groups; if not, do not merge. The distance calculation expression between rough clustering groups is: ; Wherein, is the distance between the rough clustering groups i and j ; the distance between the centers of the rough clustering groups ; and and respectively represent the distances between the farthest points in the rough clustering groups. According to the merging result, obtain the clustering center of the multi-level clustering algorithm, and match the clustering center with the existing legal storage capacity. If the straight-line distance between the clustering center and the existing legal storage capacity is within a certain threshold, it is determined as a legal point; otherwise, it is determined as an illegal point to obtain the illegal occupation area, and calculate the berthing volume of the ships in the illegal occupation area at different time periods to complete the recognition process. The calculation expression of the berthing volume is: ; In the formula, is the number of berthed ships within a certain period of time, is the number of berthed ships at a certain moment i , n represents time.
2. The method for identifying illegal occupation of ship shorelines based on AIS big data according to claim 1, wherein, The steps of the preprocessing include: Decode the AIS data of the ship. The AIS data of the ship is sent and received in the form of ASCII code, including the identifier, total number of statements, statement serial number, identification code, channel, encapsulated message, number of padding bits, and checksum; Based on the decoded AIS data, perform outlier removal and missing value filling processing in sequence to complete the preprocessing process.
3. The method for identifying illegal occupation of ship shorelines based on AIS big data according to claim 2, wherein, The steps of performing outlier removal and missing value filling processing include: Traverse and check the decoded AIS data, mark the data points with speed, longitude, and latitude outside the set range as abnormal and remove them to obtain the AIS data sequence to be filled. The calculation expression of the speed is: ; ; In the formula, is the ship speed, is the point and the distance between them, representing the distance between the latitude λ and the longitude coordinates, is the radius of the earth; Determine the time interval between adjacent data points under normal circumstances, traverse the AIS data sequence to be filled, and obtain the positions of the missing values; According to the positions of the missing values, use the cubic spline interpolation method to fill the missing values in the AIS data sequence to be filled, and perform missing value filling once every set time until the traversal of the AIS data sequence to be filled ends.
4. The method for identifying illegal occupation of ship shorelines based on AIS big data according to claim 1, wherein, The geographic information data includes VTS reporting lines, anchorages, waterways, port areas, and shoreline geographic information. The steps of constructing the geometric representation of the geographic information include: 1) For anchorages, waterways, and port areas: Based on the geographical information of the fairway, port area, and anchorage, shape feature points are used as connection control points to construct polygons respectively, obtaining the geometric features of the fairway, port area, and anchorage, which are respectively defined as: Anchorage: Anchorage { Anchorage_name , Point_List}; Waterway: Fairway { Fairway_name , Point_List}; Port area: Harbour { Harbour_name , Point_List}; Among them, for the navigable waters in the port area that are not divided into functional areas, the shape feature points of the navigable waters are connected to control points to construct a polygon, obtaining the geometric features of the navigable area outside the functional area, which are defined as: Navigation area outside the functional area: Navigable_Area { Navigable_Area_name , Point_List}; In the formula, Anchorage_name is the name of the anchorage, Fairway_name is the name of the waterway, Harbour_name is the name of the port area, Point_List is the matrix of polygon longitude and latitude control points, Navigable_Area_name is the name of the water area of the navigation area outside the functional area; 2) For the shoreline: Based on the geographical information of the shoreline, the shape feature points of the shoreline are used as connection control points, and the geometric line segment segmentation description method is adopted to express the irregular curve features, obtaining the geometric features of the shoreline, which are defined as: Shoreline: Costline { Costline_name,Point_List }; In the formula, Costline_name is the shoreline name; 3) For the VTS reporting line: Based on the geographical information of the VTS reporting line, the geometric features of the circular or rectangular VTS reporting line are formed, which are defined as: Circular VTS reporting line: VTS { r Circle ,( x cen ,y cen )}; Rectangular VTS reporting line: VTS { Point_List} In the formula, r Circle is the radius of the circular VTS reporting line, is the center control point; Among them, the closed area formed by the geometric features of the VTS reporting line and the geometric features of the shoreline determines the port water area range.
5. The method for identifying illegal occupation of ship shorelines based on AIS big data according to claim 4, wherein The steps to obtain the area where ship behavior occurs include: Based on the preprocessed AIS data and the geometric representation of the geographical information, the vector cross product method is adopted for overlay matching to determine the area where ship behavior occurs, specifically including: 1) For the geometric representation of the geographical information in the shape of each polygon: Set the longitude and latitude of the trajectory points of the ship as , and represent the polygon longitude and latitude control point matrix as ; Calculation to determine is true. If so, it indicates that a ship behavior has occurred; if not, it indicates that no ship behavior has occurred, where: ; In the formula, d i is the displacement length of the ship from the i th moment to the i +1 th moment; 2) For the geometric representation of the geographical information in the shape of a circle: Set the longitude and latitude of the trajectory point of the ship to be , with the center of the circle being ; Calculate the longitude and latitude of the trajectory point from the center of the circle and compare it with the radius of the circular area r Circle to determine whether it holds. If so, it indicates that a ship behavior has occurred; if not, it indicates that no ship behavior has occurred.
6. The method for identifying illegal occupation of ship shorelines based on AIS big data according to claim 1, wherein, The steps to obtain the ship abnormal behavior recognition result include: Set speed thresholds, including the speed threshold for vessels at anchor , the speed threshold for vessels in the port area , the speed threshold for vessels in the waterway , the speed threshold for vessels in the navigation area waters ; Set speed duration thresholds, including the time threshold for anchoring at anchorages , the time threshold for berthing in port areas , the time threshold for berthing in waterways , the time threshold for berthing of ships in the waters of the navigation area ; According to the area where ship behavior occurs and the ship navigation state, combined with the speed threshold and speed duration threshold, the ship behavior is obtained, where the ship behavior includes the following multiple types: ① Anchorage navigation: The area where the ship's behavior occurs is the anchorage and satisfies , is the ship's speed; ② Anchorage mooring: The area where the ship's behavior occurs is the anchorage and meets , T is the current stationary time of the ship; ③Navigation in the port area: The area where the ship's behavior occurs is the port area and meets ; ④Berthing in the port area: The area where the ship's behavior occurs is the port area and meets ; ⑤Channel navigation: The area where the ship's behavior occurs is the channel and meets ; ⑥ Channel anchorage: The vessel behavior occurs in the channel and meets the following conditions: ; ⑦Navigation in the navigation area: The area where the ship's behavior occurs is the navigation area, and it satisfies ; ⑧Navigation area berthing: The area where the ship's behavior occurs is the navigation area and meets ; According to the ship behavior, the abnormal behavior of the ship is determined, obtaining the ship abnormal behavior recognition result, where the abnormal behavior of the ship includes the berthing behavior in non-navigable areas, non-port areas, and non-anchorage areas.
7. A method for identifying illegal occupation of ship shorelines based on AIS big data according to claim 1, characterized in that The K-Means++ algorithm is used for rough clustering of the berthing points, and the specific steps include: Cluster center initialization: Randomly select a data point from the dataset as the first cluster center , where the dataset is constructed based on the ship abnormal behavior recognition results, and one cluster represents a group; Minimum distance calculation: For each data point X in the dataset x , calculate the minimum distance to all current cluster centers, and the calculation expression for the minimum distance is as follows: ; In the formula, is the minimum distance, is the number of cluster centers already selected, is the i th cluster center; Cluster center selection: Using as the probability distribution, randomly select the next cluster center from the remaining data points , where the selection expression is: ; In the formula, data point x the probability of being selected as the next cluster center, is for other points in the dataset, including all point information; Repeat selection: Repeat the minimum distance calculation and cluster center selection steps until k cluster centers are selected; Data point assignment: Calculate the Euclidean distance of each data point to k cluster centers, and classify each data point to the nearest cluster center, expressed as: ; In the formula, is x the final clustering of points, is the j center of the cluster; Cluster center update: By calculating the mean of the data points in the cluster, the new center of each cluster is obtained, where the calculation expression is: ; wherein, is the set of all data points belonging to the cluster j ; Iterative optimization: Repeatedly execute the data point assignment and cluster center update until the iteration end condition is reached, obtaining the rough clustering groups of multiple berthing points.
8. A method for identifying illegal occupation of ship shorelines based on AIS big data according to claim 1, characterized in that, The DBSCAN clustering algorithm is used for fine clustering of each rough clustering group, and the DBSCAN parameters of each rough clustering group are selected using the existing legal storage capacity. The specific steps include: Parameter grid setting: Set the parameters of the DBSCAN clustering algorithm and further set the parameter grid, where the parameters include the neighborhood radius and the minimum number of points that make up a cluster minPts , the minimum value of the neighborhood radius is , the maximum value is , and the step size is , the minimum value of the minimum number of points is , the maximum value is , and the step size is , and the parameter grid is: ; In the formula, is the parameter grid, j is the j th minimum point, and respectively represent the upper limits of the search steps in the two dimensions of the neighborhood radius and the minimum number of points, satisfying: ; Parameter selection: Traverse the parameter grid and select the neighborhood radius and the minimum number of points for fine clustering; Core point identification: For each parking point in the dataset, determine whether all points within its neighborhood radius are greater than or equal to the minimum number of points , if so, mark the parking point as a core point, if not, it means the parking point is a boundary point or a noise point, where the dataset is constructed based on each rough clustering group; Cluster expansion: For each of the said core points, use it as a starting point and classify all points within its neighborhood radius into the same cluster, and recursively expand the neighborhood radius for the new core points within it until no further expansion is possible; Boundary point and noise point processing: If a point is not within the neighborhood radius of any core point it is marked as a noise point. If a point is within the neighborhood of a core point but is not a core point, it is marked as a boundary point; Clustering completed: All points in the dataset are traversed, and finally one or more mooring area clusters are formed, and noise points are distinguished; Cluster evaluation: Based on the center point of each existing legal storage capacity and the cluster center of each detailed cluster , calculate the evaluation index value to obtain the cluster evaluation result, where the calculation expression of the evaluation index value is as follows: ; ; Wherein, is the evaluation result of the l th rough clustering group, n represents the number of known legal storage capacity center points in the l th rough clustering range, is the l th legal storage capacity score in the j th rough clustering range, is the straight-line distance to the closest detailed clustering center, and A, B is the set threshold value. 9. An identification system for the method of identifying illegal occupation of ship shorelines based on AIS big data according to any one of claims 1-8, characterized in that, including: Data acquisition and preprocessing module: used to acquire the AIS data of ships and perform preprocessing; Geometric characterization construction module: used to acquire geographic information data and construct the geometric characterization of geographic information; Abnormal behavior recognition module: used to superimpose and match the preprocessed AIS data and the geometric characterization of the geographic information to obtain the area where the ship behavior occurs, and combine the set threshold to obtain the ship abnormal behavior recognition result; Illegal occupation recognition module: used to perform rough clustering on the ship mooring points by using a multi-level clustering algorithm based on the ship abnormal behavior recognition result, then perform detailed clustering on each rough clustering group, and combine the preprocessed AIS data for statistical analysis to obtain the ship shoreline illegal occupation recognition result; The AIS data of the ship includes the static information data and dynamic information data of the ship. The static information data includes the maritime mobile service identity MMSI, ship master information, ship length, ship width, and ship draft information. The dynamic information data includes longitude, latitude, and speed.
Citation Information
Patent Citations
Ship anomaly identification method and system and readable storage medium
CN115774804A