Fishing port supervision method and system based on multi-source data fusion

By using multi-source data fusion technology, the fishing port supervision system can achieve accurate identification and automatic early warning in all weather conditions under low visibility conditions. This solves the problems of blind spots and data isolation in the existing supervision system, improves the comprehensiveness and reliability of the system, eliminates the problems of monitoring blind spots and data isolation, and realizes the automatic and accurate binding of visual targets with AIS identities, thereby improving the accuracy of early warning and supervision efficiency.

CN121787714APending Publication Date: 2026-04-03CRCC HARBOR & CHANNEL ENG BUREAU GRP SURVEY & DESIGN INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

The existing fishing port supervision system has blind spots in low visibility conditions, and its sensing methods are limited, data is isolated, and the level of intelligence is low, making it difficult to achieve accurate identification and automatic early warning around the clock.

Method used

A multi-source data fusion method is adopted, which synchronously collects data through optical image sensors, radio frequency positioning signal receivers and environmental parameter sensors, and performs unified conversion, processing and analysis. Combined with evidence reasoning methods, target observation information is fused and identity is verified to generate comprehensive situational information, perform real-time matching analysis and trigger alarms.

Benefits of technology

It enhances the comprehensiveness and reliability of fishing port supervision, enables effective operation under adverse weather conditions, eliminates blind spots in monitoring, achieves automatic and accurate binding of visual targets with AIS identities, and improves the accuracy of early warnings and the efficiency of supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121787714A_ABST
    Figure CN121787714A_ABST
Patent Text Reader

Abstract

The invention discloses a fishing port supervision method and system based on multi-source data fusion. The method comprises the following steps: synchronously acquiring sensing data from an optical image sensor, a radio frequency positioning signal receiver and an environmental parameter sensor; uniformly converting the multi-source sensing data to be under the same time-space reference; processing and analyzing data of each sensor, and extracting target observation information; based on the dynamic confidence coefficient weight of each sensor, adopting an evidence reasoning method to fuse the multi-source observation information of the same target, and generating fused target information with comprehensive confidence coefficient; performing cross-modal identity verification and association on the fused target information, and constructing fishing port comprehensive situation information; and performing real-time analysis on the situation information according to a predefined complex event mode rule and generating an alarm. The system comprises a sensing layer, a data fusion and processing center and an intelligent early warning module. According to the invention, through multi-source data fusion, all-weather, high-precision and intelligent active supervision of the fishing port is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of maritime supervision and public safety technology, and relates to a method and system for supervising fishing ports based on multi-source data fusion. Background Technology

[0002] Fishing ports are important places for fishing vessels to dock, load and unload catches, and for personnel to work. Their safety supervision is crucial to protecting the lives and property of fishermen and maintaining port order. Currently, fishing ports mainly employ two supervision methods:

[0003] One method relies on fixed cameras for video surveillance, with personnel reviewing the footage 24 / 7. This approach is effective in good daylight, but at night, in heavy fog, or during heavy rain, visible light cameras struggle to clearly see targets, creating blind spots. Furthermore, human monitoring is prone to fatigue, making it difficult to detect unusual behavior promptly and hindering automatic identification of vessels or personnel.

[0004] Secondly, there's the issue of using the Automatic Identification System (AIS) alone. AIS can receive information broadcast by vessels, such as their location, name, and speed. However, many small fishing boats don't have AIS equipment installed, or it may be manually turned off, making some vessels invisible. More importantly, AIS data and video footage are independent of each other. Vessels seen on the screen cannot be automatically mapped to specific vessels in the AIS list, creating a situation where you can see a vessel but can't recognize it.

[0005] The methods described above generally suffer from drawbacks such as limited sensing capabilities, isolated data, and low levels of intelligence. For example, video surveillance may fail when someone falls into the water at night. In dense fog, it is impossible to determine whether a vessel has entered a restricted area. Anomalies rely on manual detection, resulting in slow response times and a high rate of false alarms.

[0006] Therefore, a new generation of fishing port supervision system is needed that can integrate various data such as video, positioning, and environment to achieve all-weather perception, accurate target identification, and automatic intelligent early warning. Summary of the Invention

[0007] To address the problems existing in the background technology, this invention proposes a fishing port supervision method and system based on multi-source data fusion.

[0008] To achieve the above objectives, the technical solution adopted by this invention is as follows: a fishing port supervision method based on multi-source data fusion, comprising:

[0009] Simultaneously collect sensing data from multiple heterogeneous sensors, including at least image data collected by an optical image sensor, positioning and identity data collected by a radio frequency positioning signal receiver, and environmental parameter data collected by an environmental parameter sensor.

[0010] The sensing data collected by the various heterogeneous sensors are uniformly converted to the same spatiotemporal reference.

[0011] The sensed data is processed and analyzed to extract target observation information from the perspectives of each sensor;

[0012] Based on the dynamic confidence weights of each sensor, an evidence reasoning method is used to fuse observation information about the same target from different sensors to generate fused target information with comprehensive confidence.

[0013] Cross-modal identity verification and association are performed on the fused target information to construct unified comprehensive situational information of fishing ports;

[0014] Based on predefined complex event pattern rules, the comprehensive situation information of the fishing port is matched and analyzed in real time, and alarm information is generated when a matching event is detected.

[0015] Based on the aforementioned method for monitoring fishing ports using multi-source data fusion, this invention further proposes a fishing port monitoring system based on multi-source data fusion, comprising:

[0016] The perception layer includes optical image sensors, radio frequency positioning signal receivers, and environmental parameter sensors, as well as time synchronization network devices that provide a unified time reference for all sensors;

[0017] The data fusion and processing center is communicatively connected to the perception layer and configured to perform the following operations:

[0018] Transform sensor data from different sensors into a unified geographic coordinate system and time reference.

[0019] Based on the type of each sensor and the real-time parameters of its environment, the dynamic confidence weight of each sensor data is calculated during fusion.

[0020] Based on the dynamic confidence weight, the evidence reasoning method is used to fuse multi-source target observation information and output fused target information with comprehensive confidence.

[0021] Cross-modal identity association and verification are performed on the fused target information, and comprehensive situational information of the fishing port is generated;

[0022] The intelligent early warning module is communicatively connected to the data fusion and processing center, and is configured to load predefined complex event pattern rules to perform real-time analysis of the comprehensive situation information of the fishing port and trigger alarms.

[0023] Compared with existing technologies, this invention has the following advantages: by integrating multi-source data such as visible light, thermal imaging, AIS, UWB, and environmental sensors, it improves the comprehensiveness and reliability of fishing port supervision. The system can operate continuously and effectively under adverse conditions such as nighttime, heavy fog, and rain, increasing the sensing coverage and helping to eliminate blind spots in traditional video surveillance.

[0024] By using spatiotemporal alignment, motion consistency verification, and appearance feature matching, the system can automatically and accurately bind visual targets to AIS identities or UWB personnel tags, solving the problem of being able to see something but not recognize it. It can also effectively identify and track ships that have not enabled AIS.

[0025] In terms of early warning, the evidence fusion theory based on dynamic confidence is adopted to suppress false alarms from a single sensor. Combined with a complex event processing engine, it can automatically identify high-risk events such as people falling into the water, thereby improving the accuracy of early warning.

[0026] All information is presented on a unified real-time situation map, allowing regulators to intuitively grasp the overall status of the port, improving decision-making efficiency and emergency response capabilities. This enables a shift in port supervision from passive monitoring to proactive, intelligent early warning, providing an effective technical means to enhance port safety management. Attached Figure Description

[0027] Figure 1 This is a flowchart of a fishing port supervision method based on multi-source data fusion according to the present invention;

[0028] Figure 2 This is a framework diagram of a fishing port supervision system based on multi-source data fusion according to the present invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] like Figures 1-2 As shown, the technical solution adopted by this invention is as follows: A method for monitoring fishing ports based on multi-source data fusion, comprising:

[0031] S1: Synchronously acquire sensing data from multiple heterogeneous sensors. This sensing data includes at least image data acquired by an optical image sensor, location and identity data acquired by a radio frequency positioning signal receiver, and environmental parameter data acquired by an environmental parameter sensor. Synchronous acquisition does not mean that each sensor triggers sampling at the same physical moment, but rather that all raw sensing data output by the sensors are assigned a timestamp generated from a unified time base and calibrated at the microsecond level, thereby logically achieving time alignment of the multi-source data. This time alignment is achieved through a high-precision time synchronization network deployed within the fishing port's monitoring area.

[0032] Specifically, the system deploys three types of heterogeneous sensors at key locations in the fishing port: optical image sensors, radio frequency positioning signal receivers, and environmental parameter sensors. Among them, the optical image sensors include visible light cameras and thermal imaging cameras, which are used to acquire high-resolution color images under sufficient lighting conditions and to acquire target contour images through differences in infrared thermal radiation under nighttime, foggy, or low-light conditions. Together, they constitute an all-weather visual perception capability.

[0033] The radio frequency positioning signal receiver includes an Automatic Identification System (AIS) receiver and an Ultra-Wideband (UWB) positioning base station. The AIS receiver passively receives VHF band messages broadcast by ships and deciphers the ship's MMSI code, name, latitude and longitude, speed, heading, and other identification and dynamic information. The UWB base station calculates the real-time coordinates and associated identification of personnel wearing UWB tags in the port area's three-dimensional space by measuring the time-of-flight (ToF) of nanosecond-level pulse signals emitted by such personnel.

[0034] The environmental parameter sensor is a six-element meteorological and hydrological station used to monitor and output key environmental variables that affect the sensing performance in real time, including: illuminance L (unit: lux), visibility V (unit: km), rainfall intensity I (unit: mm / h), wind speed, relative humidity, and the temperature difference ΔT between the target and the background (unit: ℃) calculated with the assistance of thermal imaging equipment.

[0035] To ensure strict time alignment of data collected by the three types of sensors, the system constructs a master-slave time synchronization network based on the Precision Time Protocol (PTP, IEEE 1588 standard). This network consists of a Grandmaster Clock and multiple Slave Clocks. The Grandmaster Clock receives high-precision timing signals (such as the BDS B1I frequency) from the BeiDou Navigation Satellite System, serving as the sole time reference source for the entire system. All sensor nodes are configured as slave clocks and connected to the PTP network via industrial Ethernet.

[0036] During PTP protocol operation, the master clock and any slave clock synchronize their time by exchanging standardized PTP messages. The master clock periodically broadcasts Sync messages to transmit its local time reference. The slave clock actively sends Delay_Req messages to measure reverse link latency. These four key timestamps are defined as follows: t1: The time the master clock records the Sync message transmission on its local clock; t2: The time the slave clock records the Sync message reception on its local clock; t3: The time the slave clock records the Delay_Req message transmission on its local clock; t4: The time the master clock records the Delay_Req message reception on its local clock.

[0037] The four timestamps mentioned above are absolute time values ​​captured by the internal hardware clock of their respective devices, in seconds, with an accuracy down to the nanosecond level. Based on these four timestamps, the system uses the following formula to calculate the path delay and time offset:

[0038] offset = t2 - t1 - delay;

[0039]

[0040] Here, `delay` represents the average propagation delay of the unidirectional communication link between the master and slave clocks, assuming that the forward (master clock to slave clock) and reverse (slave clock to master clock) paths are symmetrical. `offset` represents the time deviation of the slave clock relative to the master clock. The slave clock adjusts the phase or compensates the frequency of its local clock based on the calculated offset, ultimately synchronizing its local time with the master clock at the microsecond level.

[0041] Through this synchronization mechanism, all sensors, while acquiring raw data (such as an image frame, an AIS message, or a set of meteorological values), are simultaneously stamped with a timestamp generated by a calibrated local clock, consistent with the master clock. The resulting multi-source heterogeneous sensing data stream with a unified timestamp forms the basis for spatiotemporal alignment, target fusion, identity association, and event reasoning in subsequent steps. Without this synchronization mechanism, observations of the same target by different sensors would result in state misalignment due to inconsistent sampling times, leading to cross-modal matching failures or event logic misjudgments.

[0042] S2: The sensing data collected by the various heterogeneous sensors are uniformly converted to the same spatiotemporal reference. After completing the acquisition of multi-source sensing data with a unified timestamp, the system performs spatial coordinate mapping and time alignment operations on the raw observation data from optical image sensors, radio frequency positioning signal receivers, and environmental parameter sensors, so that they are all expressed under a pre-established unified geographic coordinate system and a unified time axis, thereby providing a geometric and temporal consistency basis for subsequent multi-source fusion.

[0043] First, a unified geographic coordinate system covering the entire fishing port supervision area is established, preferably using the WGS-84 geodetic coordinate system. If the system's internal processing requires the use of planar coordinates, then the transverse Mercator projection (such as UTM Zone 50N) is used as its equivalent expression, and all spatial transformations use this as the final target coordinate system.

[0044] For visual image data output by an optical image sensor, its pixel coordinates need to be converted to world coordinates using a camera calibration model. The camera calibration model consists of three parts:

[0045] (1) Intrinsic parameter matrix K∈R 3×3 Defined as:

[0046] Among them, (f x ,f y ) represents the equivalent focal length (in pixels) along the x and y axes, (c x ,c y The matrix is ​​the coordinates of the main point (unit: pixels). It was obtained offline using a checkerboard calibration board before deployment by Zhang Zhengyou calibration method.

[0047] (2) Distortion coefficient vector: D = [k1,k2,k3,p1,p2]; where k1,k2,k3 are radial distortion coefficients, and p1,p2 are tangential distortion coefficients, used to correct lens nonlinear deformation.

[0048] (3) The extrinsic parameter matrix [Q|t], where Q∈SO(3) is the rotation matrix, and t∈R 3 The translation vector represents the pose and position of the camera coordinate system relative to the unified geographic coordinate system, and is also determined through joint optimization during the calibration process.

[0049] Based on the camera calibration model described above, the system first performs distortion correction on any pixel p(u,v) in the image. Let its corresponding normalized camera coordinates be (x,y), then the distortion-corrected coordinates (x′,y′) are calculated using the following formula:

[0050] x′=x(1+k1r 2 +k2r 4 +k3r 6 )+2p1xy+p2(r 2 +2x 2 );

[0051] y′=y(1+k1r 2 +k2r 4 +k3r 6 )+p1(r 2 +2y 2 )+2p2xy;

[0052] Where, r 2 =x 2 +y 2 This operation eliminates lens distortion and ensures accurate projection geometry. Subsequently, based on the prior assumption that fishing port targets are typically located at sea level, the system sets the target height Z in the world coordinate system. w =0, and the distorted coordinates (x′, y′) are solved by inversely solving the world horizontal coordinates (X′, y′) using the perspective projection equation. w ,Y w ):

[0053] Where s is an unknown scaling factor. Since Z w =0, the equation can be transformed into an equation concerning X. w With Y w The linear system is solved by the least squares method, and the two-dimensional position of the target in a unified geographic coordinate system is output.

[0054] For the radio frequency positioning and identity data output by the radio frequency positioning signal receiver, the system uses a coordinate system transformation algorithm to map its original location to a unified geographic coordinate system:

[0055] The latitude and longitude (φ, λ) and ellipsoidal height h output by the AIS receiver ell If conversion to planar coordinates is required, h ell This represents the vertical height of the target position relative to the WGS-84 reference ellipsoid, in meters. This is achieved through a forward projection transformation from WGS-84 to UTM, which includes standard geodetic parameters such as the central meridian, scale factor, pseudo-east, and north offset.

[0056] The local Cartesian coordinates (x) output by the UWB positioning base station uwb ,y uwb ,z uwb Then it can be converted using the following formula:

[0057] Among them, Q uwb Let X be the rotation matrix of the UWB local coordinate system relative to the geographic coordinate system. origin ,Y origin Z origin The position of the UWB origin in the geographic coordinate system is defined as follows: ) These two sets of parameters are calibrated and fixed in one go during the system deployment phase through on-site surveying using RTK-GNSS or a total station.

[0058] For environmental parameter sensors (such as six-element meteorological and hydrological stations), since they lack dynamic positioning capabilities, the system performs fixed sensor location binding during the initialization phase: the WGS-84 latitude and longitude (φ) of their installation point is determined using high-precision GNSS equipment. s ,λs ) and altitude h s The coordinates are then embedded as metadata in each uploaded environmental parameter data packet. Based on this, the data fusion and processing center correlates parameters such as illuminance L, visibility V, rainfall intensity I, and temperature difference ΔT to specific geographical areas for subsequent dynamic assessment of sensor confidence by region.

[0059] Based on spatial unification, the system further implements a time alignment mechanism: although microsecond-level time synchronization has been achieved through the PTP protocol, due to the different sampling frequencies of each sensor (e.g., 30Hz for video, about 1Hz for AIS, and once per minute for weather stations), time window matching of the data stream is still required.

[0060] Specifically, the data fusion and processing center maintains a multi-channel time buffer queue. When a sensor data packet is timestamped T... data Upon arrival, the system searches for timestamps falling within [T] in other channels. data -ΔT sync ,T data +ΔT sync The nearest neighbor data within ], where ΔT sync =50ms is the preset synchronization tolerance. If a mode is missing, linear interpolation is performed on its continuous state variables (such as position).

[0061] Among them, S prev With S next These represent the state values ​​of the two frames before and after the interpolation point. The interpolation results, together with the original data, constitute a time-aligned observation set.

[0062] In summary, by using camera calibration models, coordinate system transformation algorithms, fixed sensor location binding, and time alignment mechanisms, multi-source sensing data is transformed into target observation information expressed in a unified geographic coordinate system and a unified time axis. This eliminates spatial representation differences and temporal misalignments between sensors, providing the necessary and sufficient geometric and logical consistency prerequisites for target feature extraction and multi-source evidence fusion.

[0063] S3: Process and analyze the perceived data to extract target observation information from each sensor's perspective. Specifically, the processing of visual image data includes: analyzing the images using a target detection neural network model with spatial attention enhancement to improve the detection capability of targets against complex sea surface backgrounds. After completing the spatiotemporal benchmark unification, the system performs modality-specific parsing operations on the perceived data from the optical image sensor, RF positioning signal receiver, and environmental parameter sensor to generate structured target observation information, including target category, location, state, and auxiliary features.

[0064] For visual image data, the system employs a target detection neural network model with spatial attention enhancement. This model is based on the YOLOv5 architecture and features improvements including embedding a coordinate attention mechanism after the last C3 module of the backbone network, replacing the original path aggregation structure with a bidirectional weighted feature pyramid network, and using a combined loss function for training optimization. These improvements aim to enhance the robustness and accuracy of detecting targets such as ships, personnel, and small floating objects against complex sea surface backgrounds (e.g., strong reflections, wave interference, low contrast).

[0065] The backbone network is the convolutional neural network portion of YOLOv5 used to progressively extract multi-scale semantic features from the input image. It consists of a Focus module, multiple CBS (Conv-BN-SiLU) blocks, and several stacked C3 modules. The backbone network outputs a set of feature maps with decreasing spatial resolution and increasing semantic hierarchy. The C3 module is a cross-stage partial bottleneck structure unit in YOLOv5, containing a 1×1 convolutional dimensionality reduction branch, a backbone branch composed of multiple Bottleneck residual blocks, and a CSP (Cross Stage Partial) connection mechanism to preserve gradient flow while reducing computational cost. Each C3 module outputs a feature map with a fixed number of channels and progressively smaller spatial size.

[0066] A coordinate attention module is inserted between the last C3 module and its subsequent connected layers in the backbone network. This coordinate attention module is not placed arbitrarily, but rather at the optimal ensemble point determined after performance validation. Its output feature map possesses both high-level semantic information and retains sufficient spatial resolution (typically 1 / 8 of the input image), making it suitable for guiding the network to focus on the region containing small objects. The specific implementation of this coordinate attention module is as follows: Let the feature map output by the last C3 module be F∈R. C×H×W Where C is the number of channels, H is the height, and W is the width.

[0067] First, perform global average pooling on F along the horizontal direction to generate the feature description vector z in the vertical direction. h ∈R C×1×W The calculation for its Wth column is as follows: Where i is an integer index in the height direction, with a value range of 1≤i≤H.

[0068] Secondly, perform global average pooling on F along the vertical direction to generate the horizontal feature description vector z. v ∈R C×H×1 The calculation for its i-th row is as follows: Where j is an integer index in the height direction, with a value range of 1≤j≤H.

[0069] Then, z h With z v The vectors are concatenated along the channel dimension and then non-linearly encoded using a shared 1×1 convolutional layer (containing batch normalization and SiLU activation). They are then split into two independent vectors, denoted as a. h and a v Each channel has dimensions C. Then, the channel attention weight vector g is obtained by mapping back to the original number of channels C through two independent 1×1 convolutional layers (without activation functions). h and g v .

[0070] To further improve feature discriminative ability, a learnable gating mechanism is introduced: g h =σ(W h *a h ),g v =σ(W v *a v ); where W h ∈R C×C×1×1 With W v ∈R C×C×1×1 The parameters are learnable convolution kernels, and σ is the Sigmoid function, ensuring that the output range is [0,1].

[0071] Finally, g h Broadcasting along the height dimension extends to R C×H×W , will g v Broadcasting along the width dimension expands to R C×H×W And multiply it element-wise with the original feature map: F out =F⊙g v ⊙g h Where ⊙ represents element-wise multiplication; output F out The enhanced feature map is then passed to the subsequent feature fusion network.

[0072] The bidirectional weighted feature pyramid network replaces the path aggregation network (PANet) in the original YOLOv5. This structure includes two information fusion paths: top-down and bottom-up. In each path, feature maps from different levels are weighted and fused using learnable weights. For the output features of layer l... The calculation formula is as follows:

[0073] in, ω represents the set of adjacent level indices participating in the fusion. l,k For the learnable scalar weights of the corresponding path (optimized through backpropagation), Ψ l,k(·) indicates an upsampling or downsampling operation (such as nearest neighbor interpolation or 3×3 convolution downsampling). This mechanism enables the network to adaptively emphasize scale features that are more important to the current task, improving its ability to detect large ships in the foreground and small people and floating objects in the background.

[0074] The improved model was trained end-to-end using a dataset containing at least 100,000 images of fishing port scenes. The training data covered various weather conditions (sunny, rainy, foggy), time of day, night, and different perspectives (overhead, eye-level), as well as complex sea conditions. To alleviate the sparsity problem of small target samples, data augmentation strategies were implemented for personnel and small floating object categories, including random cropping, brightness perturbation, affine transformation, and synthetic floating object textures.

[0075] The model's loss function uses a weighted combination of CloU Loss and Focal Loss, defined as: in, To improve positioning accuracy, the regression loss of the bounding box is considered, taking into account the center point distance, aspect ratio, and overlapping areas. The classification loss to address the imbalance between positive and negative samples takes the following form: Where, p t α is the predicted probability of the true class. t λ is the category balance factor (value 0.25), γ = 2 is the focusing parameter; loc =1.0 and λ cls =0.8 is the preset loss weight coefficient, which is determined through ablation experiments.

[0076] After training, the model is deployed on edge computing units (such as Huawei Atlas 500 smart stations) in the edge processing layer to perform real-time inference on the video stream. Each frame outputs a structured detection result list, with each element containing: target category label (e.g., ship, person, vehicle, floating object), bounding box coordinates (x, y, y). min ,y min ,x max ,y max ), Detection confidence score c det ∈[0,1]. In addition, the model includes a separate ReID (Re-identification) branch, which is typically attached after the feature pyramid network to extract the 128-dimensional appearance feature vector f of the target. app ∈R 128 This vector is specifically designed for cross-camera target tracking and identity verification.

[0077] For radio frequency positioning and identity data, the system performs protocol parsing and status calculation: It parses the NMEA 0183 format message output by the AIS receiver to extract fields such as the ship's MMSI code, ship name, latitude and longitude, speed, heading, and navigation status, and converts them into a location point (X) in a unified geographic coordinate system. ais ,Y ais ) and motion vector

[0078] The least squares algorithm is used to calculate the three-dimensional coordinates (X, Y, F, Z) of the tag-wearing person based on the ranging data output from the UWB base station network. uwb ,Y uwb Z uwb ), and associate it with its unique identifier ID.

[0079] For environmental parameter data, the system directly extracts numerical observations as contextual features, including: illuminance L (unit: lux), visibility V (unit: km), rainfall intensity I (unit: mm / h), and target-background temperature difference ΔT (unit: ℃). Although these parameters do not directly describe the target, they serve as input variables for subsequent dynamic confidence assessment, quantifying the reliability of each sensor in the current environment.

[0080] All the above processing results are output in the form of structured data packets, constituting target observation information from the perspectives of each sensor. Visual observations include category, location, confidence level, and appearance features. Radio frequency observations include identity, location, and motion status. Environmental observations include meteorological and hydrological parameters that affect perception quality. This information serves as the raw input for evidence fusion, and its completeness and accuracy directly determine the reliability of the fusion decision.

[0081] Target recognition confidence refers to the quantification of the degree to which any sensor can determine that the observed target belongs to a certain preset category. Its value ranges from [0,1]. The higher the value, the more confident the sensor is that the target belongs to that category.

[0082] Furthermore, to support subsequent multi-source evidence fusion, the system uniformly represents the target recognition confidence scores output by each sensor as a target category probability vector p = [p1, p2, p3, p4, p5], where each component corresponds to the category confidence score in the recognition framework Θ = {ship, personnel, vehicle, floating object, unknown}.

[0083] The generation method of the target category probability vector p depends on the sensor type as follows:

[0084] (1) For optical image sensors (including visible light cameras and thermal imaging cameras):

[0085] The improved YOLOv5 model deployed on edge computing units outputs multi-class detection confidence scores, which are then normalized using softmax to form p, and its components p j ∈[0,1] represents the probability that the model determines the target to belong to the j-th category.

[0086] (2) For the Automatic Identification System (AIS): If a valid message with a correct checksum, valid MMSI code, and reasonable position change is received, the target is determined to belong to the ship category, and p is constructed. AIS = [1.0,0.0,0.0,0.0,0.0]; If the message is invalid or missing, no target observation will be generated, nor will it participate in subsequent fusion.

[0087] (3) For Ultra-Wideband (UWB): If the positioning solution residual is less than 0.5 meters and the tag ID is valid, then the target is determined to belong to the personnel category (because UWB tags are only issued to authorized personnel in the port area), and the following is constructed: p UWB = [0.0, 1.0, 0.0, 0.0, 0.0]; If the localization fails or the tag is offline, no target observation will be generated.

[0088] The above construction is based on the prior knowledge of the deployment of the fishing port monitoring system: AIS equipment is only installed on ships, and UWB tags are only issued to authorized personnel, thus their observations are category-exclusive. This assumption conforms to maritime regulatory business norms and is a system configuration premise known to those skilled in the art.

[0089] Although AIS and UWB outputs are hard decisions, their influence will be influenced in subsequent processes through dynamic confidence weights (such as w). AIS =0.99) to soften the evidence, thereby preserving reasonable uncertainty in evidence fusion and avoiding fusion errors caused by misjudgment of a single modality.

[0090] Ultimately, all target observation information from the sensors is output in a unified structured data packet format, containing the following fields: target category probability vector p; target location coordinates, converted to unified geographic coordinates; timestamp, time synchronized via PTP; sensor type identifier; and auxiliary features, such as visual ReID features, AIS MMSI code, and UWB tag ID.

[0091] In summary, by using a target detection neural network model with spatial attention enhancement, a radio frequency protocol parsing and localization calculation module, and an environmental parameter extraction interface, the system extracts modality-specific, semantically clear, and structurally unified target observation information from multi-source sensing data, laying a feature foundation for achieving high-precision multi-source fusion.

[0092] S4: Based on the dynamic confidence weights of each sensor, an evidence-based reasoning method is used to fuse observation information about the same target from different sensors, generating fused target information with a comprehensive confidence score. This step is performed in the data fusion and processing center, and its input is multi-source structured data that has undergone spatiotemporal benchmark unification and target observation information extraction. The same target refers to targets whose spatial distance is less than a preset threshold D under a unified geographic coordinate system. match (e.g., 50 meters) and the timestamp difference is less than the synchronization tolerance Δt. sync The physical entity pointed to by multiple sensor observations (e.g., 200 milliseconds).

[0093] Optical image sensors fall into two categories: visible light cameras, which output color RGB images and rely on external lighting; and thermal imaging cameras, which output infrared thermal radiation intensity maps and rely on the temperature difference between the target and the background.

[0094] Specifically, the determination of the dynamic confidence weight includes: for the visible light camera, dynamically calculating the evaluation factor of its data reliability based on at least one of the lighting conditions, visibility conditions and rainfall intensity in its real-time acquisition environment.

[0095] For the thermal imaging camera, an evaluation factor for its data reliability is dynamically calculated based on the temperature difference between the target and the background.

[0096] Real-time environmental data is collected by an environmental parameter sensor, which is a six-element meteorological and hydrological station deployed at a fixed location in the fishing port. Its WGS-84 coordinates are pre-bound using high-precision GNSS equipment. s ,lon s ,h s The output data packet contains the following fields: illuminance L (unit: lux), measured by a silicon photodiode sensor; visibility V (unit: km), measured by a forward-scattering visibility meter; rainfall intensity I (unit: mm / h), measured by a tipping bucket rain gauge; and target-background temperature difference ΔT (unit: ℃), calculated by the internal algorithm of the thermal imaging camera. After detecting a target area, the average temperature of the target pixels is subtracted from the average temperature of the pixels in the adjacent background area (such as the sea surface). The target area is determined by the bounding box output by the target detection model. The adjacent background area is defined as the set of non-target pixels that extend 20 pixels beyond the target bounding box and are located in the same horizontal band (±5 pixels vertical offset).

[0097] For visible light cameras, the evaluation factor C vis Calculate using the following formula: C vis (L,V,I)=β vis ·σ(α L ·L+α V ·V-α R·I); where: β vis The base reliability coefficient is ∈[0,1], reflecting the inherent reliability of the camera's hardware performance and installation angle. A typical value is 0.85, calibrated through factory testing and on-site acceptance. α L >0 represents the light sensitivity coefficient, measured in lux. -1 A typical value of 0.005 represents the marginal effect of increasing reliability by 1 lux of illumination. α V >0 represents the visibility sensitivity coefficient, measured in km. -1 A typical value of 0.2 represents the contribution of each 1km improvement in visibility to the improvement in reliability. α R >0 represents the rainfall disturbance sensitivity coefficient, with units of (mm / h). -1 The typical value is 0.1, which represents the decrease in reliability caused by each 1 mm / h increase in rainfall intensity. This is the Sigmoid function, used to compress linear combinations to the (0,1) interval, ensuring that the output is in probabilistic form and avoiding negative values ​​or exceeding the limit.

[0098] For thermal imaging cameras, the evaluation factor C ir Calculate using the following formula: Where, β ir ∈[0,1] represents the basic reliability coefficient of the thermal imaging device, with a typical value of 0.95.

[0099] T th >0 represents the temperature difference threshold, measured in °C, with a typical range of [3, 8]. It is set to 5 °C in southern fishing ports and 3 °C in northern winter fishing ports to accommodate scenarios with small temperature differences between the target and the ice surface in low-temperature environments. When ΔT ≥ T th At that time, the reliability of thermal imaging reaches its upper limit β. ir ΔT <T th At that time, the reliability decreased linearly, reflecting the weakened detection capability under low contrast.

[0100] Based on the evaluation factors and in conjunction with the data continuity index of the optical image sensor in the recent time period, its weight in the fusion decision is determined.

[0101] Data continuity index ρ cont Defined as: the number of frames N in which the optical image sensor successfully outputs valid target detection results within the past sliding time window τ = 60 seconds. valid Theoretically, the number of output frames N should be... total =f s The ratio of τ, that is: Among them, f s This represents the video frame rate (e.g., 30Hz). If consecutive frame drops occur due to lens obstruction, power outage, or edge computing unit overload, then ρcont Significantly reduced.

[0102] Final dynamic confidence weight w vis Calculated by the following formula: w vis =λ·C vis +(1-λ)·ρ cont ; where λ=0.7 is the environment-history balance coefficient, which is optimized on multiple fishing port datasets through cross-validation to ensure that environmental adaptability dominates but does not completely ignore historical stability.

[0103] For radio frequency positioning signal receivers (including AIS receivers and UWB base stations), their dynamic confidence weight m rf Determined as follows: For AIS receivers: if the received message checksum is correct, the MMSI code is valid, and the position jump is less than the threshold (e.g., 10 segments / second), then w ais =0.99, otherwise set to 0.1. For UWB base stations: if the positioning solution residual (i.e., least squares fitting error) is less than 0.5 meters, then w uwb =0.95, otherwise it will decrease linearly to 0.3 according to the residual.

[0104] Specifically, the fusion using the evidence reasoning method includes: constructing a basic confidence assignment for propositions that the target belongs to a specific category based on the dynamic confidence weights of each sensor and their output target recognition confidence. Let the recognition frame be: Θ = {θ1, θ2, θ3, θ4, θ5}; where: θ1 = ship, θ2 = personnel, θ3 = vehicle, θ4 = floating object, θ5 = unknown.

[0105] For the i-th sensor (i = 1, 2, ..., m), its target recognition confidence is a vector: p i =[p i1 ,p i2 ,p i3 ,p i4 ,p i5 ]; where: p ij ∈[0,1] represents the category θ to which the sensor determines the target belongs. j The probability. For example, the output of a visible light camera: p vis =[0.6,0.1,0.05,0.15,0.1]. The thermal imaging camera may output: p ir =[0.8,0.05,0.0,0.1,0.05]. The AIS receiver is only effective for ships, therefore the output is: p ais = [0.99, 0.0, 0.0, 0.0, 0.01].

[0106] The dynamic confidence weight of this sensor is denoted as w. i∈[0,1]. Then its Basic Reliability Assignment (BPA) function m i :2 Θ →[0,1] is defined as:

[0107]

[0108] m i (A) = 0, for all other subsets For example:

[0109] If w vis =0.7, p vis =[0.6,0.1,0.05,0.15,0.1], then: m vis ({ship}) = 0.7 × 0.6 = 0.42; m vis ({personnel}) = 0.07.

[0110] m vis (Θ)=1-(0.42+0.07+0.035+0.105+0.07)=0.3.

[0111] By using evidence synthesis rules, the basic reliability assignments from multiple sensors are combined and calculated to obtain the overall reliability of the proposition.

[0112] A recursive Dempster synthesis rule is adopted. Let the first k sensors have been fused to obtain the synthesized BPAm. (k) Now add the BPAm of the (k+1)th sensor. k+1 Then the newly synthesized BPAm (k+1) for:

[0113] The normalization constant (conflict metric) is:

[0114]

[0115] Initial state: m (1) =m1. After fusing all m sensors, the final BPAm is obtained. final For each category θ j Its overall reliability is calculated using a confidence function.

[0116] Since this scheme only assigns non-zero confidence scores to single-point sets and the entire set, therefore: the overall confidence score (θ) j )=m final ({θ j The proposition with the highest overall reliability is selected as the fusion judgment result for the stated objective. The fusion judgment category is: The corresponding overall confidence level is This value is included in the fused target information and is used for subsequent event judgment.

[0117] Specifically, when a conflict between pieces of evidence is detected during evidence synthesis calculations exceeding a preset threshold, the automatic fusion decision is paused and a manual review alarm is triggered. During each synthesis, K is calculated. (k+1) If K (k+1) >K th K th =1-10 -6 If the value is 0.999999, it is considered a severe conflict. The system will then perform the following actions:

[0118] First, the fusion process for the current target is aborted; then, no category determination is output; finally, a manual review alarm event is generated, including: the target's spatiotemporal location (converted WGS-84 coordinates), a list of involved sensors and their original observations (including p... i w i (Environmental parameters) Conflict value K (k+1) The timestamp is used. The conflict between pieces of evidence is measured by a normalization constant in the evidence synthesis rules, and the preset threshold is configured as a critical value for identifying serious conflicts between pieces of evidence.

[0119] The physical meaning of the normalization constant K is: the sum of the confidence products of all mutually exclusive hypothesis combinations. For example, if one sensor has a high confidence level of identifying a ship m1({ship}) = 0.9, and another has a high confidence level of identifying personnel m2({personnel}) = 0.85, then K ≈ 0.9 × 0.85 = 0.765. Although this does not reach the threshold, it already indicates a potential anomaly. If both confidence levels are close to 1, then K → 1, triggering an alarm.

[0120] Preset threshold K th =0.999999 was set after statistically analyzing the critical point of false alarms caused by misfusion in 10 fishing ports and a total of 2,000 hours of operation data. This ensures that manual handling is only required in cases of irreconcilable contradictions, balancing automation efficiency and safety.

[0121] S5: Perform cross-modal identity verification and association on the fused target information to construct unified comprehensive situational information for fishing ports. The fused target information refers to the output structured data object, including the target category (e.g., vessel), location in a unified geographic coordinate system (X...). w ,Y w The data includes the timestamp t, the overall confidence value, and, if it is a visual target, a 128-dimensional appearance feature vector.

[0122] Cross-modal identity verification and association refers to the consistency verification and binding of target observation results that originate from different perception modalities (i.e., visual image data and radio frequency positioning data) but point to the same physical entity at the identity level.

[0123] Unified comprehensive situation information of fishing ports refers to a dynamic digital map covering the entire fishing port supervision area. Each identified target is marked with a unique identifier (such as vessel name or personnel ID), real-time location, movement status and credibility label. The map is based on electronic nautical charts or fishing port plans and is continuously updated through a data fusion and processing center.

[0124] This step is only performed on merged targets categorized as ships. Since targets such as personnel and vehicles are not required to be equipped with AIS devices in the current system, radio frequency identity binding is not performed.

[0125] Specifically, the cross-modal identity verification and association includes: performing spatiotemporal association matching between ship targets identified based on visual image data and ship identity and location information obtained based on radio frequency positioning data. Ship targets identified based on visual image data refer to targets detected by the improved YOLOv5 model and confirmed as ship categories through fusion, and include the following fields: visual ship position p v =(X v ,Y v )∈R 2 The unit is meters, and it has been converted to the WGS-84 geographic coordinate system. Visual ship timestamp t v The unit is seconds, assigned by the PTP time synchronization network, with an accuracy down to the microsecond level. Appearance feature vector f v ∈R 128 Extracted from the ReID branch, it is used for subsequent appearance similarity calculation.

[0126] The vessel identification and location information obtained based on radio frequency positioning data refers to the structured data generated after the vessel message is parsed by the AIS receiver and processed by protocol parsing. This data includes: the vessel's MMSI code (Maritime Mobility Service Identifier), a 9-digit decimal number, globally unique; the vessel name (ASCII string); and the AIS vessel location p. a =(X a ,Y a )∈R 2 The latitude and longitude coordinates in the AIS message are converted from WGS-84 to UTM projection; the AIS timestamp t a The time is determined by the time field in the AIS message or the receiving time after PTP synchronization calibration.

[0127] Spatiotemporal correlation matching performs the following two-step filtering: Time window filtering: only retain those that satisfy |t v -t a |≤ΔT for AIS target, where ΔT=10 seconds is the preset maximum allowable time deviation, used to tolerate the asynchrony between low frequency of AIS messages (typically 1–10 seconds / message) and high frame rate of video (30Hz);

[0128] Spatial distance filtering: Based on temporal matching, calculate the Euclidean distance d = ||p v -p a ||2, only retain those satisfying d≤D th The AIS targets were selected as candidates, among which S th =50 meters is the preset spatial correlation threshold, which is set according to the density of fishing vessels and the statistics of positioning error in the fishing port.

[0129] The output is a set of one or more candidate target pairs:

[0130] Where K ≥ 0. For successfully matched candidate target pairs, a motion trajectory consistency check is performed. A successfully matched candidate target pair refers to... The non-empty element pairs mean that at least one AIS target passes the spatiotemporal screening.

[0131] Motion trajectory consistency verification is used to determine whether the motion behavior of visual vessels and AIS vessels is consistent over a period of time, in order to eliminate mismatches caused by accidental position overlap.

[0132] The specific process is as follows: Set the verification time window length T = 60 seconds; extract the historical trajectory of the visual vessel within the past T seconds from the data cache. Where N v To determine the effective visual frame count; simultaneously extract the historical message location sequence of the corresponding AIS vessel within the same time period.

[0133] Motion vectors are calculated for the two trajectories separately: Visual motion vector sequence: The velocity vector representing the displacement of the i-th segment; AIS motion vector sequence: To align the time scales, linear interpolation is used to resample the AIS motion vectors to the time points of the visual trajectory, resulting in aligned AIS motion vectors.

[0134] Calculate motion consistency similarity S mot :

[0135] Where N is the number of effective alignment segments, and the minimum value of N is taken. v -1,N a -1); the numerator is the dot product of the two velocity vectors, reflecting the consistency of direction; the denominator is the product of their respective moduli, which is normalized so that each term takes a value in [-1,1]; the whole is the time-series average of the cosine similarity, and the closer the value is to 1, the more consistent the motion direction is.

[0136] If S mot>S th S th A threshold of 0.7 (calibrated through real-ship tracking experiments) indicates that the motion trajectory consistency check has passed. If the motion trajectory consistency check passes, the ship identity information in the radio frequency positioning data is bound to the corresponding visual ship target.

[0137] The vessel identity information in radio frequency positioning data specifically refers to the vessel name and MMSI code contained in the AIS message, which together constitute the vessel's legal identity. The binding operation refers to adding the following fields to the data structure of the fused target information: identity_type = AIS; mmsi = [9-digit number]; ship_name = [string]; association_confidence = S_{\text{mot}} (optional, used to record verification strength).

[0138] Once the binding is complete, the visual vessel target acquires a traceable and searchable legal identity, can be displayed with a vessel name tag on the integrated situation map of the fishing port, and supports accurate evidence collection for subsequent violations (such as failure to activate AIS or intrusion into restricted areas). If multiple candidate AIS targets pass the motion consistency check, an appearance similarity check is further introduced as an auxiliary criterion.

[0139] Using the 128-dimensional appearance feature vector f of the visual target v Feature templates of MMSI vessels in the historical database f template Calculate cosine similarity:

[0140] Select S mot +γ·S app The largest candidate pair is bound, where γ = 0.3 is the appearance weight coefficient. If no candidate passes the verification, the vessel target is retained as unidentified, marked as a suspected vessel without AIS enabled, and a special alarm is triggered.

[0141] S6: Based on predefined complex event pattern rules, perform real-time matching analysis on the comprehensive situation information of the fishing port, and generate alarm information when a matching event is detected.

[0142] Predefined complex event pattern rules refer to a set of high-risk behavior detection templates expressed in formal logic, stored in the rule engine configuration file of the intelligent early warning module, and described using an event flow processing language (such as EPL) or a state machine model. The comprehensive situational information of the fishing port refers to the output unified dynamic map data stream, containing the real-time identity, location, speed, category, and status of all identified targets, continuously input to the intelligent early warning module in the form of a time-ordered structured message sequence.

[0143] Real-time matching analysis is performed by the ComplexEvent Processing Engine (CEP Engine) deployed in the data fusion and processing center. This engine performs a sliding window scan on the situational information stream to detect whether there is an atomic event sequence that satisfies any complex event pattern rule.

[0144] Generating alarm information means that when a match is successful, the system automatically constructs a structured alarm record containing the event type, occurrence time, geographical location, identity of the target involved, and evidence chain index (such as video clip ID, sensor list), and notifies regulatory personnel through audible and visual alarms, platform pop-ups, SMS push notifications, and other means.

[0145] Specifically, the complex event pattern rules include a personnel falling into water event pattern, which is configured to detect a personnel target moving from a first preset geographical area to an adjacent second preset geographical area. The personnel target refers to a target whose category is determined by fusion and is confirmed to have a valid UWB positioning tag or visual tracking trajectory; its identity can be a staff ID or an anonymous identifier.

[0146] The first pre-defined geographical region is denoted as These are pre-defined dock operation areas or dangerous edge areas on the shore in the electronic map of the fishing port, represented as polygonal regions in the WGS-84 coordinate system, for example, by a set of vertices. definition.

[0147] The second pre-defined geographical region is denoted as It is adjacent The nearshore waters area is typically a strip of water 0–20 meters from the edge of the pier, also in polygonal shape. It indicates that, and satisfies

[0148] Move to the target person at time t enter First entry And its previous valid position (time t) prev <t enter )lie in Within, that is: and Where p(t) = (X(t), Y(t)) represents the uniform geographical coordinates of the person at time t. Furthermore, during the subsequent preset time period, the movement speed of the person within the second preset geographical area remains below a first speed threshold.

[0149] Subsequent preset time period refers to the period from entry The moment t enter Starting from [t], the time window extending backwards [t] enter ,tenter +T obs ], where T obs =30 seconds is the preset observation time, used to eliminate false alarms from brief water entry (such as diving operations).

[0150] Movement speed refers to the speed at which the target is located. The instantaneous velocity amplitude within is calculated as follows: Among them, t i ∈[t enter ,t enter +T obs ] represents the effective location time point, and ||·||2 represents the Euclidean norm.

[0151] The condition remains below a certain value for all valid velocity sampling points within that time period, satisfying the condition: v(t) i ) <V th V th =0.3m / s is the first speed threshold, which is based on the statistical setting of typical human motion speed when floating or struggling in water (normal swimming speed >0.5m / s, static floating ≈0m / s). If at T obs Any v(t) appears in the inner i )≥V th If the condition is not met, the event matching terminates. Simultaneously, no vessel target was detected within the second preset geographical area at a distance less than the safe distance threshold from the personnel target.

[0152] A vessel target refers to a target that has completed identity binding or has at least been confirmed as a vessel category through fusion. Within the second preset geographic area, it refers to the location of the vessel target. The distance to the personnel target refers to T obs The Euclidean distance between the ship and the personnel at any time t within the time period: d sp (t)=||p ship (t)-p person (t)||2; the safety distance threshold is denoted as D. safe =10 meters, this value is set based on the lifesaving response capability and the safe handling radius of the vessel.

[0153] Not detected... less than refers to the entire T obs During the period, all ship targets For the current The collection of ships within the collection all satisfy the following: This means that no vessel comes within 10 meters of the person, ruling out normal water entry scenarios caused by rescue efforts or vessel operations. When all three conditions above are met, the CEP engine determines that a person has fallen into the water event. The system immediately generates the highest-level alarm information, including: Event Type: Person Falling into the Water; Time of Occurrence: t enter Position coordinates: p(t) enter The system identifies the person involved: UWB tag ID or visual tracking ID; evidence chain: associated thermal imaging, visible light video clips (uploaded as needed by the edge layer), UWB trajectory logs, and environmental parameters (such as visibility and illumination); alarm level: emergency (requires immediate response). Simultaneously, the system automatically highlights the person on the integrated situation map and initiates emergency response procedures (such as notifying rescue boats and retrieving nearby cameras).

[0154] Building upon the aforementioned method for monitoring fishing ports based on multi-source data fusion, this invention further proposes a fishing port monitoring system based on multi-source data fusion, comprising: a perception layer, including optical image sensors, radio frequency positioning signal receivers, and environmental parameter sensors, as well as a time synchronization network device providing a unified time reference for all sensors. The perception layer is the bottom layer of the system, composed of heterogeneous sensor nodes deployed at key locations in the fishing port (such as the wharf front, high points, and channel bends), responsible for collecting raw physical quantities. Optical image sensors refer to imaging devices used to acquire visible light or infrared band image information. Radio frequency positioning signal receivers refer to receiving devices used to receive radio signals emitted by ships or personnel and calculate their identity and location. Environmental parameter sensors refer to sensing devices used to monitor meteorological and hydrological variables affecting perception performance, typically six-element meteorological and hydrological stations, outputting data including illuminance, visibility, rainfall intensity, wind speed, humidity, and auxiliary temperature difference.

[0155] Time synchronization network equipment refers to a master-slave time synchronization network built on the Precision Time Protocol (PTP, IEEE 1588). It includes one master clock (Grandmaster Clock) and multiple slave clocks. The master clock receives the BDS B1I frequency signal from the BeiDou Navigation Satellite System, serving as the sole time reference source for the entire system. All sensor nodes have built-in slave clocks and connect to the network via industrial Ethernet to achieve microsecond-level time synchronization.

[0156] The data fusion and processing center, communicatively connected to the sensing layer, is configured to perform the following operations: convert sensing data from different sensors to a unified geographic coordinate system and time reference. The data fusion and processing center is the core computing unit of the system, deployed in the fishing port monitoring room or cloud platform, and connected in real-time to each sensor in the sensing layer via wired or wireless communication links. The unified geographic coordinate system refers to the WGS-84 geodetic coordinate system. If internal processing requires planar coordinates, a transverse Mercator projection (such as UTM Zone 50N) is used as an equivalent expression. The unified time reference refers to the PTP synchronization timestamp provided by the time synchronization network equipment; all sensing data is timestamped at the time of acquisition.

[0157] Specifically, this operation includes: converting the pixel coordinates output by the optical image sensor into world coordinates using a pre-calibrated camera intrinsic matrix, distortion coefficients, and extrinsic matrix through distortion correction and perspective back-projection models; mapping the raw location output by the RF positioning signal receiver (such as AIS latitude and longitude, UWB local coordinates) to WGS-84 using a coordinate system transformation algorithm; and binding the fixed installation location of the environmental parameter sensor as metadata to each uploaded data packet.

[0158] Based on the type of each sensor and the real-time parameters of its environment, the dynamic confidence weight of each sensor data is calculated during data fusion. The sensor type refers to three categories: optical image sensors (including visible light and thermal imaging), radio frequency positioning signal receivers (including AIS and UWB), and environmental parameter sensors, each using a different confidence model. The real-time parameters of the environment refer to values ​​uploaded in real-time by the environmental parameter sensors, such as illuminance L (lux), visibility V (km), rainfall intensity I (mm / h), and target-background temperature difference ΔT (°C). The dynamic confidence weight is a real number in the interval [0,1], reflecting the reliability of the sensor's data at the current moment. Its calculation method varies depending on the sensor type: for visible light cameras, C0 is used. vis (L,V,I)=β vis ·σ(α L ·L+α V ·V-α R •I); For thermal imaging cameras, C is used ir (ΔT)=β ir ·min(1,ΔT / T th For AIS receivers, the value is typically 0.99 or 0.1, determined by message integrity and location jump; for UWB base stations, the value is typically 0.95 to 0.3, a linear variation, determined by positioning calculation residuals.

[0159] Dynamic confidence weights serve as fundamental input parameters in the evidence fusion stage. Based on these dynamic confidence weights, an evidence reasoning method is employed to fuse multi-source target observation information, outputting fused target information with a comprehensive confidence level. The evidence reasoning method specifically refers to the Dempster-Shafer evidence theory, which includes: constructing a basic confidence assignment (BPA) for the identification framework (Θ = ship, personnel, vehicle, floating object, unknown); generating single-point proposition confidence levels using dynamic confidence weights and the category probabilities output by sensors; and combining multi-source evidence through recursive Dempster synthesis rules.

[0160] Multi-source target observation information refers to the extracted structured data, including the category, location, confidence level, and appearance features of visual targets, as well as the identity, location, and motion status of radio frequency targets. The fused target information is a structured object containing target category, unified coordinates, timestamp, and comprehensive confidence value (i.e., the maximum single-point confidence in the synthesized BPA). This operation is performed in the evidence fusion decision module of the data fusion and processing center. Cross-modal identity association and verification are performed on the fused target information, and comprehensive situational information of the fishing port is generated.

[0161] Cross-modal identity association and verification refers to performing spatiotemporal matching, motion trajectory consistency verification, and optional appearance similarity comparison on the visual observation and AIS message of the fused target of the category of ship, and finally binding the ship name and MMSI code in AIS to the visual target.

[0162] The comprehensive situational information of a fishing port refers to a dynamically updated digital map, using an electronic nautical chart or fishing port plan as a base map, overlaying the identity tags, location icons, motion vectors, and status indicators of all identified targets to form a globally unified view. This operation is completed by the identity verification and situational construction module and continuously pushed to the intelligent application and display layer. The intelligent early warning module, communicatively connected to the data fusion and processing center, is configured to load predefined complex event pattern rules, perform real-time analysis of the comprehensive situational information of the fishing port, and trigger alarms.

[0163] The intelligent early warning module is an independent software component that runs in the data fusion and processing center or on a dedicated server, establishing a subscription relationship with the situational information output interface. Predefined complex event pattern rules are stored in configuration files, including formalized logic templates for events such as people falling into the water, vessels illegally entering restricted areas, and vessels stranded without AIS activated. Real-time analysis refers to the CEP engine performing sliding window matching on the situational information stream to detect whether atomic event sequences meet the conditions for composite events. Alarm triggering involves generating structured alarm records and notifying regulatory personnel through the platform interface, audio-visual equipment, mobile terminals, and other channels.

[0164] Specifically, the optical image sensors in the perception layer include visible light cameras and thermal imaging cameras. The radio frequency positioning signal receivers include automatic identification system (AIS) receivers and ultra-wideband (UWB) positioning base stations. Visible light cameras refer to color imaging devices operating in the 400–700 nm visible light band, used to acquire high-resolution images under high-illuminance conditions during the day. Thermal imaging cameras refer to uncooled focal plane array cameras operating in the 8–14 μm long-wave infrared band, imaging by detecting differences in target thermal radiation, suitable for low-visibility scenarios such as nighttime and foggy weather. Automatic identification system (AIS) receivers refer to AIS Class B or shore-based receivers conforming to the ITU-R-RM.1371 standard, passively receiving ship broadcast messages in the VHF band (161.975 MHz, 162.025 MHz). UWB positioning base stations refer to UWB pulse radio base stations based on the IEEE 802.15.4a / z standard, calculating the three-dimensional position of personnel wearing UWB tags by measuring the time-of-flight (ToF) of nanosecond-level pulse signals, with a positioning accuracy better than 0.5 meters.

[0165] Specifically, a fishing port monitoring system based on multi-source data fusion also includes an edge processing layer, which includes an edge computing unit deployed near the optical image sensor for real-time target detection and tracking analysis of the video stream acquired by the optical image sensor, and uploading the structured analysis results to the data fusion and processing center.

[0166] The edge processing layer is an intermediate computing layer located between the perception layer and the data fusion and processing center, used to offload the video processing load from the central server. Edge computing units refer to embedded AI inference devices deployed in camera poles or cabinets, with the Huawei Atlas500 smart station being a typical example, which has GPU and NPU acceleration capabilities.

[0167] Real-time target detection and tracking analysis refers to running an improved YOLOv5 model at the edge, performing target detection on each frame of video, and combining ReID features to achieve cross-frame target tracking, outputting structured results including category, bounding box, confidence, and 128-dimensional appearance feature vector.

[0168] The structured analysis results are encapsulated in JSON format and uploaded to the data fusion and processing center via the MQTT protocol. The original video stream is not uploaded; segments are only transmitted as needed when an alarm is triggered or manual review is required, thus reducing network bandwidth consumption.

[0169] In one specific embodiment, when fog blankets a fishing port at night, reducing visibility, the system automatically enters all-weather monitoring mode. The perception layer simultaneously collects multi-source data: visible light images have a low signal-to-noise ratio due to fog; thermal imaging clearly captures personnel targets based on the temperature difference between humans and the sea; AIS receives vessel identification and location; UWB locates personnel tags; and environmental sensors report illumination, visibility, and temperature difference. All data is time-synchronized via a PTP network and stamped with a unified microsecond-level timestamp.

[0170] The data fusion center unifies the data from various sources into the WGS-84 coordinate system and uses thermal imaging frames as the time reference to align AIS, UWB, and other data within the tolerance window, with missing values ​​linearly interpolated.

[0171] The edge computing unit runs an improved YOLOv5 model (embedded with a coordinate attention module) to detect personnel targets with a confidence level of 0.85 from thermal imaging and extract 128-dimensional feature vectors; AIS parses nearby ship information; UWB calculates the wearer's three-dimensional coordinates and ID; and environmental parameters are extracted simultaneously.

[0172] The system dynamically allocates sensor weights based on the environment (thermal imaging 0.8, UWB 0.95, AIS 0.99, visible light 0.2), and fuses multi-source identification results based on Dempster evidence synthesis rules. The final comprehensive reliability of the human-proposed questions reaches 0.98, and the questions are identified as the target category.

[0173] Then cross-modal verification is performed: the spatiotemporal distance between the thermal imaging target and the UWB positioning point is less than 5 meters, and the cosine similarity of the historical trajectory direction exceeds the threshold. The match is successful, the worker's identity is bound to the visual target, and the information is updated to the comprehensive situation map of the fishing port.

[0174] The intelligent early warning module matches event rules in real time and detects that after a person enters the nearby waters from the dock, their movement speed remains below 0.3 m / s for 30 seconds, and there are no vessels within 10 meters, triggering a person falling into the water event. The system immediately activates emergency response through large-screen pop-ups, audible and visual alarms, and SMS push notifications.

[0175] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for monitoring fishing ports based on multi-source data fusion, characterized in that, include: Simultaneously collect sensing data from multiple heterogeneous sensors, including at least image data collected by an optical image sensor, positioning and identity data collected by a radio frequency positioning signal receiver, and environmental parameter data collected by an environmental parameter sensor. The sensing data collected by the various heterogeneous sensors are uniformly converted to the same spatiotemporal reference. The sensed data is processed and analyzed to extract target observation information from the perspectives of each sensor; Based on the dynamic confidence weights of each sensor, an evidence reasoning method is used to fuse observation information about the same target from different sensors to generate fused target information with comprehensive confidence. Cross-modal identity verification and association are performed on the fused target information to construct unified comprehensive situational information of fishing ports; Based on predefined complex event pattern rules, the comprehensive situation information of the fishing port is matched and analyzed in real time, and alarm information is generated when a matching event is detected.

2. The fishing port supervision method based on multi-source data fusion according to claim 1, characterized in that, Optical image sensors include visible light cameras and thermal imaging cameras, and the determination of the dynamic confidence weights includes: For the visible light camera, an evaluation factor for its data reliability is dynamically calculated based on at least one of the lighting conditions, visibility conditions, and rainfall intensity in the real-time acquisition environment. For the thermal imaging camera, an evaluation factor for its data reliability is dynamically calculated based on the temperature difference between the target and the background; Based on the evaluation factors and in conjunction with the data continuity indicators of the corresponding sensors in the recent period, their weight in the fusion decision is determined.

3. The method for monitoring fishing ports based on multi-source data fusion according to claim 1, characterized in that, The method of fusion using evidence reasoning specifically includes: The target recognition framework is defined as Θ = {ships, personnel, vehicles, floating objects, unknown}; Based on the dynamic confidence weights of each sensor and the target recognition confidence output by each sensor, a basic confidence assignment is constructed for the proposition that the target belongs to a specific category; By using evidence synthesis rules, the basic reliability assignments from multiple sensors are combined and calculated to obtain the overall reliability of the proposition. The proposition with the highest overall reliability is selected as the fusion judgment result for the stated objective.

4. The fishing port supervision method based on multi-source data fusion according to claim 1, characterized in that, The cross-modal identity verification and association specifically includes: The ship targets identified based on optical image data are spatiotemporally correlated and matched with the ship identity and location information obtained based on radio frequency positioning data. The motion trajectory consistency of the successfully matched candidate target pairs is verified. If the verification passes, the ship identity information in the radio frequency positioning data is bound to the corresponding visual ship target. The personnel targets identified based on optical image data are spatiotemporally correlated and matched with radio frequency positioning signals. The motion trajectory consistency of the successfully matched candidate target pairs is verified. If the verification passes, the personnel identity information in the UWB tag is bound to the corresponding visual personnel target.

5. A method for monitoring fishing ports based on multi-source data fusion according to claim 1, characterized in that, The complex event pattern rules include a person falling into the water event pattern, which is configured as follows: A person target is detected moving from a first preset geographical area to an adjacent second preset geographical area, and during a subsequent preset time period, the moving speed of the person target in the second preset geographical area remains below a first speed threshold, while no ship target is detected within the second preset geographical area at a distance less than a safe distance threshold from the person target.

6. A method for monitoring fishing ports based on multi-source data fusion according to claim 1, characterized in that, The processing of visual image data includes: using a target detection neural network model with spatial attention enhancement to analyze the images in order to improve the ability to detect targets against complex sea backgrounds.

7. A method for monitoring fishing ports based on multi-source data fusion according to claim 3, characterized in that, When a conflict between pieces of evidence is detected during evidence synthesis calculation that exceeds a preset threshold, the automatic fusion decision is paused and a manual review alarm is triggered. The conflict between pieces of evidence is measured by a normalization constant in the evidence synthesis rules, and the preset threshold is configured as a critical value for identifying serious conflicts between pieces of evidence.

8. A fishing port supervision system based on multi-source data fusion, characterized in that, include: The perception layer includes optical image sensors, radio frequency positioning signal receivers, and environmental parameter sensors, as well as time synchronization network devices that provide a unified time reference for all sensors; The data fusion and processing center is communicatively connected to the perception layer and configured to perform the following operations: Transform sensor data from different sensors into a unified geographic coordinate system and time reference. Based on the type of each sensor and the real-time parameters of its environment, the dynamic confidence weight of each sensor data is calculated during fusion. Based on the dynamic confidence weight, the evidence reasoning method is used to fuse multi-source target observation information and output fused target information with comprehensive confidence. Cross-modal identity association and verification are performed on the fused target information, and comprehensive situational information of the fishing port is generated; The intelligent early warning module is communicatively connected to the data fusion and processing center, and is configured to load predefined complex event pattern rules to perform real-time analysis of the comprehensive situation information of the fishing port and trigger alarms.

9. A fishing port supervision system based on multi-source data fusion according to claim 8, characterized in that, The optical image sensors in the sensing layer include a visible light camera and a thermal imaging camera; The radio frequency positioning signal receiver includes an Automatic Identification System (AIS) receiver and an ultra-wideband positioning base station.

10. A fishing port supervision system based on multi-source data fusion according to claim 8, characterized in that, It also includes an edge processing layer, which includes edge computing units deployed near the optical image sensor for real-time target detection and tracking analysis of the video stream acquired by the optical image sensor, and uploading the structured analysis results to the data fusion and processing center.